My expertise lies in deep learning theory, and yes, the "intelligence" is coming primarily from scaling up, among other things. There are good reasons for this, but essentially it comes down to taking advantage of a narrow statistical trick, where a very well-crafted model/optimizer pair that has a strong implicit bias toward simplicity can exhibit progressively increasing performance with respect to model size. Marcus Hutter's lab has shown that you can phrase this in terms of Solomonoff induction, so this bias is truly universally effective. An effective bias can continue to improve performance with larger model sizes by taking advantage of the curse of dimensionality in a way not dissimilar to how more data generally gives you a better answer (indeed, there is a duality taking place here, but I digress).
To be clear, it is an extremely narrow model class that can do this; we just got "lucky" and worked our way to it. That's why we still teach general statistical principles which often forbid this sort of behavior as a rule of thumb.
Agreed, AI is not capable at the moment of coming up with radical ideas to solve the tough problems. Sadly, I would argue many problems in math are likely to be found to be not actually tough in this sense, and those working in "comfortable" areas with fewer tough problems are having real crises of their own right now.
But even for the tough problems, it is good at executing on a particular idea with reasonable competency. It's also quite decent at verification now. That can radically speed up proof development overall, since those aspects can become quite tedious otherwise.
Agreed. The benchmark closest to my experience is FrontierMath Tier 4. Fable and Sol (90%) are very far ahead of Kimi K3 (not even 40%). Kimi is trained heavily to basic agentic tasks, like all the other open models right now.
I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here to stay in one form or another.
The fear is not about the models open weights it is the erosion of training capability in other countries. Why train models when they do it for free? Until they don't of course, or they start doing what the US is doing right now by locking out some models to government only or internal market only.
While a valid point, China also produces plenty of whitepapers going about the architecture and know how about the training and inference itself.
There’s also the fact that unless LLMs do get to AGI (which seems… doubtful, still) there comes a point where a model is good enough for what you need. Fable and gpt 5.6 are certainly pretty neat, but I’ve been happy since opus 4.6. I’d still choose a better model, obviously, but it’s not the end of the world if I was stuck with 4.6 for a while when it already lets me get the end result at acceptable quality.
It also needs to be said that the "erosion of training capability in other countries" is largely theoretical, given that Mistral hasn’t been keeping up and other countries don’t even have anything worth mentioning. You’d first need to _have_ training capability to lose it.
How exactly do you plan to pull a rug that's in my basement? The only people who are in a position to pull rugs are closed-model vendors.
And if a nation-state or other entity can't train a model that outperforms the open-weight SotA in a given respect, then they shouldn't waste electricity trying. A more-enlightened civilization would join forces and make the combined result available to all.
China has been known to set up local industry, destroy competition through subsidies, jack up prices repeatedly. US does it all the time too with tech services (uber, airbnb, are the more notorious, but all big tech is doing it now), but China is better at capex which is why they seem to be winning this race.
The main difference is that the subsidies in China usually come from the gov, while in the US it comes from VC money or anti-competition practices from established big-tech companies.
If China were setting up international funds and institutions for training with participation from other countries I would be 100% on board. Other countries could provide funding and workforce too and have a say in how the models are trained and safe-guarded. I am not saying China should bear the burden of open-weight models alone.
Ideally there would be open weight models from multiple geopolitical areas. It is not that different from telecom really, you don't want the whole world to be dependent on a single provider from a single country on this kind of stuff.
The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model, provided they continue to allow people to use it. It will be genuinely exciting when an open model is able to beat it.
I wouldn't call it a moat, but I would call it a noticeably better model. Subjectively, for my own work, I would rate the top models Fable > K3 > Sol.
But it's not like Fable is so substantially better than the other two that I would be seriously impacted if I didn't have access to it anymore. All three are amazing models, and of the three, Fable is the only one that regularly triggers refusals.
It really does depend on your application. In my domain (math research), it is substantially better. Fable can solve really hard tasks with surprising consistency. It makes mistakes, and occasionally refuses, but honestly, at the top level, ideas are the currency and the rigor is the busywork. The other models cannot come close in this domain.
If you couple Fable's idea factory with Sol's rigor, you get a real game-changer. It puts the emphasis on top-level ideas, and nearly trivialises the intermediate layers.
> saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model,
What's more interesting is that Anthropic moat shrunk to just that model. There's zero reason to use any other model from Anthropic right now. And once they take Fable off subscription there will be zero reason to have Anthropic subscription.
100% this. There's currently this [1] submission that hasn't gained much attention, but is really important. In this [2] incident report from HuggingFace, they talk about detecting an attack and not being able to analyse the logs / IoC with API models because of guardrails. If not even highly regarded reputable companies can't sort out access to SotA models for blue team use, the raw capabilities don't matter. They're useless paperweights (hah!), and nothing else. Having to resort to open models is insane!
"When we started the log analysis, we first used frontier models behind commercial APIs. This did not work [...] We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. [...] The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout [...]"
Tell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of cases that can rapidly speed up progress.
I dunno, I find Fable slops alot. Sol is my workhorse. Fable can be creative but isn't very good at doing work reliably (or without endlessly burning tokens).
Fable is still as dumb as a post. I ask it simple questions and it routinely gets things backwards, prioritises things that should be subordinate to others, etc.
An example: It just suggested that I shouldn't raise the price of my saas because it'd complicate the arithmetic if I did 0 -> $100k YT channel instead of sticking to $20 p/m.
They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for. You only see it when you ask really tough questions.
What's the point of Fable if we can't use it? I get to prompt it like 5 times on my subscription before it gets cut off, and even then I'm constantly fighting the insufferable safety classifier.
I'll switch to OpenAI soon because of this. I also can't wait for the day it becomes feasible to run these awesome open weight models on my own hardware.
That would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic being the sole arbiter of the forefront of this technology, because it is all too dangerous in the hands of anyone else.
I see no evidence that any ceo retains anything but the desire to capitalize on their marketplace of ideas for their own benefit. Like wolves inn sheep clothing, they'll put on any skin suit to convince people to keep giving them money and power.
And it has nothing to do with the individual, from what I can tell, 70% of the population placed in their position would become the same type of uberpath.
Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companies are still hanging in there.
Very confused by this comment. The older (poorer) parts of the ML literature focus on models with convex and (gradient-)Lipschitz objectives, but that's not representative of reality, not even close. Modern objectives for AI models are famously nonconvex (catastrophically, from the point of view of classical optimisation theory), and that's where the interesting research is.
I'd push back on this. Most of the core optimization techniques (eg, ADAM, stochastic gradient descent) are straight out of the convex optimization literature. Generally you need to use optimizers that work well on convex objectives because near minimizers, functions tend to be convex. (Proof by contradiction: a non-convex point has a strict descent direction.)
The fact that neural networks are highly nonconvex has encouraged a lot of research, but it's more of the kind aimed at resolving tension: these methods are probably good for convex functions, why do they continue to work for nonconvex problems, and are there tweaks we can make to improve them in that setting? It's not a lot of de novo theory; more standing on the shoulders of giants, etc etc.
No, I have to push back as well, sorry. It takes a very long time to get to the "near-minimizer" stage when training a neural network, and in practice, you never get there (see neural scaling law regimes). What you are saying is the viewpoint from 6-7 years ago. Things have changed.
The reasons why optimizers work well for neural networks in their highly nonconvex landscapes has absolutely nothing to do with their performance in convex landscapes. If that were true, everyone would be using Newton-CG. These optimizers were born in the convex optimization literature as a consequence of the genetic optimization nature of incremental publication (and because that was all we had), but their modern study is through the lens of implicit regularization (their preferences for certain minima) and their stepwise vs. continuous rates for feature learning in multilayer models.
This is completely new theory by the way, and requires painful reinvention of the field. It does not stand on the shoulders of convex optimization. The nonconvex setting is assuredly not a perturbation of the convex setting, and those that do continue to work on deep learning optimization from the convex optimization perspective are well behind the times.
It seems that we have two different stories here: in one, the new optimization theory represents a stark departure from the prior art, a sort of revolutionary new view of the understanding of optimization as applied to neural networks.
In the other story, the current understanding of optimization is a natural evolution of past work, where a new generation of researchers respond to social and technological changes, adapting and building on the work of the past, taking what's useful, downplaying the importance of some ideas, and inventing new language to describe concepts that seem most relevant to the current situation.
Both stories tell some of the truth. A revolution or evolution? Looking at the literature (eg the sibling comment here) shows that even today, convexity is used as an intuition pump for modern optimization techniques. But there are also new ideas that apply to the specific exigencies of neural nets, and downplayed ideas (eg convergence rates) that seem less relevant.
The optimizers are lifted from convex optimization, but the point above was that they are applied to highly non-convex problems. They work for finding local minima, but a lot of the deeper literature does not translate (e.g. the conjecture being discussed in this post).
I'll point out that "does not work" is not the same as "not as efficient" :) But it does seem the Adam paper had an error.
I think that Nesterov's first order method is the most efficient general first order algorithm on convex problems, so anything else is in some sense worse. (Edit: removed incorrect ADAM comment.)
Yours' "not as efficient" in [2] means that, sometimes, ADAM "does not work." Look at figure 2, ADAM literally does not work in the case of "true model."
Yes, apologies, I didn't read the articles you linked before posting this. I did update the comment.
I don't think this changes the point, which is that most optimization methods used in AI owe a substantial intellectual debt to convex optimization theory.
I love convex optimization and there are a few SciML projects I am on where I really need results from there. But in AI research with deep neural networks, it's become a liability, because people will just not let go. I'm getting tired of reviewing convex optimization theory papers in ML conferences that are still trying to wave away the obvious issues with their application to deep learning. It's harsh, but I do feel we can only start talking about an intellectual debt once that stops being the case.
Another intuition is that near a minimum you can Taylor expand the function and show that the higher order coefficients (past the square) are negligible.
It's not a matter of whether the theory "works"; it's a matter of whether one is asking the right questions. Convex optimization studies how quickly an optimizer can reach the optimum. In the non-convex case, there are many basins containing their own local minima. The more sensible questions there are "which basin is it likely to go into?" and "how do I steer it to go where I want?". Global convergence rates are largely irrelevant by comparison.
To be clear, it is an extremely narrow model class that can do this; we just got "lucky" and worked our way to it. That's why we still teach general statistical principles which often forbid this sort of behavior as a rule of thumb.
reply