I've been getting a ton done with Fable as the supervisor and astra as the implementer, with opus for adversarial reviews of the astra PRs. You can use terminal multiplexers with custom harnesses to allow Fable to start codex sessions and send instructions / read instructions / allow/deny actions. It's pretty cool!
codex has an option to expose itself as an MCP. You can also use something like OpenCodex to bring Anthropic models into Codex as any other selectable model.
So China's brand new recon satellite that was placed specifically in orbit over Iran just happened to self-destruct weeks after China was accused of giving Iran high-res satellite imagery that was used to kill American soldiers?
And the US just publicly told everyone they have weapons in space?
Small reminder that the US government rug-pulled Fable, not Dario. Lots of the safety guards that users find annoying/objectionable were the results of negotiations to get the model back online after the US government forced them to take it down.
Maybe Dario should have just "donated" $1M to Trump's inauguration fund like Altman, Meta, Amazon, Microsoft, Tim Cook, Elon, and Google. There's a reason they are the odd man out with this current Administration.
Maybe Dario shouldn’t have tried for regulatory capture. He was constantly on the news talking about how these models are so dangerous and that we need regulation to keep China from releasing open-source models without guardrails.
The US government didn't make the choices to release the worst version of Opus and label it 5.0, and then isolate portions of their subscribers to limited usage of Fable.
They may have been unfairly targeted by the US government, but they are doing more damage to themselves without government help as well.
Fable only being temporarily included in cheaper subscriptions was because anthropic is severely GPU constrained. They still are, and it impacts almost all of those unpopular decisions. They did announce from the beginning it was temporary.
Horrifying excuse, gpu constraint can be used by all of these companies to justify a shit user experience. If the user isn't properly weighed in their priorities, they have their priorities setup wrong.
Their 20$ tier currently isn't serving their best model, and they insulted their users by putting out an ill tested opus 5.0, which is the worst experience ive personally had using a model in probably 2 years(obviously adjusting for expectations at the time of release).
Yes, as a user you pick what works for you. But it is a reality for them that growth has been huge, and GPU manufacturing is bottlenecked.
People were very skeptical about how much investment most companies put into hardware/data centers two years ago, and anthropic was more conservative than OpenAI here, so it's potentially hurting them now.
(Opus is a separate story: it does seem to have improved in coding in my experience, most weirdness seems to be its human communication)
I set my env up so I can see the exact context used in CC CLI, and then once I get over about 40% ctx used I have it handoff to a new, fresh session. Nothing good comes from running above, say 60% of your context window. Coincidently, I usually have good results with CC. I never compact a session ever.
That assumes they release their models publicly. The future is leaning towards these labs air gapping their best stuff (Model 2, etc) and using it internally to snipe their competitors and charge insane amounts for monitored use in consulting environments.
You can't distill or catch up if you can't access the models. You'll basically have a situation where nation-states will need to try and steal the models Oceans 11 style.
> More importantly, what's the point of building monster-scale data centers on unprecedented amounts of debt when a more than good enough model runs on a GPU from a couple years ago?
Probably because the future "monster" models will be insane. 100T+ param models might be the type of things that can independently run a small business, which means anyone not using them is at a distinct disadvantage to their competitors.
The top model from 2025 looks silly compared to the top model of the first half of 2026. Do you feel like progress has stalled?
> The top model from 2025 looks silly compared to the top model of the first half of 2026. Do you feel like progress has stalled?
I do. Pre-training is where the industry saw the “emergent properties” of LLMs arise and, for a time, people thought you could just keep scaling up bigger and bigger models but then the incremental gains from doing this did plateau. Labs will still do bigger models (Bytedance has a 10T planned), but these are sparse architectures and they aren’t going to have mind-blowingly greater intelligence. Fable didn’t either.
Scaling pre-training tokens also flattened out. What is still delivering gains is scaling RL on verifiable tasks. But that’s not general intelligence - it’s fitting models to specific tasks, which ML has always been good at. More importantly, most tasks to which humans apply their intelligence don't have computationally verifiable answers.
they didn't plateau, the hardware was effectively saturated and it wasn't until later in 2026 that newer hardware came online to provide enough capacity to keep scaling up the models efficiently
The gains from increased parameter scaling are sublinear: there's no more hockey-stick improvement to be seen going in that direction. That doesn't mean some improvement isn't possible - it's just going to be increasingly not worth doing.
Also, I think the fact that small open-weights models are catching up to the frontier rather than the frontier rapidly pulling away is evidence of this. In fact, by far the most dramatic capability increase story over the past two years has been the gains made in the small-parameter regime.
One might think, "hey, this agentic coding thing was a pretty big deal!", but I think it's a bit of a distraction because models only recently became optimized for this specific use case. It's not like they suddenly gained so much general intelligence that they magically had the ability to use a coding harness. No, the labs started spinning up a bunch of RL environments and generating rewards over long-horizon trajectories of combining these tools. It's an excellent application of LLMs but care needs to be taken interpreting how much "progress" has been mae in terms of raw generalized capability.
Yet apparently people still dump unvetted LLM outputs onto their colleagues and expect them to thank them for the privilege. So it’s worth asking them what the consensus is of their work to find out if it’s the case.
That's the old way. The current way is nobody reviews anything. The LLM reviews it all and nobody even reads it, they just click approve, and merge blindly. I wish I were kidding.
Chinese open weight models are great for this turn, but American private models generate orders of magnitude more cashflow. This cashflow = investment in training future models. It's unclear how Chinese open weight companies are going to compete in future rounds if they can't raise the same capital for training runs.
The American business model is exceedingly efficient at building large businesses from zero. I wouldn't dismiss it as just a jobs creation thing.
It’s unclear where American labs future capital will come from. They pretty much exhausted private options at that point and it’s not clear how successful an ipo would be at the current time
Obviously to me, I express things from my point of view.
Raising too much from debt is a bit dangerous if you plan to go public relatively soon and don’t have a good story for it (I don’t believe they have one). You can continue raising from VCs, but at some point the valuation and dilution starts to become a real issue, and will make your ipo even more difficult. Their options are pretty much limited to raising money from hyperscalers (with required compute spending, so more circular funding), which is what they are doing, but you cannot do that infinitely without having a good story to tell Microsoft/Google/Amazon investors. The market is more skeptical than it was a few months ago, I’m not convinced you can do that for years to come
I think as a counter point we continue to see investment and buildout. What do you mean the market is more skeptical? Of course the market doesn’t really have an opinion per se and aren’t all of these companies growing in valuation, revenues, and profits? At least the public ones.
Kansas genuinely has a ton of paved and gravel roads all across the state. It's like a gigantic grid. Oddly high amount of street-view coverage too on google maps: https://maps.app.goo.gl/EyBFj9BNQa8nhUcE6
The Kansas landscape is very underrated. Just infinite sky in all directions.
I agree about that landscape. The state is not as flat as people think it is either. Very scenic. I worked the southwest corner of the state years ago and always loved topping a local high spot. We were surveying and out there we could set a flag a half mile away and sight off of that easily with plane table and alidade or theodolite. The atmospheric distortion would get you on a humid day or if you were shooting across a wheat field with all the transpiration but otherwise it was clear with minimal shimmer.
reply