Use a local model to produce thousands of pages worth of fake math that constantly states “I have solved the x conjecture” and methodically pump it into chat over months maybe?
I don't usually defend Apple products, but my AirPods Pro 2nd gen have been rock solid and taken insane abuse. I use them working on cars, runs, hiking, rain, shine, sauna. I drop them. They land in puddles. They smack concrete every couple of months. Sample size 1, but they are probably one of my favorite accessories and things I lug around with me. I use them a lot and if not love /really/ like them and nothing else has come close to the frictionless experience of them.
Although that may be in part due to vendor lock in and not sharing some of their connection handling API / tech.
If you are okay with waiting use GLM 5.3 max. It costs more but still cheap. It is slow, but a very strong worker. Still dollars per day (at most) with heavy concurrent agent running. I load up planning and tasks in Opus or Sol, and just have glm flash workers go to town every night. My project has never advanced more smoothly.
Which versions of flash and at what thinking levels? Which chinese flash models and at what thinking levels? What tasks? What completion rates? How was quality evaluated?
- Which versions: 3.6 vs 3.7 vs. 3.8 for Gemini Flash, and v4 0731 for Deepseek v4 Flash, and GLM 5.3 Flash
- Medium for Gemini, high for Deepseek.
- Things like find information, then understand something about it, then send a slack message or email etc.
- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini
- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.
Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.
Thanks... I have been trying to figure out some things. Been doing my own evals. Flash 3.8 does burn a lot more tokens on high. Interesting how smart and not smart it is. For personal use almost impossible to justify the cost of 3.8 Flash cost.
Deepseek also burns a lot of tokens, its output on high is 2x of Gemini on medium. But it's dirt-cheap so it still can be 60-70% cheaper.
From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.
All this really needs evals, the token prices tell nothing.
I am surprised at how well DSv4 flash does in the real world vs many benchmarks. You look at Flash 3.8 and it supposedly beats opus 5 and deepseek is far below.. but they were measuring efficiency, whatever that is…
Something doesn’t add up for me on the published benches
For what it’s worth, I am very happy with Jellyfin and the *arr suite. It took a bit of agent prodding to get them all playing nicely together and bypassing Cloudflare CAPTCHAs. However, it's pretty sweet when you get it all working.
With the *arr stack, some of the trackers it uses to search for a torrent will present a captcha when Prowlerr queries it for a search. If you integrate FlareSolverr and set up Prowlerr to submit queries through it for the trackers that have CAPTCHAs, it won't get blocked by them.
It does seem like for new code that might help. There's some really good logic and wisdom in it, but it has to be applied very contextually to the exact problem you are trying to solve. If an agent is navigating a complex codebase, this could definitely send them off on a refactoring rabbit hole. However, if you have them writing some new code, it could prevent their tendency to yak shave and write new things. So I can see some situational uses for this, but it could get out of hand as well.
Sol is my current favorite model to interact with. So much less BS than Opus 5. Fable 5.1 is okay as is Fable 5 but it has Opus like tendencies. Sol is very good at following instructions and remembering them for a session.
I still hold the line on interviewing. Maybe some don’t. I am sure it is true. My team us too small with too much responsibility to tolerate mediocrity to any real extent
Fable is okay, just slower, eats tokens and not any better at coding tasks. Maybe a little better, but not better enough. It's a lot faster to have a cheap and fast flash agent / sonnet do the implementation work with Fable tagging cleanup and divergence from spec and goals.
Flash 3.8 is genuinely my favorite all around model right now. And yeah Opus 4.6 was the last Opus model I liked. 4.8 is tolerable. Opus 5 is a terrorist. It just can't follow an instruction to save its life and regresses rapidly. Sol at least stays on track so I have to smack it's hand way less often. I am biased, but Flash 3.8 and 3.7 are the first Gemini models I just recommend to others.
Amusingly, as an autonomous coding agent, I kind of like Opus 5. But I have to bound it on tasks or it just goes off the rails. But I'm bounded tasks, it is genuinely solid. It's kind of like the new Sonnet 5. Right now my favorite model to interact with on the frontier side is Sol 5.6 so I have been using that as my coordinator. Flash 3.8 is my other favorite just because it is so fast and I use it a lot at work and know its quirks.
reply