Reminds me of letting an LLM generate a strong password in the first place. Very practical, very quick, no terminal commands needed! I'm sure some have done exactly this. But then you realise your password is always shared with your LLM provider, and it's not so random after all (unless it used tools).
Alternatively, what I tend to do after receiving generated code is a lot of asking "why?".
I've learned things I wouldn't otherwise have learned because I hadn't considered using the tools the LLM recommends. It's also a way to eliminate some hallucinating, given that critical questions are posed as unbiased as possible. For that, I also like to open a new chat with a different model and asking open-ended questions about a recommended tool I don't know much about, to double-check that the original LLM was likely correct in its recommendation in the first place.
Agent-related files should've been dot-files at least to minimize human error. In fact, I'd been discussing: shouldn't agent instructions exist on a per-developer basis, and thus be an IDE setting for each individual, removed from the repository itself? Agent instructions don't add content to the source code, after all. It's like additional READMEs, but we already have a README.
What .editorconfig has in common with CLAUDE.md and similar files, is that they should be shared among the team.
Individual guidelines can live in a dot-prefixed subdirectory or the users $HOME.
This is what they already provide.
So it's possible (but not necessary) to add personal pre-prompts, but the shared best-practises and, more importantly, learnings, can be comitted and pushed.
What's exposed on a web server is a different topic, and when that is your repo's root dir, you've done something wrong either way.
The way I see it, it's the incessant stimuli you get through these apps. There's just no point where the screen stays quiet, stimulus free. Moving elements capture our attention naturally, especially when the entire screen keeps moving. The endless scroll element just makes sure that the stimulus keeps getting renewed the moment you're done with it.
Instead, a scroll should give you a break before heading into the next video. I'm willing to bet this would help severely with addiction. People are then forced to reconsider whether they actually want to play the next video. "Done" should not always lead to "here's the next stimulus". That's what's addictive. The brain isn't made to break out of that loop easily.
The problem with using scrolling as a metric is it assumes satisfaction with the content presented and ignores the fact that, on certain apps, many scroll miles go into skipping around articles or ads or reposts to get to the content you want, and imo a punishment for seeking more content while also diluting said content with forced ads at an alarming ratio is not indicative of addiction but scarcity of satisfaction and engaging content
If you continue scrolling, that shows you think that the content presented is valuable enough that scrolling past some misses is worth it. A good scroll like TikTok carefully metes out the ads so that half of them you don't mind and the other half you enjoy. If you find a site whose scroll makes you feel this way, just stop scrolling and don't try to ban the concept for everyone else.
What I'm about to say is going to sound very 'layman', and this is coming from someone who's been building UI's for like, 20 years.
This discussion makes me both laugh and feel sad, because we all know this is bad for us and gives us zero ROI for our time... and yet there's a whole thread developing here to justify the pattern.
I don't think I have a point on that, just that observation.
Hello, what right do you have to regulate the presentation of speech? If you regulate this format because it’s now considered harmful, what stops El Presidente from moving to ban “zines” because the format is “harmful to young minds” and used by “antifa”? What stops CA from moving to ban forums because threaded formats are suddenly considered “too addicting”? Maybe we should ban VR or first person shooter video games?
There is no allowable constitutional authority for actions like this. CA is literally overstepping the 1A limits of the Constitution here.
I doubt that this will run into any 1A issues. It will probably pass the test for allowable time/place/manner restrictions.
• It is content neutral.
• The government can probably show a significant government interest in reducing the harms infinite scroll often leads to.
• It is narrowly tailored. It achieves the goal without burdening more speech than is substantially necessary to achieve the goal. Arguably it doesn't burden any speech since every word you can have on an infinite scroll page you can have on a paginated site.
• There are alternate channels. The speakers still can get their message across. In this case they can get it across to the exact same audience in the exact same place. They just have to stick in page breaks.
Time/place/manner restrictions typically apply to public property. These are private websites.
While the court has once or twice extended protections to people using private property as a public forum, to my knowledge they have never done so with time/place/manner restrictions.
"No loud music between 10pm and 7am." That's on private property when it can be heard by others. Laws have all kind of restrictions on the time/place/manner of self-expression when what is being expressed has a negative effect on others. What you can't do is have laws that would say loud opera after 10pm is OK but loud Country & Western isn't.
“No writing at home between 10pm and 7am.” Would that be allowed? I imagine you agree the answer is “no”. What, specifically, is the difference? Because the answer lies there.
Hint: the physical world is very different from the virtual world, and has different limitations.
Hint #2: if I crank up the amp to 1000db and shout into it, it’s obviously not a question of speech anymore. This is obviously an extreme example (the energetic release just destroyed the planet), so dial it back to where it’s reasonable and concerns are balanced. Are you still facing actual physical discomfort? Did you dial it back enough?
Hint #3: is my nighttime writing keeping you awake in your home?
Speaking as someone who agrees with you, it's much better to just make your point than to lay it out in hints like this. It's condescending and annoying. No offense intended, just a note.
Fair enough, I just get bored of being the lone voice refuting these obvious talking points.
To be honest I suspect much of the support for this bill here is inorganic, and I do feel extreme contempt for the people pushing it.
In the end I’m not really trying to convince these posters—they have obviously made up their mind—but rather to entertain and educate the nonaligned audience.
Honestly this is the big reason I have stopped participating in threads like this. Infinite scroll is not itself addictive. Hell, there are studies that suggest that us calling it "addictive" could itself create the very problem we're trying to "solve". There are also studies that completely refute the "social media is harmful" narrative too, and I'm talking on the order of millions of participants across more than 50 countries. Hell, even the kids don't agree with the narrative. I bring this up because all of this is so interconnected. I'll just leave this here for those curious: https://www.youtube.com/watch?v=Vzbz--aPQLE
Thanks for bringing this up. “Infinite scroll is literally cancer” is a new one that I’ve only seen in this thread, and I’m not even bothering to respond to. (Is infinite scroll annoying? Absolutely. But it’s not a “public health” threat, and even if it were that doesn’t trump 1A concerns!)
> There are also studies that completely refute the "social media is harmful" narrative too, and I'm talking on the order of millions of participants across more than 50 countries.
Because the youtube video essentially says what the studies do (and it's a developmental psychologist who's giving that talk too). I felt it would be a bit easier to digest than more than 6 separate studies that you'd have to read. But if you really want them here are at least 6 of them all essentially saying the same thing: https://www.techdirt.com/2023/12/18/yet-another-massive-stud...
Thanks; it's much easier to examine assertions and weigh evidence when I can actually see the evidence instead of someone talking at me, regardless of their credentials. I appreciate the link.
These are private websites accessible over the public internet, to be clear. Also, there are known relationships between the intelligence community and the big platforms, lets not pretend they have free reign.
Should/do we allow foreign propaganda radio stations? If we accept that the government can (and very much does) impose itself on content platforms for "national security", what exactly is the difference between deliberately insidious information warfare, and collateral damage from market incentives?
I agree that its better to find solutions that involve protections instead of restrictions though. I think it means forced decoupling of indices/curation from advertising. This would make advertising funded addiction feeds compete with paid feed applications.
Courts already do lawfully regulate the “presentation of speech” as you’re calling it. Say facebook was to present each post surrounded by pornography for example. That’s clearly a “presentation of speech” in your framing. However courts have decided that it is possible to regulate the circumstances under which that would or would not be ok and 1A arguments have not prevailed in that case.
That’s more likely to be publication in itself, not presentation, but regardless porn is one of the very few areas where the courts still listen to speech arguments. However, the remaining decisions allowing (limited) regulation of porn rely on complicated, twisted reasoning—it’s clear that the justices don’t like touching this subject and feel that there are still constitutional issues here that may eventually need to be resolved via amendment. Usually this involves classifying porn in some other context, so it’s no longer “just” speech. Then the non-speech part can be regulated. Whenever they do this, they like to draw a very tight line biased towards favoring speech wherever possible, and they have consistently made it clear that they are not looking for more areas to do this kind of tightrope walking.
We already do limit harmful speech, at presence it's limited to speech and will immediately cause harm (the whole "shouting fire" thing) and the demonstrable addiction properties can be reasonably shown as harmful.
It's also telling that only corporations seem to be the ones demanding the right to infinite scroll; what's the scenario where an individual can only express themselves and their ideas through implementing infinite scroll on a social media?
We draw lines in the sand all the time for the sake of public safety, I'd like to hear a specific case of harm here.
“Shouting fire” was a bad decision denying the right to protest the draft, and it’s since been overturned. (Thankfully, as we may need that right soon!)
The First Amendment is clear: there shall be no law abridging freedom of speech. Courts have bent around that in the past, in earlier eras, but they were wrong to do so. Their mistakes have mostly been corrected although there’s still a few left.
The document that governs this country spells it out: it can’t be done. Public safety be damned. There’s no public safety exemption in the Constitution. If you want it done, pass an amendment. There’s a process for it.
I personally dislike infinite scroll, but I dislike the camel’s nose in the tent even more. No speech laws.
Hang on, let's go back - clarify for me how we're calling an addictive feature in a product built by the wealthiest corporations on the planet a matter of individual free speech? Precisely whose free speech would be harmed here?
Seriously, this diffusion of individual liberties into corporations has no presence in the constitution, and courts have fabricated this wholesale. There is no idea, no concept, no notion that infinite scroll provides. We regulate the size, location, and brightness of billboards; is this also a matter of speech?
Oh is this law’s scope limited to only the world’s largest corporations, and not smaller competitors, new entrants, individual developers, or nonprofits? I didn’t realize that.
Oh is the presentation of text and images not “speech” because it’s “addictive”? I didn’t realize that.
Your strategy with billboards is more clever than I’ve usually seen from you lot; I’ll give you credit for that. A billboard is actually a physical structure. The message on the billboard is the speech. If I stopped here you’d have a “gotcha”; the software must be like the billboard! But no, because first of all, code is speech, and secondly, the layout of items on the screen and how they interact is also just speech. It’s just graphic and UX design! There is no physical structure here. You’re attempting to regulate the presentation of information—design.
The 1A jurisprudence, to my understanding, basically results in the courts virtually never finding that the government has a legitimate, competing interest in limiting political speech.
But courts are willing to find that certain speech that is apolitical can be limited (the previous "fire in a crowded theatre" example). Basically the courts have recognized 1A established freedom of speech to protect political dissent and political ideas. Porn, for example, has limitations that would never apply to political ideas.
Again, the fire in a crowded theater example was actually political, and the decision was overturned. It no longer stands as precedent.
Limitations on porn still exist in a few areas, but they are gradually being rolled back—obscenity laws were once widespread and highly restrictive. Most still standing carveouts are pretzel twists that probably need to be corrected with a clarifying amendment; they are on very shaky ground.
The court has recognized speech protections outside of politics many times, including protections for authors and creators who were not explicitly aiming for political statements. For example, Brown v. Entertainment Merchants Association established that video games are protected expressive speech, even if they are violent trash that aren’t attempting any political point whatsoever.
Isn't it fascinating that the people making the most extensive use of infinite feeds and A/B testing for maximum user engagement are also the massive platforms with dominating network effects and captive audiences? It's like _specifically regulating large social media conglomerates with outsized impact, capacity for harm, and demonstrated propensity to maximize user addiction might provide an ideal balance of societal improvement without harming smaller actors_.
Re, source code: you can print out an implementation of your infinite feed and put it on GitHub. Go nuts. That's your freedom of speech. Likewise, I can write DDoS control software and clients. However I can't run said software as a service because that specific act is illegal. Same thing applies to the application feeds we're discussing; hosting content and offering software as a service has different semantics.
If you think that UX is a matter of free speech then I have an illuminated freeway sign running at 3000 nits to sell you.
We can have nice things. We can push corporations to act in pro-social manners. We can put individuals at a better footing with respect to large corporations while ensuring the liberty of individuals and small businesses. This libertarian idea that we cannot constrain obviously harmful behavior from massive corporations without immediately turning into an authoritarian both flies in the face of historical precedent and basic reason.
Sophistry. The question is not whether or not regulation is authoritarian, it’s whether or not it’s constitutional. As in, whether or not the government is even allowed to make such a law.
A law doesn’t just get a constitutional bypass because it’s addressing known harms or “anti-social” behavior or whatever. This is not the UK.
Illuminated signs exist in the real, physical world. They can beam bright light into your home, involuntarily. Design and presentation exists in the realm of a printed page, or on the display of your device. Can we regulate how a book lays out its type?
The First Amendment is quite possibly the most uniquely American thing about our Constitution, and its most defining feature. It’s worth defending.
Can I buy a 40mm grenade launcher without a FFL? No, I can't. Can I legally manufacture and install an auto sear on my AR? Also no. Would it be sick if I could? Hell yes. Do I own a delightful selection of firearms, including AR pattern rifles? Yes, and the cardinality of that set is only going up.
Does society benefit from mass ownership and unlimited access to fully automatic rifles and grenade launchers? If it does, what country allows it?
Are the above constraints explicitly decided as constitutional though years of legal decisions at all levels of the courts? Yes? Then we can observe that we can reasonably constrain constitutional rights through law and legal opinions. The line may be hard to draw and may shift, see the AR ban, but it is accepted that constitutional guaranteed rights have bounds that can be articulated and clarified through the legal and political system.
We put upper bounds on the rights and freedoms of individuals and corporations because we all must live within proximity to each other. These bounds may be authoritarian at times, and of course that's bad. But we collectively can limit freedoms because the alternative is actively and disproportionally harmful to society.
When it comes to the rights and the freedoms of the largest and wealthiest corporations, we already live in an era where these entities are shaping major aspects of our lives. Infinite scroll is one small mechanism by which they're hacking our biology; this is more than just pixels on a screen but a component in a system that was A/B tested to maximize behavior modification.
Help me understand - do you believe that it is possible to regulate these entities in any form? Or do we need to say that the folks that yeeted tea into a harbor were fine with infinite corporate power and regulatory capture?
Not that I think this significantly alters the point, but it's pretty common in the US to regulate or ban signage. e.g. billboards are illegal in my city and there are specific regulations about what kind of elements can be present on buildings to signal business names. I'm pretty certain illuminated signs beaming into people's homes would be illegal here. Actually I don't think light-up signs are allowed at all; I believe they have to be lit via projected light pointed at the building the're on.
Yes, my point is that things like illuminated signs or loudspeakers can actually physically affect neighbors, so speech concerns have to be balanced against other concerns. Often the speech still wins, but not always.
We’re talking apples and oranges because a website is more like a book than an illuminated sign. You have to decide to view it, and it doesn’t shine through your window at night, disturbing the peaceful enjoyment of your home.
A website like mygeotechnicalblog.example.com is like a book that you have to seek out. But websites like Facebook and Twitter may be so ubiquitous that they are more akin to a street that you walk down for many purposes and shouldn't be bombarded by obnoxious advertisements on the way.
While we’re just stretching metaphors to fit our preferences, comments like yours are so odious that they are akin to an open sewer, and should be regulated for public health reasons. Am I doing this right?
> A law doesn’t just get a constitutional bypass because it’s addressing known harms or “anti-social” behavior or whatever. This is not the UK.
First, the harm arguments are regularly made in front of the supreme court. And sometimes, when it suits them, justices make their own harm or sociality arguments. No, USA is worst. It gets to be constitutional if it advances conservative right wing agenda and unconstitutional otherwise.
> The First Amendment is quite possibly the most uniquely American thing about our Constitution, and its most defining feature.
You dont defend it by redefining its meaning to unrecognizable to encompass things non-speech of corporations. All the while making it so that in practice, poorer people have no defense anyway.
There's the constitution (basically a piece of toilet paper with scribbles on it) and then there's the actual reality of how the country operates, and the actual reality is that speech is restricted in many ways.
I must also mention that courts are not Congress and states are not Congress. The first amendment does not say "there shall be no law" - that is your poor paraphrasing - it says "Congress shall make no law"
Well just throw it all out the window then, if we’re not going to pay attention to the constitution. First things first, let’s make a law to ban you.
If you’re just going to pick and choose what rights you apply, then it’s not much of a governing document, is it? Is this just “Parliament is Sovereign” with extra fluff? Might makes right?
Too bad all the old “rightful” standbys have gone rogue, while rapidly losing their capacity to effect change.
It’s almost like we need a robust system of checks and balances, governed some kind of rigid framework to ensure that everyone plays by the rules. Or we could just continue to ignore that and see what happens.
That specific case is why I brought this up, no they aren’t illegal, this is Making Shit Up and exaggerating beyond what actually happened. But if people like you got your way, they could be made illegal.
First they came for the infinite scroll, and I said nothing...
That was a little hyperbolic. The government already can regulate "speech" to some extent in limited, targeted ways, as this is. For example: they can (and increasingly will) require ADA accessibility standards on web and mobile sites and apps-- even private sector-- that deal with the public.
It may be that this isn’t as settled as you think when speech concerns are present. The existence of alternative accessible formats, or sufficient assistive technology in the marketplace, may be just as compliant. It’s likely that these will be favored over mandating changes that affect design or presentation, given the Court’s prior decisions on balancing speech concerns in other areas.
While I would agree, I think many are missing the even more fundamental issue. Even if you get rid of every single predatory thing imaginable, social media itself is full of incessant stimuli, FOMO and other features that make it destructive for everybody, but especially young people. For instance image crafting is going to harmful for everybody, because many adults have a tendency to engage in keeping up with the Joneses, but for children it's especially harmful because of much greater personal insecurities and proclivities towards envy.
And things like image crafting are not even necessarily intentional or malicious. Somebody who posts a pic of their filet mignon dinner probably isn't posting much in the way of their microwaved leftovers mashed up in a bowl. And adults already get mistaken perceptions of others because of this bias, let alone children.
And those installed programs could have vulnerabilities that just a non-root user account could still take advantage of. Perhaps not likely for an LLM to do, but more so if you let them loose on the internet and they end up coming across prompt injection that instructs exactly that :)
Though I'm in the camp "people should really know to sandbox by now and be careful", I'd say we should also be mindful of how far from everyone has deep knowledge of the systems and tools they use. This behaviour of a tool is just malicious. You have to take into account the human factor, of how people likely end up using a system. And in this case, the consequences of exfiltrating so many secrets this way are really quite unacceptable.
This is a fight I deal with every day. We have folks in the technology group at work who use AI to write code and do so without issue. But now folks in supply chain or in the executive suite are using it to generate web pages that they want published or apps they want on the internet, and while Claude can generate an HTML file how that gets published, how authentication works, etc is just glossed completely over, and generates a ton of work for the IT team to build up around this stuff as it comes in
macOS largely _does_ bake this into the OS, and it is annoying. They also provide a way to turn it off for specific applications (including, for example, Terminal.app).
I've found that I can usually write apps that respect it. MacOS is a free-love hippie, compared to iOS. In many cases, we have no choice.
It's annoying, and often rather infuriating, but I understand that one of the motivations for people buying Apple stuff, is for that very reason, so I'm really sawing off the branch that I'm sitting on, by trying to work around it.
Not to mention the very wide push to "Use AI NOW, for EVERYTHING!" in marketing ans many companies, with hardly any though given to safety or where does all the data end up.
Given the long history of even the most egregious data breaches with millions of affected people never having the slightest consequences, what level of care are you expecting here?
I specifically chose a Mac Studio 128GB as my home server that's also running LLMs to be always online, in part due to the minimal idle power consumption and mostly fan-less operation. It's definitely expensive, especially nowadays, but I can still recommend Mac Minis as a cheaper alternative for someone to just get started with an affordable, always-on home server that won't annoy any housemates. I think both are in some sweet spot in terms of value for money, depending on what you're looking for in a home server. If image or video generation is your thing, look further though, definitely look into a proper GPU then. Macs are quite slow at that. They're just great at MoE LLMs because it's mostly a matter of (V)RAM size.
I have! I care about data privacy and LLMs being free. I'm using the Pi coding harness but containerized and sandboxed, to make sure it's running completely offline. On my Mac Studio with 128GB RAM (or MacBook with 36GB RAM) I'm using Qwen3.6 35b, with only 3b active parameters so that it runs really fast. I've done a complete redesign for my website's homepage and blog with Django + Wagtail. The latter is interesting, because Wagtail is a bit less well-known, so the agent, without giving it internet access, doesn't always know how to develop for Wagtail. I've used Qwen3.5 122b for when things get more complex. At 10b active parameters, it's significantly slower though.
I've noticed a few things compared to large models like Claude. For starters, you really need to know what you're asking, and be precise; it doesn't do much thinking for you. Any assumptions left open, and it'll take the easiest route to reach the goal (e.g. CSS in HTML), often not the best in terms of architecture.
It gets into loops quite often, and surprisingly often gets the edit tool call wrong, after which it will spend lots of thinking tokens and re-read files instead of retrying (despite the system prompt suggesting so).
Comparing agentic Qwen3.6 35b to Claude Opus is like a junior with knowledge across the board, that you really need to guide, versus a senior that thinks with you on architecture. If Opus gives a 15x speedup, local and fully offline Qwen gives a 5x speedup. Which, given that it's completely free, is still mind-boggling to me :)
This is very similar to my setup. Pi in a container (I do let it have network access, just no access to creds or anything, only the one directory that I'm working on at the time and my ~/.pi directory), talking to llama.cpp in another container. I'm on a Strix Halo 128 GiB unified memory laptop.
I've never used the frontier models in earnest, I don't believe in using proprietary tools for my programming, so I can't really compare.
And I'm still a AI skeptic, so I'm doing more testing and kicking the tires than I am actually using it. That means I spend a lot of time trying to break various models, probe them for strengths and weaknesses, etc.
But I find that when I do try to use it for real for agentic coding, Qwen 3.6 35B-A3B is definitely the one I reach for the most often.
For other chat tasks and translation, I'll frequently use Gemma 4 31B.
For audio, I'll use Gemma 4 12B.
I keep a bunch of other models around to try out every once in a while (Qwen 3.5 122B-A10B, Qwen 3.6 27B, Nemotron 3 Super 122B-A12B, Step 3.7 Flash and Minimax M2.7 both at somewhat more aggressive quants, and GPT-OSS 120B if I want super fast but not terribly smart), but so far Qwen 3.6 35B-A3B is really the sweet spot for coding on a setup like this.
Hopefully this isn't off-topic, but your setup sounds just like mine, Strix Halo and (I'm assuming) llama.cpp on ROCm, and I'm finding that the Qwen hybrid models don't handle prompt caching and instead re-process the context in full on every turn. I'm wondering if you were able to solve this and how?
I use Vulkan mostly instead of ROCm. Vulkan is actually a bit faster, paradoxically. I do switch out and try them both out, and it's not a huge difference, but I've been mostly saying on Vulkan.
The re-processing context every turn problem is definitely something I've hit. Some of the causes have been solved upstream in llama.cpp; make sure you're up to date.
But another cause of the issue that has a big effect is that older Qwen models didn't support preserving thinking. This means that each time you have a long sequence of tool calls with interleaved thinkging, as soon as you had your next turn in the chat, it would have to re-process all of that as it would drop all of the reasoning.
Qwen 3.6, however, now supports preserving thinking. This can use a bit more context, becasue you're not dropping the thinking every turn, but it re-uses the cache better, not causing you to have to reprocess a whole turn at a time each time.
In my models.ini, I have this for the Qwen3.6 models:
Thanks for sharing have been running ROCm primarily with Qwen 3.6 and Qwen Coder, on the runs much better statement is that a stability, performance or other capability your experiencing?
I'm a little surprised that preserve_thinking would matter here for cache purposes. for actual capabilities/intelligence, yes, I'd imagine it helps to have past reasoning traces in multi-turn setups.
but for caching, all you are doing is leaving off a fraction of the most recent assistant message generation, which will have little/no impact on cache hit rate.
> all you are doing is leaving off a fraction of the most recent assistant message generation
True, but not a tiny fraction, qwen is very verbose in its thinking traces. And it basically means that for every (nonthinking) generated token you have to compute the KV twice (once as tg, the second one as pp).
I was able to solve this for my setup, 7900XTX and llama.cpp on ROCM in the oh-my-pi fork of pi.dev harness. I documented my setup on github, check under my username/omp-config, but the important thing is making sure the context is strictly append-only, and starting llama.cpp with
If you're hitting this you have a bug, this is not related to the model. Either your harness is editing the messages between turns incorrectly (i.e. it is not append-only), or sometimes this is because of llama.cpp bugs, but bet on the former. Setting up something like Tailscale's Aperture will let you capture the requests and then you can diff them.
What harness are you using? Some of them (e.g. OpenCode) mutate the system prompt every turn, and therefore can't work with a KV cache.
I've had the best luck with Pi so far, but it comes without some bells and whistles you might be used to (e.g. plan mode, subagents, MCP client support)
> Qwen hybrid models don't handle prompt caching and instead re-process the context in full on every turn. I'm wondering if you were able to solve this and how?
Isn't this the nature of how LLMs work? Or do you mean that it recalculates the entire KV cache instead of saving the old KV cache, in which case the problem is likely in your executor (llama.cpp, vllm, e.g.) configuration or capabilities?
So, one of the ways that this problem manifests is that most local models aren't trained on preserving the full reasoning between turns. Every turn, they skip passing the reasoning trace from previous turns to the the LLM. So if on one turn you have a long interleaved chain of reasoning and tool calls, then it responds to you, and then you give a new prompt to fix something, it has to re-process all of those tools calls now with the reasoning stripped out.
Qwen 3.6 has finally been trained both with and without preserving thinking, so you can optionally enable preserving thinking. This will use up a bit more context, but it will avoid having to do this re-processing of long agentic turns, and also the preserved thinking can avoid having to re-do some of the same reasoning over again in later turns.
Besides that, modern LLMs don't only use full attention (apparently, attention is not all you need). Full attention is very expensive to compute and store (0(n^2)). But additionally, full attention is actually bad at certain kinds of reasoning; keeping track of some value that gets replaced over the course of time, for example. So most models these days use various forms of local attention which is fixed length and gets updated as you go; sliding window attention, Mamba-2 state space models, etc.
But one advantage of attention is that you can go back and reprocess by truncating the KV cache and starting over. You can't do that with other forms of local attention; you've lost the state earlier in the sequence.
So to allow you to go back without fully recomputing the cache all over again, your engine will save snapshots of the local attention state at various times, so if you need to go back to recompute the cache, you can start from the last snapshot. However, these snapshots can get large, you can't keep too many of these, so sometimes you need to go back quite far to get to one, or they're all past the point you need to go back to and you need to start over again from the beginning.
There have been particular bugs in llama.cpp that have caused this to be triggered more often than it should; for instance, it wouldn't take snapshots before turns that included images at one point, so if you had an image heavy agentic workflow, that issue plus the lack of preserving thinking would mean you would frequently have to go back and start over from scratch.
Some of these issue have been fixed, some are addressed by preserving thinking. There are still some issues sometimes; for instance, one that's hard to fix is that the tokens generated autoregressively don't always parse the same when doing prefill. For instance, you could generate something as two tokens "pre" and "fill", but it turns out that "prefill" is also a single token so the tokenizer will use that, so when you send that back again on the next turn, it will see a divergence and have to recompute from that point. It might be possible to ignore that and use the not fully greedy tokenization that's in the cache, but I've definitely seen llama.cpp have to do some cache recomputation due to that.
use that to install assuming you have whatevers needed to set up.
ive a much fancier next gen thing i hope to make available as a sass mid to late summer, but if you have any feedback or questions on mine do reach out
Not a harness issue. The harness (pi in my case) passes back the cot for all previous turns.
The jinja template is what renders the openai-format request sent by the harness, into the actual string of text that will be tokenized and fed to the model. For models without preserve thinking support, the jinja template drops the reasoning from all but the current turn.
{#- Render reasoning/reasoning_content as thinking channel -#}
{%- set thinking_text = message.get('reasoning') or message.get('reasoning_content') -%}
{%- if thinking_text and loop.index0 > ns_turn.last_user_idx and message.get('tool_calls') -%}
{{- '<|channel>thought\n' + thinking_text + '\n<channel|>' -}}
{%- endif -%}
You see that it only preserves the thinking for indexes that are later than the last user message; thinking is only preserved for a single turn (which can include a lot of interleaved thinking and tool calls), once it goes back to the user and the user replies, it will replay the tool calls but not the thinking between them.
{%- if (preserve_thinking is defined and preserve_thinking is true) or (loop.index0 > ns.last_query_index) %}
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
{%- else %}
{{- '<|im_start|>' + message.role + '\n' + content }}
{%- endif %}
It additionally has a preserve_thinking flag that you can set. If that's set, it will include all turns thinking in the text passed to the model. But you do have to set that, it's not the default.
It's possible to modify the jinja file that you're using with a model. Some people do that with models that haven't been specifically trained for it, and report good results; but some report that because it wasn't trained for that, they get worse results if they include thinking from previous turns.
So for models like Gemma, you would have to modify the default jinja to enable this. For Qwen, you can just set the preserve_thinking flag to get this behavior; and apparently they have trained in this mode so you get better results than models that have not trained this way.
How could the harness fix this? It's the jinja template used by the inference engine to render the API requests into the raw text that gets tokenized and completed by the model. Unless you're using something like the raw completions API instead of the `/v1/chat/completions` API, and effectively applying the template yourself. In which case, you could also just modify the jinja template on your server.
Anyhow, I've heard mixed results on any method of supplying reasoning traces beyond the current turn to models not trained on them. For some models, I've heard that it works fine this way, for others I've heard it degrades performance. But I don't know of anyone who has any kind of reliable benchmark for how well this works.
I'm a housekeeper skeptic. While I concede that a professional housekeeper would probably do a better job than me on most domestic tasks, I still think everyone should clean their own home, cook their own dinner, and write their own code.
For me the distinction is that your rice only needs to be edible once, while your code may need to last for decades. Using AI to code anything I could comfortably throw away if needed is a lot less fraught than letting it make choices that I and anybody who inherits the code is gonna have to live with, especially if by outsourcing those choices I reduce my understanding of the implications of those choices.
I don't let the AI make any choices. I have a lot of instructions and sample code for it to follow. It is basically a glorified code generator at that point.
I think the idea that code should last decades is now questionable, if not problematic. If we can now produce code at 10x the rate, that means we can have 10x more code (probably not desirable) or we can have 10x as many revisions. Whoever inherits the code can have it rewritten to their liking and understanding. Nothing helps better in understanding a system than to rebuild it, even if just by handholding an LLM.
Exactly, but if I start from working code with a lot of tests I don't need to remember the requirements. I just need to know my current requirement and figure out the ones I'm changing with my new requirement. It doesn't catch everything, but in most cases if I break some other requirement I find out about it and can figure out just that one more requirement and not the millions of others that still work.
I'm not getting it. OP said they are wary of letting the agent make choices for them, and outsourcing those choices lessens their understanding of them. They could interrogate the agent on why those choices were made until they have sufficient understanding, and they can also change the solution if they want to.
The thing about this is that you can choose how high level you go.
For example you can just tell it to make a website for a business with a webshop and it'll just generate thousands of lines of code and you have no control over anything. Or you can spend hours/days writing the specification and then have it generate it.
Or you can do what I do and work iteratively one feature at a time making sure everything is exactly the way you want it. I generally solve the problem myself then tell it what to do, or if I'm not sure what the best solution is I might discuss with the AI until we agree on a plan and then have it execute it. Often this leads to me learning useful things, like it will suggest a tool/feature that I didn't know about that's perfect for my usecase or it will identify a problem in my plan that I wouldn't have found until after spending hours on the implementation.
I've always been very detail oriented and I care a lot about code quality, I want my solutions to be clean, consistent and as simple as possible while solving the problem. To me, AI tools let me do that more quickly and better, it's not a compromise it's just flat out better in every dimension. It's about how you use it.
A lot of people seem to think that it's a binary choice, either hand craft a high quality bespoke solution or just vibe code a pile of trash. There's a whole spectrum in between those two, and I think there's a sweet spot where you still maintain control and understanding, it's just much faster and the result is actually better because it's not just you and the knowledge in your brain it's also the AI that practically knows everything - it will teach you things and suggest solutions you wouldn't have thought about, it makes you a better developer. It's a force multiplier and the smarter you are the better you will be at using it.
It's not a replacement it's an enhancement. It's like imagine a developer with Google vs one without, obviously the one with Google will be better because they have access to more information. The AI is like automatic google that just googles everything all the time, things you wouldn't have even thought to Google or things you couldn't possibly formulate a good search term for. With AI you can just show it a screenshot or describe an issue in detail and get a really solid answer a lot of the time. It's like having an expert on standby all the time, sure it's sometimes wrong but most of the time it's not and if you're smart you'll recognize when it isn't.
I'd say anyone who isn't using AI today aren't using their full potential. I don't see how anyone could possibly perform better without this tool than with it. I do see how someone who doesn't care could produce a lot of slop, but the people who refuse to use it aren't that guy. That guy has been using it to produce slop for years already. You can use it to produce top quality code if you choose to.
Much more complex than that. Even if it does give you a speedup at certain tasks, is it worth the cost and risks? You go faster, but now you have more code that you don't understand and so won't be as good at maintaining. There's the engergy use, the water use, the scrapers destroying the internet, the massive piles of slop, the hallucinations and bullshit, etc.
It means that even if it works for certain tasks, I think that the problems caused by use of LLMs outweigh their benefits. I think it's a bad idea to generate large piles of code that you don't understand, but due to competitive pressures, it's too tempting for people to pass up, leading to a world in which software is getting worse by the day, while pumping CO2 into the atmosphere and boiling scarce water supplies to do so, DDOSing websites to scrape the data, and polluting the internet with mountains of slop.
This isn't about using rice cookers or not, that's a personal choice for how you cook your food, and choosing to do so or not really only affects the person cooking and cleaning. A rice cooker probably uses a similar amount of energy as cooking it by hand, possibly even less.
But when people using LLMs are causing active harm, and are making it more difficult to collaborate on a team, it's a lot harder to accept that it's just a personal preference.
If you wanted to use the rice cooker analogy, imagine if rice cookers let you cook rice in just one minute. Faster, don't have to wait for the rice to be done, great! But in order to do so, you have to cook 50 pounts of rice, but throw out the majority of it, and use a thousand kilowatt hours of energy to do so. You'd better believe I'm going to be skeptical of everyone deciding that they suddenly have to use these 1-minute rice cookers that burn so much energy and generate so much waste.
Haven't used for actual coding but was testing locally - for example running some swebench instances - whether qwen-3.6-35b-a3b@Q8 was better than qwen-3.5-122b-a10b@Q4. With MTP the former runs at around 55t/s and the latter at around 30t/s meaning the latter is also usable. It looked like qwen-3.5-122b-a10b@Q4 performed a bit better.
For the edit tool, you should consider implementing a hash-based approach where each line of code is hashed and referenced by it when doing replacements. You can read up on the approach here: https://blog.can.ac/2026/02/12/the-harness-problem/
I didn't do much benchmarking, but anecdotally, I found it to be making less edit errors. YMMV
Yup, I used this for a while and IME it may get you a few percentages more of useful context initially, so quality feels a bit higher, but things start breaking down in funnier ways when you do run out of that quality for any reason later, so definitely caveat emptor.
I can use Gemini 3 Flash with the harness I built for around 8 years and still not exceed the cost of a Mac Studio with 128GB, the price for privacy is very high. Agentic flows that get stuck can be worked around but I prefer developer velocity.
For corporate use, if the corporation would break the law sending anything to the open internet or to the US, then you can't use any model that's not hosted in house. And there are many such cases.
Sure, but Gemini subscription gives you just that - Gemini subscription, but new computer allows you to do other stuff with it as well. When you're upgrading anyway for other reasons then it's not fair to compare full Studio price to just one subscription.
Right. Tokens/s decode isn't the most important thing to me: wall clock time for task completion is. And tracking all of that, on my GB10-based Asus box, Step 3.7 Flash at IQ4_XS beats Qwen 3.6 27B despite the latter having MTP, on all of my actual coding task evaluations in real codebases.
Qwen seems better at one-shotting things based on vague prompts to an acceptable degree, but thats literally not what I use these things for!
One thing if people do play with it, is it seems very very sensitive to quantisation of the K part of the KV cache. F16 K and Q8 V got rid of a lot of the loops that it was otherwise hitting.
There's also a regression in llama.cpp wrt. Step Flash, where quantisation is getting worse KLD and Perplexity than it otherwise was previously, for the exact same quants. Very odd, but it's being looked into at least!
Do you think the choice of quantization matters that much for other models? I've seen a lot of discussion about different quantization and FP formats but I feel totally unequipped to make an informed decision about what to try.
What's your evaluation setup like? It sounds like maybe the best thing to do is have a realistic evaluation that resembles your actual intended workload and workflow, and then just try everything.
>What's your evaluation setup like? It sounds like maybe the best thing to do is have a realistic evaluation that resembles your actual intended workload and workflow, and then just try everything.
That is quite literally what I have setup :)
I have a few codebases I've written over the years that I attempt a suite of specific tasks: code analysis/bug finding, bug fixing, adding features, that kind of thing. I keep track of the results, including wall clock time
>Do you think the choice of quantization matters that much for other models
It hugely matters. Lots more than r/LocalLlama would have you believe, sadly. Some model architectures can handle more aggressive quantisation than others, and it's hard to know ahead of time.
Step handles it surprisingly well (sparse MoE models seem to generally, when the particular layers are chosen to be quantised carefully). Qwen 3.6 27B handles it okay, but FP8 was better... except annoyingly Qwen's official FP8 has worse KLD/perplexity numbers/accuracy than it otherwise should. RedHat's one was better in my testing, though not by a huge amount.
It isn’t though, I’ve run both through a bunch of coding evals. You nearly certainly didn’t have the right sampling parameters or quantised the KV cache?
Ds4 is impressive for what it is, but it loops and over thinks even more, burning massive wall clock time to not even get great outcomes. It’s also limited to a slow speed on my Spark
I tried a bunch of stuff with step 3.5 and step 3.7 maybe not as much as you. Could you tell me what parameters and launched you’re using ? Antirez ds4 flash q2-q4 works almost out of the box for me
To be fair: if you're happy with ds4 then IMO stick with it!
Step 3.7 is notably better than 3.5
1. Use the official StepFun GGUF, IQ4_XS - theirs is better tuned in my experience than the other quants
2. Temp 1.0 top_p 0.95 sampling parameters for reasoning/agentic coding
3. It's really quite important that you don't quantise the KV cache: it made a surprising amount of difference to the looping and over thinking I found, at least for the quantised version of the model. I'm using the full F16 for K, and Q8 for V
4. Note that it now supports `reasoning_effort: low|medium|high` in your chat_template_kwargs; this is super useful :)
I've got a tool that sits in between the harness and inference engine called petsitter. It is a middleman validator to avoid just these kinds of issues. You can stack the fixes as needed (they're called tricks in the petsitter parlance)
> Comparing agentic Qwen3.6 35b to Claude Opus is like a junior with knowledge across the board, that you really need to guide, versus a senior that thinks with you on architecture.
that's why i use the frontier models because its a senior co-worker vs a junior. if you use the junior for the sake of privacy i think you're missing out on the best insights for a specific task.
Consumer-grade subscriptions of the frontier models give you superb capabilities per dollar, them being heavily subsidized. But if you're working in an enterprise setting, that won't work. You need to upgrade, and that gets significantly more expensive.
Furthermore, basing the SDLC on leveraging the bargain subscriptions risks falling apart in the future, both from a cost perspective as well as the question of availability (e.g. Mythos).
So from a strategic perspective, going local on the LLM and still achieving great results with the right approach is very relevant.
Or you can get the best of both worlds--use frontier models to build a spec/plan, and use cheap models (open source or not) for implementation. Your max or team plan can go a lot further this way without giving up much for quality. Play with something like Superpowers to make this really approachable.
Best insights can be over rated due to bandwith limitation of the brain. Even if Einstein is sitting next to you the whole day and helping out Theory of Bounded Rationality applies.
What kind of coding do you do?
Do you keep track of frontier models to vibe check the differences and re-evaluate constantly or are you ok with having a nerfed model forever?
(not being judmental, just really wanto to know your framework here)
Some of the work I do, I do for an (EU) organisation that doesn't have clear rules or guidelines on the use of AI yet. Though I have seen colleague-developers blatantly putting source code into external Claude-like models, I stay true to my principles and don't. I know for certain that everything that I run through my local, offline Pi Container Sandbox cannot leave the machine, and thus can't result in a data breach. I do this for the peace of mind.
I do (unscientifically) experiment whenever a new capable local LLM (<=130b) releases with a license that permits commercial use. As for knowing my models require more work than Opus, I don't mind still having to puzzle on getting the architecture right. In any case, it forces me to stay in the loop of what's being built, which is a good thing.
Could you give more details on how to make such a set up?
I'm not familiar with Pi, and not sure which kind of container you are referring to. Something mainstream like docker, or more classic like a BSD jail?
I started to experiment with locale LLMs, through ollama and Lemonade. Enough to throw simple prompts with code excerpts and get small scope code refactors. Though I still struggled to make them work with external tools, like my IDE, so they can be leveraged on to an agentic level with access to a full repository.
That's mainly for work, as they push for using LLMs, though with the new copilote license they provide it doesn't take me even a week to burn the whole token credit.
The tool can be useful, but in my experience without heavy guard rails and loops over tests. I suspect late models to also burn many token into rabbit hole of nonsense hypothesis, instead of doing straight forward correct implemention as you would expect from any entity with such a huge cumulated resources eaten and experimental playground to leverage on. Maybe incentives don't help model provider to minimize sold token, maybe it's just so hard to tame the beast all these bright minds with virtually infinite resources are not good enough.
Anyway, sorry for digression, but I would be extremely interested with a step by step tutorial to make a local LLM work in agentic level, including which kind of hardware is required to make it work properly.
So you don't really trust the data policy (non-retention) of the big companies like Anthropic/OpenAI + regulations in EU.
This is very interesting. I myself have been blindly trusting these organizations with my data and still not sure if I am trading code/trajectories for productivity.
Another POV is that most of the code written in most of my codebases were generated by Codex/Claude, so they would be "stealing data from themselves" in a sense.
I've been working with Transformers/LLM training in 2018-2021 and then now, more recently again. Things are far different. I think they would be more interested in the "how" you got your code to be satisfactory with your guidance than the actual code generated.
But mostly I personally trust that they are not really using my trajectories for that (unless I explicitly allow it in the configs)
Why I do like Qwen 3.6 35B A3B, I have found that the difference improvement of Qwen 3.6 27B is massive. Sure, it is 3x slower (https://github.com/stared/benching-local-llms-on-apple-silic...), but for the total development time it felt that still 27B is faster to get the goal.
My experience is almost identical. I have found that I need to be very careful with planning, breaking things down into small isolated steps (I can have qwen do this); and also (me) writing a very clear design. Relying on qwen to fill in a lot of those precise details results in those about-to-write loops.
Yeah, that edit inability is weird. I’ve updated AGENTS.md to limit editing (as opposed to rewriting) and that helps a little.
Maybe even more useful than Opus when I have all the constraints to an issue. There is less "knowledge" in the model (I get by with 48GB of RAM allocated to an 8b quant), so it has fewer things to hallucinate about.
I've been getting to know its limits pretty well over the last few weeks and would say it's an excellent code search/replacement/generation* engine.
It's got the "in-context script generation" flow down as well, so it will easily help automate tasks that you describe with text and perhaps example commands, or tools, or skills* that you provide.
*Think of it + Pi as an NLP abstraction layer over grep, or a shell, rather than a jack of all trades + world knowledge all-in-one.
How are you sandboxing your Pi coding harness? Directly only mounting certain folders, using capabilities to kill the network and not giving it all your shell env vars, that sort of thing? Or do you use a tool?
And, is the sandboxing for security (avoid RCE on the host) or merely guardrails for the models?
I've wanted the latter quite a bit for Pi, because weaker models like Deepseek V4 have extreme issues with obeying prompts (e.g. I'll instruct it to find a bug but not fix it, and it'll "helpfully" try to fix it anyway), so having a "read-only mode" actually backed by the OS would be very useful.
Haha, yes! Last time I asked it for options how to tackle a task and only do the research without touching any code. With xhigh rasoning, it echoed the options that many times until it was convinced that option A is the better choice and started implementing it.
The harness and the LLM parameters are pretty essential to getting better results and reducing loops. Tweak the parameters and you can mostly eliminate loops without negatively affecting performance (it's a bit complex but ask a SOTA AI to guide you and it's not hard). The harness should also react more intelligently to failures; it can do things like return additional context or hints as it tracks error rates and avg duration of calls. Pi can be easily extended, and it's suggested by the author you modify it to perform better for your use case.
I've noticed the same about the edit tool, in both Gemma and Qwen. Maybe I'm not running them with the right sampler settings, but I'm happy to hear I'm not the only one. Lots of mismatched whitespace and stuff, the model ends up doing hex dumps and maybe 5 or 6 attempts at editing a 5-line function into a 250-line Python file.
All of these models also seem to get stuck in long thinking loops, sometimes tripling the tokens of a frontier closed model which is really painful when inference is already on the slow side (on my Macbook).
I am right there with you. Mind-boggling. It's a indistinguishable from magic technology!! I tried running some basic tasks through Qwen with Opencode on a 10 year old dual Xeon server for shits and giggles. I gave it a simple task like "use ffprobe first but convert this webm to mp4" and it was able to complete the task with zero network calls outside my network. On 10 year old hardware. It took about 3 minutes to complete the task. Now you may be saying 3 minutes? pfft. But I dare you to do it yourself. You're gonna be googling the CLI switches for at least 10 minutes and setting up your command. I had it actually optimize all the switches on the fly for me based on an initial ffprobe to see what is optimal.
> 10 year old dual Xeon server...On 10 year old hardware.
Hold on, what are the specs of your rig? How much RAM?
I've been considering getting an old refurbished 2018 Mac Mini with 64Gb of DDR4 RAM but everything I've read suggests this will be way slower than my 16Gb M1 Pro Macbook.
You can absolutely still use this to do some basic stuff like tell opencode to convert a video file from one format to another. But frankly you're better off getting two AMD GPUs. Say a dual 7900XT would get way better performance.
And that would be a much better source for a phone number than Googling. Similarly, the docs that ship with software are a better source for command line switches for that software than a search engine or LLM.
My lived experience right now is a lot of super talented people around me using these tools all day every day to build awesome things and then there's the randos like you on HN who think they know better. Protip: You don't know squat.
The docs that ship with it are a great source for the LLM who will be running the command and monitoring its output, fixing or adjusting whatever in order to complete my goal. Why on earth would I be calling it by hand?
You are right, but I think you miss the whole point of the agentic workflows that are being discussed in this post comments.
Yes, you surely can read man, docs, whatever, then DIY. The point is that in many areas people don’t really want to become an expert, like in ffmpeg cli arguments, they just want the work to be done. Above is an example of agent being able to do it locally, and I think it’s great
This is the part of LLMs I may just be too old to "get" because I can't imagine what people are doing that looking up command line switches is some kind of time suck they worry enough about to find a way to automate
also the context window sizes are too low. I can't operate in 65,000 windows any more because even just reading the code's file structure overruns it and gets me nowhere. Definitely its own art form.
200k context windows and above for me now
I saw a paper last night that should help this a lot though
I get that it's a deal breaker to some; it definitely requires patience.
In Pi, /new is my best friend and most-used command for sure. For simple tasks (I decompose complex ones anyway since I don't trust small local LLMs to do this for me), the model doesn't need much context, given that I'm proficient in my codebase myself: "I'd like Feature X. Look into files 1, 2 and 3 to make your edits."
> you really need to know what you're asking, and be precise
Any chance that you could share some recent prompts to give other HNers a head start on his to approach Qwen? If you are uncomfortable posting them here, my Gmail username is the same as my HN username.
I'm glad you're asking. I already started writing a blog post on how to best make use of local models. I'll share it as soon as I have a complete enough list. If anyone else reading this would like to chime in with their tips & tricks, let us know!
For the time being, off the top of my head, I'd say:
- Prompt Engineering tips & tricks apply here (like being complete in the relevant context you provide in your question, and the specific task(s) the agent should do like reasoning, modifying one file, or trying to fix a complex task all at once (not recommended)).
- If you already know which files the agent should look into, mention them to save time and potentially context.
- In my personal workflow, I write down lots of atomic TODOs needed to solve a problem. As I write it down, I'll notice assumptions I'm making, or the fact that the TODO could still be decomposed further into (atomic) subtasks.
- It's best to get a feeling yourself for how Qwen handles your repository. I noticed if I don't specify an architecture for development, it'll make quick & dirty fixes. If I don't tell it to remove debug statements, it won't. This is what was meant with "be precise" – Claude Opus might think for you and act in your best interest. Smaller Qwen models will just do what you ask them to, and no more. They have design knowledge, but you have to explicitly ask them to "activate" that part of their knowledge.
Given your knowledge on this - do you think we'll see an open source model with Opus levels of capability? IMO if/when this happens - I would 100% stop using Anthropic.
Let me put it like this. I started with local LLMs when ChatGPT still used GPT-3.5. I was amazed how my MacBook with 8GB RAM could run openhermes2.5-mistral: a 7b parameter model that could generate short stories that sort of made sense. Incredible!
Two years later, and I'm running Qwen3.6 35b agentically to develop the start of a repository and automatically run tests to then improve on itself. I never thought we'd get here so quickly with LLMs back then.
I'm pretty sure in two years we'll have current Opus-like quality in the 30-100b parameter model range. But at that point, Opus 6.3 will reason along for us so much better still, that we'll still look at those models in awe. It's great to look ahead, but let's not forget to appreciate how effective the current local models already are :)
Haha well I ask because I don't really want/need anything beyond Opus most of the time. And I'm paranoid that Anthropic is going to be forced to charge the true cost of all this before too long.
The other upside of running local LLMs is that there's no cloud provider to suddenly charge more for the same, or even less, model use.
It's personal, but I prefer CapEx over OpEx for this. If you can purchase a device upfront that runs a decent local LLM, you get the peace of mind that your setup won't suddenly change over time and can only get better.
If you believe the benchmarks, Qwen 3.6 35B-A3B already outperforms Claude 4 Opus.
Now, there's a bit of a degree to which some of the open source models do some benchmaxxing, and bigger models with more params may always feel like they have more depth. But anyhow, right now you have something that is arguably comparable to Claude 4 Opus on your laptop. I can't really compare myself because I never used it. It looks like Claude 4 Opus is still available on OpenRouter, so you could try it out and compare yourself if you're interested.
It will likely always be the case that there are proprietary cloud models that are more powerful than what you can run on a laptop. You can just do a whole lot more with terabytes of VRAM on multi-GPU clusters than you can do on a laptop. So for folks who must have the most capable, you're probably not going to want to leave Anthropic.
But right now, the models you can run on your laptop are comparable to the cloud models that were popular when vibecoding and Claude Code first took off.
You really need to take the benchmarks with a massive pinch of salt. I’ve been testing local LLMs since the original llama and there’s nothing I’ve tried that is in the same category as Opus.
Which Opus? They certainly outperform Claude 3 Opus.
Anyhow, feel free to try them out head to head on OpenRouter. I'd love to see someone write up their results, of a modern local sized open source model vs. frontier models from ~a year ago, on something other than the standard benchmarks.
There's a guy on Youtube named Bijan Bowen who tests all the models (open and frontier) on a series of one/few shot programming exercises and has been for a long while now. You can pretty much watch him compare the results for any two models you're likely to be interested in.
I'm not affiliated, I just like his style and have found it handy. I know it's not very rigorous, but it's good enough for me and I've found his examples to pretty closely match the results I see in real life.
Qwen 3.6 produced far more working functionality than Claude 4 Opus did.
Obviously, just one test of a single one-shot prompt of a silly toy OS, but yeah, this particular test shows Qwen 3.6 running locally dramatically outperforming Claude 4 Opus, which was a frontier model a year ago.
I’m normally comparing frontier open/cheap models against frontier closed source. I use deepseek/glm regularly, they’re fine and you can get real work done with them but it’s super obvious when you switch back to opus or even sonnet. A 3B active param MoE model is not comparable.
Yeah. I was pointing out that local 3b active models outperform frontier models from a year ago.
Will this trend continue? Who knows. Both the frontier and local model will probably continue to get better. Which one will hit the top of the S-curve first? Hard to say, really. But what you can do right now locally is better than what you could do a year ago on the frontier, and lots of people were already using it pretty heavily a year ago.
Hoever, November is when most folks agree that the frontier models got good enough for much of their work. Local models aren't quite there yet (where by "local" I mean "can run at reasonable speed and quant on a system less that $10,000 with today's RAM and GPU prices"). The biggest open weights models are getting there, but those require something like an 8x H100 server to reasonably run.
It's likely that there will always be a gap between frontier and local if you're comparing models at the same time, you can just do a lot more with terabytes of HBM than gigabytes of DDR. But will local models get good enough to be usable for useful work? For many folks, they already are.
Agreed, but at their current prices Deepseek + GLM are clear winners in my book. This weekend I spent $5 between the two where as I'd probably have to pay $20-30 to Anthropic (and that's still with the massive VC subsidies).
For web development (or anything else with an extreme amount of training data) it's number one for sure. You can't beat it at its costs. US companies will not be able to compete on a competitive market, which is why they rely on so much US government protection + corporate welfare.
There is no Claude 4 Opus model... It's a series of model, of which the strongest is Opus 4.8, and Qwen 3.6 35B-A3b gets 51.5% on Swe-bench pro to Opus 4.8's 69.2%
But it is still available on Google Vertex according to OpenRouter (though it's possible that info is just out of date, it's currently quoting 3tps which is unusably slow): https://openrouter.ai/anthropic/claude-opus-4
People can't seem to agree on what "Opus class" even means (the latest Opus is apparently pretty weak) but DeepSeek Pro, Kimi and GLM all are quite capable.
Nothing compares to Opus when it comes to "taste" in web design in my experience. Nothing compares to opus in very difficult HPC/model inference development. I worked on this with opus: https://github.com/computerex/dlgo
OpenAI was offering 2x usage at one point and I still used opus just because it's so much more effective.
Right. Local models haven't quite hit that level yet. The biggest open models, which you need tens of thousands of dollars of hardware to run at reasonable speed, have pretty much hit that level of capability, but most models you can reasonably run at home aren't quite there yet. But given the gap, if local models keep improving, you'd expect to maybe see that level by this November.
My understanding is that we could in fact run the largest models on "reasonable" home hardware by focusing on throughput rather than raw speed and having them do unattended inference in large batches. The big proprietary suppliers have no interest in this because their own incentive is to fill all the physical space available with top-performing hardware and doing huge amounts of inference as quickly as possible. A home user with limited hardware investment has very different constraints.
To me totally yes, even further, if they keep their existing route, over time people will stop using Anthropic.
More and more specialized and ultra-performant chips are going to flood the consumer market. Especially once new hardware foundries will start producing (well if we don't die from WW3 in the interval).
In 10 years from now, when even basic computers will have 128 GB of memory, and phones will have super optimized tuned models, then what will be the point of Anthropic ?
Just use Gemma/Gemini/Siri or whatever.
Pornography and uncensored models is also pushing toward local models.
It's not like needs of people grows exponentially, the needs follow an asymptote instead (they are capped).
The real revolution is offline robots and self-driving cars, but LLMs are already quite maxed.
For programmers, now, what Anthropic offers is like 3% improvement on a known test (like this pelican riding a bicycle), or on questions leaked from benchmark insiders.
It's ok but not like revolutionary (Fable was better but it was unusable, easy 20 minutes per one prompt due to overthinking).
@greenpants, "Pi coding harness but containerized and sandboxed" care to address some specifics and/or reference implementation for this. may be a GITHub URL?
What IDE do you use? How do you integrate it? I was using Continue but it exited its funding round to the Titler octopus and the Chat function in VSCode is choking on the Ollama responses.
That's precisely why my agent use is IDE-agnostic: I run Pi in any terminal. Often use it with the terminal inside VSCodium, though sometimes in a terminal outside an IDE if I don't expect to edit any files myself (e.g. for small one-shot projects).
Based on your explanation, it doesn't sound feasible for me, a complete non-engineer, to switch to fully offline? I do a lot of back and forth discussion with LLMs as someone who reads and writes 0 code.
about the edit tool it is almost always trailing white spaces. if you give it a skill with a sed 's/( )*$//g' or something like that it speeds up things
The thing is, to do a proper fix it would really need all of the context (maybe the tool call that failed was for an edit to a file that was last touched way at the beginning of the context), so you'd need to either keep that smaller model running doing prompt processing all the time, or have a very long wait while it does prompt processing on your whole session.
And then also, sometimes the tool call errors are because of something like a file was changed out from under it; the larger model is probably going to do a better job of figuring that out and fixing it up.
Finally, in Pi, you can always just use the /tree command to skip back to before a series of failed tool calls, with a summary if you want to let the model know what happened. The Pi /tree command is pretty powerful in managing your context
An illustrative example I've seen a lot is creating Jira tickets in projects with custom fields marked as mandatory. It tries to create the ticket without the field and the tool call fails. The LLM needs access to the full context so that it can generate text to put in the "Why couldn't this meeting be an email?" field.
I'm actually quite sure that directly retrying the tool call would often fix the edit-call already. But these models have been trained to "think" for a while for any problem solving, so they'll presume the problem of the edit is more fundamental and spend unnecessary tokens filling up the context.
I'll experiment more with the effectiveness of AGENTS.md rules for local Pi agents. I feel like smaller (local) LLMs just lack in attentiveness to elements in the context window, like precise instructions, compared to e.g. Claude models.
4-5 bit quants would probably fit pretty well on your rig. Check HuggingFace for Qwen3.6-35B-A3B-MTP-GGUF [1]. They've also got a cool UI thing these days to help indicate which quants of a model will run on your hardware.
Full octane isn't gonna fit on much of anything south of a 128GB machine once adding KV cache.
> is like a junior with knowledge across the board, that you really need to guide, versus a senior that thinks with you on architecture
I don't want to be rude, but your linkedin has a sumtotal (generous) of like 8 months of programming as a profession (job title is AI Engineer). The rest is at best programming adjacent. How would you know what either of these situations are really like?
I haven't logged in to LinkedIn or looked at it since a former employer demanded that everyone create a profile. So mine is now about 20 years out of date.
I've dabbled a bit in GitHub Copilot using Claude Opus and Sonnet models via work, but I couldn't shake the thought that we weren't allowed to use this on any of our clients' codebases. Having been a fan of Ollama, I wanted to try something truly local.
First I tried OpenCode but they unexpectedly make external requests (!) even when using Ollama (I noticed when Ollama wasn't properly connected and I still got a title generated).
So I settled for Pi, but I strongly disliked the idea that the agent could, at any point, decide to delete files or exfiltrate .env secrets. So I created Picosa (https://github.com/GreenpantsDeveloper/Picosa), containerizing and sandboxing Pi, with firewall rules such that it could only ever reach the local network (for Ollama), scoped by just the current working directory, and nothing else. Combined with Qwen3.6:35b, it works surprisingly well, and I could ask it to improve itself when run on its own repository.