Hacker Newsnew | past | comments | ask | show | jobs | submit | nvch's commentslogin

The harness instructs them to behave this way. Also this approach saves tokens. The scripts allow to edit files in bulk, and most of the session cost is in cache reads (e.g. for 300K context each command costs the same as 30K input tokens).

> The harness instructs them to behave this way. Also this approach saves tokens. The scripts allow to edit files in bulk, and most of the session cost is in cache reads (e.g. for 300K context each command costs the same as 30K input tokens).

I understand the reasoning, but at that point wouldn't the LLM be better off creating `sed` commands and executing those? I mean, if it's already executing Python, it can literally do anything to the environment, so using `sed` is at least as safe, with a bonus that it (or a subagent, or a human) can double-check the intention with the sed script and flag incorrect or missing changes.


I've experimented quite a bit with giving agents python vs sed + awk. They make mistakes with both, a lot. The only thing that has stood out is that agents reach for python too quickly if it's available, and that awk causes the least problems, while sed might take several attempts to get results, similar to python.

Also it's the only way that makes sense when you need to work with big files, or large amount of files, or documents that look small when fetched through a RAG tool, but then you read one and get hit with couple megabytes of base64-encoded binary data you didn't expect because RAG tool stripped out embedded images...

Ask me how I know. Or don't. I have a standing rule for all agents warning about that failure mode (and related, doing `ls` in `/tmp` and few other directories that like to accumulate files by the hundreds..)


My harness forbids it, they end up spending time debugging their scripts

I was wondering why Fable's 5.1 writing in Claude Code became even more unreadable, and found that they added "No em-dashes, no parentheticals, no arrows" to its system prompt.

For starters, if you repeat a specific prompt multiple times per day, you may save it as a skill.

Claude will do this for you after a few times. But yes, I have a skill called plan-to-epic which creates a Jira epic and ticket per milestone. It helps my agents persist context and, because I’m terrible at competing with my coworkers for “visibility,” means I can point to all my work if asked.

What if the app asks one question per swipe?

Advanced version: and shows a match with the same answer.


When I was not 17 at the times of GPT2, I decided to not bother with learning how to build LLMs because it’s too expensive for an individual. This escalated quickly.


Looking how much agents like to use and push to have functional a11y trees, we may accidentally solve accessibility as well


Haha, one day, one day...


12 tok/s and almost instant response on M1 Max Mac Studio (with faster SSD than laptops) are impressive – gives hope that large models may run locally from SSDs instead of memory.


Thanks for sharing! SSD read speed is the biggest limiting factor here, unfortunately


We have speed limits for vehicles because speed kills.

In computing, waiting kills (indirectly, by wasting time). Speed is life.

Some roads have minimum speed limits. If we're talking about limits, that's the kind of limit we want.


I am partial to this speed-is-life sentiment, and it made me think of this beautiful and wonderfully poignant demoscene production by Farbrausch and Haujobb: "Time Index". It is made in the memory of a friend of theirs that died.

The softsynth soundtrack includes lyrics and one of them is "we slow down" which I always interpreted as a kind of demoscener's lament since making things go fast is sorta the whole point!

https://youtu.be/fngv1dCFrdo?si=tR-1uQ4vKIPLHZu3

I am not arguing for either side of the fastness debate, while I certainly adore fast computers I don't like the mental image of our computers all blazing away doing stuff mainly humans care about, while the the world 'outside' is steadily getting hotter and more polluted.


I had learnt that trick, so now I explicitly disallow Fable subagents.

Yesterday, I wanted to review a complex piece after a large refactoring, and requested a review plan beforehand. The first step was 8 agents + one more to verify the findings (all Fable). Looks good, approved.

The verification step turned into an attempt to throw a party with 41 Fable verifiers.

It will find a way.


Don't do that; limit concurrency


"LGTM"

That'll be $50 — please.


I'm waiting for the day when Claude will figure out to use em dashes, en dashes or dashes depending on whether the user is nice or unpleasant, or write notes in the unallocated disk space.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: