These are all interrelated points and the sibling comment is correct: it's a skill.
The key thing with these voice systems is that you do not need to edit anything. You can literally just stream of consciousness into them, no editing, include the backtracking, the live-revisions, etc., and it will actually all produce vastly better context for the LLM than the written thing you took even 30 seconds to edit for clarity or brevity.
I had a very similar disposition towards this idea just 6 months ago. I highly recommend trying it out. The key thing is that you do not need to edit. Just keep talking. Try it for a few weeks!
I mean we don't need to do any epidemiological studies here or anything.
If someone hasn't tried it, they should. It's probably quite different from how they're expecting, might be great, and costs basically nothing. Try it for a few days and if it doesn't work in your workflow, obviously don't do it.
But I have encountered many many people who raised these exact same arguments against trying it, then tried it, and were hooked within days. Exactly 0% of people I've ever convinced to try it decided it just wasn't for them and went back to typing full-time.
What do you mean how you never write? You can just write what for want to say instead of saying it out loud. It is not a special skill, it's no different than how you write a message, just that you don't hit backspace to go back and correct things.
> it's no different than how you write a message, just that you don't hit backspace to go back and correct things.
Yes, that.
If you do that when writing to another person, you come off as blabbering, incoherent moron. Fortunately, very few people do that, because communicating with people who write like they talk is incredibly hard.
I'd worry about trying to do this on purpose; feels like the kind of "learning" that can easily spill over to non-LLM communication and make your life much harder.
There's a middle ground, though.
When shooting IM messages, people sometimes make typps
ypos*
typos**
there's an append-only procedure for fixing that, which I just demonstrated.
Also they don't always finish everything in one message
in fact a nice thing about IMs is being able to not use full sentences
that would work with LLMs, if not for the annoying "feature" that persists in most harnesses:
conversations take turns, and you can just send message after message - you have to wait for LLM to finish.
You already have to juggle multiple writing styles between different people and use cases. It is easy to avoid submitting a pure stream of thought if you are writing a formal letter.
>there's an append-only procedure for fixing that, which I just demonstrated.
LLMs can figure out most typos on your own.
>you have to wait for LLM to finish.
I don't think any coding harnesses work like this. They let you send more messages to steer the model while it's working.
When you're making a big deal out of it being "much harder" because it's "how you never write" and they're saying that's just "not hitting backspace"? No, you don't agree.
How frequently do you write in a stream of consciousness and not hit backspace?
Maybe give a ballpark estimate of characters typed per week in this manner versus characters typed where you are doing some combination of: 1) thinking about what you're writing before you write it, 2) punctuating and formatting correctly, or 3) correcting your writing output?
Ridiculous proposition. And I type correctly at 110+ wpm.
How often I do it doesn't matter because it's such a trivial thing to switch. If you're gonna "try it for a few weeks" the part of you that has to learn the typing-specific parts of that method is about 1% of the difficulty.
It's really easy to ignore typos. And the way you have to approach thinking and correcting is the same whether you're typing or voicing.
If you can't just type the way you would just talk, and you find it notably hard, it's you that's being ridiculous.
"never" schmever. I never ramble unrestrained either. Deciding not to edit at all is a skill either way. If you just want the equivalent of voice in a normal way, you remove backspace and that's it.
> Have you tried the voice-based prompting, as I'm describing?
I've never prompted a thing. I can just see your distinction is nonsense. If you think it's hard you're doing it wrong.
And wow I did that post without revising a thing. Wow.
Maybe you lost the thread, but this is a conversation about using dictation (specifically modern dictation tools like Whisper derivatives) to prompt coding agents.
Do you think it suddenly becomes harder to avoid backspace when you're in a different text box?
The thing you're claiming is hard, writing exactly the way you would speak, is not hard.
There's also some additional benefit to going with the flow and not thinking about words much before saying them, but that's equally hard with text or voice.
I hadn't before. Then I started doing it just to show how easy it was.
It doesn't matter what text box you're typing in. The ability to type as you'd speak is easy. Without any extra delays or issues.
I hope you're not trying to argue that typing the same way into an AI prompt is harder than doing it into HN. It's just not hard in any situation. Voice isn't special.
I see. You've fallen into some weird pedantry to think what I'm saying isn't relevant to your argument. I hope you figure out my very simple meaning later, have a good night too!
Cute snark but no you don't, nor does anyone else who speaks to you. People self-repair their speech in 10% to nearly 40% of speech turns. The upper end is for cognitively demanding communication but the lower bound is very normal.
Good evidence of my point though on how natural this is. People literally don't even notice it as either the listener or the speaker.
Even when a listener is told specifically to listen for and detect errors or self-repairs in speech, listeners will not even detect 50% to 80% of minor repairs. Your brain literally doesn't even perceive them.
The key thing with these voice systems is that you do not need to edit anything. You can literally just stream of consciousness into them, no editing, include the backtracking, the live-revisions, etc., and it will actually all produce vastly better context for the LLM than the written thing you took even 30 seconds to edit for clarity or brevity.
I had a very similar disposition towards this idea just 6 months ago. I highly recommend trying it out. The key thing is that you do not need to edit. Just keep talking. Try it for a few weeks!