Writing for the ear: turning long agent replies into something you can listen to

By Kevin Hofmann · Updated

A good spoken summary of an agent's reply puts the truth first, says every number and symbol the way a person would say it, and is exactly as long as you asked. Those three rules, plus asking for the length twice, are most of the prompt Pintail uses to turn a Claude Code reply into something you can listen to.

Reading and listening are different. When you read, you skim, jump back, look at the code block. When you listen, you hear it once, in order, at the speed of speech. So the summary can't be a shorter version of the written reply. It has to be written for the ear.

Why not ask Claude to write a spoken version?

It was the first idea. Add "also give me a one-line spoken version" to the system prompt, and read that part aloud. I dropped it quickly:

  • It clutters every written reply with a line you only need when you're listening.
  • It costs tokens on every turn.
  • Turning it on and off means editing your setup.

So a separate small model writes the spoken version after the reply is done. The agent doesn't know it exists.

The four lengths

Pintail has four levels. Here's one reply, about users getting logged out after every deploy, at each:

  • Brief (one sentence): "Fixed, just set SESSION_SECRET before the next deploy."
  • Summary (two to four sentences): "Deploys changed the session secret. Fixed. Set SESSION_SECRET before the next deploy."
  • Detailed (five to ten sentences): "Every deploy made a new session secret, so people got logged out. Local dev never rotates it, which is why you only saw it after deploys. The secret now stays the same across releases, and the old one still works for a day. Two new tests cover it, and all 214 pass. Set SESSION_SECRET in production before the next deploy."
  • Full reply: the whole thing, read as written.

Notice what survives at every length: what happened, and what you need to do. The Brief version drops the why. The Detailed version keeps it.

Rule 1: truth first

The listener can't check. If the summary says "done", they'll believe it. So:

  • A plan stays a plan. If the agent proposed something and didn't do it, the summary must not say it's done.
  • No invented questions. If the reply didn't ask you anything, the summary doesn't either.
  • No promises the agent didn't make.

This sounds obvious. It wasn't. In one test, Mistral Small called finished work a plan and invented a question at the end. Both are worse than no summary, because you act on them.

Rule 2: ask for the length twice

I tell the model the length before the reply, and then again after it. With the instruction only at the top, Claude Haiku regularly ignored it on long replies: the reply is thousands of characters, and by the end the instruction is far away. Repeating it after the reply fixed that.

Rule 3: write it the way you'd say it

Text to speech models read what's there. So the summary has to spell out anything that isn't a plain word:

  • "46/46" becomes "46 of 46".
  • "−16" becomes "minus 16".
  • "⌃⌥" becomes "Control Option".
  • snake_case_names become words.
  • File names and symbols are said like a person would say them, or left out.

This matters more than it seems. One voice model I use doesn't just mispronounce symbols, it can babble for a minute on them. I wrote up what broke when a TTS model read long replies.

Which models followed it?

I ran the same replies through several models on OpenRouter:

  • Claude Haiku 4.5 followed the prompt best, once the length was repeated.
  • Claude Sonnet 5.5 sounded the most natural. It's slower and costs more.
  • Mistral Small was fast and cheap, but it's the one that turned done work into a plan and invented a question.
  • Gemini 2.5 Flash-Lite was the fastest to its first words.
  • Reasoning models were too slow to use: 8 to 30 seconds before the first word, because they think first.

On a Mac with macOS 26, Pintail uses Apple Intelligence by default, so the summary is written on the Mac. The cloud models are there for people who want a different voice in the writing, or don't have macOS 26 yet. The cost post has the numbers.

What I'd tell anyone doing this

Write the prompt for a listener, not a reader, and test it by listening, not by reading the output. Too many numbers in a row, a file path read out character by character, "the following changes:" followed by nothing you can picture: all of those look fine on screen and sound wrong. If it sounds like a colleague telling you what they did, it's right.

Hear it first

Pintail isn’t out yet. Leave your email and you’ll get the signed Mac build first.