Text to speech models are tuned for sentences, and coding-agent replies aren't sentences. When I pointed Voxtral at my Claude Code replies, it babbled for up to 70 seconds on symbols, drifted on long input, and different voices came out 13 dB apart. The fixes were boring and they work: send whole sentences in small chunks, drop symbols, level the loudness, and let the player, not a clock, drive the captions.
This is what I learned building Pintail, a Mac app that reads coding agents' replies aloud.
Why Voxtral?
I tried Apple's built-in say first. It sounds robotic, and I didn't want to listen to it all day. Then I compared Qwen3-TTS, Voxtral, Supertonic 3 and Kokoro on an M5 Pro, reading the same replies. Voxtral won by ear. Everything below is about Voxtral, but I'd expect the same failure modes from any model trained on normal speech.
Problem 1: it babbles on single words and symbols
Give Voxtral a bare word and it sometimes doesn't stop. I sent it the German word "zu" six times. The audio came back between 0.5 and 41 seconds long. Same input, same voice.
Symbols were worse. The string "⟲5 ■ ⟳5" produced 70 seconds of audio. And long input drifted.
On the Mac, the local server stops every request at 96 seconds as a safety net. A cap stops the worst case, but it doesn't fix it.
Fix: whole sentences, small chunks, no symbols
- Split at sentence ends. Never send a fragment. A fragment is where the babbling starts.
- Start early. The first chunk goes out as soon as it has 40 characters and a sentence end, so the voice can start while the rest is still being written.
- Keep the rest at 400 characters or less. Short enough that drift doesn't get going.
- Cap each chunk's length. If the audio is much longer than the text could possibly take, cut it.
- Drop symbols. The summary prompt already writes things out ("46 of 46", "Control Option"); anything left that isn't a word gets removed before it reaches the voice.
For scale: a 2,600-character reply became 182 seconds of audio in 9 chunks. On the Mac it generated 3 times faster than real time; the hosted version was 6.6 times faster. Either way the listener never waits for the next chunk.
Problem 2: voices differ by about 13 dB
Switch from one voice to another and the volume jumps. Across the voices I use, the difference was about 13 dB, and the German voices were louder. That's the difference between "comfortable" and "reach for the volume key".
Fix: level everything to one loudness
The player measures each chunk and levels it to −24 dBFS RMS, with a limiter after it so peaks don't clip.
One gotcha with Apple's AUPeakLimiter: its decay parameter accepts values outside the documented range, and applies them. I had 0.2 seconds in there. The result was the gain held down for about 300 ms after every peak, which you hear as the voice ducking after loud consonants. Check every value you pass it against the documented range; it won't.
Problem 3: captions drift
Pintail shows live captions in a pill at the top of the screen, with the current word highlighted. My first version assumed a fixed speaking rate. That's wrong: the voices I measured range from about 12 to about 21 characters per second. With a fixed rate the captions ran ahead or fell behind within a sentence.
Fix: the player drives the captions
The player reports how many bytes of audio it has actually played. The captions map that to a position in the text and follow it. If playback stalls, the captions stall. If you change the speed to 1.5×, they keep up.
Problem 4: the first word gets eaten
The hosted voice's first bytes arrive after about 0.5 to 0.7 seconds, and up to 1.4 seconds when it's cold. If the pill appears and starts the caption before the sound, it looks broken. If the audio starts at full volume, the first syllable clicks.
Fix: wait for the first sound, then fade in a little
The pill opens with the first sound, not before. The player puts 250 ms of silence in front and fades in over 100 ms. I tried a longer fade; it ate the first word.
Problem 5: German with an English accent
Hosted Voxtral has no German voice. Its English voices read German with a strong accent. Microsoft's Klaus, Deepgram's Julius and Fish speak German properly, so German summaries go to one of those.
What I'd do again
Treat the voice model as something that needs clean, sentence-shaped input and will misbehave on anything else. Most of the work wasn't the voice at all: it was the summary that feeds it (how I write for the ear) and the player that plays it. The voice itself costs about half a cent to 2 cents a reply.