What it costs to hear your coding agent (measured, September 2026)

By Kevin Hofmann · Updated

Hearing a coding agent's reply costs about half a cent to 2 cents on OpenRouter, measured in September 2026. The voice is 1.5 to 3 cents per 1,000 spoken characters, a spoken summary is a few hundred characters, and the summary itself is free on a Mac with Apple Intelligence or a fraction of a cent in the cloud.

I built Pintail, a Mac app that reads Claude Code replies out loud, and it runs on the user's own OpenRouter account. So "what will this cost me?" is the first question, and "it depends" is a bad answer. Here are the numbers I measured.

What does the voice cost per reply?

OpenRouter bills text to speech by the character. These are the voices I tested, with OpenRouter's prices on 29 September 2026:

Voice Per 1,000 spoken characters
Fish about 1.5 ¢
Microsoft Flash about 1.5 ¢
Voxtral (hosted) 1.6 ¢
Microsoft HD 2.2 ¢
Deepgram 3 ¢

The trick is that you don't speak the whole reply. A long agent reply is easily 3,000 characters; the spoken version at a normal length is a few hundred. At 1.6 ¢ per 1,000 characters, 400 spoken characters cost about 0.6 ¢.

If you do want the whole reply read, a 3,000-character reply is about 5 cents with Voxtral. That adds up fast, which is one reason Pintail reads a summary by default.

Are there free voices?

Yes, with limits. OpenRouter lists free voices with daily caps. Fish S2.1 Pro Free allows 1,000 requests a day, or 50 a day on accounts with less than $10 of credit, and there's a free Deepgram voice for English. Good for trying things out; for a full day of agent work you'll probably hit the cap.

What does the summary cost?

Before anything is spoken, a small model condenses the reply. On macOS 26 Pintail does that on the Mac with Apple Intelligence, which costs nothing. In the cloud, here's what I measured for a reply of about 3,000 characters at the "Detailed" level (five to ten sentences):

Summary model First words after Cost per reply
Gemini 2.5 Flash-Lite 0.4–0.7 s $0.0002
Mistral Small 0.5 s $0.0003
Gemini 3.1 Flash-Lite 0.8–1.0 s $0.0005
GPT-5.4 mini 0.7–1.2 s $0.002
Claude Haiku 4.5 1.1–1.4 s $0.003
Claude Sonnet 5.5 1.8–1.9 s $0.007

Sonnet sounded the most natural, but for a listener who hears it once, Haiku and Gemini 2.5 Flash-Lite are both fine. I wrote about which models followed the prompt and which invented things separately.

The "first words after" column matters more than the price. Speech starts as soon as the summary starts streaming, so a model that takes two seconds to say anything feels slow even if it's cheap.

Why not use a reasoning model?

They think before they answer, and you hear the thinking as silence. In my tests Grok 4.3 took about 8 seconds to its first word, Gemini 3.8 Flash about 11 seconds and DeepSeek V4 Flash about 30. By then you've already switched tabs and read the reply yourself.

Worked examples

Put the two halves together for one reply at the Detailed level:

  • Gemini 2.5 Flash-Lite + Voxtral: about 1.5 ¢.
  • Claude Haiku 4.5 + Voxtral: about 3 ¢.
  • Apple Intelligence + Voxtral: only the voice, usually under a cent for a short summary.

Per day it depends on how much you listen. Fifty replies at a cent each is 50 cents.

Does it use my Claude tokens?

No. The summary runs outside the agent, on its own model, and Pintail's voice commands are caught by the hook before they reach Claude Code. Your Claude plan pays for the coding, OpenRouter pays for the listening.

How to check the current prices

Prices move. Everything above was measured on 29 September 2026 with OpenRouter's prices at the time. Check the model pages on OpenRouter before you trust my table, and if you see it's out of date, tell me and I'll update this post.

Pintail's own price is planned at €15 once, with 100 free read-outs to try it. The pricing page has the same tables and stays current.

Hear it first

Pintail isn’t out yet. Leave your email and you’ll get the signed Mac build first.