· via dev.to (home feed)
OpenRouter adds ElevenLabs speech-to-text and text-to-speech to its model API
OpenRouter's October 7 launch puts ElevenLabs transcription and speech synthesis alongside its LLM catalogue under one API, simplifying voice features — though only for file-based audio.
What launched
OpenRouter, best known as a single API sitting in front of many language models, has added ElevenLabs audio models to its catalogue. According to a dev.to article covering the October 7 launch, apps can now call ElevenLabs speech-to-text and text-to-speech endpoints alongside the chat models they already reach through OpenRouter, which can simplify routing and billing for teams already sending traffic there.
The additions span nine ElevenLabs text-to-speech models and two speech-to-text models. The dev.to walkthrough highlights three of them as the main building blocks: elevenlabs/scribe-v2 for transcription, elevenlabs/eleven-v4-turbo for lower-latency spoken replies, and elevenlabs/eleven-v4 for more expressive narration. Per the announcement it describes, v4 Turbo is positioned for real-time-style use and priced at half the per-character rate of v4.
There is an important boundary: this is not a realtime voice connection. The launch includes no streaming version of Scribe v2 over WebSocket. The endpoints process audio files in a request/response pattern, so continuous back-and-forth audio remains a separate product problem requiring a different realtime path.
How the endpoints work
Transcription accepts either JSON input with base64-encoded audio and an explicit source format, or a multipart upload carrying file, model and response-format fields — the two shapes are alternatives, and the choice depends on how a client stores or uploads audio. Uploaded audio is capped at 25 MB, which the article notes covers roughly 27 minutes of 128 kbps MP3, but requests can still time out upstream after 180 seconds. The practical advice is to enforce tighter app-level size and duration limits, keep API keys on the server, and have clients upload to an authenticated application endpoint rather than calling OpenRouter directly.
Scribe v2 returns a transcript and can also provide word timestamps, speaker labels and audio-event tags. Its responses include a usage object reporting audio seconds and a dollar cost. For multi-speaker recordings, the walkthrough shows requesting verbose_ with word-level timestamps and enabling diarization, optionally passing a known speaker count.
On the output side, the speech endpoint takes text, a model and a voice. Setting response_format to mp3 produces 44.1 kHz, 128 kbps audio that most browser and mobile players can open, while omitting the field returns raw 16-bit little-endian mono PCM at 24 kHz — better suited to a custom audio pipeline than a downloadable reply. Eleven v4 Turbo has no speed control and rejects non-default speed settings; supported delivery tags exist, but because tags are text sent to the model, they count toward billing.
Two cost meters, not one
The two sides of the pipeline are billed on different units: transcription by audio duration, speech by input characters counted as Unicode code points, tags included. Those meters diverge in practice. A short command can sit inside a long recording full of pauses, and a short question can trigger a long spoken answer. The dev.to article recommends tracking the two sides separately, summarising transcription cost by duration bands that match the product rather than a single blended average, and recording failed, retried and abandoned requests in product metrics so provider dashboards don't hide real usage.
A temporary launch discount is also described, with an end time of October 19, 2026 at 8 a.m. Pacific Time. The article treats that as a dated promotion rather than a standing price, and advises rechecking live rates before projecting monthly spend.
Labels and latency need scrutiny
Speaker labels are model output, not verified identity. A speaker_0 tag should not be surfaced as a named person; if an application must name speakers, the mapping should go through an explicit, user-confirmed step. The article also stresses testing on audio that reflects real conditions — accents, room noise, microphones, vocabulary — and prioritising errors that change intent, such as a mistaken "can" versus "can't" or a wrong date, over average word-error scores.
Latency deserves measurement too. The path is serial — upload, transcribe, generate a chat reply, synthesise speech — and each stage adds delay, so end-to-end timing should be checked before anything is described as conversational.
Why it matters
For apps already routing LLM calls through OpenRouter, voice input and output can now live behind the same API key and billing account, removing a layer of integration overhead from voice features. The scope is the caveat: this launch handles recorded audio, not live sessions, so genuinely realtime products still need another path. And because costs are metered in audio seconds on one side and characters on the other, teams should model voice spending per feature rather than collapsing it into a single "voice usage" number.
- #openrouter
- #elevenlabs
- #speech-to-text
- #text-to-speech
- #ai-api