Catalog

Voice AI models and benchmarks

Every model Speko has measured, and the 73 of 73 catalog entries a request can pin today. Accuracy, latency and price as each board published them — nothing on this page is estimated to fill a column.

Provider

66 models

Speech-to-text19

#ModelProviderId
1PulseSmallest AI
5.1%8.0%178ms~$0.0050
2Nova-3Deepgram
12.0%12.9%106ms$0.0048
3FluxDeepgram
not measured6.6%406ms$0.0065
4Ink-2Cartesia
11.0%9.9%102ms$0.0090
5Qwen3-ASRAlibaba
2.8%4.0%424ms$0.0054
6Chirp 3Google
3.9%7.4%581ms$0.0160
7Scribe v2 RealtimeElevenLabs
not measured3.4%233ms$0.0065
8Realtime STT-1Inworld
3.3%3.6%139ms$0.0025
9Universal-3.5 ProAssemblyAI
2.0%2.0%66ms$0.0075
10Velma 2Modulate
4.4%5.4%1.11s$0.0010
11stt-rt-v5Soniox
7.5%7.3%78ms$0.0020
12GPT-4o TranscribeOpenAI
2.3%5.8%572ms$0.0060
13GPT Live TranscribeOpenAI
not measured4.5%1.12s$0.0170
14GPT-4o-mini TranscribeOpenAI
2.7%6.4%460ms$0.0030
15Grok STTxAI
4.8%10.9%305ms$0.0033
16Gradium ASRGradium
8.4%11.7%334ms$0.0104
17Solaria-1Gladia
5.0%11.4%596ms$0.0125
18Muse Voice TranscribeMeta
not measured4.6%not measured$0.0030
19Gemini 3.5 Transcribe LiveGoogle
not measurednot measured1.22s~$0.0090

LLM14

#ModelProviderId
1gpt-5.6-lunaOpenAI
94.21%659ms70%0%$1.20
2gpt-4.1OpenAI
91.40%640ms59%10%$8.00
3gpt-5.6-terraOpenAI
87.37%701ms66%0%$12.00
4gpt-4.1-miniOpenAI
87.19%708ms40%57%$1.60
5gpt-oss-120b (Baseten)Baseten
86.32%393ms71%27%$0.50
6gpt-oss-120b (Cerebras)Cerebras
85.00%195ms68%30%$0.75
7Claude Haiku 4.5Anthropic
82.46%532ms2%0%$5.00
8Claude Sonnet 5Anthropic
81.72%1.21s40%0%$10.00
9DeepSeek-V4-Flash-0731Baseten
80.18%361ms5%10%$0.26
10gemma-4-31bCerebras
77.72%192ms53%0%$1.49
11gemini-3.8-flashGemini
77.37%not measured69%0%$3.75
12GLM-4.7Baseten
72.46%275ms2%17%$2.20
13Llama-3.3-70BTogether
64.21%698ms74%0%$1.04
14inkling-smallBaseten
57.72%177ms72%0%$1.20

Text-to-speech23

#ModelProviderId
1eleven_v3_conversationalElevenLabs
1590266ms$100.0
2gemini-3.1-flash-tts-previewGemini
1591978ms~$33.3
3aura-2Deepgram
1584125ms$30.0
4fluxDeepgram
~1550106ms$45.0
5sonic-3.5Cartesia
1574121ms$50.0
6sonic-3.6Cartesia
~1578120ms$50.0
7tts-rt-v1Soniox
1569362ms~$13.0
8tts-rt-v2Soniox
~1606381ms~$13.0
9simba-3.2Speechify
1573345ms$10.0
10inworld-tts-2Inworld
1561116ms$25.0
11inworld-tts-2-flashInworld
~150182ms$15.0
12s2.1-proFish Audio
~1566185ms$15.0
13lightning_v3.1Smallest
1544173ms$25.0
14grok-ttsxAI Grok
1495272ms$15.0
15Gradium TTSGradium
~1585272ms$57.8
16gpt-4o-mini-ttsOpenAI
1424691ms~$20.0
17arcanav3Rime
1429238ms$40.0
18speech-2.8-hdMiniMax
1431294ms$100.0
19octave-2Hume
1377448ms$100.0
20qwen3-tts-flashQwen
1310472ms$10.0
21palabra-tts-v1Palabra
~152072ms$30.0
22bland-speechBland
~1569303ms$15.0
23Maya 2 NativeMaya
not measured100ms~$4.0

Speech-to-speech10

#ModelProviderIdPriceas published
1grok-voice-think-fast-2.0xAI
0.800.67820ms$0.08 · min
2gemini-3.1-flash-liveGemini
0.770.781.11s$3.00 / $12.00 · 1M tok
3gpt-realtimeOpenAI
0.760.83494ms$32 / $64 · 1M tok
4gpt-realtime-2.1-miniOpenAI
0.750.831.01s$10 / $20 · 1M tok
5grok-voice-fastxAI
0.720.441.11s$0.05 · min
6gpt-realtime-2OpenAI
0.680.871.10s$32 / $64 · 1M tok
7gpt-realtime-2.1OpenAI
0.630.561.10s$32 / $64 · 1M tok
8gpt-realtime-miniOpenAI
0.620.70614ms$10 / $20 · 1M tok
9grok-voice-think-fast-1.0xAI
0.560.221.01s$0.05 · min
10gpt-4.1-nano (cascade)Inworld
0.451.00not measurednot measured

Measured per language

24 studies across 12 languages.

Arabic speech-to-text

#ModelCERbatch
1GPT-4o Transcribe2.7%
2Universal-3.5 Pro3.3%
3Qwen3-ASR3.5%
4Qwen3-ASR Realtimenot measured
5Scribe v2 Realtimenot measured
6Grok STTnot measured
7GPT-4o-mini Transcribe4.7%
8S35.0%
9stt-rt-v55.1%
10Nova-37.1%
11Ink-Whisper10.2%

Arabic text-to-speech

#ModelNaturalnessElo
1grok-tts1736.00
2sonic-3.51644.00
3tts-rt-v21612.00
4eleven_v3_conversational1564.00
5inworld-tts-21454.00

Published stacks

What benchmarks.speko.ai picks for each job off the boards above, and the measurement that decided it. A stack whose legs are identical to another’s is one row: the publisher distinguishes four jobs and its boards distinguish two stacks. A stack with a leg Router cannot currently call is not shown at all — a measurement may stay visible in that state, an instruction you would copy may not.

Use case
Real-time phone agentdecided on 66ms · 192ms · 116ms
Speech-to-textUniversal-3.5 Proassemblyai:universal-3-5-pro
LLMgemma-4-31bcerebras:gemma-4-31b
Text-to-speechinworld-tts-2inworld:inworld-tts-2
Accuracy-criticaldecided on 2.0% WER · 0% fabrication · 0.93 robustnessTool-heavy agentdecided on 2.0% WER · 2.7% tool silence · 116ms
Speech-to-textUniversal-3.5 Proassemblyai:universal-3-5-pro
LLMClaude Haiku 4.5anthropic:claude-haiku-4-5
Text-to-speechinworld-tts-2inworld:inworld-tts-2
Natural conversationdecided on 2.0% WER · 1.6% dead-air · MOS 1590
Speech-to-textUniversal-3.5 Proassemblyai:universal-3-5-pro
LLMClaude Haiku 4.5anthropic:claude-haiku-4-5
Text-to-speecheleven_v3_conversationalelevenlabs:eleven_v3_conversational

How to read this

Every model id is copyable. Click an id to copy it. An id offered by the current Router catalog can be sent as model; other ids identify benchmark evidence without claiming that the route is currently available. A direct benchmark label means the provider published no Router model id for that row.

Not measured is not zero. A dash in a column of word error rates reads as a perfect score, so an unmeasured cell says so in words. A leading ~ is the board’s own mark for an estimate.

Nothing here re-ranks a board. The # column indexes the order you sorted into. It is not a verdict.

The catalog endpoint did not answer this request, so the routable column comes from the bundled snapshot. The boards are unaffected.