---
title: "Voice AI models and benchmarks"
description: "Measured voice AI benchmarks across speech-to-text, LLM, text-to-speech, and speech-to-speech, with current Router route availability kept separate from historical evidence."
canonical: "https://speko.ai/models"
last-updated: "2026-08-24"
---

> ## Speko page index
> The complete index of every page on this site is at: https://speko.ai/llms.txt
> Read it before exploring further. It lists the exact Markdown URL for every canonical HTML page.

# Voice AI models and benchmarks

Catalog

Every model Speko has measured, and the 73 of 73 catalog entries a request can pin today. Accuracy, latency and price as each board published them — nothing on this page is estimated to fill a column.

Provider

78 models

## Speech-to-text19

\# Model Provider Id

1 Pulse Smallest AI

5.1% 8.0% 178ms ~$0.0050 (Streaming Pulse. The vendor never publishes an exact rate, and its own pricing page contradicts this model page with ~$0.009/min.)

2 Nova-3 Deepgram

12.0% 12.9% 106ms $0.0048 (Streaming pay-as-you-go. Vendor flags it as a limited-time promotional rate; pre-recorded is $0.0043.)

3 Flux Deepgram

not measured 6.6% 406ms $0.0065 (Flux English streaming pay-as-you-go. Vendor flags it as a limited-time promotional rate; flux-general-multi is $0.0078.)

4 Ink-2 Cartesia

11.0% 9.9% 102ms $0.0090 (Ink-2 bills 3 credits/sec on Pro, the cheapest paid plan; $0.0071 at Startup, $0.0067 at Scale. The 1 credit/sec rate is Ink-Whisper, a different model.)

5 Qwen3-ASR Alibaba

2.8% 4.0% 424ms $0.0054 (qwen3-asr-flash-realtime, Singapore. The file/sync SKU is $0.0021.)

6 Chirp 3 Google

3.9% 7.4% 581ms $0.0160 (Speech-to-Text V2 Standard, first tier — one rate for streaming, sync and batch. Dynamic Batch Recognition is $0.003.)

7 Scribe v2 Realtime ElevenLabs

not measured 3.4% 233ms $0.0065 (Scribe v2 Realtime at $0.39/hr, flat across every tier. Batch Scribe v2 is $0.0037.)

8 Realtime STT-1 Inworld

3.3% 3.6% 139ms $0.0025 (On-demand $0.15/hr. The $0.10/hr headline needs a paid Creator subscription.)

9 Universal-3.5 Pro AssemblyAI

2.0% 2.0% 66ms $0.0075 (Universal-3.5 Pro Realtime at $0.45/hr. Async pre-recorded is $0.0035. Diarization and other add-ons bill on top.)

10 Velma 2 Modulate

4.4% 5.4% 1.11s $0.0010 (Published streaming rate, $0.06/hr. Batch is $0.03/hr. Not a measurement.)

11 stt-rt-v5 Soniox

7.5% 7.3% 78ms $0.0020

12 GPT-4o Transcribe OpenAI

2.3% 5.8% 572ms $0.0060

13 GPT Live Transcribe OpenAI

not measured 4.5% 1.12s $0.0170 (gpt-live-transcribe, the streaming path this row is recommended for.)

14 GPT-4o-mini Transcribe OpenAI

2.7% 6.4% 460ms $0.0030

15 Grok STT xAI

4.8% 10.9% 305ms $0.0033 (grok-stt streaming at $0.20/hr. REST is $0.0017.)

16 Gradium ASR Gradium

8.4% 11.7% 334ms $0.0104 (3 credits/sec on the XS tier, the cheapest paid plan; $0.0068 at the L tier.)

17 Solaria-1 Gladia

5.0% 11.4% 596ms $0.0125 (Starter pay-as-you-go real-time ($0.75/hr). Pre-recorded is $0.0102/min, and committed Growth pricing goes to $0.0042.)

18 Muse Voice Transcribe Meta

not measured 4.6% not measured $0.0030 (Vendor list price: $3.00 per 1,000 minutes, equivalently $0.18/hr. Streaming and batch are priced identically -- the table does not differentiate -- and there is a single tier, no plan or volume rates.)

19 Gemini 3.5 Transcribe Live Google

not measured not measured 1.22s ~$0.0090 (ESTIMATE, not a published per-minute rate. Google prices this model per token -- $3.50/1M audio in, $21.00/1M text out -- and the ~$0.005/min + ~$0.004/min figures on the pricing page are Google's own estimate at 25 audio tokens/sec input and 175 text tokens/min output, for a blended ~$0.009/min. Speech density moves it, so read it as an order of magnitude rather than a rate card. Every other price on this column is the vendor's published per-minute rate.)

## LLM14

\# Model Provider Id

1 gpt-5.6-luna OpenAI

94.21% (84–100% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 659ms 70% 0% $1.20

2 gpt-4.1 OpenAI

91.40% (78–100% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 640ms 59% 10% $8.00

3 gpt-5.6-terra OpenAI

87.37% (74–98% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 701ms 66% 0% $12.00

4 gpt-4.1-mini OpenAI

87.19% (78–95% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 708ms 40% 57% $1.60

5 gpt-oss-120b (Baseten) Baseten

86.32% (72–98% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 393ms 71% 27% $0.50

6 gpt-oss-120b (Cerebras) Cerebras

85.00% (70–96% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 195ms 68% 30% $0.75

7 Claude Haiku 4.5 Anthropic

82.46% (67–95% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 532ms 2% 0% $5.00

8 Claude Sonnet 5 Anthropic

81.72% (65–95% 95% CI bootstrap over items · n=93 · measured 2026-08-25) 1.21s 40% 0% $10.00

9 DeepSeek-V4-Flash-0731 Baseten

80.18% (71–90% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 361ms 5% 10% $0.26

10 gemma-4-31b Cerebras

77.72% (62–92% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 192ms 53% 0% $1.49

11 gemini-3.8-flash Gemini

77.37% (61–92% 95% CI bootstrap over items · n=95 · measured 2026-09-03) not measured 69% 0% $3.75

12 GLM-4.7 Baseten

72.46% (59–85% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 275ms 2% 17% $2.20

13 Llama-3.3-70B Together

64.21% (43–84% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 698ms 74% 0% $1.04

14 inkling-small Baseten

57.72% (43–73% 95% CI bootstrap over items · n=95 · measured 2026-08-25) 177ms 72% 0% $1.20

## Text-to-speech35

\# Model Provider Id

1 eleven_v3_conversational ElevenLabs

1590 266ms $100.0

2 gemini-3.1-flash-tts-preview Gemini

1591 978ms ~$33.3

3 aura-2 Deepgram

1584 125ms $30.0

4 flux Deepgram

~1550 106ms $45.0

5 sonic-3.5 Cartesia

1574 121ms $50.0

6 sonic-3.6 Cartesia

~1578 120ms $50.0

7 tts-rt-v1 Soniox

1569 362ms ~$13.0

8 tts-rt-v2 Soniox

~1606 381ms ~$13.0

9 simba-3.2 Speechify

1573 345ms $10.0

10 inworld-tts-2 Inworld

1561 116ms $25.0

11 inworld-tts-2-flash Inworld

~1501 82ms $15.0

12 s2.1-pro Fish Audio

~1566 185ms $15.0

13 lightning_v3.1 Smallest

1544 173ms $25.0

14 grok-tts xAI Grok

1495 272ms $15.0

15 Gradium TTS Gradium

~1585 272ms $57.8

16 gpt-4o-mini-tts OpenAI

1424 691ms ~$20.0

17 arcanav3 Rime

1429 238ms $40.0

18 speech-2.8-hd MiniMax

1431 294ms $100.0

19 octave-2 Hume

1377 448ms $100.0

20 qwen3-tts-flash Qwen

1310 472ms $10.0

21 palabra-tts-v1 Palabra

~1520 72ms $30.0

22 bland-speech Bland

~1569 303ms $15.0

23 Maya 2 Native Maya

not measured 100ms ~$4.0

24 eleven_flash_v2_5 ElevenLabs

~1494 not measured $50.0

25 eleven_turbo_v2_5 ElevenLabs

~1466 not measured $50.0

26 eleven_multilingual_v2 ElevenLabs

~1472 not measured $100.0

27 eleven_flash_v2 ElevenLabs

~1507 not measured $50.0

28 eleven_turbo_v2 ElevenLabs

~1492 not measured $50.0

29 sonic-3 Cartesia

~1551 not measured $50.0

30 coda Rime

~1535 not measured $50.0

31 mistv3 Rime

~1428 not measured $30.0

32 octave-1 Hume

~1455 not measured $100.0

33 speech-2.6-hd MiniMax

~1450 not measured ~$100.0

34 speech-2.6-turbo MiniMax

~1408 not measured ~$60.0

35 lightning_v3.1_pro Smallest

~1447 not measured not measured

## Speech-to-speech10

\# Model Provider Id Price as published

1 grok-voice-think-fast-2.0 xAI

0.85 0.67 820ms $0.08 · min

2 gemini-3.1-flash-live Gemini

0.77 0.78 1.11s $3.00 / $12.00 · 1M tok

3 gpt-realtime OpenAI

0.79 0.84 494ms $32 / $64 · 1M tok

4 gpt-realtime-2.1-mini OpenAI

0.76 0.83 1.01s $10 / $20 · 1M tok

5 gpt-live-1 OpenAI

0.86 0.78 1.22s $0.05 · min

6 grok-voice-fast xAI

0.74 0.44 1.11s $0.05 · min

7 gpt-realtime-2 OpenAI

0.67 0.87 1.10s $32 / $64 · 1M tok

8 gpt-realtime-2.1 OpenAI

0.63 0.56 1.10s $32 / $64 · 1M tok

9 gpt-realtime-mini OpenAI

0.65 0.77 614ms $10 / $20 · 1M tok

10 grok-voice-think-fast-1.0 xAI

0.60 0.22 1.01s $0.05 · min

## Measured per language

26 studies across 13 languages.

### Spanish speech-to-text

\# Model Provider Id WER batch

1 GPT-4o Transcribe OpenAI

3.8%

2 Universal-3.5 Pro AssemblyAI

4.0%

3 Qwen3-ASR Alibaba

4.4%

4 GPT-4o-mini Transcribe OpenAI

4.6%

5 Chirp 3 Google

5.4%

6 Gradium ASR Gradium

6.7%

7 Ink-Whisper Cartesia

7.3%

8 stt-rt-v5 Soniox

7.3%

9 Nova-3 Deepgram

8.4%

10 Pulse Pro Smallest AI

8.4%

11 Nova-2 Deepgram

11.3%

12 Scribe v2 Realtime ElevenLabs

not measured

13 Qwen3-ASR Realtime Alibaba

not measured

14 Pulse Smallest AI

not measured

### German speech-to-text

\# Model Provider Id WER batch

1 GPT-4o Transcribe OpenAI

2.1%

2 Universal-3.5 Pro AssemblyAI

2.6%

3 Qwen3-ASR Alibaba

3.0%

4 GPT-4o-mini Transcribe OpenAI

4.1%

5 Scribe v2 Realtime ElevenLabs

not measured

6 Chirp 3 Google

5.0%

7 Qwen3-ASR Realtime Alibaba

not measured

8 Gradium ASR Gradium

5.7%

9 Ink-Whisper Cartesia

7.4%

10 Pulse Smallest AI

not measured

11 Pulse Pro Smallest AI

9.2%

12 Nova-3 Deepgram

11.2%

13 Nova-2 Deepgram

12.0%

14 stt-rt-v5 Soniox

13.4%

### French speech-to-text

\# Model Provider Id WER batch

1 Universal-3.5 Pro AssemblyAI

3.3%

2 GPT-4o Transcribe OpenAI

4.0%

3 Qwen3-ASR Realtime Alibaba

not measured

4 Qwen3-ASR Alibaba

4.4%

5 Scribe v2 Realtime ElevenLabs

not measured

6 GPT-4o-mini Transcribe OpenAI

6.8%

7 GPT Live Transcribe OpenAI

not measured

8 Chirp 3 Google

not measured

9 stt-rt-v5 Soniox

8.6%

10 Nova-3 Deepgram

9.9%

11 Gradium ASR Gradium

10.0%

12 Ink-Whisper Cartesia

10.4%

13 Pulse Smallest AI

not measured

14 Pulse Pro Smallest AI

11.8%

15 Nova-2 Deepgram

12.9%

### Arabic speech-to-text

\# Model Provider Id CER batch

1 GPT-4o Transcribe OpenAI

2.7%

2 Universal-3.5 Pro AssemblyAI

3.3%

3 Qwen3-ASR Alibaba

3.5%

4 Qwen3-ASR Realtime Alibaba

not measured

5 Scribe v2 Realtime ElevenLabs

not measured

6 Grok STT xAI

not measured

7 GPT-4o-mini Transcribe OpenAI

4.7%

8 S3 Hamsa

5.0%

9 stt-rt-v5 Soniox

5.1%

10 Nova-3 Deepgram

7.1%

11 Ink-Whisper Cartesia

10.2%

### Filipino speech-to-text

\# Model Provider Id WER batch

1 GPT-4o Transcribe OpenAI

6.1%

4 GPT-4o-mini Transcribe OpenAI

10.8%

5 Scribe v2 Realtime ElevenLabs

not measured

6 GPT Live Transcribe OpenAI

not measured

7 Velma 2 multilingual Modulate

12.5%

8 Qwen3-ASR Alibaba

18.8%

9 Qwen-ASR Alibaba

18.8%

10 Ink-Whisper Cartesia

20.7%

11 stt-rt-v5 Soniox

21.9%

12 Nova-3 Deepgram

25.6%

13 Chirp 3 Google

27.3%

### Norwegian speech-to-text

\# Model Provider Id WER batch

2 Scribe v1 ElevenLabs

5.7%

3 GPT-4o Transcribe OpenAI

6.4%

5 Scribe v2 Realtime ElevenLabs

not measured

6 Velma 2 multilingual Modulate

8.4%

7 stt-rt-v5 Soniox

8.6%

8 Universal-3.5 Pro AssemblyAI

10.1%

9 Chirp 3 Google

11.2%

10 Qwen3-ASR Alibaba

12.0%

11 Qwen-ASR Alibaba

12.0%

12 GPT-4o mini Transcribe OpenAI

14.4%

13 Qwen3-ASR Realtime Alibaba

not measured

14 Ink-Whisper Cartesia

14.5%

15 Nova-3 Deepgram

14.6%

16 Nova-2 Deepgram

19.1%

### Hindi speech-to-text

\# Model Provider Id WER batch

1 Chirp 3 Google

6.1%

2 Qwen3-ASR Alibaba

9.8%

3 stt-rt-v5 Soniox

12.3%

4 GPT-4o-mini Transcribe OpenAI

12.7%

5 Pulse Pro Smallest AI

13.4%

6 Universal-3.5 Pro AssemblyAI

17.2%

7 Nova-3 Deepgram

20.3%

8 GPT-4o Transcribe OpenAI

21.0%

9 Nova-2 Deepgram

25.1%

10 Ink-Whisper Cartesia

41.9%

### Tamil speech-to-text

\# Model Provider Id CER batch

1 Chirp 3 Google

7.4%

2 GPT-4o Transcribe OpenAI

10.1%

3 Pulse Pro Smallest AI

12.1%

4 Nova-3 Deepgram

14.2%

5 stt-rt-v5 Soniox

14.6%

6 GPT-4o-mini Transcribe OpenAI

16.1%

7 Ink-Whisper Cartesia

30.6%

### Telugu speech-to-text

\# Model Provider Id CER batch

1 Chirp 3 Google

3.2%

2 GPT-4o Transcribe OpenAI

7.5%

3 Pulse Pro Smallest AI

8.5%

4 stt-rt-v5 Soniox

9.8%

5 GPT-4o-mini Transcribe OpenAI

10.2%

6 Nova-3 Deepgram

13.4%

### Korean speech-to-text

\# Model Provider Id CER stream

1 stt-rt-v5 Soniox

not measured

2 Scribe v2 Realtime ElevenLabs

not measured

3 Grok STT xAI

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

not measured

4 Qwen3-ASR Alibaba

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

not measured

5 GPT-4o Transcribe OpenAI

not measured

6 Solaria-1 Gladia

not measured

8 Nova-3 Deepgram

not measured

### Chinese (Mandarin) speech-to-text

\# Model Provider Id CER stream

1 stt-rt-v5 Soniox

not measured

2 Qwen3-ASR Alibaba

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

not measured

3 Scribe v2 Realtime ElevenLabs

not measured

4 Grok STT xAI

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

not measured

6 GPT-4o Transcribe OpenAI

not measured

7 Nova-3 Deepgram

not measured

8 Solaria-1 Gladia

not measured

### Japanese speech-to-text

\# Model Provider Id CER batch

1 GPT-4o Transcribe OpenAI

1.6%

3 stt-rt-v5 Soniox

2.0%

4 Scribe v1 ElevenLabs

2.1%

5 Qwen3-ASR Alibaba

2.1%

6 Qwen-ASR Alibaba

2.1%

8 GPT-4o mini Transcribe OpenAI

2.8%

9 Universal-3.5 Pro AssemblyAI

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

not measured

10 Scribe v2 Realtime ElevenLabs

not measured

11 Velma 2 multilingual Modulate

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

4.8%

12 Pulse Pro Smallest AI

5.5%

13 Pulse Smallest AI

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

not measured

14 GPT Live Transcribe OpenAI

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

not measured

15 Solaria-1 Gladia

11.5%

16 Nova-3 Deepgram

20.7%

17 Ink-Whisper Cartesia

30.3%

18 Grok STT xAI

42.9%

### Thai speech-to-text

\# Model Provider Id CER batch

1 Qwen-ASR Alibaba

3.5%

3 Qwen3-ASR Alibaba

3.5%

5 Scribe v1 ElevenLabs

4.3%

6 Qwen3-ASR Realtime Alibaba

not measured

7 GPT Live Transcribe OpenAI

not measured

8 stt-rt-v5 Soniox

4.7%

9 Scribe v2 Realtime ElevenLabs

not measured

10 Velma 2 multilingual Modulate

5.1%

11 GPT-4o mini Transcribe OpenAI

5.2%

12 GPT-4o Transcribe OpenAI

6.1%

13 Chirp 3 Google

6.4%

14 Solaria-1 Gladia

7.9%

15 Nova-3 Deepgram

8.4%

16 Universal (batch tier) AssemblyAI

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

11.2%

17 Grok STT xAI

16.6%

18 Ink-Whisper Cartesia

24.1%

### Spanish text-to-speech

\# Model Provider Id Naturalness Elo

1 eleven_v3_conversational ElevenLabs

1951.00

2 sonic-3.5 Cartesia

1751.00

3 tts-rt-v2 Soniox

1696.00

4 grok-tts xAI Grok

1634.00

5 inworld-tts-2 Inworld

1598.00

6 octave-2 Hume

1568.00

7 tts-rt-v1 Soniox

1525.00

8 lightning_v3.1 Smallest

1487.00

9 speech-2.8-hd MiniMax

1460.00

10 qwen3-tts-flash Qwen

1437.00

11 gpt-4o-mini-tts OpenAI

1408.00

12 arcanav3 Rime

1172.00

13 aura-2 Deepgram

1055.00

### German text-to-speech

\# Model Provider Id Naturalness Elo

1 eleven_v3_conversational ElevenLabs

1798.00

2 grok-tts xAI Grok

1762.00

3 tts-rt-v1 Soniox

1717.00

4 tts-rt-v2 Soniox

1700.00

5 octave-2 Hume

1687.00

6 sonic-3.5 Cartesia

1665.00

7 inworld-tts-2 Inworld

1576.00

8 qwen3-tts-flash Qwen

1573.00

9 speech-2.8-hd MiniMax

1528.00

10 gpt-4o-mini-tts OpenAI

1497.00

11 arcanav3 Rime

1257.00

12 lightning_v3.1 Smallest

1254.00

13 Gradium TTS Gradium

1245.00

14 aura-2 Deepgram

1062.00

### French text-to-speech

\# Model Provider Id Naturalness Elo

1 eleven_v3_conversational ElevenLabs

1775.00

2 tts-rt-v2 Soniox

1695.00

3 grok-tts xAI Grok

1692.00

4 inworld-tts-2 Inworld

1684.00

5 octave-2 Hume

1671.00

6 tts-rt-v1 Soniox

1628.00

7 sonic-3.5 Cartesia

1622.00

8 gpt-4o-mini-tts OpenAI

1535.00

9 qwen3-tts-flash Qwen

1529.00

10 speech-2.8-hd MiniMax

1419.00

11 lightning_v3.1 Smallest

1416.00

12 arcanav3 Rime

1316.00

13 Gradium TTS Gradium

1233.00

14 aura-2 Deepgram

1114.00

### Arabic text-to-speech

\# Model Provider Id Naturalness Elo

1 grok-tts xAI Grok

1736.00

2 sonic-3.5 Cartesia

1644.00

3 tts-rt-v2 Soniox

1612.00

4 eleven_v3_conversational ElevenLabs

1564.00

5 inworld-tts-2 Inworld

1454.00

### Filipino text-to-speech

\# Model Provider Id Naturalness Elo

1 gemini-3.1-flash-tts-preview Gemini

2012.00

2 tts-rt-v2 Soniox

1790.00

3 eleven_v3_conversational ElevenLabs

1693.00

4 gpt-4o-mini-tts OpenAI

1676.00

5 grok-tts xAI Grok

1657.00

6 speech-2.8-hd MiniMax

1598.00

7 sonic-3.5 Cartesia

1597.00

8 inworld-tts-2 Inworld

1251.00

### Norwegian text-to-speech

\# Model Provider Id Naturalness Elo

1 sonic-3.5 Cartesia

1836.00

2 gemini-3.1-flash-tts-preview Gemini

1705.00

3 eleven_v3_conversational ElevenLabs

1628.00

4 tts-rt-v2 Soniox

1621.00

5 Azure AI Speech (native) Azure

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

1467.00

6 gpt-4o-mini-tts OpenAI

1420.00

7 speech-2.8-hd MiniMax

1386.00

### Hindi text-to-speech

\# Model Provider Id Naturalness Elo

1 sonic-3.5 Cartesia

1736.00

2 gemini-3.1-flash-tts-preview Gemini

1563.00

3 speech-2.8-hd MiniMax

1522.00

4 eleven_v3_conversational ElevenLabs

1490.00

5 Maya 2 Native Maya

1478.00

6 chirp-3-hd Google Chirp

1468.00

7 Azure AI Speech (native) Azure

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

1412.00

8 simba-multilingual Speechify

1310.00

### Tamil text-to-speech

\# Model Provider Id Naturalness Elo

1 gemini-3.1-flash-tts-preview Gemini

1676.00

2 sonic-3.5 Cartesia

1620.00

3 Maya 2 Native Maya

1603.00

4 chirp-3-hd Google Chirp

1555.00

5 eleven_v3_conversational ElevenLabs

1543.00

6 Azure AI Speech (native) Azure

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

1462.00

7 simba-multilingual Speechify

1144.00

### Telugu text-to-speech

\# Model Provider Id Naturalness Elo

1 sonic-3.5 Cartesia

1639.00

2 gemini-3.1-flash-tts-preview Gemini

1625.00

3 chirp-3-hd Google Chirp

1548.00

4 Maya 2 Native Maya

1540.00

5 Azure AI Speech (native) Azure

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

1419.00

6 simba-multilingual Speechify

1269.00

### Chinese (Mandarin) text-to-speech

\# Model Provider Id Naturalness Elo

1 inworld-tts-2 Inworld

1688.00

2 eleven_v3_conversational ElevenLabs

1675.00

3 gemini-3.1-flash-tts-preview Gemini

1649.00

4 sonic-3.5 Cartesia

1564.00

5 Azure AI Speech (native) Azure

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

1552.00

6 speech-2.6-hd MiniMax

1467.00

7 gpt-4o-mini-tts OpenAI

1365.00

### Korean text-to-speech

\# Model Provider Id Naturalness Elo

1 gemini-3.1-flash-tts-preview Gemini

1917.00

2 sonic-3.5 Cartesia

1859.00

3 inworld-tts-2 Inworld

1722.00

4 eleven_v3_conversational ElevenLabs

1628.00

5 gpt-4o-mini-tts OpenAI

1498.00

6 Azure AI Speech (native) Azure

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

1438.00

7 speech-2.6-hd MiniMax

1407.00

8 coda Rime

1187.00

9 generative AWS Polly

953.00

### Japanese text-to-speech

\# Model Provider Id Naturalness Elo

1 gemini-3.1-flash-tts-preview Gemini

1756.00

2 eleven_v3 ElevenLabs

1646.00

3 sonic-3.5 Cartesia

1600.00

4 inworld-tts-2 Inworld

1597.00

5 DragonHDLatestNeural Azure

1539.00

6 tts-rt-v2 Soniox

1447.00

7 speech-2.6-hd MiniMax

1441.00

8 chirp-3-hd Google Chirp 3 HD

1407.00

9 gpt-4o-mini-tts OpenAI

1338.00

10 octave-2 Hume

1228.00

### Thai text-to-speech

\# Model Provider Id Naturalness Elo

1 gemini-3.1-flash-tts-preview Gemini

1722.00

2 tts xAI Grok

1719.00

3 speech-2.6-hd MiniMax

1699.00

4 sonic-3.5 Cartesia

1688.00

5 eleven_v3 ElevenLabs

1605.00

6 tts-rt-v2 Soniox

1591.00

7 chirp-3-hd Google Chirp 3 HD

1442.00

8 Azure AI Speech (native) Azure

direct benchmark (Measured directly; no Router model id exists for this benchmark row.)

1280.00

9 inworld-tts-2 Inworld

1202.00

10 gpt-4o-mini-tts OpenAI

1084.00

## Published stacks

What benchmarks.speko.ai picks for each job off the boards above, and the measurement that decided it. A stack whose legs are identical to another’s is one row: the publisher distinguishes four jobs and its boards distinguish two stacks. A stack with a leg Router cannot currently call is not shown at all — a measurement may stay visible in that state, an instruction you would copy may not.

Use case Speech-to-text LLM Text-to-speech

Real-time phone agent decided on 66ms · 192ms · 116ms

Universal-3.5 Pro assemblyai:universal-3-5-pro

gemma-4-31b cerebras:gemma-4-31b

inworld-tts-2 inworld:inworld-tts-2

Accuracy-critical decided on 2.0% WER · 0% fabrication · 0.93 robustness Tool-heavy agent decided on 2.0% WER · 2.7% tool silence · 116ms

Universal-3.5 Pro assemblyai:universal-3-5-pro

Claude Haiku 4.5 anthropic:claude-haiku-4-5

inworld-tts-2 inworld:inworld-tts-2

Natural conversation decided on 2.0% WER · 1.6% dead-air · MOS 1590

Universal-3.5 Pro assemblyai:universal-3-5-pro

Claude Haiku 4.5 anthropic:claude-haiku-4-5

eleven_v3_conversational elevenlabs:eleven_v3_conversational

## How to read this

**Every model id is copyable.** Click an id to copy it. An id offered by the current Router catalog can be sent as`model`; other ids identify benchmark evidence without claiming that the route is currently available. A direct benchmark label means the provider published no Router model id for that row.

**Not measured is not zero.** A dash in a column of word error rates reads as a perfect score, so an unmeasured cell says so in words. A leading `~` is the board’s own mark for an estimate.

**Nothing here re-ranks a board.** The `#` column indexes the order you sorted into. It is not a verdict.

The catalog endpoint did not answer this request, so the routable column comes from the bundled snapshot. The boards are unaffected.
