---
title: "The Router for Voice AI"
description: "Hosted Router for managed STT, LLM and TTS routing, plus an open customer-side Gateway for provider-direct voice workloads."
canonical: "https://speko.ai/"
last-updated: "2026-08-24"
---

> ## Speko page index
> The complete index of every page on this site is at: https://speko.ai/llms.txt
> Read it before exploring further. It lists the exact Markdown URL for every canonical HTML page.

# The Router for Voice AI

[Backed by Y Combinator](https://www.ycombinator.com/companies/speko)

Every speech model, benchmarked language by language, wired into one API.

[Get API key](https://platform.speko.ai/sign-in)

<a id="router"></a>

## Router

A hosted, provider-neutral STT, LLM and TTS data plane at router.speko.dev, with typed contracts and managed routing.

<a id="coverage"></a>

### Benchmark coverage by language

Model EN (English) ES (Spanish) DE (German) FR (French) AR (Arabic) FIL (Filipino) NB (Norwegian) HI (Hindi) TA (Tamil) TE (Telugu) KO (Korean) ZH (Chinese (Mandarin)) JA (Japanese) TH (Thai)

deepgram:nova-3

openai:gpt-4o-transcribe

soniox:stt-rt-v5

cartesia:ink-whisper

Alibaba:Qwen3-ASR bench

google:chirp_3

OpenAI:GPT-4o-mini Transcribe bench

assemblyai:universal-3-5-pro

openai:gpt-transcribe

Smallest AI:Pulse Pro bench

Deepgram:Nova-2 bench

xai:grok-stt

Alibaba:Qwen-ASR bench

ElevenLabs:Scribe v2 bench

Gladia:Solaria-1 bench

gradium:default

modulate:velma-2-stt-streaming

ElevenLabs:Scribe v1 bench

OpenAI:GPT-4o mini Transcribe bench

alibaba:qwen3-asr-flash-realtime

elevenlabs:scribe_v2_realtime

AssemblyAI:Universal (batch tier) bench

azure:MAI-Transcribe-2

cartesia:ink-2

fish:transcribe-1

fish:transcribe-1-pro

Hamsa:S3 bench

modulate:velma-2-stt-streaming-english-v2

nari:qwen3-asr

nari:qwen3-asr-fast

smallest:pulse

speechmatics:enhanced

speechmatics:standard

Hover a cell for its number.

worse better

20 of 47

are only measured in English

Their rank in any other language is unknown — including the model that sits at the top of the English table.

5

different models win across 14 languages

No single model is best everywhere, so the right pick changes with the language your users speak.

<a id="benchmarks"></a>

### Score against cost, per stage

Accuracy (WER · conversational)

25ms 4.57s 12.0% 35.2%

Qwen3-ASR Fast 12.0% · 33ms

Qwen3-ASR 12.0% · 25ms

Finalize (ms, p50)

Model

Qwen3-ASR Fast nari:qwen3-asr-fast

12.0% 33ms

Qwen3-ASR nari:qwen3-asr

12.0% 25ms

Qwen3-ASR Realtime alibaba:qwen3-asr-flash-realtime

12.2% 889ms

Scribe v2 Realtime elevenlabs:scribe_v2_realtime

12.5% 1.75s

Gemini 3.5 Transcribe Live gemini:gemini-3.5-transcribe-live

13.1% 1.28s

GPT Live Transcribe openai:gpt-live-transcribe

13.5% 620ms

Grok Voice Transcribe 1.0 xai:grok-stt

13.7% 1.13s

Universal-3.6 Pro assemblyai:universal-3-6-pro

13.7% 505ms

Muse Voice Transcribe meta:muse-voice-transcribe-1.0

14.2% —

Velma 2 modulate:velma-2-stt-streaming-english-v2

14.4% 1.21s

Pulse smallest:pulse

14.5% 839ms

Universal-3.5 Pro assemblyai:universal-3-5-pro

14.5% 408ms

Enhanced speechmatics:enhanced

15.1% 4.49s

stt-rt-v5 soniox:stt-rt-v5

16.6% 340ms

Ink-2 cartesia:ink-2

16.7% 587ms

GPT Realtime Whisper openai:gpt-realtime-whisper

16.8% 1.14s

GPT Transcribe openai:gpt-transcribe

16.8% 817ms

Nova-3 deepgram:nova-3

17.0% 189ms

Standard speechmatics:standard

18.1% 4.57s

Chirp 3 google:chirp_3

18.4% —

Flux deepgram:flux-general-en

19.7% 712ms

flux-general-multi deepgram:flux-general-multi

21.4% 854ms

Gradium ASR gradium:default

24.2% 1.13s

Universal-Streaming English assemblyai:universal-streaming-english

26.2% 452ms

GPT-4o Transcribe openai:gpt-4o-transcribe

32.9% 848ms

Universal-Streaming Multilingual assemblyai:universal-streaming-multilingual

35.2% 495ms

transcribe-1-pro fish:transcribe-1-pro

— —

transcribe-1 fish:transcribe-1

— —

MAI-Transcribe-2 azure:MAI-Transcribe-2

— —

ink-whisper cartesia:ink-whisper

— —

gemini-3.5-transcribe gemini:gemini-3.5-transcribe

— —

velma-2-stt-streaming modulate:velma-2-stt-streaming

— —

default palabra:default

— —

<a id="gateway"></a>

## Gateway

The open customer-side runtime for LiveKit and Pipecat: provider-direct streaming, local BYOK credentials and optional Speko-managed routes.

<a id="integrate"></a>

### Run and observe your voice workers

Use the native Gateway integrations in your framework, or call the hosted Router through its public OpenAPI and AsyncAPI contracts.

```python
from livekit.agents import AgentSession
from speko_gateway.livekit import LLM, STT, TTS

session = AgentSession(
    stt=STT(credential_source="auto"),
    llm=LLM(model="auto", objective="balanced"),
    tts=TTS(credential_source="auto"),
)
```

## Point your agent at Speko

```bash
$ claude mcp add --transport http speko https://mcp.speko.ai/mcp
```

[Read the docs](https://docs.speko.ai) [Get API key](https://platform.speko.ai/sign-in)
