Streaming speech-to-text, compared.
For live calls, transcription price is only half the story - turn-taking latency is what callers feel. Here's price next to the providers, plus the realtime piece most leaderboards miss.
Price per 1,000 minutes (USD)
Lower is better. PyAI in teal.
PyAI from our rate card. Competitor list prices: Artificial Analysis, as of June 2026; realtime/streaming tiers can differ from prerecorded - verify before relying.
Same fast class as Deepgram.
Lower is better. PyAI in teal. Bands show each provider’s published range.
Hear’s ~185 ms in-region first partial (revisable - stable_textfollows) sits inside Deepgram’s published 150-300 ms interim band. We don’t claim to beat Deepgram on latency - same fast class, verify in-region. (First-partial latency is separate from batch throughput, which Hear runs at 92-247x real-time.)
Turn-taking beats raw price.
On a phone call, the win is replying the instant the caller stops - not shaving a fraction of a cent. Start with Hear for transcription (~185 ms first partial, in-region), then go end-to-end with Omni (~390 ms voice-to-voice, in-region).
We don’t grade ourselves.
For an independent streaming accuracy and latency leaderboard across providers, see Artificial Analysis.
Streaming ASR leaderboardRealtime transcription that keeps up.
Start free with $50 in free credits - stream partials from one OpenAI-compatible socket.
Agents is live in beta. Sandbox keys have daily limits and never touch billing.