Skip to content

Comparison

Turn detection for voice agents: built into Omni vs DIY

Omni packages turn detection and grounded retrieval inside the voice agent. You do not buy a separate turn-detection product or tune silence timers per accent and line.

Built into Omni

Included
  • Turn detection (endpointing) the moment speech ends
  • Grounded KB context retrieved inline
  • Included in the Omni voice-agent stack
  • No separate product to price and integrate

Build it yourself

Your time + infra
  • Tune your own VAD / silence thresholds per accent and line
  • Stand up + host a vector DB and retrieval service
  • Glue transcription, endpointing, and retrieval yourself
  • Own the latency, scaling, and on-call

Use Omni for the whole loop, or build each component yourself

Omni ($0.05/min for speech + brain) gives you the whole agent end to end: hearing, turn-taking, grounding, reasoning, and speaking. If you keep your own LLM and TTS, use Hear where English transcription fits and own the remaining turn and retrieval logic in your application. Cue grounding is not active on the serving Hear stream, is unavailable, and has no published launch date.

Get an API key. Make a phone agent.

Sandbox keys never bill. Fund a live key when you go to production.