Comparison
Turn detection for voice agents: built into Omni vs DIY
Omni packages turn detection and grounded retrieval inside the voice agent. You do not buy a separate turn-detection product or tune silence timers per accent and line.
Built into Omni
Included- Turn detection (endpointing) the moment speech ends
- Grounded KB context retrieved inline
- Included in the Omni voice-agent stack
- No separate product to price and integrate
Build it yourself
Your time + infra- Tune your own VAD / silence thresholds per accent and line
- Stand up + host a vector DB and retrieval service
- Glue transcription, endpointing, and retrieval yourself
- Own the latency, scaling, and on-call
Use Omni for the whole loop, or build each component yourself
Omni ($0.05/min for speech + brain) gives you the whole agent end to end: hearing, turn-taking, grounding, reasoning, and speaking. If you keep your own LLM and TTS, use Hear where English transcription fits and own the remaining turn and retrieval logic in your application. Cue grounding is not active on the serving Hear stream, is unavailable, and has no published launch date.
Get an API key. Make a phone agent.
Sandbox keys never bill. Fund a live key when you go to production.