Skip to content
All migration guides
Speech-to-text

Switch from Deepgram to PyAI Hear

Streaming transcription tuned for 8 kHz call audio.

Hear is telephony-native speech-to-text with eager streaming partials over a single WebSocket. You map Deepgram's interim/is_final results to PyAI's partial / partial_stable / speech_final / final frames and force-finalize a turn with {"type":"commit"}. Cue grounding is unavailable on the serving Hear stream.

Why teams move from Deepgram

Tuned for the phone

Hear is tuned on narrowband 8 kHz call audio, not studio takes - accuracy holds up on real lines.

Half-price batch

Queue recordings to the async jobs API at $0.0005/min - 50% off the realtime rate - with webhooks on completion.

Explicit turn boundaries

Send {"type":"commit"} when your own VAD decides the turn is complete, or close the socket to flush buffered audio.

Before and after

Before - Deepgram
import { createClient, LiveTranscriptionEvents } from "@deepgram/sdk";

const dg = createClient(process.env.DEEPGRAM_API_KEY);
const conn = dg.listen.live({ model: "nova-3", interim_results: true });
conn.on(LiveTranscriptionEvents.Transcript, (d) => {
  const text = d.channel.alternatives[0].transcript;
  if (d.is_final) commit(text); else render(text);
});
// pipe PCM16 audio frames: conn.send(chunk)
After - PyAI Hear
// Pass the key as a WebSocket subprotocol (browser-safe).
const ws = new WebSocket(
  "wss://api.pyai.com/v1/audio/transcriptions/stream" +
    "?protocol=pyai-hear-v1&model=pyai-hear&encoding=pcm16&sample_rate=16000&interim_results=true",
  ["pyai.v1", "pyai-key." + apiKey],
);
ws.onmessage = (e) => {
  const f = JSON.parse(e.data); // bare JSON frames
  if (f.type === "partial" || f.type === "partial_stable") render(f.text);
  else if (f.type === "speech_final" || f.type === "final") commit(f.text);
};
// stream PCM16 frames up; end a turn with:
// ws.send(JSON.stringify({ type: "commit" }))

Migration checklist

1

Swap the connection

Change the base URL or WebSocket URL, pass a PyAI key, and keep the old client where the API shape is compatible.

2

Map models, voices, and formats

Use the table below to replace model ids, voice ids, response formats, sample rates, and auth headers without rewriting the product flow.

3

Replay customer traffic

Run real prompts, recordings, and phone-call samples through both systems. Compare latency, quality, completion rate, and all-in cost.

4

Launch with guardrails

Start on free credits or test keys, add usage alerts, then enable Trace, Recap, or the Agents feature when calls need production review.

If the system you are leaving already uses OpenAI-compatible transcription, speech, or realtime APIs, start with the smallest compatible swap: base URL, key, model, and voice. The real test is whether PyAI improves cost, latency, and caller outcomes on your own traffic.

What maps to what

DeepgramPyAI
Authorization: Token <key>Subprotocols pyai.v1, pyai-key.<key> (or ?api_key=)
wss://api.deepgram.com/v1/listenwss://api.pyai.com/v1/audio/transcriptions/stream?protocol=pyai-hear-v1
interim_results=trueinterim_results=true (same)
Interim Results framepartial / partial_stable
is_final / speech_finalspeech_final, then corrected final
Finalize / CloseStream{"type":"commit"} (or close the socket)

Good to know

  • Frames are bare JSON ({type, text, utterance_id, t_ms, ...}); there is no session.started or usage frame.
  • {"type":"end"} is ignored - use {"type":"commit"} or close the socket to flush a final.
  • Close codes: 1000 normal, 1008 auth/policy, 1011 engine error (benign right after a final), 4429 over the concurrency cap.

FAQ

How do interim results map?

Deepgram interim results map to PyAI partial / partial_stable; is_final maps to speech_final, followed by a corrected full-context final.

How do I force end-of-turn?

Send the text frame {"type":"commit"} (e.g. when your VAD detects end-of-turn). Closing the socket also flushes a final for buffered audio.

Can I add knowledge-base grounding?

Not on the serving Hear stream. Cue grounding fields are reserved but inactive, with no published availability date. Use Omni when you need grounded realtime conversations.

Stream partials from one socket.

Start free with $50 in free credits - telephony-tuned streaming STT.