Skip to content

PyAI Hear / Speaker diarization

Speaker diarization: know who said what.

Turn recorded conversations into speaker-labelled transcripts with timestamps. PyAI Hear supports speaker diarization for mono audio and channel separation for stereo call recordings.

Choose the right speaker separation for your recording

Mono recordings

Set diarize: true on a Hear async transcription job. Speaker labels are inferred from the audio and identify turns within that recording. Check that the returned segments have speaker labels before relying on them.

Dual-channel calls

Set channel: true when each participant has a separate stereo channel. Channel 0 maps to speaker_1 and channel 1 to speaker_2. Keep your call metadata to identify participant roles.

From a recording to a useful conversation record

  1. Submit an audio URL or file to POST /v1/transcription/jobs with one speaker separation option.
  2. Poll the job until it is completed, failed, or cancelled. A queued response means processing has been accepted.
  3. Read the result, checking for speaker-labelled segments and timestamps. Use the transcript for call review, searchable conversation records, or Recap input after mapping participant roles.

This is an async recording workflow. Read the transcription jobs guide for formats, retention and webhook handling.

Speaker diarization, speech generation and call summaries

Hear transcribes recordings and labels speakers. Speak turns text into spoken audio. Recap analyzes a supplied transcript. Combine these capabilities when your application needs both conversation understanding and spoken output.

Speaker diarization FAQ

What is speaker diarization?

Speaker diarization groups speech by speaker and marks when each speaker talks. Combined with transcription, it helps answer who said what and when in a recording. It does not verify a person's identity.

Does PyAI offer a speaker diarization API?

Yes. Submit a recording to Hear using POST /v1/transcription/jobs with diarize: true for mono audio. Poll the returned job ID until processing finishes, then read the returned speaker-labelled segments.

Should I use diarization or stereo channel separation?

Use channel: true when each participant already occupies a separate stereo channel. Use diarize: true for mono recordings. Do not set both options. Channel labels describe the source channel; they do not infer agent or customer roles.

Is speaker diarization part of Speak?

Speak converts text to speech. Speaker diarization belongs to Hear's async recording transcription workflow. Use Hear for speaker-labelled transcripts and Speak when your application needs to generate spoken audio.

Can I use it for call analysis?

Speaker-labelled transcripts help reviewers follow calls and locate speaker turns by timestamp. Your application can prepare the transcript for Recap, but must establish participant roles from call context before mapping them to agent or customer.

What limitations should I handle?

Mono speaker labels are model-derived and are not stable identities across jobs. A job can complete with a transcript but no speaker labels when diarization cannot be produced. Check that segments have speaker labels before using them. The Trace transcription option cannot be combined with diarize or channel in the same job.

How do I migrate a diarization workflow?

Map your recording submission to Hear async jobs, choose mono diarization or stereo channel separation, and adapt your result handling to PyAI's job statuses, speaker labels and timestamps. Compare outputs on representative recordings before switching production traffic.

See current pricing for Hear usage and plan availability.