Text-to-speech that starts in milliseconds.
Speak is PyAI's text-to-speech (TTS) API: low-latency streaming synthesis with 200 stock voices, free instant voice cloning, and prompt-to-voice design. First audio bytes arrive in tens of milliseconds over the same /v1/audio/speech contract your OpenAI client already speaks.
- Price
- $0.04/min
- Endpoint
POST /v1/audio/speech- Scope
speak:synthesize- Model id
pyai-speak
Hear it before you build with it.
The box on the right calls the live Speak API. Pick a voice, type a line, and listen.
Open the PlaygroundWant the full thing - Hear, Omni, code export?Open the Playground
Start in minutes
curl https://api.pyai.com/v1/audio/speech \
-H "Authorization: Bearer $PYAI_KEY" \
-d '{"model":"pyai-speak","input":"Hello from PyAI.","voice":"stock_ava_en_us","response_format":"mp3"}' \
--output hello.mp3Speaker diarization: know who said what
Turn recorded conversations into speaker-labelled transcripts with PyAI Hear.
Speak creates spoken audio. For speaker diarization of a recording, pair it with Hear: identify speaker turns and their timestamps before using the transcript in your application.
Language support
Speak across every supported language
Use one PyAI product surface across the language your customers speak.
English
English
Explore Speak
Español
Spanish
Explore Speak
Français
French
Explore Speak
हिन्दी
Hindi
Explore Speak
Deutsch
German
Explore Speak
Italiano
Italian
Explore Speak
Português
Portuguese
Explore Speak
Русский
Russian
Explore Speak
日本語
Japanese
Explore Speak
한국어
Korean
Explore Speak
ગુજરાતી
Gujarati
Explore Speak
తెలుగు
Telugu
Explore Speak
ਪੰਜਾਬੀ
Punjabi
Explore Speak
ಕನ್ನಡ
Kannada
Explore Speak
বাংলা
Bengali
Explore Speak
اردو
Urdu
Explore Speak
Kreyòl ayisyen
Haitian Creole
Explore Speak
普通话
Mandarin Chinese
Explore Speak
العربية
Arabic
Explore Speak
What you get
Snappy by default
Audio streams from the first byte (~53 ms warm / ~98 ms cold time-to-first-audio-byte, in-region), so playback starts almost instantly.
Your voice, free to clone
Cloning enrollment and prompt-to-voice design are free, and a cloned voice streams its first audio byte in ~32 ms warm (in-region).
Drop-in OpenAI shape
The same /v1/audio/speech contract, with OpenAI preset voice names mapped to PyAI voices, so existing TTS code runs unchanged.
Capabilities
- ~53 ms warm TTFB (in-region)
- 200 stock voices
- Voice cloning (free)
- Designed voices (free)
- Telephony formats (G.711)
- OpenAI drop-in
Time-to-first-audio-byte ~53 ms warm / ~98 ms cold on the streaming path, measured in-region (first audio byte, not first spoken word); progressive playback the whole way through. Server-side formats (mp3, wav, opus, pcm, and 8 kHz G.711 µ-law/A-law) mean telephony callers skip the resampler and µ-law encoder.
FAQ
Can I use speaker diarization with Speak?
PyAI supports speaker diarization through Hear async transcription jobs. Speak generates audio from text; Hear identifies speaker turns in recorded audio. Use Hear to produce a speaker-labelled transcript, then use Speak when your application needs spoken output.
What is Speak?
Speak is PyAI's text-to-speech (TTS) API. It turns text into natural speech over a low-latency streaming path, with 200 stock voices, free voice cloning, and prompt-to-voice design. It speaks the OpenAI /v1/audio/speech contract, so it's a drop-in replacement for OpenAI TTS.
How much does Speak TTS cost?
Voice cloning enrollment and prompt-to-voice design are free - you pay only for the audio you synthesize, billed per minute. Check our pricing page for the latest rates.
How do I migrate from OpenAI TTS?
Point your existing OpenAI SDK at https://api.pyai.com/v1 with your PyAI key. The /v1/audio/speech request shape matches, and the OpenAI preset voice names (alloy, echo, fable, onyx, nova, shimmer) are accepted and map to PyAI stock voices - so most code runs unchanged.
How many voices are there?
200 stock voices, plus your own cloned voices and prompt-designed voices. Browse the live catalog at GET /v1/voices, or pass an OpenAI preset name for drop-in compatibility.
What does voice cloning cost?
Cloning enrollment is free, as is prompt-to-voice design. You pay only for the audio you later synthesize with the voice, at the standard per-minute Speak rate.
Can I get telephony-ready audio?
Yes. Speak encodes server-side into mp3, wav, opus, pcm, or 8 kHz G.711 µ-law/A-law, so phone callers get exactly the format the line needs without a client-side resampler or µ-law encoder.
Build with Speak today.
An instant sandbox key lets you test without billing. Create an account when you want saved projects, and fund a live key for production traffic.