Skip to content
AMD API

PyAI AMD API: know who answered in one line of TwiML

An AI answers the phone now. Google Call Assist picks up before the person does, says one short line, and hands back the turn, so person-or-machine detection waves it through as a human and your agent pitches into a transcript. PyAI AMD returns screening as its own outcome, alongside voicemail, IVR, hold music and dead numbers, with the reason attached to every decision. Keep your carrier, keep your code.

First 5,000 answered calls free every month. No card.

The whole integration
<Response>
  <Start>
    <Stream url="wss://api.pyai.com/v1/amd/stream">
      <Parameter name="api_key" value="YOUR_PYAI_KEY"/>
      <Parameter name="aggressiveness" value="0.25"/>
      <Parameter name="webhook" value="https://you/amd-events"/>
    </Stream>
  </Start>
  <!-- Standalone test: replace this pause with your existing call flow. -->
  <Pause length="30"/>
</Response>

That is the migration. The key rides a <Parameter> and is verified before any audio is processed. Decisions push mid-call on the same socket and to your webhook.

Works with what you already dial from

TwilioOne line of TwiML
SignalWireTwilio-format streams
TelnyxTwilio-format streams
PlivoStream bridge
VonageStream bridge
GenesysStream bridge
RingCentralStream bridge
Amazon ConnectKVS bridge

The endpoint speaks Twilio’s Media Streams wire format natively, so platforms that speak it connect directly. Platforms with their own stream format fork call audio to a wss:// bridge that forwards the frames, about a page of server code, authenticated with ?api_key=. All trademarks belong to their owners; listing means protocol compatibility, not endorsement.

45/45
Google Call Assist screeners caught, on real harvested calls
0/182
false screening fires across 32 gatekeepers and 150 human calls
~940 ms
earlier than Twilio AMD at the median, same 200 real calls
$0.004
per answered call after 5,000 free each month

Our own measurements on real customer traffic, not an SLA and not a vendor benchmark. The screening figures come from calls harvested out of live dialing and kept with audio and transcript, so the set can be re-scored whenever the model changes. The latency figure is one side-by-side run of 200 answered calls replayed at real-time pace against Twilio AMD’s recorded decision on the same calls.

The category shift

An AI answers the phone now. Legacy AMD calls it a human.

iPhone Call Screening and Google Call Assist pick up unknown calls with an on-device assistant before the person does. It sounds like a person because it is speaking, not beeping. Person-or-machine detection from 2015 has no label for it, so it picks the closest one it has: human. Your agent pitches into a transcript, and the spam flag lands on your caller ID. PyAI AMD returns it as screening so your dialer can behave and your numbers stay clean.

What comes back

Every outcome is a different next move

One person-or-machine label forces one crude branch. The class plus the subtype lets the dialer act: respect the screener, drop the voicemail after the beep, navigate the menu, scrub the dead number, keep the human.

human

Connect the agent.

machine / screening

An AI screener answered. Say who you are and why. Do not run the pitch into a transcript.

machine / voicemail

Wait for the beep, drop the message, move on.

machine / ivr

An auto-attendant or phone tree. Navigate it or route to the right queue; do not leave a message on a menu.

machine / music

Hold music or a musical greeting with no words. Treat as a machine; it cannot fire on anyone who spoke.

sit_invalid / dead_number

The carrier's intercept says the number is dead. Scrub it from the list instead of retrying it.

fax

Fax tone on the line. Drop and flag the record.

unknown

No decisive evidence inside the window. Your default, and you set it.

The receipt

Every decision says why

The reason field quotes what was heard and when. When a campaign goes sideways at 2am, you read the reason, not a black-box label. answered_by_twilio speaks Twilio's exact vocabulary, so existing routing logic runs unchanged while you decide what to do with the richer fields.

Pushed mid-call on the socket and to your webhook
{
  "event": "amd",
  "call_id": "CA7d0…",
  "answered_by": "machine",
  "subtype": "screening",
  "answered_by_twilio": "machine_start",
  "confidence": 0.97,
  "decision_ms": 720,
  "reason": "AI call-screener: 'call assist' @720ms"
}

FAQ

How do you know the screening detection works?

We harvest real screened calls out of live traffic and keep the audio and transcript. On the Google Call Assist captures the screening outcome fires on 45 of 45, with zero false fires across 32 human gatekeepers and 150 ordinary human calls. Google screeners announce themselves in a near-closed vocabulary, which is why this is a high-precision problem rather than a guess. iPhone Call Screen detection ships too, matched on the prompts Apple's assistant uses, but we have not caught enough organic iPhone screeners in the wild to publish a recall number for it, so we do not.

Will it hang up on a live person?

That is the one failure that costs you a customer, so the design is bent around it. The thresholds are asymmetric on purpose, and the rules that fire machine from audio alone are gated so they cannot fire on someone who spoke. In our side-by-side on 1,341 real calls where both engines committed to an answer, Twilio AMD marked a live person as a machine 6 times; PyAI AMD did 0. We publish that number rather than claim never. When the window expires with nothing decisive, the human-safe default hands you unknown and lets the agent listen instead of guessing machine.

Which telephony platforms does PyAI AMD work with?

Any platform that can fork call audio to a WebSocket. The endpoint speaks Twilio's Media Streams wire format natively, so Twilio is one line of TwiML, and SignalWire and Telnyx speak the same stream format. Plivo, Vonage, Genesys, and RingCentral fork call audio from their own media streaming features through a thin server-side bridge that forwards the frames; Amazon Connect reaches it the same way via Kinesis Video Streams. Server-side integrations authenticate with ?api_key= on the URL or the pyai-key subprotocol.

How is PyAI AMD billed?

Per answered call. The first 5,000 answered calls each month are free, then $0.004 per answered call. Calls that never get answered are not billed, and AMD is included at no charge when the call already runs on PyAI telephony.

Do I have to change my routing logic to leave Twilio AMD?

No. Every decision carries answered_by_twilio, which maps to Twilio's exact AnsweredBy enum, next to our own answered_by class and subtype. Route on the Twilio field on day one, then read subtype when you are ready to act on screening, IVRs and dead numbers.

How fast is the decision?

We do not publish an absolute latency SLA. What we measured: on 200 real answered calls replayed at real-time pace through the live endpoint, PyAI's decision landed about 940 ms earlier at the median than Twilio AMD's recorded decision on the same calls, and about 1.5 s earlier on the calls a person answered. Our own measurement on one run, not a guarantee. decision_ms measures processed audio; when you set a decision timeout, decision_elapsed_ms measures elapsed time to prepare the result. Webhook delivery takes additional time.

Can I limit how long AMD waits?

Yes. Set decision_timeout_ms on the stream to 3000 for a three-second decision budget or 5000 for five seconds; the supported range is one to fifteen seconds. The clock starts when PyAI accepts the authenticated start. Earlier decisions return immediately; if the evidence available at the cutoff is inconclusive, AMD returns unknown instead of guessing. Shorter budgets can increase unknowns. Allow additional time for webhook delivery.

Point one call at it. Read the reason.

A free key, one line of TwiML, and the first 5,000 answered calls each month on us. Dial a Pixel with Call Assist on and watch it come back as screening instead of human. That is the whole evaluation.