AI voice agent platforms provide the runtime for building real-time AI voice agents. They coordinate speech-to-text (STT), a large language model (LLM), and text-to-speech (TTS) into a single conversational loop, handling turn detection, interruption, and audio transport over telephony or WebRTC so that a spoken exchange stays within conversational latency.
These products are consumed as infrastructure rather than as finished applications. A team adopting one is building voice into a product they own: an in-app assistant, a drive-through ordering system, a healthcare intake line, a voice interface in a consumer device. The platform supplies the conversational plumbing rather than the business logic.
The qualifying capability must be the product's primary reason to exist. A speech-to-text or text-to-speech service that added an agent wrapper belongs in its component category; a CPaaS provider that terminates calls belongs in CPaaS; a finished phone-answering product sold to a business belongs in AI voice assistants. Whether the platform is configured through code, a dashboard, or both is a packaging decision and does not determine membership.
Buyers are engineering teams embedding voice into their own products, building inbound and outbound phone agents, in-app voice assistants, drive-through and quick-service ordering, healthcare intake, appointment scheduling, and voice interfaces in consumer devices. This distinguishes the category from packaged AI voice assistant products by both buyer and form factor: a developer composing or configuring a voice agent, not a business buying a finished phone-answering tool.
To qualify for inclusion in the AI Voice Agent category, a product must:
- Provide real-time STT, LLM, and TTS orchestration in a single conversational loop with sub-second latency, as the product's primary function
- Offer a developer API or SDK for building voice agents programmatically, not solely a fixed no-code configuration UI aimed at non-technical business buyers
- Provide real-time audio transport via WebRTC, WebSocket, or telephony (SIP/PSTN), built in or through documented integration, as core to the product
- Support programmable conversation control, where the developer defines the agent's behavior, prompts, tools, and flow, rather than configuring a fixed template
- Support function calling or tool use so the agent can take action during a live conversation