This is an early release preview. You may encounter bugs.

Give agents a voice

Speech recognition, synthesis and the parts of a spoken conversation.

37 tools · showing 1–24

Framework

LiveKit Agents

Voice

livekit

Framework for building realtime, server-side voice agents with pluggable STT, LLM, and TTS and WebRTC/telephony support

webrtctelephonyreal-timestt

A+
Tool

LiveKit

VoiceDeployment

livekit

WebRTC transport and infrastructure for real-time audio/video agents.

webrtcsfureal-timetelephony

A+
Framework

OpenAI Agents SDK

CodingVoiceSecurity

openai

Python framework for multi-agent workflows with handoffs, guardrails, sessions, sandbox and voice agents, and built-in tracing

multi-agentguardrailsobservabilitysandbox

A+
Framework

Pipecat

Voice

pipecat-ai

Python framework for real-time voice and multimodal conversational agents, with composable pipelines over WebSockets or WebRTC

real-timemultimodalwebrtcvoice-agent

A
Tool

Dograh

VoiceGenerative Media

dograh-hq

Self-hostable open-source voice AI platform for building inbound and outbound phone agents with a drag-and-drop workflow builder

local-firsttelephonyorchestrationpipecat

A
Tool

Moss

Memory

usemoss

Knowledge-retrieval service that injects semantic search results into an agent's LLM context.

ragsemantic-searchhybrid-searchlow-latency

A
Tool

HomeRail

VoiceCodingMonitoring

xiaotianfotos

Voice-first local agent orchestration runtime for auditable DAG workflows.

local-firstdagorchestrationhome-lab

B
Tool

Jarvis

VoiceMemoryInterface

isair

A 100% private AI voice assistant that lives on your computer (works offline). Talk naturally as if Jarvis is a third person in the room, an

productivityvoice-agentwake-wordsttlocal-first

B
Tool

Neuphonic

VoiceGenerative Media

neuphonic

Text-to-speech via Neuphonic's API.

ttsvoice-cloningon-devicegguf

B
Tool

Finchvox

MonitoringVoice

finchvox

Voice-AI observability that unifies conversation audio, logs, traces, and metrics for Pipecat agents.

observabilitypipecattelemetrysession-replay

B
Tool

Chatterbox

VoiceGenerative Media

resemble-ai

Open zero-shot voice-cloning TTS family with multilingual, low-latency Turbo, and CPU-scale Nano variants

ttsvoice-cloningmultilingualzero-shot

C
Tool

Inworld

VoiceGenerative MediaTraining

inworld-ai

Realtime speech-to-speech and expressive TTS-2 text-to-speech.

ttsfine-tunespeech-to-speechreal-time

C
App

Airi

VoiceGenerative Media

moeru-ai

Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds.

avatarlive2dai-companioncharacter-ai

Tool

AssemblyAI

Voice

Real-time and file speech-to-text via AssemblyAI's Universal models.

sttreal-timevoice-agent

Score unavailable
Tool

Async

VoiceGenerative Media

Text-to-speech via Async's WebSocket and HTTP APIs.

ttswebsocketreal-timelow-latency

Score unavailable
Tool

Awaaz AI

Voice

Telephony media-stream serializer for hosting voice agents over the Awaaz AI websocket.

financetelephonywebsocketmultilingualvoice-agent

Score unavailable
Tool

Beyond Presence

VoiceGenerative Media

Real-time video-avatar transport for voice agents.

videoavatarreal-timelivekitpipecat

Score unavailable
Tool

Cartesia

VoiceGenerative Media

Real-time speech-to-text and text-to-speech via Cartesia's streaming APIs.

ttssttreal-timestreaming

Score unavailable
Tool

Daily

Voice

WebRTC transport for real-time audio/video agent communication.

videowebrtcpipecatreal-time

Score unavailable
Tool

Deepdub

VoiceGenerative Media

Text-to-speech via Deepdub AI's streaming WebSocket API.

ttsvoice-cloningspeech-to-speechstreaming

Score unavailable
Tool

Deepgram

VoiceGenerative Media

Real-time speech-to-text (Nova/Flux) and Aura text-to-speech.

sttttsreal-timevoice-agent

Score unavailable
Tool

ElevenLabs

VoiceGenerative Media

Expressive text-to-speech with word-level timing plus file-based transcription.

ttssttvoice-cloningvoice-agent

Score unavailable
Tool

Exotel

Voice

Serializer for the Exotel WebSocket media-streaming protocol.

telephonywebsocketvoice-agent

Score unavailable
Tool

Gradium

VoiceGenerative Media

Low-latency streaming speech-to-text and text-to-speech via Gradium.

sttttsreal-timevoice-cloning

Score unavailable

More ways in