Collection
Give agents a voice
Speech recognition, synthesis and the parts of a spoken conversation.
37 tools · showing 1–24
LiveKit Agents
Voice
livekit
Framework for building realtime, server-side voice agents with pluggable STT, LLM, and TTS and WebRTC/telephony support
webrtctelephonyreal-timestt
LiveKit
VoiceDeployment
livekit
WebRTC transport and infrastructure for real-time audio/video agents.
webrtcsfureal-timetelephony
OpenAI Agents SDK
CodingVoiceSecurity
openai
Python framework for multi-agent workflows with handoffs, guardrails, sessions, sandbox and voice agents, and built-in tracing
multi-agentguardrailsobservabilitysandbox
FrameworkPipecat
Voice
pipecat-ai
Python framework for real-time voice and multimodal conversational agents, with composable pipelines over WebSockets or WebRTC
real-timemultimodalwebrtcvoice-agent
ToolDograh
VoiceGenerative Media
dograh-hq
Self-hostable open-source voice AI platform for building inbound and outbound phone agents with a drag-and-drop workflow builder
local-firsttelephonyorchestrationpipecat
ToolMoss
Memory
usemoss
Knowledge-retrieval service that injects semantic search results into an agent's LLM context.
ragsemantic-searchhybrid-searchlow-latency
ToolHomeRail
VoiceCodingMonitoring
xiaotianfotos
Voice-first local agent orchestration runtime for auditable DAG workflows.
local-firstdagorchestrationhome-lab
ToolJarvis
VoiceMemoryInterface
isair
A 100% private AI voice assistant that lives on your computer (works offline). Talk naturally as if Jarvis is a third person in the room, an
productivityvoice-agentwake-wordsttlocal-first
ToolNeuphonic
VoiceGenerative Media
neuphonic
Text-to-speech via Neuphonic's API.
ttsvoice-cloningon-devicegguf
ToolFinchvox
MonitoringVoice
finchvox
Voice-AI observability that unifies conversation audio, logs, traces, and metrics for Pipecat agents.
observabilitypipecattelemetrysession-replay
ToolChatterbox
VoiceGenerative Media
resemble-ai
Open zero-shot voice-cloning TTS family with multilingual, low-latency Turbo, and CPU-scale Nano variants
ttsvoice-cloningmultilingualzero-shot
ToolInworld
VoiceGenerative MediaTraining
inworld-ai
Realtime speech-to-speech and expressive TTS-2 text-to-speech.
ttsfine-tunespeech-to-speechreal-time
AppAiri
VoiceGenerative Media
moeru-ai
Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds.
avatarlive2dai-companioncharacter-ai
AssemblyAI
Voice
Real-time and file speech-to-text via AssemblyAI's Universal models.
sttreal-timevoice-agent
Async
VoiceGenerative Media
Text-to-speech via Async's WebSocket and HTTP APIs.
ttswebsocketreal-timelow-latency
Awaaz AI
Voice
Telephony media-stream serializer for hosting voice agents over the Awaaz AI websocket.
financetelephonywebsocketmultilingualvoice-agent
Beyond Presence
VoiceGenerative Media
Real-time video-avatar transport for voice agents.
videoavatarreal-timelivekitpipecat
Cartesia
VoiceGenerative Media
Real-time speech-to-text and text-to-speech via Cartesia's streaming APIs.
ttssttreal-timestreaming
Daily
Voice
WebRTC transport for real-time audio/video agent communication.
videowebrtcpipecatreal-time
Deepdub
VoiceGenerative Media
Text-to-speech via Deepdub AI's streaming WebSocket API.
ttsvoice-cloningspeech-to-speechstreaming
Deepgram
VoiceGenerative Media
Real-time speech-to-text (Nova/Flux) and Aura text-to-speech.
sttttsreal-timevoice-agent
ElevenLabs
VoiceGenerative Media
Expressive text-to-speech with word-level timing plus file-based transcription.
ttssttvoice-cloningvoice-agent
Exotel
Voice
Serializer for the Exotel WebSocket media-streaming protocol.
telephonywebsocketvoice-agent
Gradium
VoiceGenerative Media
Low-latency streaming speech-to-text and text-to-speech via Gradium.
sttttsreal-timevoice-cloning