Collection
Give agents a voice
Speech recognition, synthesis and the parts of a spoken conversation.
94 tools · showing 25–48
Confer
VoiceConnectors
A complete AI assistant — writing, images, web search, memory. Every conversation is end-to-end encrypted, so no one else can read it.
productivitymeetingssttaction-items
OpenMontage
Generative MediaCoding
calesthio
Agentic video production system that drives an AI coding assistant through scripting, asset generation, editing, and rendering
videoffmpegremotionttsstable-diffusion
ToolMiniMax-MCP
VoiceGenerative Media
MiniMax-AI
Official MCP server exposing MiniMax text-to-speech, voice cloning, image and video generation APIs to MCP clients.
mcp-serverttsvoice-cloning
ToolNeuphonic
VoiceGenerative Media
neuphonic
Text-to-speech via Neuphonic's API.
ttsvoice-cloningon-devicegguf
ToolChatterbox
VoiceGenerative Media
resemble-ai
Open zero-shot voice-cloning TTS family with multilingual, low-latency Turbo, and CPU-scale Nano variants
ttsvoice-cloningmultilingualzero-shot
ToolFinchvox
MonitoringVoice
finchvox
Voice-AI observability that unifies conversation audio, logs, traces, and metrics for Pipecat agents.
observabilitypipecattelemetrysession-replay
ToolVideo Use
Generative MediaVoice
browser-use
Edit videos with coding agents — drop raw footage in a folder, chat, get final.mp4.
videoffmpegremotionelevenlabs
ToolWhisper
Voice
openai
Transformer model for multilingual speech recognition, translation, and language identification
sttmultilingualtranslationtransformer
ToolChatTTS
VoiceGenerative Media
2noise
Generative text-to-speech for dialogue in English and Chinese, with multi-speaker output and prosody control
ttsmultilingualprosody-controldialogue
ToolMeetily
Voice
Zackriya-Solutions
Self-hosted AI meeting assistant that transcribes, diarizes, and summarizes meetings locally with Whisper/Parakeet and Ollama
productivitywhispersttdiarizationollama
ToolInworld
VoiceGenerative MediaTraining
inworld-ai
Realtime speech-to-speech and expressive TTS-2 text-to-speech.
ttsfine-tunespeech-to-speechreal-time
ToolHermes M5Stick Firmware
VoiceInterface
Syax89
M5StickC Plus 2 firmware for a Wi-Fi desk companion that shows Hermes Agent status and sends voice input via Groq STT
m5stickcgroqsttfirmware
ToolVideo Report Nemotron
Data WranglingVoice
AetherX-Technologies
Hermes skill that turns videos into Markdown, HTML, and PDF reports, using subtitles first and local Nemotron ASR as fallback
videoproductivitynemotronsttmlx
AI Coustics
Voice
Real-time speech enhancement and standalone voice-activity detection for agent audio streams.
audiospeech-enhancementvoice-activity-detectionreal-timeaudio-processing
AppAiri
VoiceGenerative Media
moeru-ai
Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds.
avatarlive2dai-companioncharacter-ai
Anam
VoiceGenerative Media
Real-time video-avatar service synchronising audio and video.
videoavatarlip-syncreal-timechatbot
AssemblyAI
Voice
Real-time and file speech-to-text via AssemblyAI's Universal models.
sttreal-timevoice-agent
Async
VoiceGenerative Media
Text-to-speech via Async's WebSocket and HTTP APIs.
ttswebsocketreal-timelow-latency
Awaaz AI
Voice
Telephony media-stream serializer for hosting voice agents over the Awaaz AI websocket.
financetelephonywebsocketmultilingual
AWS Polly
VoiceGenerative Media
Text-to-speech via Amazon Polly.
ttscloudmultilingual
AWS Transcribe
Voice
Real-time speech-to-text via Amazon Transcribe streaming.
sttcloudstreaming
Bandwidth
Voice
Telephony serializer for Bandwidth Programmable Voice WebSocket media streams.
telephonywebsocketmedia-streaming
Beyond Presence
VoiceGenerative Media
Real-time video-avatar transport for voice agents.
videoavatarreal-timelivekit
Cartesia
VoiceGenerative Media
Real-time speech-to-text and text-to-speech via Cartesia's streaming APIs.
ttssttreal-timestreaming