Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
36 tools · showing 1–24
LiveKit Agents
Voice
livekit
Framework for building realtime, server-side voice agents with pluggable STT, LLM, and TTS and WebRTC/telephony support
webrtctelephonyreal-timestt
ToolNeMo
VoiceGenerative MediaTraining
NVIDIA-NeMo
NVIDIA framework for building speech AI models, including ASR, TTS, translation, and speaker diarization
sttttsdiarizationtranslation
PlatformOpenMAIC
Generative MediaVoice
THU-MAIC
Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
educationvideomulti-agentcourse-generationttspowerpoint
ToolVoiceStudio
CodingVoice
debpalash
The Open-Source Elevenlabs alternative AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.
voice-cloningttssttdubbing
Fish Audio
VoiceGenerative Media
fishaudio
Real-time text-to-speech via Fish Audio's WebSocket API.
ttsvoice-cloningwebsocket
OpenLive
VoiceGenerative Media
katipally
Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) r
sttttsbarge-invoice-cloning
Voicebox
VoiceGenerative Media
jamiepine
Local-first voice studio to clone voices, generate speech in 23 languages, dictate into any app, and give MCP agents a voice
ttsvoice-cloningsttlocal-first
ToolHermes for Home Assistant
VoiceConnectors
rusty4444
Bundle connecting Hermes Agent to Home Assistant with a custom integration, service-call tools, and an optional wake-word/STT/TTS voice loop
home-assistantwake-wordttsstt
OpenMontage
Generative MediaCoding
calesthio
Agentic video production system that drives an AI coding assistant through scripting, asset generation, editing, and rendering
videoffmpegremotionttsstable-diffusion
ToolMiniMax-MCP
VoiceGenerative Media
MiniMax-AI
Official MCP server exposing MiniMax text-to-speech, voice cloning, image and video generation APIs to MCP clients.
mcp-serverttsvoice-cloning
ToolNeuphonic
VoiceGenerative Media
neuphonic
Text-to-speech via Neuphonic's API.
ttsvoice-cloningon-devicegguf
ToolChatterbox
VoiceGenerative Media
resemble-ai
Open zero-shot voice-cloning TTS family with multilingual, low-latency Turbo, and CPU-scale Nano variants
ttsvoice-cloningmultilingualzero-shot
ToolChatTTS
VoiceGenerative Media
2noise
Generative text-to-speech for dialogue in English and Chinese, with multi-speaker output and prosody control
ttsmultilingualprosody-controldialogue
ToolInworld
VoiceGenerative MediaTraining
inworld-ai
Realtime speech-to-speech and expressive TTS-2 text-to-speech.
ttsfine-tunespeech-to-speechreal-time
Async
VoiceGenerative Media
Text-to-speech via Async's WebSocket and HTTP APIs.
ttswebsocketreal-timelow-latency
AWS Polly
VoiceGenerative Media
Text-to-speech via Amazon Polly.
ttscloudmultilingual
Cartesia
VoiceGenerative Media
Real-time speech-to-text and text-to-speech via Cartesia's streaming APIs.
ttssttreal-timestreaming
Deepdub
VoiceGenerative Media
Text-to-speech via Deepdub AI's streaming WebSocket API.
ttsvoice-cloningspeech-to-speechstreaming
Deepgram
VoiceGenerative Media
Real-time speech-to-text (Nova/Flux) and Aura text-to-speech.
sttttsreal-time
ElevenLabs
VoiceGenerative Media
Expressive text-to-speech with word-level timing plus file-based transcription.
ttssttvoice-cloningvoice-agent
Google Cloud Speech
VoiceGenerative Media
Google Cloud's Speech-to-Text and Text-to-Speech services.
sttttscloudapi
Gradium
VoiceGenerative Media
Low-latency streaming speech-to-text and text-to-speech via Gradium.
sttttsreal-timevoice-cloning
Hume
VoiceGenerative Media
Expressive text-to-speech using Hume AI's Octave models with word timestamps.
ttsoctaveemotionalword-timestamps
LMNT
VoiceGenerative Media
Text-to-speech via LMNT's streaming API.
ttsstreamingapivoice-cloning