Collection
Give agents a voice
Speech recognition, synthesis and the parts of a spoken conversation.
36 tools · showing 1–24
Toolwhisper.cpp
Voice
ggml-org
C/C++ port of OpenAI's Whisper speech-recognition model, dependency-free and optimized for on-device inference
stton-devicemetalquantization
LiveKit Agents
Voice
livekit
Framework for building realtime, server-side voice agents with pluggable STT, LLM, and TTS and WebRTC/telephony support
webrtctelephonyreal-timestt
ToolNeMo
VoiceGenerative MediaTraining
NVIDIA-NeMo
NVIDIA framework for building speech AI models, including ASR, TTS, translation, and speaker diarization
sttttsdiarizationtranslation
PlatformOpenMAIC
Generative MediaVoice
THU-MAIC
Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
educationvideomulti-agentcourse-generationttspowerpoint
ToolDograh
VoiceGenerative Media
dograh-hq
Self-hostable open-source voice AI platform for building inbound and outbound phone agents with a drag-and-drop workflow builder
local-firsttelephonyorchestrationpipecat
ToolFluidVoice
Voice
altic-dev
Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. A local Wispr Flow alternative. ⭐ helps a t
sttmacoson-devicewhisper
ToolVoiceStudio
CodingVoice
debpalash
The Open-Source Elevenlabs alternative AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.
voice-cloningttssttdubbing
Voicebox
VoiceGenerative Media
jamiepine
Local-first voice studio to clone voices, generate speech in 23 languages, dictate into any app, and give MCP agents a voice
ttsvoice-cloningsttlocal-first
ToolMeetily
Voice
Zackriya-Solutions
Self-hosted AI meeting assistant that transcribes, diarizes, and summarizes meetings locally with Whisper/Parakeet and Ollama
productivitywhispersttdiarizationollama
ToolOGAM
Generative MediaVoiceInterface
off-grid-ai
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text,
on-deviceggufwhisperstable-diffusion
OpenLive
VoiceGenerative Media
katipally
Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) r
sttttsbarge-invoice-cloning
ToolJarvis
VoiceMemoryInterface
isair
A 100% private AI voice assistant that lives on your computer (works offline). Talk naturally as if Jarvis is a third person in the room, an
productivityvoice-agentwake-wordsttlocal-first
ToolHermes for Home Assistant
VoiceConnectors
rusty4444
Bundle connecting Hermes Agent to Home Assistant with a custom integration, service-call tools, and an optional wake-word/STT/TTS voice loop
home-assistantwake-wordttsstt
Confer
VoiceConnectors
A complete AI assistant — writing, images, web search, memory. Every conversation is end-to-end encrypted, so no one else can read it.
productivitymeetingssttaction-items
ToolWhisper
Voice
openai
Transformer model for multilingual speech recognition, translation, and language identification
sttmultilingualtranslationtransformer
ToolVideo Use
Generative MediaVoice
browser-use
Edit videos with coding agents — drop raw footage in a folder, chat, get final.mp4.
videoffmpegremotionelevenlabsdiarization
ToolInworld
VoiceGenerative MediaTraining
inworld-ai
Realtime speech-to-speech and expressive TTS-2 text-to-speech.
ttsfine-tunespeech-to-speechreal-time
ToolHermes M5Stick Firmware
VoiceInterface
Syax89
M5StickC Plus 2 firmware for a Wi-Fi desk companion that shows Hermes Agent status and sends voice input via Groq STT
m5stickcgroqsttfirmware
ToolVideo Report Nemotron
Data WranglingVoice
AetherX-Technologies
Hermes skill that turns videos into Markdown, HTML, and PDF reports, using subtitles first and local Nemotron ASR as fallback
videoproductivitynemotronsttmlxffmpeg
AI Coustics
Voice
Real-time speech enhancement and standalone voice-activity detection for agent audio streams.
audiospeech-enhancementvoice-activity-detectionreal-timeaudio-processing
AssemblyAI
Voice
Real-time and file speech-to-text via AssemblyAI's Universal models.
sttreal-timevoice-agent
AWS Transcribe
Voice
Real-time speech-to-text via Amazon Transcribe streaming.
sttcloudstreaming
Cartesia
VoiceGenerative Media
Real-time speech-to-text and text-to-speech via Cartesia's streaming APIs.
ttssttreal-timestreaming
Deepgram
VoiceGenerative Media
Real-time speech-to-text (Nova/Flux) and Aura text-to-speech.
sttttsreal-timevoice-agent