Collection
Give agents a voice
Speech recognition, synthesis and the parts of a spoken conversation.
8 tools
ToolNeMo
VoiceGenerative MediaTraining
NVIDIA-NeMo
NVIDIA framework for building speech AI models, including ASR, TTS, translation, and speaker diarization
sttttsdiarizationtranslation
ToolVoiceStudio
CodingVoice
debpalash
The Open-Source Elevenlabs alternative AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.
voice-cloningttssttdubbing
ToolWhisper
Voice
openai
Transformer model for multilingual speech recognition, translation, and language identification
sttmultilingualtranslationtransformer
Deepdub
VoiceGenerative Media
Text-to-speech via Deepdub AI's streaming WebSocket API.
ttsvoice-cloningspeech-to-speechstreaming
Gradium
VoiceGenerative Media
Low-latency streaming speech-to-text and text-to-speech via Gradium.
sttttsreal-timevoice-cloning
OpenWhispr
Voice
OpenWhispr
Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.
talkingsttwhisperdiarizationtranslation
Pinch
VoiceGenerative Media
Real-time speech-to-speech translation for voice agents.
translationdubbingreal-timespeech-to-speech
Soniox
VoiceGenerative Media
Real-time WebSocket speech-to-text and streaming text-to-speech.
sttttsreal-timetranslation