This is an early release preview. You may encounter bugs.

Give agents a voice

Speech recognition, synthesis and the parts of a spoken conversation.

36 tools · showing 1–24

Tool

whisper.cpp

Voice

ggml-org

C/C++ port of OpenAI's Whisper speech-recognition model, dependency-free and optimized for on-device inference

stton-devicemetalquantization

A+
Framework

LiveKit Agents

Voice

livekit

Framework for building realtime, server-side voice agents with pluggable STT, LLM, and TTS and WebRTC/telephony support

webrtctelephonyreal-timestt

A+
Tool

NeMo

VoiceGenerative MediaTraining

NVIDIA-NeMo

NVIDIA framework for building speech AI models, including ASR, TTS, translation, and speaker diarization

sttttsdiarizationtranslation

A+
Platform

OpenMAIC

Generative MediaVoice

THU-MAIC

Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click

educationvideomulti-agentcourse-generationttspowerpoint

A+
Tool

Dograh

VoiceGenerative Media

dograh-hq

Self-hostable open-source voice AI platform for building inbound and outbound phone agents with a drag-and-drop workflow builder

local-firsttelephonyorchestrationpipecat

A
Tool

FluidVoice

Voice

altic-dev

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. A local Wispr Flow alternative. ⭐ helps a t

sttmacoson-devicewhisper

A
Tool

VoiceStudio

CodingVoice

debpalash

The Open-Source Elevenlabs alternative AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.

voice-cloningttssttdubbing

A
Tool

Voicebox

VoiceGenerative Media

jamiepine

Local-first voice studio to clone voices, generate speech in 23 languages, dictate into any app, and give MCP agents a voice

ttsvoice-cloningsttlocal-first

A
Tool

Meetily

Voice

Zackriya-Solutions

Self-hosted AI meeting assistant that transcribes, diarizes, and summarizes meetings locally with Whisper/Parakeet and Ollama

productivitywhispersttdiarizationollama

A
Tool

OGAM

Generative MediaVoiceInterface

off-grid-ai

The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text,

on-deviceggufwhisperstable-diffusion

B
Tool

OpenLive

VoiceGenerative Media

katipally

Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) r

sttttsbarge-invoice-cloning

B
Tool

Jarvis

VoiceMemoryInterface

isair

A 100% private AI voice assistant that lives on your computer (works offline). Talk naturally as if Jarvis is a third person in the room, an

productivityvoice-agentwake-wordsttlocal-first

B
Tool

Hermes for Home Assistant

VoiceConnectors

rusty4444

Bundle connecting Hermes Agent to Home Assistant with a custom integration, service-call tools, and an optional wake-word/STT/TTS voice loop

home-assistantwake-wordttsstt

B
App

Confer

VoiceConnectors

A complete AI assistant — writing, images, web search, memory. Every conversation is end-to-end encrypted, so no one else can read it.

productivitymeetingssttaction-items

B
Tool

Whisper

Voice

openai

Transformer model for multilingual speech recognition, translation, and language identification

sttmultilingualtranslationtransformer

B
Tool

Video Use

Generative MediaVoice

browser-use

Edit videos with coding agents — drop raw footage in a folder, chat, get final.mp4.

videoffmpegremotionelevenlabsdiarization

B
Tool

Inworld

VoiceGenerative MediaTraining

inworld-ai

Realtime speech-to-speech and expressive TTS-2 text-to-speech.

ttsfine-tunespeech-to-speechreal-time

C
Tool

Hermes M5Stick Firmware

VoiceInterface

Syax89

M5StickC Plus 2 firmware for a Wi-Fi desk companion that shows Hermes Agent status and sends voice input via Groq STT

m5stickcgroqsttfirmware

C
Tool

Video Report Nemotron

Data WranglingVoice

AetherX-Technologies

Hermes skill that turns videos into Markdown, HTML, and PDF reports, using subtitles first and local Nemotron ASR as fallback

videoproductivitynemotronsttmlxffmpeg

C
Tool

AI Coustics

Voice

Real-time speech enhancement and standalone voice-activity detection for agent audio streams.

audiospeech-enhancementvoice-activity-detectionreal-timeaudio-processing

Score unavailable
Tool

AssemblyAI

Voice

Real-time and file speech-to-text via AssemblyAI's Universal models.

sttreal-timevoice-agent

Score unavailable
Tool

AWS Transcribe

Voice

Real-time speech-to-text via Amazon Transcribe streaming.

sttcloudstreaming

Score unavailable
Tool

Cartesia

VoiceGenerative Media

Real-time speech-to-text and text-to-speech via Cartesia's streaming APIs.

ttssttreal-timestreaming

Score unavailable
Tool

Deepgram

VoiceGenerative Media

Real-time speech-to-text (Nova/Flux) and Aura text-to-speech.

sttttsreal-timevoice-agent

Score unavailable

More ways in