This is an early release preview. You may encounter bugs.

Catalogue

Submit a tool

Harnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.

30 tools · showing 1–24

#stt ×
Tool

whisper.cpp

Voice

ggml-org

C/C++ port of OpenAI's Whisper speech-recognition model, dependency-free and optimized for on-device inference

stton-devicemetalquantization

A+
Framework

LiveKit Agents

Voice

livekit

Framework for building realtime, server-side voice agents with pluggable STT, LLM, and TTS and WebRTC/telephony support

webrtctelephonyreal-timestt

A+
Tool

NeMo

VoiceGenerative MediaTraining

NVIDIA

NVIDIA framework for building speech AI models, including ASR, TTS, translation, and speaker diarization

sttttsdiarizationtranslation

A+
Tool

FluidVoice

Voice

altic-dev

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. A local Wispr Flow alternative. ⭐ helps a t

sttmacoson-devicewhisper

A
Tool

VoiceStudio

CodingVoice

debpalash

The Open-Source Elevenlabs alternative AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.

voice-cloningttssttdubbing

A
Tool

OpenLive

VoiceGenerative Media

katipally

Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) r

sttttsbarge-invoice-cloning

A
Tool

Voicebox

VoiceGenerative Media

jamiepine

Local-first voice studio to clone voices, generate speech in 23 languages, dictate into any app, and give MCP agents a voice

ttsvoice-cloningsttlocal-first

A
Tool

Jarvis

VoiceMemoryInterface

isair

A 100% private AI voice assistant that lives on your computer (works offline). Talk naturally as if Jarvis is a third person in the room, an

productivityvoice-agentwake-wordsttlocal-first

B
Tool

Hermes for Home Assistant

VoiceConnectors

rusty4444

Bundle connecting Hermes Agent to Home Assistant with a custom integration, service-call tools, and an optional wake-word/STT/TTS voice loop

home-assistantwake-wordttsstt

B
App

Confer

VoiceConnectors

A complete AI assistant — writing, images, web search, memory. Every conversation is end-to-end encrypted, so no one else can read it.

productivitymeetingssttaction-items

B
Tool

Whisper

Voice

openai

Transformer model for multilingual speech recognition, translation, and language identification

sttmultilingualtranslationtransformer

B
Tool

Meetily

Voice

Zackriya-Solutions

Self-hosted AI meeting assistant that transcribes, diarizes, and summarizes meetings locally with Whisper/Parakeet and Ollama

productivitywhispersttdiarizationollama

C
Tool

Hermes M5Stick Firmware

VoiceInterface

Syax89

M5StickC Plus 2 firmware for a Wi-Fi desk companion that shows Hermes Agent status and sends voice input via Groq STT

m5stickcgroqsttfirmware

C
Tool

Video Report Nemotron

Data WranglingVoice

AetherX-Technologies

Hermes skill that turns videos into Markdown, HTML, and PDF reports, using subtitles first and local Nemotron ASR as fallback

videoproductivitynemotronsttmlx

C
Tool

AssemblyAI

Voice

Real-time and file speech-to-text via AssemblyAI's Universal models.

sttreal-timevoice-agent

Score unavailable
Tool

AWS Transcribe

Voice

Real-time speech-to-text via Amazon Transcribe streaming.

sttcloudstreaming

Score unavailable
Tool

Cartesia

VoiceGenerative Media

Real-time speech-to-text and text-to-speech via Cartesia's streaming APIs.

ttssttreal-timestreaming

Score unavailable
Tool

Deepgram

VoiceGenerative Media

Real-time speech-to-text (Nova/Flux) and Aura text-to-speech.

sttttsreal-time

Score unavailable
Tool

ElevenLabs

VoiceGenerative Media

Expressive text-to-speech with word-level timing plus file-based transcription.

ttssttvoice-cloningvoice-agent

Score unavailable
Tool

Gladia

Voice

Real-time speech-to-text via Gladia's live API.

sttapimultilingualreal-time

Score unavailable
Tool

Google Cloud Speech

VoiceGenerative Media

Google Cloud's Speech-to-Text and Text-to-Speech services.

sttttscloudapi

Score unavailable
Tool

Gradium

VoiceGenerative Media

Low-latency streaming speech-to-text and text-to-speech via Gradium.

sttttsreal-timevoice-cloning

Score unavailable
Tool

Lokutor

VoiceGenerative Media

CPU-only text-to-speech synthesis via Lokutor.

ttscpu-onlyapistt

Score unavailable
App

Openwhispr

Voice

OpenWhispr

Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.

talkingstt