This is an early release preview. You may encounter bugs.
VoiceStudio logo
Unclaimed

Tool coding voice

VoiceStudio

The Open-Source Elevenlabs alternative AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.

A 85/100 GitHub score ? This grade is derived from GitHub signals, not user votes. Open for the full breakdown.
No votes yet

01 / About

What VoiceStudio is.

VoiceStudio (previously OmniVoice-Studio) is a desktop application for voice cloning, voice design, video dubbing, dictation, and long-form audio such as stories and audiobooks. It runs the models on your own machine by default: voices, projects, settings, and outputs stay local, and the features that send audio off the machine, such as remote workers and OpenAI-compatible ASR, are explicit opt-ins.

The app bundles 16 text-to-speech engines and 11 speech-recognition engines that you add, remove, and switch from a Model Catalogue. The default TTS engine is built on k2-fsa/OmniVoice and covers 600+ languages; others include CosyVoice 3, GPT-SoVITS, VoxCPM2, IndexTTS 2.5, Supertonic 3, KittenTTS, MLX-Audio, and Sherpa-ONNX, with per-engine differences in language coverage, cloning, instruction following, and platform support. ASR engines include WhisperX (default), Faster-Whisper, MLX Whisper, Parakeet TDT, Moonshine, FunASR, and streaming sherpa-onnx for live dictation. Compute routes to CUDA, Apple Silicon MPS/MLX, ROCm on Linux, or CPU, with optional enrolled remote workers.

Architecturally it is a Tauri v2 shell around a React UI and a FastAPI backend on localhost:3900. That backend also exposes a local speech platform: REST, SSE, and WebSocket routes, an OpenAI-compatible audio API, and an MCP server, so an OpenAI audio client can be repointed at the local base URL and agents can call synthesis and transcription as tools. Skills for Claude Code, Codex, Cursor, and other skills.sh-compatible agents are published for the same purpose. It runs on macOS 13.3+ (Apple Silicon), Windows 10/11 x64, and Linux x86_64, with Docker images for CUDA, ROCm, and CPU.

Features

  • Voice cloning: zero-shot synthesis from a 3–15 second reference clip
  • Voice design: create a voice from age, accent, pitch, style, and delivery instructions
  • Video dubbing: transcribe, translate, preserve speakers, synthesize, and export video
  • Stories and audiobooks: multi-voice scripts, EPUB/PDF import, chapter rendering, and .m4b export
  • Dictation widget: system-wide shortcut with live transcription and optional local-LLM cleanup
  • Audio processing: Demucs vocal isolation, Pyannote and WhisperX speaker diarization, and AudioSeal watermark embedding and detection
  • Batch queue: large sets of audio and video jobs with per-job progress
  • Model Catalogue: add, remove, select, and route TTS, ASR, and LLM models, including downloads onto remote workers
  • OpenAI-compatible audio API: /v1/audio/speech, /v1/audio/transcriptions, a streaming transcription WebSocket, and a voice listing endpoint
  • MCP server and agent skills: synthesis and transcription tools for MCP clients, plus published skills for coding agents
  • Diagnostics: self-checks, an error journal, logs, and scrubbed support bundles
  • Extensible engines: registry-based TTS, ASR, and plugin interfaces

02 / Discussion CREDIBILITY-GATED

Discussion

Reading is open to everyone. Posting and voting need a verified identity or a GitHub grade of B or higher.

  • No discussions yet.

03 / Build

Build with VoiceStudio.

Browse the catalogue for frameworks, tools, and harnesses, each scored on real GitHub credibility.

Get VoiceStudio →

Browse the catalogue