ElevenLabs offers speech synthesis, speech recognition, and a conversational-agent platform through one API. The product lines are ElevenAPI for developers, ElevenAgents for deploying voice agents, and ElevenCreative for content production, and the underlying models cover text-to-speech, speech-to-text, voice cloning and design, dubbing, music, and sound effects.
Text-to-speech models are chosen per request. Eleven v3 targets expressive delivery across 70+ languages with a 5,000-character limit, and a v3 Conversational variant streams at about 280 ms for agents. Eleven Multilingual v2 covers 29 languages with a 10,000-character limit. Eleven Flash v2.5 covers 32 languages, allows 40,000 characters, and returns audio at about 75 ms at half the per-character price. Streaming and WebSocket endpoints can return word-level timestamps for aligning audio with text.
Speech-to-text runs on Scribe v2, which transcribes 90+ languages with speaker diarization for up to 32 speakers, word-level timestamps, and detection of 65 entity types; Scribe v2 Realtime streams at about 150 ms. Official Python and TypeScript SDKs wrap the REST and WebSocket APIs, and usage is metered in credits (one credit per input character for TTS) that reset monthly.
Features
- Text-to-speech models: Eleven v3, v3 Conversational, Multilingual v2, and Flash v2.5 with different language counts, character limits, and latencies
- Streaming with timestamps: WebSocket synthesis that returns word-level timing alongside audio
- Speech-to-text: Scribe v2 for files and Scribe v2 Realtime for live audio in 90+ languages
- Diarization and entities: up to 32 speakers and 65 entity types per transcript
- Voice cloning and design: create voices from samples or from a description
- Dubbing: translate audio and video while keeping the speaker's voice
- Forced alignment: map an existing transcript onto its audio
- Music and sound effects: generative audio beyond speech
- ElevenAgents: hosted platform for building and deploying conversational voice agents
- SDKs: official Python and TypeScript clients
Integrated by
Airi
Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds.
Autonovel
Agent pipeline that drafts, revises, typesets, illustrates, and narrates a complete novel from a seed concept
Hermes
Self-improving AI agent with a learning loop that creates and refines skills, recalls past sessions, and runs across chat platforms
-
LemonSlice
AI-avatar video service for adding LemonSlice avatars to real-time rooms.
Alternatives
-
Async
Text-to-speech via Async's WebSocket and HTTP APIs.
-
Gladia
Real-time speech-to-text via Gladia's live API.
OpenLive
Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) r
-
Pinch
Real-time speech-to-speech translation for voice agents.