This is an early release preview. You may encounter bugs.
Unclaimed

Tool voice

AssemblyAI

Real-time and file speech-to-text via AssemblyAI's Universal models.

No votes yet

01 / About

What AssemblyAI is.

AssemblyAI provides speech-to-text and speech-understanding models behind a single API, aimed at products that need to transcribe or analyse voice: voice agents, AI notetakers, call analytics, medical transcription, dictation, and agent assist. Transcription is offered in three modes: pre-recorded (asynchronous) for files, real-time streaming over a WebSocket for live audio, and a synchronous endpoint for short clips.

The current flagship speech-to-text model is Universal-3.5 Pro, available for both real-time and pre-recorded audio; streaming also offers Universal-Streaming (English), Universal-Streaming Multilingual, and Whisper-Streaming for wider language coverage. Streaming sessions are billed on WebSocket duration. The platform lists coverage for 99 languages.

On top of transcripts, the Speech Understanding API adds analysis such as speaker identification, sentiment, entities, topics, chapters, summaries, and translation, and an LLM Gateway plus guardrails layer supports building downstream reasoning over audio. Official Python and JavaScript SDKs are available, alongside adapters for the LiveKit and Pipecat voice-agent frameworks and a self-hosted Voice AI Cloud option.

Features

  • Pre-recorded transcription: asynchronous speech-to-text for uploaded or hosted audio files
  • Real-time streaming: WebSocket transcription with Universal-3.5 Pro Realtime, Universal-Streaming, Universal-Streaming Multilingual, and Whisper-Streaming models
  • Sync endpoint: synchronous speech-to-text for short audio
  • Speaker identification: labels each utterance with the speaker who said it
  • Speech understanding: sentiment per sentence, entity detection, IAB topic detection, auto chapters, key phrases, action items, and summaries in several formats
  • Translation: transcripts translated into other languages
  • LLM Gateway and guardrails: route transcript-based prompts to LLMs with safety controls
  • Voice-agent integrations: LiveKit and Pipecat adapters, plus guides for coding assistants such as Cursor and Claude Code
  • SDKs: official Python and JavaScript clients
  • Self-hosted option: Voice AI Cloud deployment for running the models in your own environment

02 / Discussion CREDIBILITY-GATED

Discussion

Reading is open to everyone. Posting and voting need a verified identity or a GitHub grade of B or higher.

  • No discussions yet.

03 / Build

Build with AssemblyAI.

Browse the catalogue for frameworks, tools, and harnesses, each scored on real GitHub credibility.

Get AssemblyAI →

Browse the catalogue