This is an early release preview. You may encounter bugs.

Create images, video and music

Generation and editing for imagery, animation, video and audio.

37 tools · showing 1–24

Tool

NeMo

VoiceGenerative MediaTraining

NVIDIA-NeMo

NVIDIA framework for building speech AI models, including ASR, TTS, translation, and speaker diarization

sttttsdiarizationtranslation

A+
Platform

OpenMAIC

Generative MediaVoice

THU-MAIC

Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click

educationvideomulti-agentcourse-generationttspowerpoint

A+
Tool

LocalAI

InferenceGenerative Media

mudler

Self-hosted engine that runs LLM, vision, voice, image, and video models on any hardware behind OpenAI-compatible APIs

local-firstopenai-compatiblellama-cppmultimodal

A
Tool

ArcReel

Generative MediaInterface

ArcReel

Open-source AI video workspace that turns a novel into characters, storyboards, and video clips with cross-shot consistency

videostoryboardmulti-providerdockerffmpeg

A
Tool

Fish Audio

VoiceGenerative Media

fishaudio

Real-time text-to-speech via Fish Audio's WebSocket API.

ttsvoice-cloningwebsocket

A
Tool

Voicebox

VoiceGenerative Media

jamiepine

Local-first voice studio to clone voices, generate speech in 23 languages, dictate into any app, and give MCP agents a voice

ttsvoice-cloningsttlocal-first

A
Tool

OGAM

Generative MediaVoiceInterface

off-grid-ai

The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text,

on-deviceggufwhisperstable-diffusion

B
Tool

OpenLive

VoiceGenerative Media

katipally

Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) r

sttttsbarge-invoice-cloning

B
Framework

OpenMontage

Generative MediaCoding

calesthio

Agentic video production system that drives an AI coding assistant through scripting, asset generation, editing, and rendering

videoffmpegremotionttsstable-diffusion

B
Tool

MiniMax-MCP

VoiceGenerative Media

MiniMax-AI

Official MCP server exposing MiniMax text-to-speech, voice cloning, image and video generation APIs to MCP clients.

mcp-serverttsvoice-cloning

B
Tool

Neuphonic

VoiceGenerative Media

neuphonic

Text-to-speech via Neuphonic's API.

ttsvoice-cloningon-devicegguf

B
Tool

Chatterbox

VoiceGenerative Media

resemble-ai

Open zero-shot voice-cloning TTS family with multilingual, low-latency Turbo, and CPU-scale Nano variants

ttsvoice-cloningmultilingualzero-shot

C
Tool

ChatTTS

VoiceGenerative Media

2noise

Generative text-to-speech for dialogue in English and Chinese, with multi-speaker output and prosody control

ttsmultilingualprosody-controldialogue

C
Tool

Inworld

VoiceGenerative MediaTraining

inworld-ai

Realtime speech-to-speech and expressive TTS-2 text-to-speech.

ttsfine-tunespeech-to-speechreal-time

C
Tool

AI Coustics

Voice

Real-time speech enhancement and standalone voice-activity detection for agent audio streams.

audiospeech-enhancementvoice-activity-detectionreal-timeaudio-processing

Score unavailable
Tool

Async

VoiceGenerative Media

Text-to-speech via Async's WebSocket and HTTP APIs.

ttswebsocketreal-timelow-latency

Score unavailable
Tool

AWS Polly

VoiceGenerative Media

Text-to-speech via Amazon Polly.

ttscloudmultilingual

Score unavailable
Tool

Cartesia

VoiceGenerative Media

Real-time speech-to-text and text-to-speech via Cartesia's streaming APIs.

ttssttreal-timestreaming

Score unavailable
Tool

Deepdub

VoiceGenerative Media

Text-to-speech via Deepdub AI's streaming WebSocket API.

ttsvoice-cloningspeech-to-speechstreaming

Score unavailable
Tool

Deepgram

VoiceGenerative Media

Real-time speech-to-text (Nova/Flux) and Aura text-to-speech.

sttttsreal-timevoice-agent

Score unavailable
Tool

ElevenLabs

VoiceGenerative Media

Expressive text-to-speech with word-level timing plus file-based transcription.

ttssttvoice-cloningvoice-agent

Score unavailable
Tool

Google Cloud Speech

VoiceGenerative Media

Google Cloud's Speech-to-Text and Text-to-Speech services.

sttttscloudapi

Score unavailable
Tool

Gradium

VoiceGenerative Media

Low-latency streaming speech-to-text and text-to-speech via Gradium.

sttttsreal-timevoice-cloning

Score unavailable
Tool

Hume

VoiceGenerative Media

Expressive text-to-speech using Hume AI's Octave models with word timestamps.

ttsoctaveemotionalword-timestamps

Score unavailable

More ways in