This is an early release preview. You may encounter bugs.

Create images, video and music

Generation and editing for imagery, animation, video and audio.

76 tools · showing 1–24

Tool

NeMo

VoiceGenerative MediaTraining

NVIDIA

NVIDIA framework for building speech AI models, including ASR, TTS, translation, and speaker diarization

sttttsdiarizationtranslation

A+
Tool

MNN

Coding

alibaba

Lightweight deep learning engine for on-device inference and training, with runtimes for local LLMs and diffusion models

on-devicemobileembeddedquantization

A+
Tool

InvokeAI

Generative MediaInterface

invoke-ai

Locally hosted creative engine for diffusion image generation with a unified canvas, node workflows, and gallery management

stable-diffusiondiffusionfluxlocal-first

A+
Tool

Supervision

DeploymentCodingMonitoring

roboflow

Reusable computer-vision utilities for detection, tracking and annotation pipelines.

videocomputer-visionobject-detectionobject-trackingannotation

A+
Tool

ComfyUI

Generative MediaInterface

comfyanonymous

Node-graph engine for generating images, video, 3D, and audio, locally or via API

stable-diffusionnode-graphaudio-generation

A+
Tool

LocalAI

InferenceGenerative Media

mudler

Self-hosted engine that runs LLM, vision, voice, image, and video models on any hardware behind OpenAI-compatible APIs

local-firstopenai-compatiblellama-cppmultimodal

A+
Platform

OpenMAIC

Generative MediaVoice

THU-MAIC

Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click

educationvideomulti-agentcourse-generationttspowerpoint

A+
Tool

HyperFrames

Interface

heygen-com

HyperFrames renders video from plain HTML.

videohtml-to-videomotion-graphicsffmpegpuppeteer

A
Tool

Diffusers

Generative MediaInferenceTraining

huggingface

Library of pretrained diffusion models for generating images, audio, and 3D structures, for inference or training

diffusionstable-diffusionpytorchflux

A
Platform

Roboflow Inference

Deployment

roboflow

Server and SDK for running vision models locally or in the cloud behind a simple API.

videocomputer-visionobject-detectionon-devicedocker

A
Tool

Dograh

VoiceGenerative Media

dograh-hq

Self-hostable open-source voice AI platform for building inbound and outbound phone agents with a drag-and-drop workflow builder

local-firsttelephonyorchestrationpipecat

A
Tool

Mlx Serve

InferenceGenerative MediaVoice

ddalcu

Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mo

mlxggufapple-siliconopenai-compatible

A
Tool

ODS

DeploymentInterfaceInference

Osmantic

Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.

local-firstollamacomfyuin8n

A
Tool

Pi Web Access

CodingResearch

nicobailon

Web search and content extraction extension for Pi coding agent

videosearchcontent-extractionexatavily

A
Tool

PPT Master

Generative MediaVoice

hugohe3

AI generates a real, editable PowerPoint from any document — native shapes & animations, speaker notes voiced as audio narration, and the op

productivitypowerpoint-generationdocument-generationpresentationaudio-narration

A
Tool

ArcReel

Generative MediaInterface

ArcReel

Open-source AI video workspace that turns a novel into characters, storyboards, and video clips with cross-shot consistency

videostoryboardmulti-providerdocker

A
Tool

OpenPencil

Generative MediaCoding

open-pencil

AI-native design editor. Open-source Figma alternative.

design-editorfigmamcp-serverdesign-to-code

A
Tool

Fish Audio

VoiceGenerative Media

fishaudio

Real-time text-to-speech via Fish Audio's WebSocket API.

ttsvoice-cloningwebsocket

A
Tool

ComfyUI MCP

Generative MediaInterface

artokun

The local-first, agent-native control plane for ComfyUI — MCP server + Claude Code plugin. 108 tools, 29 AI skills (Flux · WAN · LT2.3 · Qwe

comfyuistable-diffusionfluxlocal-first

A
Tool

Fal

InferenceGenerative MediaDeployment

fal-ai

Fast generative-media inference (image, audio, video) via fal serverless models.

serverlessgpu

A
Tool

OpenLive

VoiceGenerative Media

katipally

Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) r

sttttsbarge-invoice-cloning

A
Tool

Voicebox

VoiceGenerative Media

jamiepine

Local-first voice studio to clone voices, generate speech in 23 languages, dictate into any app, and give MCP agents a voice

ttsvoice-cloningsttlocal-first

A
Tool

iPolloWork

CodingInterfaceGenerative Media

Devin-AXIS

Next-gen, source-available alternative to Codex and Claude Code — one local-first, self-hostable agent workspace for code, office work, edit

productivitymulti-agentpresentationsguilocal-first

A
Tool

OGAM

Generative MediaVoiceInterface

off-grid-ai

The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text,

on-deviceggufwhisperstable-diffusion

B

More ways in