This is an early release preview. You may encounter bugs.

Run agents on your own machine

Runtimes, models and tooling that work on hardware you control.

58 tools · showing 1–24

Tool

llama.cpp

Deployment

ggml-org

LLM inference in C/C++ across CPU and GPU backends, using the GGUF format with quantization, a REST server, and a WebUI

ggufquantizationguilocal-first

A+
Tool

whisper.cpp

Voice

ggml-org

C/C++ port of OpenAI's Whisper speech-recognition model, dependency-free and optimized for on-device inference

stton-devicemetalquantization

A+
Tool

MNN

Coding

alibaba

Lightweight deep learning engine for on-device inference and training, with runtimes for local LLMs and diffusion models

on-devicemobileembeddedquantization

A+
Tool

InvokeAI

Generative MediaInterface

invoke-ai

Locally hosted creative engine for diffusion image generation with a unified canvas, node workflows, and gallery management

stable-diffusiondiffusionfluxlocal-first

A+
Tool

LiteRT

Coding

google-ai-edge

On-device runtime (formerly TensorFlow Lite) for running models at the edge.

on-devicegpunpuquantization

A+
Tool

Manifest

Inference

mnfst

Open-source LLM router exposing one OpenAI-compatible endpoint across API keys, subscriptions, and local models, with cost tracking

gatewaytoken-optimizationopenai-compatiblebyok

A+
Tool

Open WebUI

Interface

open-webui

Self-hosted, offline-capable AI platform for Ollama and OpenAI-compatible models with RAG, tools, and multi-user access control

ollamaraglocal-firstopenai-compatible

A+
Tool

OpenJarvis

InferenceTraining

open-jarvis

Personal AI, On Personal Devices

productivitylocal-firstollamaclicron

A+
Tool

LocalAI

InferenceGenerative Media

mudler

Self-hosted engine that runs LLM, vision, voice, image, and video models on any hardware behind OpenAI-compatible APIs

local-firstopenai-compatiblellama-cppmultimodal

A
Tool

Nanocoder

Coding

Nano-Collective

An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your

productivitycoding-agentcliollamaopenrouter

A
Tool

Unsloth

TrainingInference

unslothai

Run and fine-tune open LLMs, vision, and audio models locally, with training claimed at up to 2x faster and up to 70% less VRAM

fine-tuneloragguflocal-first

A
Harness

Codewhale

CodingInterface

Hmbown

Open-source, community-driven agent harness

coding-agentclilocal-firstmulti-agent

A
Tool

Jan

InterfaceInference

janhq

Desktop app for running open-weight LLMs locally or connecting to cloud providers, exposing an OpenAI-compatible local API

local-firstopenai-compatiblellama-cppdesktop

A
Tool

Mlx Serve

InferenceGenerative MediaVoice

ddalcu

Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mo

mlxggufapple-siliconopenai-compatible

A
Tool

Mesh LLM

Deployment

Mesh-LLM

Distributed LLM inference that pools GPUs across machines and serves one OpenAI-compatible API, splitting large models across nodes

distributed-inferenceopenai-compatiblegpulocal-first

A
Tool

oMLX

Deployment

jundot

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

mlxlocal-firstmacosdesktop

A
Harness

Ante Preview

Coding

AntigmaLabs

Ghost in your shell. Ante is a self-contained agent harness with a highly optimized core. It works like Claude Code or Codex, with none of t

clilocal-firstlightweightself-organizing

A
Tool

OpenCodex

Coding

lidge-jun

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, an

proxygatewaycodexcli

A
Harness

Osaurus

InferenceMemory

osaurus-ai

Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity

productivitylocal-firstmacosdesktopmlx

A
Tool

Local Deep Research

Research

LearningCircuit

Self-hosted assistant that runs iterative, cited research across search engines with local or cloud LLMs

productivitylocal-firstragcitationsollama

A
Platform

Superlinked Inference Engine (SIE)

InferenceDeploymentData Wrangling

superlinked

Self-hosted Kubernetes inference cluster serving LLMs, embeddings, rerankers, OCR and vision models with cluster-wide batching.

inference-serverkubernetesrerankingocr

A
Tool

Apfel

InterfaceInference

Arthur-Ficial

The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API key

apple-intelligenceon-devicemacosopenai-compatible

A
Tool

Rapid-MLX

Inference

raullenchai

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache,

apple-siliconmlxlocal-firstollama-alternative

A
Platform

Exo

InferenceDeployment

exo-explore

Open-source distributed inference runtime that pools everyday devices into one local cluster to serve frontier models behind an OpenAI-compatible API.

distributed-inferencedevice-clusteropenai-compatiblelocal-first

A

More ways in