Collection
Run agents on your own machine
Runtimes, models and tooling that work on hardware you control.
58 tools · showing 1–24
Toolllama.cpp
Deployment
ggml-org
LLM inference in C/C++ across CPU and GPU backends, using the GGUF format with quantization, a REST server, and a WebUI
ggufquantizationguilocal-first
Toolwhisper.cpp
Voice
ggml-org
C/C++ port of OpenAI's Whisper speech-recognition model, dependency-free and optimized for on-device inference
stton-devicemetalquantization
MNN
Coding
alibaba
Lightweight deep learning engine for on-device inference and training, with runtimes for local LLMs and diffusion models
on-devicemobileembeddedquantization
ToolInvokeAI
Generative MediaInterface
invoke-ai
Locally hosted creative engine for diffusion image generation with a unified canvas, node workflows, and gallery management
stable-diffusiondiffusionfluxlocal-first
ToolLiteRT
Coding
google-ai-edge
On-device runtime (formerly TensorFlow Lite) for running models at the edge.
on-devicegpunpuquantization
ToolManifest
Inference
mnfst
Open-source LLM router exposing one OpenAI-compatible endpoint across API keys, subscriptions, and local models, with cost tracking
gatewaytoken-optimizationopenai-compatiblebyok
Open WebUI
Interface
open-webui
Self-hosted, offline-capable AI platform for Ollama and OpenAI-compatible models with RAG, tools, and multi-user access control
ollamaraglocal-firstopenai-compatible
ToolOpenJarvis
InferenceTraining
open-jarvis
Personal AI, On Personal Devices
productivitylocal-firstollamaclicron
ToolLocalAI
InferenceGenerative Media
mudler
Self-hosted engine that runs LLM, vision, voice, image, and video models on any hardware behind OpenAI-compatible APIs
local-firstopenai-compatiblellama-cppmultimodal
ToolNanocoder
Coding
Nano-Collective
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your
productivitycoding-agentcliollamaopenrouter
ToolUnsloth
TrainingInference
unslothai
Run and fine-tune open LLMs, vision, and audio models locally, with training claimed at up to 2x faster and up to 70% less VRAM
fine-tuneloragguflocal-first
HarnessCodewhale
CodingInterface
Hmbown
Open-source, community-driven agent harness
coding-agentclilocal-firstmulti-agent
ToolJan
InterfaceInference
janhq
Desktop app for running open-weight LLMs locally or connecting to cloud providers, exposing an OpenAI-compatible local API
local-firstopenai-compatiblellama-cppdesktop
ToolMlx Serve
InferenceGenerative MediaVoice
ddalcu
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mo
mlxggufapple-siliconopenai-compatible
ToolMesh LLM
Deployment
Mesh-LLM
Distributed LLM inference that pools GPUs across machines and serves one OpenAI-compatible API, splitting large models across nodes
distributed-inferenceopenai-compatiblegpulocal-first
oMLX
Deployment
jundot
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
mlxlocal-firstmacosdesktop
HarnessAnte Preview
Coding
AntigmaLabs
Ghost in your shell. Ante is a self-contained agent harness with a highly optimized core. It works like Claude Code or Codex, with none of t
clilocal-firstlightweightself-organizing
ToolOpenCodex
Coding
lidge-jun
Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, an
proxygatewaycodexcli
HarnessOsaurus
InferenceMemory
osaurus-ai
Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity
productivitylocal-firstmacosdesktopmlx
ToolLocal Deep Research
Research
LearningCircuit
Self-hosted assistant that runs iterative, cited research across search engines with local or cloud LLMs
productivitylocal-firstragcitationsollama
PlatformSuperlinked Inference Engine (SIE)
InferenceDeploymentData Wrangling
superlinked
Self-hosted Kubernetes inference cluster serving LLMs, embeddings, rerankers, OCR and vision models with cluster-wide batching.
inference-serverkubernetesrerankingocr
ToolApfel
InterfaceInference
Arthur-Ficial
The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API key
apple-intelligenceon-devicemacosopenai-compatible
ToolRapid-MLX
Inference
raullenchai
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache,
apple-siliconmlxlocal-firstollama-alternative
PlatformExo
InferenceDeployment
exo-explore
Open-source distributed inference runtime that pools everyday devices into one local cluster to serve frontier models behind an OpenAI-compatible API.
distributed-inferencedevice-clusteropenai-compatiblelocal-first