Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
80 tools · showing 25–48
Rtk
Inference
rtk-ai
CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
productivitytoken-optimizationcliproxycoding-agent
PlatformSuperlinked Inference Engine (SIE)
InferenceDeploymentData Wrangling
superlinked
Self-hosted Kubernetes inference cluster serving LLMs, embeddings, rerankers, OCR and vision models with cluster-wide batching.
inference-serverkubernetesrerankingocr
ToolApfel
InterfaceInference
Arthur-Ficial
The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API key
apple-intelligenceon-devicemacosopenai-compatible
Atomic Agent
InterfaceMemoryInference
AtomicBot-ai
Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.
local-firstllama-cppcomputer-usebrowser-automation
PlatformExo
InferenceDeployment
exo-explore
Open-source distributed inference runtime that pools everyday devices into one local cluster to serve frontier models behind an OpenAI-compatible API.
distributed-inferencedevice-clusteropenai-compatiblelocal-first
ToolRapid-MLX
Inference
raullenchai
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache,
apple-siliconmlxlocal-firstollama-alternative
Fal
InferenceGenerative MediaDeployment
fal-ai
Fast generative-media inference (image, audio, video) via fal serverless models.
serverlessgpu
ToolLLM
InterfaceInference
simonw
CLI tool and Python library for running prompts, embeddings and tools against many local and remote LLMs, with responses logged to SQLite
clisqlitevector-dbplugin
ToolOpenSquilla
InferenceMemory
OpenSquilla
Token-efficient microkernel AI agent with an on-device model router, persistent memory and pluggable multi-provider support
model-routeron-devicemulti-providercli
ToolWebLLM
Inference
mlc-ai
In-browser LLM inference engine running open models client-side on WebGPU with an OpenAI-compatible API
webgpubrowseropenai-compatiblejson-mode
ToolFlama
InferenceInterfaceConnectors
vortico
Ollama is fine for trying a model. vLLM becomes interesting when that model has to serve traffic: many users, agent work
openai-compatiblechatbotasgi
ToolMagnitude
Inference
magnitudedev
Your fully local, private agent. Runs models on your machine with its built-in inference engine. Works out of the box, on any hardware.
local-firstmodel-managementcli
PlatformModal
DeploymentInference
modal-labs
Serverless GPU/CPU cloud for inference, sandboxes and agent workloads.
serverlessgpusandboxautoscaling
ToolGoModel
InferenceMonitoring
ENTERPILOT
AI gateway written in Go. Lightweight unified OpenAI-compatible API for OpenAI, Anthropic, Gemini, Groq, xAI & Ollama. LiteLLM alternative w
gatewayproxyopenai-compatibletoken-optimization
ToolNadirClaw
Inference
NadirRouter
Open-source LLM router & AI cost optimizer. Routes simple prompts to cheap/local models, complex ones to premium — automatically. Drop-in Op
gatewaytoken-optimizationproxylocal-first
ToolRatel
MemoryInference
ratel-ai
Context engineering for agents: in-process BM25 + semantic retrieval over tools, skills and memory with progressive disclosure, no vector DB.
token-optimizationtool-callingbm25progressive-disclosure
ToolLuminal
Inference
luminal-ai
Rust inference compiler that lowers static graphs of ~15 primitive ops straight to CUDA/Metal, with a torch.compile backend.
inference-compilercudametalpytorch
ToolMuna
InferenceDeployment
muna-ai
Compiles Python AI functions into self-contained native binaries and serves open models via an OpenAI-compatible client across cloud, edge and device.
openai-compatibleon-devicegpumodel-compilation
ToolBitNet
Inference
microsoft
Microsoft’s bitnet.cpp is a new CPU-first engine for running 1‑bit LLMs based on the BitNet b1.58 architecture, which us
1-bit-llmquantizationcpu-inferencevector-db
ToolLynkr
Inference
Fast-Editor
Streamline your workflow with Lynkr, a CLI tool that acts as an HTTP proxy for efficient code interactions using Claude Code CLI.
gatewaysemantic-cachetoken-optimizationproxy
HarnessOdysseus
InterfaceResearchInference
pewdiepie-archdaemon
Odysseus is a self-hosted workspace with powerful local tools. Keep auth enabled, keep private data out of Git, and do not expose raw model/service ports publicly.
productivitylocal-firstchatbotdockeremail
PlatformTinyagentos
MemoryDeploymentInference
jaylfc
Self-hosted, framework-agnostic AI agent platform that runs on your own hardware with a browser desktop
local-firstframework-agnosticmulti-frameworkknowledge-graph
ToolEmbedAnything
Data WranglingMemoryInference
StarlightSearch
Rust-based inference, ingestion and indexing library for embeddings and retrieval, with Python bindings.
local-firstcloudragsdk
ToolMLC LLM
Inference
mlc-ai
Machine-learning compiler and engine for deploying LLMs across AMD, NVIDIA, Apple, and Intel GPUs, browsers, iOS, and Android
cross-platformwebgpuon-devicequantization