This is an early release preview. You may encounter bugs.

Catalogue

Submit a tool

Harnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.

80 tools · showing 25–48

inference ×
Tool

Rtk

Inference

rtk-ai

CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies

productivitytoken-optimizationcliproxycoding-agent

A
Platform

Superlinked Inference Engine (SIE)

InferenceDeploymentData Wrangling

superlinked

Self-hosted Kubernetes inference cluster serving LLMs, embeddings, rerankers, OCR and vision models with cluster-wide batching.

inference-serverkubernetesrerankingocr

A
Tool

Apfel

InterfaceInference

Arthur-Ficial

The free AI already on your Mac. CLI tool, OpenAI-compatible server, and interactive chat — all on-device via Apple Intelligence. No API key

apple-intelligenceon-devicemacosopenai-compatible

A
Tool

Atomic Agent

InterfaceMemoryInference

AtomicBot-ai

Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.

local-firstllama-cppcomputer-usebrowser-automation

A
Platform

Exo

InferenceDeployment

exo-explore

Open-source distributed inference runtime that pools everyday devices into one local cluster to serve frontier models behind an OpenAI-compatible API.

distributed-inferencedevice-clusteropenai-compatiblelocal-first

A
Tool

Rapid-MLX

Inference

raullenchai

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache,

apple-siliconmlxlocal-firstollama-alternative

A
Tool

Fal

InferenceGenerative MediaDeployment

fal-ai

Fast generative-media inference (image, audio, video) via fal serverless models.

serverlessgpu

A
Tool

LLM

InterfaceInference

simonw

CLI tool and Python library for running prompts, embeddings and tools against many local and remote LLMs, with responses logged to SQLite

clisqlitevector-dbplugin

A
Tool

OpenSquilla

InferenceMemory

OpenSquilla

Token-efficient microkernel AI agent with an on-device model router, persistent memory and pluggable multi-provider support

model-routeron-devicemulti-providercli

A
Tool

WebLLM

Inference

mlc-ai

In-browser LLM inference engine running open models client-side on WebGPU with an OpenAI-compatible API

webgpubrowseropenai-compatiblejson-mode

A
Tool

Flama

InferenceInterfaceConnectors

vortico

Ollama is fine for trying a model. vLLM becomes interesting when that model has to serve traffic: many users, agent work

openai-compatiblechatbotasgi

A
Tool

Magnitude

Inference

magnitudedev

Your fully local, private agent. Runs models on your machine with its built-in inference engine. Works out of the box, on any hardware.

local-firstmodel-managementcli

A
Platform

Modal

DeploymentInference

modal-labs

Serverless GPU/CPU cloud for inference, sandboxes and agent workloads.

serverlessgpusandboxautoscaling

A
Tool

GoModel

InferenceMonitoring

ENTERPILOT

AI gateway written in Go. Lightweight unified OpenAI-compatible API for OpenAI, Anthropic, Gemini, Groq, xAI & Ollama. LiteLLM alternative w

gatewayproxyopenai-compatibletoken-optimization

A
Tool

NadirClaw

Inference

NadirRouter

Open-source LLM router & AI cost optimizer. Routes simple prompts to cheap/local models, complex ones to premium — automatically. Drop-in Op

gatewaytoken-optimizationproxylocal-first

A
Tool

Ratel

MemoryInference

ratel-ai

Context engineering for agents: in-process BM25 + semantic retrieval over tools, skills and memory with progressive disclosure, no vector DB.

token-optimizationtool-callingbm25progressive-disclosure

A
Tool

Luminal

Inference

luminal-ai

Rust inference compiler that lowers static graphs of ~15 primitive ops straight to CUDA/Metal, with a torch.compile backend.

inference-compilercudametalpytorch

B
Tool

Muna

InferenceDeployment

muna-ai

Compiles Python AI functions into self-contained native binaries and serves open models via an OpenAI-compatible client across cloud, edge and device.

openai-compatibleon-devicegpumodel-compilation

B
Tool

BitNet

Inference

microsoft

Microsoft’s bitnet.cpp is a new CPU-first engine for running 1‑bit LLMs based on the BitNet b1.58 architecture, which us

1-bit-llmquantizationcpu-inferencevector-db

B
Tool

Lynkr

Inference

Fast-Editor

Streamline your workflow with Lynkr, a CLI tool that acts as an HTTP proxy for efficient code interactions using Claude Code CLI.

gatewaysemantic-cachetoken-optimizationproxy

B
Harness

Odysseus

InterfaceResearchInference

pewdiepie-archdaemon

Odysseus is a self-hosted workspace with powerful local tools. Keep auth enabled, keep private data out of Git, and do not expose raw model/service ports publicly.

productivitylocal-firstchatbotdockeremail

B
Platform

Tinyagentos

MemoryDeploymentInference

jaylfc

Self-hosted, framework-agnostic AI agent platform that runs on your own hardware with a browser desktop

local-firstframework-agnosticmulti-frameworkknowledge-graph

B
Tool

EmbedAnything

Data WranglingMemoryInference

StarlightSearch

Rust-based inference, ingestion and indexing library for embeddings and retrieval, with Python bindings.

local-firstcloudragsdk

B
Tool

MLC LLM

Inference

mlc-ai

Machine-learning compiler and engine for deploying LLMs across AMD, NVIDIA, Apple, and Intel GPUs, browsers, iOS, and Android

cross-platformwebgpuon-devicequantization

B