Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
80 tools · showing 1–24
vLLM
Inference
vllm-project
LLM inference and serving library using PagedAttention and continuous batching, with an OpenAI-compatible API server
pagedattentionopenai-compatiblequantization
ToolRay
InferenceTrainingDeployment
ray-project
Distributed runtime and ML libraries for scaling Python and AI workloads from a laptop to a cluster
distributed-computingreinforcement-learninghyperparameter-tuning
Toolany LLM
InferenceCoding
mozilla-ai
One interface to talk to 40+ LLM providers.
multi-providergatewayprovider-agnostic
ToolTruss
DeploymentInference
basetenlabs
Open-source format and CLI for packaging a model as a production API.
gpudockerclivllm
ToolLiteLLM
Inference
BerriAI
Open-source AI gateway exposing 100+ LLM providers through one OpenAI-compatible interface, as a Python SDK or self-hosted proxy
openai-compatibleproxymulti-providergateway
ToolLocalAI
InferenceGenerative Media
mudler
Self-hosted engine that runs LLM, vision, voice, image, and video models on any hardware behind OpenAI-compatible APIs
local-firstopenai-compatiblellama-cppmultimodal
ToolOmniRoute
Inference
diegosouzapw
AI gateway fronting many model providers behind one endpoint, with token-saving compression, auto-fallback, and routing for coding CLIs
gatewaytoken-optimizationmulti-providerauto-fallback
ToolAPISIX
Inference
apache
Dynamic API gateway on NGINX and etcd that also serves as an AI gateway proxying and rate-limiting LLM traffic
gatewaykubernetesproxy
ToolOpenJarvis
InferenceTraining
open-jarvis
Personal AI, On Personal Devices
productivitylocal-firstollamaclicron
ToolHeadroom
Inference
chopratejas
Content-aware context-compression layer that reversibly shrinks tool outputs, logs, and history before they reach the LLM
token-optimizationproxyraglangchain
ToolManifest
Inference
mnfst
Open-source LLM router exposing one OpenAI-compatible endpoint across API keys, subscriptions, and local models, with cost tracking
gatewaytoken-optimizationopenai-compatiblebyok
ToolUnsloth
TrainingInference
unslothai
Run and fine-tune open LLMs, vision, and audio models locally, with training claimed at up to 2x faster and up to 70% less VRAM
fine-tuneloragguflocal-first
ToolBifrost
Inference
maximhq
High-throughput AI gateway unifying many model providers behind one OpenAI-compatible API with failover, load balancing, and caching
gateway
Diffusers
Generative MediaInferenceTraining
huggingface
Library of pretrained diffusion models for generating images, audio, and 3D structures, for inference or training
diffusionstable-diffusionpytorchflux
ToolJan
InterfaceInference
menloresearch
Desktop app for running open-weight LLMs locally or connecting to cloud providers, exposing an OpenAI-compatible local API
local-firstopenai-compatiblellama-cppdesktop
ToolKTransformers
InferenceTraining
kvcache-ai
Framework for LLM inference and fine-tuning on CPU-GPU heterogeneous hardware, tuned for large mixture-of-experts models
moecpu-gpufine-tune
Replicate
InferenceDeployment
replicate
Run and host open-source models via API — image generation and beyond.
dockercudaapi
ToolCodex Chatgpt Web
InferenceInterface
miuuyy
Use ChatGPT Web (including Pro) as a native model in the Codex app — with context, tools, streaming and images beyond Codex usage limits.
playwrightbrowser-automationproxyresponses-api
Caveman
CodingInference
JuliusBrussee
Skill for Claude Code and 30+ agents that compresses replies into terse caveman-speak to cut output tokens
token-optimizationprompt-compressionskill
ToolMlx Serve
InferenceGenerative MediaVoice
ddalcu
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mo
mlxggufapple-siliconopenai-compatible
HarnessOsaurus
InferenceMemory
osaurus-ai
Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity
productivitylocal-firstmacosdesktopmlx
ToolSwitchyard
Inference
NVIDIA-NeMo
Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility -
gatewaylitellmopenai-compatibletoken-optimization
ToolODS
DeploymentInterfaceInference
Osmantic
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
local-firstollamacomfyuin8n
Pruna
Inference
PrunaAI
Python package that quantises, prunes, caches and compiles models to cut inference cost and latency.
quantizationpruningmodel-compressionlatency-optimization