Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
80 tools · showing 49–72
ToolMlx Dspark
Inference
ARahim3
Up to 3× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-
speculative-decodingmlxapple-siliconinference-optimization
ToolRouter
Inference
workweave
Model router for agentic systems. Routes every prompt to the right model in <50ms. Cut costs 40-70% with just an endpoint change.
gatewayproxyopenai-compatibletoken-optimization
ToolCollective Intelligence
Inference
ailinone
Point an AI agent at your company database and it writes SQL that looks right and gives the wrong answer, because it doe
gatewaymulti-providerconsensusopenai-compatible
ToolContext Gateway
CodingInference
Compresr-ai
Interesting idea: context management as infrastructure, not application logic. Context Gateway runs as a transparent pro
token-optimizationproxyprompt-compression
ToolCascadeFlow
InferenceMonitoring
lemony-ai
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
model-cascadingtoken-optimizationgatewaylangchain
ToolAisuite
CodingInference
andrewyng
Simple, unified interface to multiple Generative AI providers
multi-providergatewaytool-callingagents-api
Helicone
InferenceMonitoring
Acquired
Helicone
AI gateway to 100+ models with routing and fallbacks, plus request logging, tracing, and cost analytics
gatewayfallbackscost-analyticsobservability
PlatformTogether AI
InferenceTrainingDeployment
togethercomputer
Inference, fine-tuning and GPU-cluster platform serving open-weight models.
gpu-clusterfine-tuneapisdk
ToolZinc
Inference
zolotukhin
Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon
ggufvulkanrocmmetal
ToolHermes Tool Router
Inference
AtlasOmnia
Experimental Hermes plugin that prunes the first-turn tool surface to cut tool-schema token overhead, failing open
gatewaytoken-optimizationpluginopenrouter
Cerebras Inference
Inference
Cerebras
OpenAI-compatible inference cloud serving open models on wafer-scale hardware
openai-compatiblelow-latencycloudsdk
PlatformCoAI
InferenceData WranglingInterface
coaidev
Unified LLM gateway and chat platform with multi-provider routing, billing, load balancing, file parsing, and web search
financeproductivitygatewaymulti-providerocrsearch
Caveman Code
CodingInference
JuliusBrussee
Terminal coding agent that answers in a compressed caveman register to cut per-turn output tokens
multi-providertoken-optimizationcliautopilot
ToolInter-1
Inference
Signals API over an omni-modal model that detects confidence, frustration, engagement and nine more cues from video
multimodalsocial-signalsemotion-detection
ToolMoondream
Inference
vikhyat
Small open-source vision-language model with a local server and client SDKs for VQA, captioning, pointing and object detection in agent pipelines.
vision-language-modelimage-captioningobject-detectionmultimodal
ToolPal MCP Server
CodingInference
BeehiveInnovations
MCP server that lets one coding CLI orchestrate multiple models and spawn isolated CLI subagents within a shared context
multi-providercli-bridgemulti-agentmcp-server
ToolQuantum Free Router
Inference
spacepirate15
Pre-configured Bifrost router: one OpenAI-compatible endpoint over free-tier LLM providers, with auto-failover and model certification
gatewayopenai-compatiblebifrostfailover
ToolCcflare
InferenceMonitoring
snipeship
Rate limits and opaque failures are a primary scaling bottleneck for Claude workloads. ccflare sits in front of Anthropi
proxygatewaymulti-account
ToolModel Router
Inference
open-world-project
Cost-aware per-turn model routing plugin for Hermes Agent with five tiers, automatic escalation, and manual pins
gatewayopenroutertoken-optimizationplugin
ToolLLM Keypool
Inference
piyush-tyagi-13
Local OpenAI-compatible proxy that pools free-tier LLM keys with round-robin rotation, 429 cooldowns, and an audit log
openai-compatibleproxygatewaymulti-provider
Adaptive Engine
TrainingInferenceQA
RLOps platform for continuously fine-tuning, evaluating and serving specialised LLMs from production feedback (company acquired by Datadog).
rlopsfine-tuneproduction-feedback
Aperture
SecurityInference
Identity-aware AI gateway on Tailscale's network that centralizes provider credentials, model tokens and MCP access control for agents.
gatewayauthcredential-managementtailscale
Atomic Hermes
InterfaceConnectorsInference
AtomicBot-ai
Native macOS AI assistant bundling chat, terminal, file browser, computer-use, and local models in one app
productivitymacoslocal-firstdesktopcomputer-use
BaseRT
Inference
6.4x faster than llama.cpp, 3.9x faster than MLX
apple-siliconlocal-firstmacosllm-runtime