This is an early release preview. You may encounter bugs.

Catalogue

Submit a tool

Harnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.

80 tools · showing 49–72

inference ×
Tool

Mlx Dspark

Inference

ARahim3

Up to 3× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-

speculative-decodingmlxapple-siliconinference-optimization

B
Tool

Router

Inference

workweave

Model router for agentic systems. Routes every prompt to the right model in <50ms. Cut costs 40-70% with just an endpoint change.

gatewayproxyopenai-compatibletoken-optimization

B
Tool

Collective Intelligence

Inference

ailinone

Point an AI agent at your company database and it writes SQL that looks right and gives the wrong answer, because it doe

gatewaymulti-providerconsensusopenai-compatible

B
Tool

Context Gateway

CodingInference

Compresr-ai

Interesting idea: context management as infrastructure, not application logic. Context Gateway runs as a transparent pro

token-optimizationproxyprompt-compression

B
Tool

CascadeFlow

InferenceMonitoring

lemony-ai

Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.

model-cascadingtoken-optimizationgatewaylangchain

B
Tool

Aisuite

CodingInference

andrewyng

Simple, unified interface to multiple Generative AI providers

multi-providergatewaytool-callingagents-api

B
Platform

Helicone

InferenceMonitoring

Acquired

Helicone

AI gateway to 100+ models with routing and fallbacks, plus request logging, tracing, and cost analytics

gatewayfallbackscost-analyticsobservability

B
Platform

Together AI

InferenceTrainingDeployment

togethercomputer

Inference, fine-tuning and GPU-cluster platform serving open-weight models.

gpu-clusterfine-tuneapisdk

B
Tool

Zinc

Inference

zolotukhin

Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon

ggufvulkanrocmmetal

B
Tool

Hermes Tool Router

Inference

AtlasOmnia

Experimental Hermes plugin that prunes the first-turn tool surface to cut tool-schema token overhead, failing open

gatewaytoken-optimizationpluginopenrouter

C
Platform

Cerebras Inference

Inference

Cerebras

OpenAI-compatible inference cloud serving open models on wafer-scale hardware

openai-compatiblelow-latencycloudsdk

C
Platform

CoAI

InferenceData WranglingInterface

coaidev

Unified LLM gateway and chat platform with multi-provider routing, billing, load balancing, file parsing, and web search

financeproductivitygatewaymulti-providerocrsearch

C
Harness

Caveman Code

CodingInference

JuliusBrussee

Terminal coding agent that answers in a compressed caveman register to cut per-turn output tokens

multi-providertoken-optimizationcliautopilot

C
Tool

Inter-1

Inference

Signals API over an omni-modal model that detects confidence, frustration, engagement and nine more cues from video

multimodalsocial-signalsemotion-detection

C
Tool

Moondream

Inference

vikhyat

Small open-source vision-language model with a local server and client SDKs for VQA, captioning, pointing and object detection in agent pipelines.

vision-language-modelimage-captioningobject-detectionmultimodal

C
Tool

Pal MCP Server

CodingInference

BeehiveInnovations

MCP server that lets one coding CLI orchestrate multiple models and spawn isolated CLI subagents within a shared context

multi-providercli-bridgemulti-agentmcp-server

C
Tool

Quantum Free Router

Inference

spacepirate15

Pre-configured Bifrost router: one OpenAI-compatible endpoint over free-tier LLM providers, with auto-failover and model certification

gatewayopenai-compatiblebifrostfailover

C
Tool

Ccflare

InferenceMonitoring

snipeship

Rate limits and opaque failures are a primary scaling bottleneck for Claude workloads. ccflare sits in front of Anthropi

proxygatewaymulti-account

C
Tool

Model Router

Inference

open-world-project

Cost-aware per-turn model routing plugin for Hermes Agent with five tiers, automatic escalation, and manual pins

gatewayopenroutertoken-optimizationplugin

C
Tool

LLM Keypool

Inference

piyush-tyagi-13

Local OpenAI-compatible proxy that pools free-tier LLM keys with round-robin rotation, 429 cooldowns, and an audit log

openai-compatibleproxygatewaymulti-provider

C
Platform

Adaptive Engine

TrainingInferenceQA

RLOps platform for continuously fine-tuning, evaluating and serving specialised LLMs from production feedback (company acquired by Datadog).

rlopsfine-tuneproduction-feedback

Score unavailable
Platform

Aperture

SecurityInference

Identity-aware AI gateway on Tailscale's network that centralizes provider credentials, model tokens and MCP access control for agents.

gatewayauthcredential-managementtailscale

Score unavailable
App

Atomic Hermes

InterfaceConnectorsInference

AtomicBot-ai

Native macOS AI assistant bundling chat, terminal, file browser, computer-use, and local models in one app

productivitymacoslocal-firstdesktopcomputer-use

App

BaseRT

Inference

6.4x faster than llama.cpp, 3.9x faster than MLX

apple-siliconlocal-firstmacosllm-runtime