Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
15 tools
ToolLMCache
DeploymentCoding
lmcache
A KV Cache Management Layer for Scalable LLM Inference.
kv-cachevllmgpupytorch
ToolTruss
DeploymentInference
basetenlabs
Open-source format and CLI for packaging a model as a production API.
gpudockerclivllm
ToolLLM D
DeploymentCoding
llm-d
A high-performance distributed inference serving stack optimized for production deployments on Kubernetes
kubernetesvllmkv-cachegpu
ToolLiteRT
Coding
google-ai-edge
On-device runtime (formerly TensorFlow Lite) for running models at the edge.
on-devicegpunpuquantization
ToolServer
Coding
triton-inference-server
Inference-serving software that deploys models from multiple frameworks across GPU and CPU on cloud, data center, and edge
gputensorrtonnxnvidia
ToolBentoML
Coding
bentoml
Python framework for turning AI models into inference APIs and multi-model serving systems
mlopsdockergpu
ToolMesh LLM
Deployment
Mesh-LLM
Distributed LLM inference that pools GPUs across machines and serves one OpenAI-compatible API, splitting large models across nodes
distributed-inferenceopenai-compatiblegpulocal-first
Fal
InferenceGenerative MediaDeployment
fal-ai
Fast generative-media inference (image, audio, video) via fal serverless models.
serverlessgpu
PlatformModal
DeploymentInference
modal-labs
Serverless GPU/CPU cloud for inference, sandboxes and agent workloads.
serverlessgpusandboxautoscaling
ToolTuFT
TrainingDeployment
agentscope-ai
Multi-tenant fine-tuning for LLMs with Tinker-compatible API
fine-tunemulti-tenantgpusdk
ToolLuminal
Inference
luminal-ai
Rust inference compiler that lowers static graphs of ~15 primitive ops straight to CUDA/Metal, with a torch.compile backend.
inference-compilercudametalpytorch
ToolMuna
InferenceDeployment
muna-ai
Compiles Python AI functions into self-contained native binaries and serves open models via an OpenAI-compatible client across cloud, edge and device.
openai-compatibleon-devicegpumodel-compilation
Baseten
Deployment
Managed inference platform for deploying and autoscaling models behind APIs.
autoscalinggpumulti-cloud
Hyperbolic
Inference
Decentralised GPU marketplace plus serverless OpenAI-compatible inference API over 25+ open models, with on-demand and reserved bare-metal compute.
gpuopenai-compatibleserverlessdecentralized
RunPod
DeploymentTrainingInference
GPU cloud offering Pods, autoscaling Serverless workers, multi-node Clusters and a repo-backed Hub for one-click model endpoints.
gpucloudserverlessautoscaling