This is an early release preview. You may encounter bugs.

Catalogue

Submit a tool

Harnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.

15 tools

#gpu ×
Tool

LMCache

DeploymentCoding

lmcache

A KV Cache Management Layer for Scalable LLM Inference.

kv-cachevllmgpupytorch

A+
Tool

Truss

DeploymentInference

basetenlabs

Open-source format and CLI for packaging a model as a production API.

gpudockerclivllm

A+
Tool

LLM D

DeploymentCoding

llm-d

A high-performance distributed inference serving stack optimized for production deployments on Kubernetes

kubernetesvllmkv-cachegpu

A+
Tool

LiteRT

Coding

google-ai-edge

On-device runtime (formerly TensorFlow Lite) for running models at the edge.

on-devicegpunpuquantization

A+
Tool

Server

Coding

triton-inference-server

Inference-serving software that deploys models from multiple frameworks across GPU and CPU on cloud, data center, and edge

gputensorrtonnxnvidia

A
Tool

BentoML

Coding

bentoml

Python framework for turning AI models into inference APIs and multi-model serving systems

mlopsdockergpu

A
Tool

Mesh LLM

Deployment

Mesh-LLM

Distributed LLM inference that pools GPUs across machines and serves one OpenAI-compatible API, splitting large models across nodes

distributed-inferenceopenai-compatiblegpulocal-first

A
Tool

Fal

InferenceGenerative MediaDeployment

fal-ai

Fast generative-media inference (image, audio, video) via fal serverless models.

serverlessgpu

A
Platform

Modal

DeploymentInference

modal-labs

Serverless GPU/CPU cloud for inference, sandboxes and agent workloads.

serverlessgpusandboxautoscaling

A
Tool

TuFT

TrainingDeployment

agentscope-ai

Multi-tenant fine-tuning for LLMs with Tinker-compatible API

fine-tunemulti-tenantgpusdk

B
Tool

Luminal

Inference

luminal-ai

Rust inference compiler that lowers static graphs of ~15 primitive ops straight to CUDA/Metal, with a torch.compile backend.

inference-compilercudametalpytorch

B
Tool

Muna

InferenceDeployment

muna-ai

Compiles Python AI functions into self-contained native binaries and serves open models via an OpenAI-compatible client across cloud, edge and device.

openai-compatibleon-devicegpumodel-compilation

B
Platform

Baseten

Deployment

Managed inference platform for deploying and autoscaling models behind APIs.

autoscalinggpumulti-cloud

Score unavailable
Platform

Hyperbolic

Inference

Decentralised GPU marketplace plus serverless OpenAI-compatible inference API over 25+ open models, with on-demand and reserved bare-metal compute.

gpuopenai-compatibleserverlessdecentralized

Score unavailable
Platform

RunPod

DeploymentTrainingInference

GPU cloud offering Pods, autoscaling Serverless workers, multi-node Clusters and a repo-backed Hub for one-click model endpoints.

gpucloudserverlessautoscaling

Score unavailable