Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
12 tools
Toolllama.cpp
Deployment
ggml-org
LLM inference in C/C++ across CPU and GPU backends, using the GGUF format with quantization, a REST server, and a WebUI
ggufquantizationguilocal-first
vLLM
Inference
vllm-project
LLM inference and serving library using PagedAttention and continuous batching, with an OpenAI-compatible API server
pagedattentionopenai-compatiblequantization
Toolwhisper.cpp
Voice
ggml-org
C/C++ port of OpenAI's Whisper speech-recognition model, dependency-free and optimized for on-device inference
stton-devicemetalquantization
MNN
Coding
alibaba
Lightweight deep learning engine for on-device inference and training, with runtimes for local LLMs and diffusion models
on-devicemobileembeddedquantization
ToolAxolotl
Training
OpenAccess-AI-Collective
Open-source framework for fine-tuning and post-training large language models across many architectures and hardware setups
fine-tuneloramoequantization
ToolLiteRT
Coding
google-ai-edge
On-device runtime (formerly TensorFlow Lite) for running models at the edge.
on-devicegpunpuquantization
ToolUnsloth
TrainingInference
unslothai
Run and fine-tune open LLMs, vision, and audio models locally, with training claimed at up to 2x faster and up to 70% less VRAM
fine-tuneloragguflocal-first
ToolSoup
Training
MakazhanAlpamys
Fine-tuning a small model on a few hundred rows costs a fraction of sending the same task through an API.
loradpoquantizationconsumer-gpu
Pruna
Inference
PrunaAI
Python package that quantises, prunes, caches and compiles models to cut inference cost and latency.
quantizationpruningmodel-compressionlatency-optimization
Toolmistral.rs
Coding
EricLBuehler
Rust LLM inference engine with multimodal support, automatic quantization, and OpenAI- and Anthropic-compatible serving
quantizationopenai-compatiblemultimodallora
ToolBitNet
Inference
microsoft
Microsoft’s bitnet.cpp is a new CPU-first engine for running 1‑bit LLMs based on the BitNet b1.58 architecture, which us
1-bit-llmquantizationcpu-inferencevector-db
ToolMLC LLM
Inference
mlc-ai
Machine-learning compiler and engine for deploying LLMs across AMD, NVIDIA, Apple, and Intel GPUs, browsers, iOS, and Android
cross-platformwebgpuon-devicequantization