This is an early release preview. You may encounter bugs.

Catalogue

Submit a tool

Harnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.

12 tools

#quantization ×
Tool

llama.cpp

Deployment

ggml-org

LLM inference in C/C++ across CPU and GPU backends, using the GGUF format with quantization, a REST server, and a WebUI

ggufquantizationguilocal-first

A+
Tool

vLLM

Inference

vllm-project

LLM inference and serving library using PagedAttention and continuous batching, with an OpenAI-compatible API server

pagedattentionopenai-compatiblequantization

A+
Tool

whisper.cpp

Voice

ggml-org

C/C++ port of OpenAI's Whisper speech-recognition model, dependency-free and optimized for on-device inference

stton-devicemetalquantization

A+
Tool

MNN

Coding

alibaba

Lightweight deep learning engine for on-device inference and training, with runtimes for local LLMs and diffusion models

on-devicemobileembeddedquantization

A+
Tool

Axolotl

Training

OpenAccess-AI-Collective

Open-source framework for fine-tuning and post-training large language models across many architectures and hardware setups

fine-tuneloramoequantization

A+
Tool

LiteRT

Coding

google-ai-edge

On-device runtime (formerly TensorFlow Lite) for running models at the edge.

on-devicegpunpuquantization

A+
Tool

Unsloth

TrainingInference

unslothai

Run and fine-tune open LLMs, vision, and audio models locally, with training claimed at up to 2x faster and up to 70% less VRAM

fine-tuneloragguflocal-first

A
Tool

Soup

Training

MakazhanAlpamys

Fine-tuning a small model on a few hundred rows costs a fraction of sending the same task through an API.

loradpoquantizationconsumer-gpu

A
Tool

Pruna

Inference

PrunaAI

Python package that quantises, prunes, caches and compiles models to cut inference cost and latency.

quantizationpruningmodel-compressionlatency-optimization

A
Tool

mistral.rs

Coding

EricLBuehler

Rust LLM inference engine with multimodal support, automatic quantization, and OpenAI- and Anthropic-compatible serving

quantizationopenai-compatiblemultimodallora

A
Tool

BitNet

Inference

microsoft

Microsoft’s bitnet.cpp is a new CPU-first engine for running 1‑bit LLMs based on the BitNet b1.58 architecture, which us

1-bit-llmquantizationcpu-inferencevector-db

B
Tool

MLC LLM

Inference

mlc-ai

Machine-learning compiler and engine for deploying LLMs across AMD, NVIDIA, Apple, and Intel GPUs, browsers, iOS, and Android

cross-platformwebgpuon-devicequantization

B