This is an early release preview. You may encounter bugs.
Unsloth logo
Unclaimed

Tool training inference

Unsloth

Run and fine-tune open LLMs, vision, and audio models locally, with training claimed at up to 2x faster and up to 70% less VRAM

A 89/100 GitHub score ? This grade is derived from GitHub signals, not user votes. Open for the full breakdown.
No votes yet

01 / About

What Unsloth is.

Unsloth runs and fine-tunes open models on your own machine, through a native desktop app, a web interface called Unsloth Studio, or a Python package. It covers text, vision, audio, diffusion, and embedding models in Hugging Face, GGUF, and MLX formats, and works on Windows, Linux, WSL, and macOS across multi-GPU NVIDIA, AMD, and Intel setups, CPUs, and the Vulkan backend.

On the training side it supports full fine-tuning, LoRA, QLoRA, pretraining, FP8, and reinforcement learning methods including GRPO and DPO, with reported figures of 2x faster training at 70% less VRAM and no accuracy loss, mixture-of-experts training at 12x faster with 35% less VRAM, 1.8-3.3x faster embedding fine-tuning, and a 20B model trained at over 500K context on an 80GB GPU. Free Colab notebooks cover Gemma 4, Qwen 3.5, gpt-oss, Llama 3.1 and 3.2, EmbeddingGemma, and Orpheus text-to-speech, each listing its speed and memory saving. Trained models export to GGUF, NVFP4, FP8, and other formats, and datasets can be built from PDFs, CSVs, and DOCX files with Data Recipes.

On the serving side, models are exposed through an OpenAI-compatible API, over the local network, or remotely through a Cloudflare HTTPS tunnel. unsloth start points a coding agent at a local model in one command — Claude Code, OpenAI Codex, Hermes Agent, OpenClaw, OpenCode, and DeepSeek Harness each have a subcommand — using Unsloth's OpenAI- and Anthropic-compatible APIs. MCP servers connect local models to files, applications, databases, and external tools, and connections can mix local models with API providers such as OpenAI and Anthropic or servers such as vLLM and Ollama in one interface. Web search, deep research, auto-compaction, and retrieval-augmented generation are built in.

Features

  • Three surfaces: a native desktop app, the Unsloth Studio web interface, and the code-based Unsloth Core package
  • Model breadth: LLMs plus MLX, GGUF, diffusion, embedding, and audio models, including Qwen3.8, GLM-5.3-Flash, Kimi K3, MiniMax-H3, DeepSeek-V4, and Gemma 4
  • Fine-tuning methods: LoRA, QLoRA, full fine-tuning, pretraining, FP8, and reinforcement learning with GRPO and DPO
  • Reported efficiency: 2x faster training with 70% less VRAM, 12x faster mixture-of-experts training with 35% less VRAM
  • Agent connections: unsloth start wires Claude Code, Codex, Hermes Agent, OpenClaw, OpenCode, and DeepSeek Harness to a local model
  • OpenAI-compatible serving: an OpenAI- and Anthropic-compatible API, plus LAN access and a Cloudflare HTTPS tunnel for remote use
  • MCP and tools: local models call files, applications, databases, and external tools over the Model Context Protocol
  • Mixed connections: local models, API providers, and vLLM or Ollama servers in one interface
  • Search and RAG: built-in web search, deep research, rolling-context auto-compaction, and retrieval-augmented generation
  • Export formats: GGUF, NVFP4, FP8, and other deployment formats
  • Dataset building: Data Recipes builds training sets from PDFs, CSVs, and DOCX files
  • Hardware coverage: NVIDIA, AMD, and Intel GPUs, CPUs, the Vulkan backend, multi-GPU setups, DGX Spark, and Blackwell

02 / Discussion CREDIBILITY-GATED

Discussion

Reading is open to everyone. Posting and voting need a verified identity or a GitHub grade of B or higher.

  • No discussions yet.

03 / Build

Build with Unsloth.

Browse the catalogue for frameworks, tools, and harnesses, each scored on real GitHub credibility.

Get Unsloth →

Browse the catalogue