Sign in

Unsloth AI

@unsloth.ai
1.5K followers 0 following 88 posts

Making open-source AI more accessible! 🦥 Github: github.com/unslothai/unsloth

PostsRepliesMedia
Unsloth AI @unsloth.ai · 28/09/2026
You can now run Laya Decision models locally on just 4GB RAM! 🔥 Works on CPU, Mac, Windows, Linux and GPU setups. Serve Laya through a Jev-compatible API via Unsloth Desktop. GitHub: github.com/unslothai/un... Guide: unsloth.ai/docs/models/...
0568
Unsloth AI @unsloth.ai · 23/09/2026
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗 Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever. Thanks for all your support!
0343
Unsloth AI @unsloth.ai · 22/09/2026
Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! The 7B parameter model performs on par with Nano Banana 2.0. Run our Dynamic FP8 or GGUFs for higher quality via diffusers, Unsloth Desktop & more. GGUF: huggingface.co/unsloth/Qwen... Guide: unsloth.ai/docs/models/...
1407
Unsloth AI @unsloth.ai · 17/09/2026
You can now train and run 500+ models locally with our Unsloth Docker image! 🐳 Use our new GUI or notebooks workflow. No setup required. Works on NVIDIA and AMD. Guide: unsloth.ai/docs/get-sta... GitHub: github.com/unslothai/un...
1221
Unsloth AI @unsloth.ai · 08/09/2026
Qwen3.8-27B Unsloth GGUF is now the #1 most-liked GGUF of all time! The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥 GGUF: huggingface.co/unsloth/Qwen... Guide: unsloth.ai/docs/models/...
0524
Unsloth AI @unsloth.ai · 04/09/2026
We made GLM-5.3-Flash run 3.3x faster locally! Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction. Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/GLM-...
0637
Unsloth AI @unsloth.ai · 02/09/2026
Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: huggingface.co/unsloth/Qwen... Guide: unsloth.ai/docs/models/...
1331
Unsloth AI @unsloth.ai · 28/08/2026
GLM-5.3 can now be run locally! The 2-bit model retains ~81% accuracy after we shrunk it from 1.51TB to 239GB (-83% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.3 is the strongest open model to date. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/GLM-...
0454
Unsloth AI @unsloth.ai · 27/08/2026
GLM-5.3-Flash can now be run locally! ✨ Run 3-bit on 128GB RAM via Unsloth GGUF. GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/GLM-...
0364
Unsloth AI @unsloth.ai · 26/08/2026
Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Qwen...
0274
Unsloth AI @unsloth.ai · 26/08/2026
GLM-5.3-Flash, also known as ox-alpha, is out now! Zai's GLM-5.3-Flash is a new 320B parameter multimodal open model with 18B active parameters. GLM-5.3-Flash approaches Claude 4.8 Opus on coding and agentic benchmarks. GGUF coming soon!
111914
Unsloth AI @unsloth.ai · 25/08/2026
You can now fine-tune Qwen3.8-27B for free with our notebook! 🔥 Local training works on 24GB VRAM. Unsloth trains Qwen3.8-27B 1.5x faster with 50% less VRAM. GitHub: github.com/unslothai/un... Qwen3.8-27B Notebooks + Guide: unsloth.ai/docs/models/...
0378
Unsloth AI @unsloth.ai · 25/08/2026
Qwen announces Qwen3.8-Flash-Next, a new open-weight multimodal MoE model. 💜 The model will be released tomorrow and we are working on Unsloth day zero support. Qwen3.8-Flash-Next is built on the new Qwen4 architecture.
110011
Unsloth AI @unsloth.ai · 19/08/2026
We’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy. Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks. We also release 1-bit quants that retain 77% accuracy. Run on 8GB RAM. Blog: unsloth.ai/docs/basics/... GGUF: huggingface.co/unsloth/Qwen...
2629
Unsloth AI @unsloth.ai · 18/08/2026
Qwen3.8-27B Unsloth GGUF is now the #2 trending model on Hugging Face with 2.7M downloads! 💗 Unsloth also reached #3 trending on GitHub! Thanks so much for the love! Model: huggingface.co/unsloth/Qwen... GitHub: github.com/unslothai/un...
0343
Unsloth AI @unsloth.ai · 14/08/2026
Qwen3.8-27B can now be run locally! ✨ Run on 17GB RAM via Unsloth Dynamic GGUFs. Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants. GGUF: huggingface.co/unsloth/Qwen... Guide: unsloth.ai/docs/models/...
0448
Unsloth AI @unsloth.ai · 13/08/2026
You can now fine-tune Meta Muse Glimmer 30B for free! 🔥 Our free notebook also supports GRPO RL training. Unsloth trains Muse Glimmer 1.5× faster with 50% less VRAM vs FA2 setups. Train locally with 24GB VRAM. Guide: unsloth.ai/docs/models/... Notebooks: unsloth.ai/docs/models/...
0122
Unsloth AI @unsloth.ai · 12/08/2026
Qwen3.8 can now be run locally! 🔥 We shrank Qwen3.8-2.4T-A95B from 4.9TB to 397GB (-91% size) via Dynamic 1-bit by selectively quantizing layers. Run on 410GB+ RAM/VRAM via Unsloth Desktop. Qwen3.8 rivals GPT-5.6 Sol. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Qwen...
0264
Unsloth AI @unsloth.ai · 11/08/2026
Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally. • Open-source. Runs on Mac, Windows, Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code + Codex to local LLMs • 50% more accurate self-healing tool calls Download on unsloth.ai + GitHub
4616
Unsloth AI @unsloth.ai · 10/08/2026
2-bit Muse Glimmer GGUF made 100+ tool calls on just 14GB RAM. 🔥 Muse Glimmer did a complete repo bug hunt for 5 mins nonstop with: evidence, repro, fix, tests and a PR writeup. Run and train it in Unsloth. GitHub repo: github.com/unslothai/un...
1455
Unsloth AI @unsloth.ai · 10/08/2026
Meta releases Muse Glimmer, a new 30B open model that runs on 18GB RAM. Apache 2.0 licensed and built for agentic coding, it supports vision and is the strongest agentic model for its size. Run and train the model via Unsloth. GGUF: huggingface.co/unsloth/Muse... Guide: unsloth.ai/docs/models/...
1428
Unsloth AI @unsloth.ai · 06/08/2026
DeepSeek-V4-Flash can now run 2× faster locally with DSpark! ⚡️ DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change. DeepSeek-V4-Flash-0731 can reach at 120 tokens/s. GGUFs: huggingface.co/unsloth/Deep... Guide: unsloth.ai/docs/models/...
0282
Unsloth AI @unsloth.ai · 03/08/2026
Qwen just announced Qwen3.8-27B along with Qwen3.8-Max! 🔥 Qwen3.8-27B will run locally on 17GB RAM/VRAM setups and is expected to be the best performing model for its size.
711612
Unsloth AI @unsloth.ai · 31/07/2026
DeepSeek V4 Flash 0731 can now be run locally! 🐳 Run DeepSeek V4 Flash lossless 4-bit on 168GB RAM and 3-bit on 110GB RAM. V4 Flash 0731 outperforms V4 Pro. Run via Unsloth or llama.cpp. Smaller quants coming today. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Deep...
1465
Unsloth AI @unsloth.ai · 30/07/2026
You can now run Inkling-Small, a new 276B model by Thinking Machines. Inkling-Small is the strongest open model for its size and runs local on 128GB RAM. Apache-2.0 Licensed, it has image, audio + 1M context support. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Inkl...
2391
Unsloth AI @unsloth.ai · 29/07/2026
We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts. 1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s. GitHub repo: github.com/unslothai/un...
0221
Unsloth AI @unsloth.ai · 29/07/2026
Kimi K3 can now be run locally! ✨ The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size). Run on a Mac Studio + 128GB RAM device. Kimi K3 is the strongest open model to date. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Kimi...
21009
Unsloth AI @unsloth.ai · 27/07/2026
We signed the Open Weights letter because we believe the future of AI should be shaped by everyone, not controlled by a select few. That belief has always been at the heart of Unsloth: everyone should be able to train and run models on their own local device.
2875
Unsloth AI @unsloth.ai · 20/07/2026
Introducing Unsloth for AMD 🚀 You can now train & run LLMs on your AMD hardware • We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs • Works on Windows, WSL, Linux • Train Qwen, Gemma on just 3GB VRAM GitHub: github.com/unslothai/un... Blog: unsloth.ai/docs/basics/...
05710
Unsloth AI @unsloth.ai · 17/07/2026
Gemma 4 is now faster and much more accurate! 🚀 Google made huge improvements to tool-calling and chat accuracy, reliability + speed. To get fixes, re-download our updated GGUF, MLX, NVFP4 quants! Unsloth quants: huggingface.co/collections/... Gemma 4 Guide: unsloth.ai/docs/models/...
0563
Unsloth AI @unsloth.ai · 15/07/2026
Inkling, a new 975B parameter open model is here! From Thinking Machines, Inkling supports image, audio, text & 1M context. We quantized Inkling to Dynamic 1-bit (-86% size) and retained 74.2% of top-1% accuracy. Run on 280GB. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/inkl...
0661
Unsloth AI @unsloth.ai · 14/07/2026
We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: unsloth.ai/docs/basics/... Gemma NVFP4: huggingface.co/collections/...
0342
Unsloth AI @unsloth.ai · 13/07/2026
We collaborated with AWS on a complete guide to LLM Quantization and Deployment. Learn about: • Model formats, dynamic quants & making your own • Choosing GGUF, NVFP4 or FP8 • Picking the right tools & deploy on AWS SageMaker • Benchmark quality, latency & cost Read: aws.amazon.com/blogs/machin...
1171
Unsloth AI @unsloth.ai · 10/07/2026
We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. ⚡ Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/... Qwen3.6 NVFP4: huggingface.co/collections/...
1442
Unsloth AI @unsloth.ai · 07/07/2026
DeepSeek-V4 can now run locally with Unsloth GGUFs! 🐳 Run lossless DeepSeek-V4-Flash on 168GB RAM. 3-bit works on 110GB Mac, RAM, VRAM setups. Run via Unsloth Studio or llama.cpp. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Deep...
0272
Unsloth AI @unsloth.ai · 25/06/2026
What’s your go-to local model right now?
9256
Unsloth AI @unsloth.ai · 23/06/2026
1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5 We gave 3 models the same prompt and compared one-shot outputs. The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra 256GB RAM at ~21.6 tok/s. Which do you like best? GGUF: huggingface.co/unsloth/GLM-... Guide: unsloth.ai/docs/models/...
0484
Unsloth AI @unsloth.ai · 18/06/2026
GLM-5.2 can now be run locally! 🔥 The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.2 is the strongest open model to date. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/GLM-...
1717
Unsloth AI @unsloth.ai · 15/06/2026
You can now run Kimi K2.7 Code locally! 🌘 We shrank the 1T model to 325GB (-48%) via Dynamic 2-bit where important layers are upcasted. Run at >40 tok/s on 330GB RAM/VRAM setups. Run full precision on 610 GB. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Kimi...
41205
Unsloth AI @unsloth.ai · 12/06/2026
DiffusionGemma can now run at 2000+ tokens/sec! ⚡ We made local DiffusionGemma inference 1.8× faster. Run it on 18GB RAM via Unsloth Studio. GitHub: github.com/unslothai/un... Guide: unsloth.ai/docs/models/...
2815
Unsloth AI @unsloth.ai · 11/06/2026
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide: unsloth.ai/docs/models/...
46715
Unsloth AI @unsloth.ai · 10/06/2026
Google releases DiffusionGemma.✨ The new 26B-A4B diffusion text model runs locally on 18GB RAM. It supports high-speed text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio. GGUF: huggingface.co/unsloth/diff... Guide: unsloth.ai/docs/models/...
2626
Unsloth AI @unsloth.ai · 05/06/2026
Google releases Gemma 4 QAT. ✨ You can now run Gemma 4 at 3x less memory with near original performance. Quantization-Aware Training (QAT) makes it possible to run Gemma 4 26B-A4B on 16GB RAM. GGUFs: huggingface.co/collections/... QAT Guide: unsloth.ai/docs/models/...
2684
Unsloth AI @unsloth.ai · 04/06/2026
NVIDIA releases Nemotron 3 Ultra, a new 550B model. 💚 Nemotron-3-Ultra-550B-A55B is NVIDIA's largest LLM yet, with 1M context, frontier coding & chat. Run 2-bit on 200GB RAM, 3-bit on 256GB, 8-bit on 600GB. GGUF: huggingface.co/unsloth/NVID... Guide: unsloth.ai/docs/models/...
0212
Unsloth AI @unsloth.ai · 03/06/2026
Google releases Gemma 4 12B, a new model that can run locally on 8GB RAM. Gemma 4 12B Unified model supports image, audio and 256K context. Run and train the model via Unsloth Studio. GGUF: huggingface.co/unsloth/gemm... Guide: unsloth.ai/docs/models/...
2736
Unsloth AI @unsloth.ai · 01/06/2026
We made a guide on using MCP with local LLMs. Connect Qwen3.6 and Gemma 4 for controlled access to tools, files, APIs, enabling private automated workflows. Learn to use OAuth, Exa, Context7, Hugging Face & more. Guide: unsloth.ai/docs/basics/... GitHub: github.com/unslothai/un...
0121
Unsloth AI @unsloth.ai · 18/05/2026
Qwen3.6 now runs 2x faster with MTP GGUFs! Run locally on just 18GB RAM. ⚡️ MTP enables Qwen3.6 to generate ~1.4–2.2× faster with no accuracy change. Qwen3.6-27B MTP runs at 160 tokens/s. 35B-A3B reaches 240 t/s. GGUFs: huggingface.co/unsloth/Qwen... Guide: unsloth.ai/docs/models/...
111214
Unsloth AI @unsloth.ai · 11/05/2026
We’re excited to share that Unsloth has joined the PyTorch Ecosystem! Unsloth is an open-source project that makes training & running models faster, more accurate with less compute. We want AI to be accessible to everyone. Blog: unsloth.ai/blog/pytorch GitHub: github.com/unslothai/un...
0272
Unsloth AI @unsloth.ai · 06/05/2026
We collaborated with NVIDIA to teach you how we made LLM training ~25% faster! 🚀 Learn how 3 optimizations help your home GPU train models faster: 1. Packed-sequence metadata caching 2. Double-buffered checkpoint reloads 3. Faster MoE routing Guide: unsloth.ai/blog/nvidia-...
0195
Unsloth AI @unsloth.ai · 05/05/2026
We made a guide on how to run open LLMs in Claude Code, Codex and OpenClaw. Use Gemma 4 and Qwen3.6 GGUFs for local agentic coding on 24GB RAM Run with self-healing tool calls, code execution, web search via the Unsloth API endpoint and llama.cpp Guide: unsloth.ai/docs/basics/...
2364