Sign in

Xuan Son Nguyen

@ngxson.hf.co
810 followers 137 following 80 posts

Software Engineer @ Hugging Face 🤗

PostsRepliesMedia
Xuan Son Nguyen @ngxson.hf.co · 10/04/2026
llama.cpp now supports various small OCR models that can run on low-end devices. These models are small enough to run on GPU with 4GB VRAM, and some of them can even run on CPU with decent performance. In this post, I will show you how to use these OCR models with llama.cpp 👇
blog.ngxson.com
Using OCR models with llama.cpp
An easy-to-follow guide on how to use OCR models with llama.cpp.
030
Xuan Son Nguyen @ngxson.hf.co · 05/10/2025
Very nice touch, Gmail 😅
030
Xuan Son Nguyen @ngxson.hf.co · 29/08/2025
Part 2 of my journey building a smart home! 🚀 In this part: > ESPHome & custom component > RF433 receiver & transmitter > Hassio custom addon
100
Xuan Son Nguyen @ngxson.hf.co · 27/08/2025
Just published a new article on my blog 🏃‍♂️ Building My Smart Home - Part 1: Plan, Idea & Home Assistant Check it out!
100
Xuan Son Nguyen @ngxson.hf.co · 14/08/2025
Kudos to Google and the llama.cpp team! 🤝 GGUF support for Gemma 270M right from day-0
150
Xuan Son Nguyen @ngxson.hf.co · 21/07/2025
Richy Mini and SmolLM3 are featured in Github's weekly news! 🚀 🚀
100
Xuan Son Nguyen @ngxson.hf.co · 26/06/2025
Gemma 3n has arrived in llama.cpp 👨‍🍳 🍰 Comes in 2 flavors: E2B and E4B (E means "effective/active parameters")
010
Xuan Son Nguyen @ngxson.hf.co · 11/06/2025
See you this Sunday at AI Plumbers conference: 2nd edition! 📍 Where: GLS Event Campus Berlin, Kastanienallee 82 | 10435 Berlin 👉 Register here: lu.ma/vqx423ct
000
Xuan Son Nguyen @ngxson.hf.co · 03/06/2025
✨✨ AIFoundry is bringing you the AI Plumbers Conference: 2nd edition — an open source meetup for low-level AI builders to dive deep into "the plumbing" of modern AI 📍 Where: GLS Event Campus Berlin, Kastanienallee 82 | 10435 Berlin 📅 When: June 15, 2025 👉 Register now: lu.ma/vqx423ct
011
Xuan Son Nguyen @ngxson.hf.co · 15/05/2025
Hugging Face Inference Endpoints now officially support deploying **vision** models via llama.cpp 👀 👀 Try it now: endpoints.huggingface.co/catalog
000
Xuan Son Nguyen @ngxson.hf.co · 12/05/2025
Real-time webcam demo with @huggingface.bsky.social SmolVLM and llama.cpp server. All running locally on a Macbook M3
120
Xuan Son Nguyen @ngxson.hf.co · 25/04/2025
Although we have A100, H200, M3 Ultra, etc Still can't match the power of that Casio FX 😆
020
Xuan Son Nguyen @ngxson.hf.co · 21/04/2025
llama.cpp vision support just got much better! 🚀 Traditionally, models with complicated chat template like MiniCPM-V or Gemma 3 requires a dedicated binary to run. Now, you can use all supported models via a "llama-mtmd-cli" 🔥 (Only Qwen2VL is not yet supported)
040
Xuan Son Nguyen @ngxson.hf.co · 21/04/2025
Finally have time to write a blog post about ggml-easy! 😂 ggml-easy is a header-only wrapper for GGML, simplifies development with a cleaner API, easy debugging utilities, and native safetensors loading ✨ Great for rapid prototyping!
100
Xuan Son Nguyen @ngxson.hf.co · 20/04/2025
Someone at Google definitely had a lot of fun making this 😆 And if you don't know, it's available in "Starter apps" section on AI Studio. The app is called "Gemini 95"
010
Xuan Son Nguyen @ngxson.hf.co · 20/04/2025
Telling LLM memory requirement WITHOUT a calculator? Just use your good old human brain 🧠 😎 Check out my 3‑step estimation 🚀
031
Xuan Son Nguyen @ngxson.hf.co · 19/04/2025
Google having a quite good sense of humor 😂 Joke aside, 1B model quantized to Q4 without performance degrading is sweet 🤏
021
Xuan Son Nguyen @ngxson.hf.co · 31/03/2025
Cooking a fun thing today, I can now load safetensors file directly to GGML without having to convert it to GGUF! Why? Because this allow me to do experiments faster, especially with models outside of llama.cpp 😆
100
Xuan Son Nguyen @ngxson.hf.co · 30/03/2025
No vibe coding. Just code it ✅ Visit my website --> ngxson.com
020
Xuan Son Nguyen @ngxson.hf.co · 20/03/2025
On Monday, the 24th, I'm proud to give a talk at sota's webinar. My main talk will last for an hour to deep dive into the current state of on-device LLMs, exploring their advantages, trade-offs, and limitations. The session will end with an Q&A, where you can ask me anything about this subject.
130
Xuan Son Nguyen @ngxson.hf.co · 19/03/2025
Had a fantastic chat today with Georgi Gerganov, the brilliant mind behind ggml, llama.cpp, and whisper.cpp! We discussed about: 🚀 The integration of vision models into llama.cpp 🚀 The challenges of maintaining a smooth UX/DX 🚀 The exciting future of llama.cpp Big things ahead - stay tuned!
010
Xuan Son Nguyen @ngxson.hf.co · 13/03/2025
OK now you are the best, Gememe 2.0
000
Xuan Son Nguyen @ngxson.hf.co · 12/03/2025
Wanna try Gemma 3 vision with llama.cpp? There is a playground for that! More in 🧵
120
Xuan Son Nguyen @ngxson.hf.co · 12/03/2025
Day-zero Gemma 3 support in llama.cpp 🤯 👉 4 model sizes: 1B, 4B, 12B, 27B 👉 Vision capability (except for 1B) with bi-direction attention 👉 Context size: 32k (1B) and 128k (4B, 12B, 27B) 👉 +140 languages support (except for 1B) 👉 Day-zero support on many frameworks 🚀
251
Xuan Son Nguyen @ngxson.hf.co · 10/03/2025
Aya Vision is now the number one trending OCR model on Hugging Face 🚀 👉 Comes in 2 sizes, 8B and 32B 👉 Supports 32 languages 👉 Day-zero support with HF Transformers
120
Xuan Son Nguyen @ngxson.hf.co · 08/03/2025
Did you know? A number of 🤗 Hugging Face's blog posts now feature AI-created podcasts 🎙️ This offers an alternative way to absorb extensive and intricate articles 🔍
020
Xuan Son Nguyen @ngxson.hf.co · 06/03/2025
Qwen/QwQ-32B has just arrived on Hugging Chat! Try it now: huggingface.co/chat/models/...
000
Xuan Son Nguyen @ngxson.hf.co · 06/03/2025
CogView-4 is out 🔥🚀 The SoTa OPEN text to image model by ZhipuAI Demo: huggingface.co/spaces/THUDM... ✨ 6B with Apache2.0 ✨ Supports Chinese & English Prompts by ANY length ✨ Generate Chinese characters within images ✨ Creates images at any resolution within a given range
010
Xuan Son Nguyen @ngxson.hf.co · 05/03/2025
Wondering how much RAM is needed to run a given GGUF? Try: npx @huggingface/gguf [model].gguf This also work with remote file, for example: npx @huggingface/gguf https: //huggingface.co/bartowski/Qwen_QwQ-32B-GGUF/resolve/main/Qwen_QwQ-32B-Q4_K_M.gguf
141
Xuan Son Nguyen @ngxson.hf.co · 05/03/2025
Apple unveils the M3 Ultra chip, support up to 512GB unified money, oops sorry, unified memory. Perfect for pro workflows and AI development 👀 Read more: www.apple.com/newsroom/202...
010
Xuan Son Nguyen @ngxson.hf.co · 05/03/2025
Baby wake up! The Hugging Face Reasoning Course is out 🚀 Huge thanks to Maxime Labonne for building the first practical example in the reasoning course. Link: huggingface.co/reasoning-co...
021
Xuan Son Nguyen @ngxson.hf.co · 05/03/2025
DiffRhythm's Revolutionary Music Generation 🎵🎵 🚀 Lightning-Fast Production: Create full-length songs with vocals in under 10 seconds! ⚡ Non-Autoregressive Structure: build on top of variational autoencoder (VAE) 🤏 Small: Both VAE + Base model combined is < 2.5GB 🌍 Open-Source model code + weights
110
Xuan Son Nguyen @ngxson.hf.co · 05/03/2025
With the new 🐸 JFrog 's model scanner on the 🤗 Hugging Face hub, we're making running AI models even more secured for everyone!
020
Xuan Son Nguyen @ngxson.hf.co · 03/03/2025
EgoLife: An AI-Powered Egocentric Life Assistant Key details: > Open-source dataset: 300+ hrs egocentric, multimodal data > 3K long-context QAs for daily insights > Open-source models: EgoGPT & EgoRAG for smart recall Turning real-life moments into personalized AI help!
130
Xuan Son Nguyen @ngxson.hf.co · 03/03/2025
Disassemble Phi-4-multimodal-instruct: > Minimal Vision encoder: 440M/460M respectively > Projector: 2-layer MLP for both modalities > Language model: Phi-4-mini 3.3B parameters > LoRA adapters for Vision/Audio decoder, applied on top of Phi-4-mini
010
Xuan Son Nguyen @ngxson.hf.co · 02/03/2025
This weekend, I found a very fun project from the community: TinyLM 🤏 Built around transformers.js 🚀 , this project aims to provide to the developer a straight-forward API to work with LLM And most important, it can run inference on-browser using WebGPU or wasm 🚀 No server is needed!
121
Xuan Son Nguyen @ngxson.hf.co · 27/02/2025
What is GGUF, Safetensors, PyTorch, ONNX? In this blog post, let's discover common formats for storing an AI model. huggingface.co/blog/ngxson/...
huggingface.co
Common AI Model Formats
A Blog post by Xuan-Son Nguyen on Hugging Face
053
Xuan Son Nguyen @ngxson.hf.co · 26/02/2025
My AI Podcast generator space is featured on top spaces of the week 🔥 🔥 Try it now --> huggingface.co/spaces/ngxso...
100
Xuan Son Nguyen @ngxson.hf.co · 22/02/2025
The new Data Studio is a fun way to play with datasets. Very useful feature for me! Kudos to @hf.co team ❤️
020
Xuan Son Nguyen @ngxson.hf.co · 20/02/2025
Nice explanation! I learned something today 😍
020
Xuan Son Nguyen @ngxson.hf.co · 30/01/2025
Interesting, DeepSeek mitigates DDOS attack by doing a challenge-response on-browser, asking it to calculate SHA3 of a random string using wasm module. Image: (1) the challenge (2) extract from wasm byte code (3) the handler code in javascript
142
Xuan Son Nguyen @ngxson.hf.co · 29/01/2025
We release yet another config to deploy DeepSeek-R1 on HF inference endpoints! It may looks expensive, but you get 32K context length and a bigger, better quality quantization. Thanks @unsloth.bsky.social for providing the IQ2_XXS dynamic quant!
110
Reposted by Xuan Son Nguyen
Adina Yakup @adinayakup.bsky.social · 27/01/2025
🔥So many exciting releases coming from the Chinese community this month! huggingface.co/collections/...
huggingface.co
2025 January - a zh-ai-community Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
063
Xuan Son Nguyen @ngxson.hf.co · 10/01/2025
Meet MiniThinky, the ultimate AI powerhouse! With only 1 billion parameters, MiniThinky delivers rapid insights and accurate solutions. Perfect for tackling complex problems swiftly! Huge thanks to @xenova.bsky.social who made this awesome WebGPU demo. Try it here: huggingface.co/spaces/webml...
040
Xuan Son Nguyen @ngxson.hf.co · 08/01/2025
Can 1B model **think** 🤔 🤔 ? Check this out --> ollama run hf(.)co/ngxson/MiniThinky-v2-1B-Llama-3.2-Q8_0-GGUF
462
Xuan Son Nguyen @ngxson.hf.co · 06/12/2024
Very useful code suggestion
010
Xuan Son Nguyen @ngxson.hf.co · 02/12/2024
The new "chat with your database" feature on @hf.co is a game-changer for me! 🎉 I can now simply ask it to write queries like "show me top N users having ..."
050
Reposted by Xuan Son Nguyen
𝚖𝚘𝚘𝚍𝚋𝚘𝚊𝚛𝚍. @moodboard.bsky.social · 27/11/2024
#moodboard
236107291003
Xuan Son Nguyen @ngxson.hf.co · 27/11/2024
Hugging Face inference endpoints now support CPU deployment for llama.cpp 🚀 🚀 Why this is a huge deal? Llama.cpp is well-known for running very well on CPU. If you're running small models like Llama 1B or embedding models, this will definitely save tons of money 💰 💰
3266
Xuan Son Nguyen @ngxson.hf.co · 26/11/2024
It's drama time! Someone leaked a well-known text-to-video API and blaming a well-known company for unfair wages.
000