Sign in

Duc Nguyen Huu

@ducnh279.bsky.social
45 followers 19 following 57 posts

Data Science in ♥️ Home in 🇻🇳 Kaggle Competitions Master 🥇 1 Solo Gold 🥈 2 Silvers (1 Solo, 1 Team) 🌍 Ranked 272 / 202K globally (Top 0.14%)

PostsRepliesMedia
Reposted by Duc Nguyen Huu
Gaël Varoquaux @gaelvaroquaux.bsky.social · 26/05/2026
In this high-level yet detailed blog post, I expend on the statistics that hold together AI and data science: blog.probabl.ai/data-science...
blog.probabl.ai
Data science is not AI, but it is its genesis and its new frontier
Gaël Varoquaux explains why and how statistical thinking is the heart of AI.
1136
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 13/03/2026
My book is the #1 New Release in NLP! 🥳 Amazon US put it on sale... for $0.95 off 😂 Get the paperback: geni.us/MasterML Or read online (free!): mlbook.dataschool.io #MachineLearning #Python @scikit-learn.org
064
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 27/02/2026
My new book - on sale NEXT WEEK! 🎉 Sign up to get notified when it's available: dataschool.kit.com/mlbook #MachineLearning #Python @scikit-learn.org
031
Duc Nguyen Huu @ducnh279.bsky.social · 20/02/2026
When I first started learning data science, I often got lost in the @scikit-learn.org documentation. After taking the course that became this book, I understood it much better and gained confidence using it. Scikit-learn is a powerful library,but without guidance, it can feel overwhelming at first.
110
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 11/09/2025
Dream unlocked: I'm publishing my first book! 🎉🎉🎉 It's called "Master Machine Learning with scikit-learn: A Practical Guide to Building Better Models with Python" Download the first 3 chapters right now: 👉 dataschool.kit.com/mlbook 👈 Thanks for your support 🙏
1266
Reposted by Duc Nguyen Huu
David Holzmüller @dholzmueller.bsky.social · 29/07/2025
I got 3rd out of 691 in a tabular kaggle competition – with only neural networks! 🥉 My solution is short (48 LOC) and relatively general-purpose – I used skrub to preprocess string and date columns, and pytabkit to create an ensemble of RealMLP and TabM models. Link below👇
2112
Reposted by Duc Nguyen Huu
Gaël Varoquaux @gaelvaroquaux.bsky.social · 09/07/2025
This work is presented at ICML next week. • The paper arxiv.org/html/2502.05... • The python package: pypistats.org/packages/tab... (try it out 🐍) • The source code github.com/soda-inria/t... (100% open source, including pre-training 💞) Longer read (5mn): gael-varoquaux.info/science/tabi... 8/9
arxiv.org
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
0122
Reposted by Duc Nguyen Huu
Gaël Varoquaux @gaelvaroquaux.bsky.social · 09/07/2025
👨‍🎓🧾✨#icml2025 Paper: TabICL, A Tabular Foundation Model for In-Context Learning on Large Data With Jingang Qu, @dholzmueller.bsky.social, and Marine Le Morvan TL;DR: a well-designed architecture and pretraining gives best tabular learner, and more scalable On top, it's 100% open source 1/9
15115
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 28/05/2025
My thoughts on the current state of AI progress and the most important developments in 2025: www.dataschool.io/ai-progress-...
dataschool.io
AI progress in 2025 📈
Thoughts on the current state of AI progress and the most important developments in 2025
011
Duc Nguyen Huu @ducnh279.bsky.social · 20/04/2025
000
Duc Nguyen Huu @ducnh279.bsky.social · 05/04/2025
Are you familiar with Token Pooling? Models that use late interaction, like ColBERT, ColPali, and ColQwen, gain significant benefits from this pooling technique! By integrating token pooling methods, the number of vectors to store can be reduced. Blog: www.answer.ai/posts/colber...
answer.ai
A little pooling goes a long way for multi-vector representations – Answer.AI
Practical AI R&D
000
Duc Nguyen Huu @ducnh279.bsky.social · 04/04/2025
Efficiently scale long CoT models like DeepSeek when using Best-of-N or Majority Voting by early pruning reasoning chains. Kaggle Discussion: www.kaggle.com/competitions...
kaggle.com
AI Mathematical Olympiad - Progress Prize 2
Solve national-level math challenges using artificial intelligence models
000
Duc Nguyen Huu @ducnh279.bsky.social · 31/03/2025
I find making your agents safe is just as important as making them smart. 🔒 A good read for building secure AI! arxiv.org/pdf/2503.18813
010
Duc Nguyen Huu @ducnh279.bsky.social · 30/03/2025
There will be one day ... in 🇺🇸 or 🇻🇳
210
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 27/03/2025
Claude finally integrated web search into its results... But with LangChain & LangGraph, you can build a chatbot that integrates web search into ANY model you like! You'll learn how to do that (and much more) in my new AI course... Sign up for EARLY ACCESS: 👉 dataschool.kit.com/agents 👈
022
Duc Nguyen Huu @ducnh279.bsky.social · 24/03/2025
A practical way for students to secure jobs and earn money is by developing real-world projects. Researching or engineering LLMs often seems like a field dominated by the big tech! It's still important to learn fundamentals from scratch for growth and problem-solving (e.g be able to fix things)! 😁
020
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 23/03/2025
My next tutorial on pretraining an LLM from scratch is now out. It starts with a step-by-step walkthrough of understanding, calculating, and optimizing the loss. After training, we update the text generation function with temperature scaling and top-k sampling: www.youtube.com/watch?v=Zar2...
06112
Reposted by Duc Nguyen Huu
Rashmi Banthia @rashmib.bsky.social · 20/03/2025
cuDF-pandas (%load_ext cudf.pandas) with Rapids ... work similarly and super cool to see we will be able to speed up scikit-learn
111
Duc Nguyen Huu @ducnh279.bsky.social · 20/03/2025
Scikit-learn accelerated 🚀 My company has a bunch of unused T4 GPUs because the LLMs are too big for AI teams run exps. Now the data science team finally has a reason to ask for them! 🤣 developer.nvidia.com/blog/nvidia-...
developer.nvidia.com
NVIDIA cuML Brings Zero Code Change Acceleration to scikit-learn | NVIDIA Technical Blog
Scikit-learn, the most widely used ML library, is popular for processing tabular data because of its simple API, diversity of algorithms, and compatibility with popular Python libraries such as pandas...
120
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 17/03/2025
In honor of March Madness 🏀, I've got a new blog post: www.dataschool.io/pandas-strea... Learn how to identify & analyze scoring streaks using pandas operations: - shift() - cumsum() - boolean math - groupby()
dataschool.io
How to calculate "scoring streaks" with pandas 🏀
Learn how to identify & analyze consecutive events in your data using advanced DataFrame methods!
011
Duc Nguyen Huu @ducnh279.bsky.social · 18/03/2025
Many good advices/best practices for missing value imputation in the paper! I now have a much deeper appreciation for Data School's course and regard it as the best scikit-learn course. Master Machine Learning with scikit-learn: courses.dataschool.io/master-machi...
courses.dataschool.io
Course: Master Machine Learning with scikit-learn
Get ready for your dream job in Machine Learning!
121
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 13/03/2025
"Some people today are discouraging others from learning programming on the grounds AI will automate it. This advice will be seen as some of the worst career advice ever given." -- Andrew Ng, legendary AI researcher Source: www.deeplearning.ai/the-batch/is...
deeplearning.ai
DeepSeek-R1 Uncensored, QwQ-32B Puts Reasoning in Smaller Model, and more...
The Batch AI News and Insights: Some people today are discouraging others from learning programming on the grounds AI will automate it.
011
Reposted by Duc Nguyen Huu
Gaël Varoquaux @gaelvaroquaux.bsky.social · 14/03/2025
A recent talk, fully in a vscode: 100% code on data wrangling for machine learning with @skrub-data.bsky.social www.youtube.com/watch?v=hdWW... super powerful to easily assemble production-ready pipelines in easy syntax
youtube.com
The Future of AI & Machine Learning | The Python Exchange February 2025
YouTube video by Don't Use This Code • James Powell
2235
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 13/03/2025
Yesterday, Google released Gemma 3, their latest open-weight LLM. Finally, a new addition to the "Big 5" of open-weight models (Gemma, Llama, DeepSeek, Qwen, and Mistral). I just went through the Gemma 3 report and experimented a bit with the models, and there are plenty of interesting tidbits:
35811
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 11/03/2025
Just uploaded my "Coding Attention Mechanisms" tutorial. A 2h15m session on coding attention mechanisms to understand how the engine of LLMs works: self-attention → parameterized self-attention → causal self-attention → multi-head self-attention www.youtube.com/watch?v=-Ll8...
youtube.com
Build an LLM from Scratch 3: Coding attention mechanisms
YouTube video by Sebastian Raschka
0364
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 08/03/2025
I just shared a new article, "The State of Reasoning Models", where I am exploring 12 new research articles on improving the reasoning capabilities of LLMs (all published after the release of DeepSeek R1): magazine.sebastianraschka.com/p/state-of-l... Happy reading!
magazine.sebastianraschka.com
The State of LLM Reasoning Models
Part 1: Inference-Time Compute Scaling Methods
16114
Reposted by Duc Nguyen Huu
Trey Hunner @trey.io · 03/03/2025
A couple months ago @dataschool.io wrote about a tool he uses to chat with different LLM models without paying a monthly subscription to all of them. The tool is called Typing Mind and I decided to pay $30 for lifetime access. It was well worth it. Kevin's post 👇 www.dataschool.io/save-money-o...
dataschool.io
Use premium AI models for pennies 💰
Learn how to access ChatGPT, Claude, and more for pennies per conversation rather than paying for expensive subscriptions!
442
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 04/03/2025
19 professionals (in a variety of fields) evaluated OpenAI's Deep Research vs Google's Deep Research. OpenAI was the clear winner 🏆 Neat study by @binarybits.bsky.social, read more here: www.understandingai.org/p/these-expe...
043
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 02/03/2025
A new tutorial in my “Build A Large Language Model From Scratch” series is now live (www.youtube.com/watch?v=341R...) - Tokenizing raw text and converting tokens into token IDs - Applying byte pair encoding - Setting up data loaders in PyTorch for efficient training
youtube.com
Build an LLM from Scratch 2: Working with text data
YouTube video by Sebastian Raschka
1446
Reposted by Duc Nguyen Huu
Rashmi Banthia @rashmib.bsky.social · 26/02/2025
Comprehensive !! mlcontests.com/state-of-mac...
mlcontests.com
The State of Machine Learning Competitions | ML Contests
We summarise the state of the ML competitions landscape and analyse the hundreds of competitions that took place in 2024. Plus an overview of winning solutions and commentary on techniques used.
021
Duc Nguyen Huu @ducnh279.bsky.social · 20/02/2025
Another great read on reasoning models! 🧠 Small LMs struggle to learn from long or complex CoTs from larger teachers. 🔍 Why? The reasoning complexity may be too overwhelming. 🚀 Solution: Mix simple & complex CoTs! 📈 Results: Clear gains over complex CoTs alone! Arxiv: arxiv.org/pdf/2502.121...
arxiv.org
000
Duc Nguyen Huu @ducnh279.bsky.social · 19/02/2025
OMGGG!!!!!!!!!!! 100+ pages book to fully understand distributed training. Author: "🎯 Our goal: democratize knowledge about LLM training at scale. Whether you're starting with 1 GPU or orchestrating thousands, this guide walks you through the journey step by step." huggingface.co/spaces/nanot...
huggingface.co
The Ultra-Scale Playbook - a Hugging Face Space by nanotron
The ultimate guide to training LLM on large GPU Clusters
020
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 16/02/2025
Here's my personal guide to Bluesky happiness 🌈 1. Be very selective about who I follow (42 people at the moment) 2. Read everything in the "Following" feed, and nothing else 3. Trust that the people I follow will surface other good people to follow 4. Unfollow anyone who no longer makes the cut
011
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 15/02/2025
It's 2025, and I’ve finally updated my Python setup guide to use uv + venv instead of conda + pip! Here's my go-to recommendation for uv + venv in Python projects for faster installs, better dependency management: github.com/rasbt/LLMs-f... (Any additional suggestions?)
1115920
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 12/02/2025
Wondering about the differences between "Jupyter", "Jupyter Notebook", "JupyterLab", "IPython", "Colab", and other related terms? I'll explain these terms (and more) in 5 minutes: ▶️ www.youtube.com/watch?v=TDlG...
youtube.com
Jupyter & IPython terminology explained
YouTube video by Data School
021
Duc Nguyen Huu @ducnh279.bsky.social · 12/02/2025
Reading this paper, you feel the power of RL. RL is great, but not all we need. A blend of RL, inference time scaling, good pretraining, and domain-specific tuning is often the way to go! arxiv.org/pdf/2502.06807
arxiv.org
000
Duc Nguyen Huu @ducnh279.bsky.social · 12/02/2025
LLM Trick - Finding the Near Optimal Temperature for Multi-Sample Inference (e.g., Majority Vote, Best-of-N) - Notebook (implement paper): 🤖https://github.com/ducnh279/Optimize-Temperature-for-LLM-Inference/blob/main/optimize-temperature-for-llm-inference.ipynb
110
Duc Nguyen Huu @ducnh279.bsky.social · 11/02/2025
When I first started learning about LLMs, I used this cheat sheet (community.openai.com/t/cheat-shee...) to choose sampling "temperature." Now, I’ve found another useful technique to add to my bag of tricks: arxiv.org/pdf/2502.05234.
bit.ly
Cheat Sheet: Mastering Temperature and Top_p in ChatGPT API
Hello everyone! Ok, I admit had help from OpenAi with this. But what I “helped” put together I think can greatly improve the results and costs of using OpenAi within your apps and plugins, specially ...
030
Duc Nguyen Huu @ducnh279.bsky.social · 11/02/2025
Have you ever wondered what are the differences between 2 types of knowledge distillation (for LLMs) these days? Link: developer.nvidia.com/blog/how-to-...
010
Duc Nguyen Huu @ducnh279.bsky.social · 10/02/2025
My learning project: Implement a Reasoning Model with GRPO from scratch Code: github.com/ducnh279/grp... - Base model: Qwen2.5-1.5B - Dataset: GSM8K (math) - Reward functions: Format, Accuracy - Objective: GRPO - PEFT: LoRA
150
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 09/02/2025
Maybe a hot take, but what about the following advice to the next gen: Don't get an AI degree; the curriculum will be outdated before you graduate. Instead, study math, stats, or physics as your foundation, and stay current with AI through code-focused books, blogs, and papers.
1214722
Duc Nguyen Huu @ducnh279.bsky.social · 08/02/2025
𝐂𝐨𝐬𝐢𝐧𝐞 𝐫𝐞𝐰𝐚𝐫𝐝 stabilizes training & incentivizes the model to generate responses that are both accurate and reasonably lengthy. To prevent reward hacking, a repetition penalty is added to the cosine reward. Paper: arxiv.org/pdf/2502.03373 Hackable Code: github.com/eddycmu/demy...
010
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 07/02/2025
Comparing: - Software Developers - ML Engineers - LLM Developers - Prompt Engineers Read here: www.louisbouchard.ai/llm-develope... Image & article by Louis-François Bouchard
011
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 07/02/2025
Just read the s1: Simple Test-Time Scaling paper. Super interesting approach to improving reasoning models! TL;DR: 1. SFT on 1k curated examples w/ reasoning traces. 2. Control response length w/ budget forcing: "Wait" tokens → longer reasoning/self-correction. "Final Answer:" → enforce stopping.
2376
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 05/02/2025
I just finished writing up my take on reasoning models: magazine.sebastianraschka.com/p/understand... Here, I 1. Discuss the advantages & disadvantages of reasoning models 2. Of course, describe and discuss DeepSeek R1 3. Describe the 4 main ways to building & improving reasoning models
magazine.sebastianraschka.com
Understanding Reasoning LLMs
Methods and Strategies for Building and Refining Reasoning Models
39020
Duc Nguyen Huu @ducnh279.bsky.social · 05/02/2025
The best way to understand GRPO is to code it from scratch without HF Trainer. 😂
100
Duc Nguyen Huu @ducnh279.bsky.social · 05/02/2025
Want to start with RL basics for training reasoning LLMs using GRPO? Check out: superb-makemake-3a4.notion.site/group-relati...
superb-makemake-3a4.notion.site
group relative policy optimization (GRPO) | Notion
intro
010
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 31/01/2025
This blog post is now available in video form: www.youtube.com/watch?v=rOKX... (Hopefully I did a good job pronouncing Andrew Ng & Sebastian Raschka 🤷‍♂️)
youtube.com
How to keep up with AI in 2025
YouTube video by Data School
021
Reposted by Duc Nguyen Huu
Kevin Markham @dataschool.io · 30/01/2025
Want to keep up with the latest developments in AI? Check out the 7 newsletters I personally read & recommend for AI news, insights, and analysis: 👉 www.dataschool.io/best-ai-news... 👈 Includes: @simonwillison.net @emollick.bsky.social @sebastianraschka.com @binarybits.bsky.social
dataschool.io
How to keep up with AI in 2025 🏃‍♂️
These 7 AI experts will guide you through the most important developments in Artificial Intelligence.
3143
Reposted by Duc Nguyen Huu
Sebastian Raschka (rasbt) @rasbt.bsky.social · 24/01/2025
But before I get to the reasoning model space... if you are looking to do some focused offline reading this weekend, I just re-compiled my take on the "noteworthy AI research papers of 2024" into one PDF-export-friendly 47-page mega-post with TOC and all: sebastianraschka.com/blog/2025/ll...
sebastianraschka.com
Noteworthy LLM Research Papers of 2024
This article covers 12 influential AI research papers of 2024, ranging from mixture-of-experts models to new LLM scaling laws for precision.
56113