Jia-Bin Huang @jbhuang0604.bsky.social · 02/09/2026Speculative Decoding is the coolest trick for speeding up LLM inference! Check out the video and learn why rejection sampling preserves quality, and how methods such as draft trees, Medusa, MTP, EAGLE, and DFlash further accelerate LLM inference. youtu.be/l8gWQlrVOKQ 161
Jia-Bin Huang @jbhuang0604.bsky.social · 24/08/2026Two API Calls Exposed AI's Hidden Reasoning Such a cool attack exploits a vulnerability in frontier models. Just replay encrypted reasoning traces through weaker models and recover the hidden reasoning! 🤯 Many fun reasoning examples! New video: youtu.be/P1v1-2CCKD0youtu.beTwo API Calls Exposed AI's Hidden ReasoningYouTube video by Jia-Bin Huang 270
Jia-Bin Huang @jbhuang0604.bsky.social · 17/07/2026How Small Models Learn to Think Like Giants On-policy distillation has emerged as a powerful way to transfer knowledge between models. BUT, how do they work? 🤔 In this video, let's explore knowledge distillation from first principles. youtu.be/YH0YXgDWZXAyoutu.beHow Small Models Learn to Think Like GiantsYouTube video by Jia-Bin Huang 0224
Jia-Bin Huang @jbhuang0604.bsky.social · 25/06/2026LoRA, low-rank adaptation, is arguably the most popular parameter-efficient fine-tuning method for LLMs. But how does it actually work? Check out the video to learn LoRA and friends (LoRA+, QLoRA, VeRA, and DoRA)! youtu.be/U80tjcThl9Q 091
Jia-Bin Huang @jbhuang0604.bsky.social · 15/06/2026I didn't know what JEPA is, and at this point I am too afraid to ask ...😬 so I made a video tracing the last 30+ years of self-supervised learning, covering ideas from contrastive learning, distillation, masked modeling, JEPA, and world models. youtu.be/gVEr2cnDE_8 1272
Jia-Bin Huang @jbhuang0604.bsky.social · 18/05/2026The 60-Year Hunt for AI's Most Important Function I was trying to understand how SwiGLU works, but I couldn’t find an explanation that clicked for me. So I made this video to explain it from first principles. Check it out: youtu.be/JRaPNrpsQ9s 071
Jia-Bin Huang @jbhuang0604.bsky.social · 17/05/2026**Modern Transformer - Complete Guide** Interested in learning the recent advances in transformers? After 14 videos, I've finally completed this series! 🥳🥳🥳 Check out the course here: www.youtube.com/playlist?lis... 0325
Jia-Bin Huang @jbhuang0604.bsky.social · 06/05/2026The Most Underrated Layer Inside Every AI Model Virtually every AI model has normalization layers. BUT, what makes them so essential? 🤔 New video on learning the role of normalization in stabilizing training and alternatives like DyT and Derf. youtu.be/JHl_gwVoh-k 0304
Jia-Bin Huang @jbhuang0604.bsky.social · 30/04/2026How is DeepSeek V4 so INSANELY cheap? 🤔 Compared to a GQA baseline, it's new *compressed attention* mechanism (CSA and HCA) slashes the KV cache memory cost by 98% 🤯 at a 1M-token context! Here’s how: youtu.be/q8holiIirgo 040
Jia-Bin Huang @jbhuang0604.bsky.social · 14/04/2026How do we make attention actually capture context? Exclusive Self Attention (XSA) is an interesting variant that improves attention with minimal cost in speed & memory. Check out the video here: youtu.be/2eZKT4H9_iQ 0164
Jia-Bin Huang @jbhuang0604.bsky.social · 10/04/2026**Modern Transformer architecture explained** I compiled a list of videos on the Transformer architecture into a short "YouTube course". www.youtube.com/playlist?lis... Hopefully, this would be helpful for beginners in the community. Happy learning! 😎 0388
Jia-Bin Huang @jbhuang0604.bsky.social · 07/04/2026Finally got some time to read the DeepSeek Engram paper! Idea: Replace repeated reconstruction with direct lookup of common knowledge. It’s so intuitive that it feels strange this wasn’t part of the design from the start. Video summary here: youtu.be/87Q8nf1XHKA 060
Jia-Bin Huang @jbhuang0604.bsky.social · 09/02/2026New video! How do LLMs grow outrageously large yet blazingly fast? The secret: Mixture of Experts (MoE) In this video, we cover the role of FFNs, how to scale them without slowing down, and how to maintain load balance and training stability. Full video here: youtu.be/0QQlYR1r6pQyoutu.beHow LLMs Get Outrageously Large Yet Blazingly Fast [MoE]YouTube video by Jia-Bin Huang 0140
Jia-Bin Huang @jbhuang0604.bsky.social · 20/01/2026Beyond softmax attention Linear attention and its variants enable faster inference without growing the KV cache. Let’s learn the core ideas behind efficient sequence modeling. youtu.be/pUCWwGR5WmQyoutu.beBeyond Softmax: The Future of Attention MechanismsYouTube video by Jia-Bin Huang 1290
Jia-Bin Huang @jbhuang0604.bsky.social · 12/01/2026Residual connections are adopted in virtually every deep learning model. BUT, can we further improve it? Hyper-connections is an exciting recent exploration to generalize residual connections. Check out the video explaining manifold constraints hyper-connections! youtu.be/jYn_1PpRzxIyoutu.beHow Residual Connections Are Getting an Upgrade [mHC]YouTube video by Jia-Bin Huang 1171
Jia-Bin Huang @jbhuang0604.bsky.social · 05/12/2025Introducing *Edit-by-Track*, a framework that enables precise video motion editing via 3D point tracks. Our method supports a wide range of motion editing applications, including object removal, shape deformation, dynamic view synthesis, and many more! Video explainer: youtu.be/aAj6PIgx20o 0101
Jia-Bin Huang @jbhuang0604.bsky.social · 01/12/2025Wondering how DeepSeek v3.2 rivals SOTA models (e.g., GPT5/Gemini 3 pro) while being ~30x cheaper? 🤔 Let's learn how the base model works! We'll focus on attention, the need for KV caching, and key ideas for improving attention (MQA/GQA/MLA/DSA). youtu.be/Y-o545eYjXM 0122
Jia-Bin Huang @jbhuang0604.bsky.social · 21/11/2025Sharing the slides for a talk on faculty job search Hope it's helpful to people exploring and preparing for the process. Feedback is welcome! www.dropbox.com/scl/fi/p7xdt... 0112
Jia-Bin Huang @jbhuang0604.bsky.social · 29/10/2025How to organize your talk? I used to present like this, thinking that I was being "academic", "organized", and "professional". BUT, from the audience's viewpoints, this sucks. 😱 Look how far they need to hold a long-term context to just make sense of what you're saying! 1131
Jia-Bin Huang @jbhuang0604.bsky.social · 24/10/2025Muon is a (relatively) new optimizer that powered large-scale training of recent foundation models, e.g., Kimi K2 and GLM 4.5. Interested in learning how it works? Check out the video here: youtu.be/bO5nvE289ecyoutu.beThis Simple Optimizer Is Revolutionizing How We Train AI [Muon]YouTube video by Jia-Bin Huang 1395
Jia-Bin Huang @jbhuang0604.bsky.social · 17/09/2025How AI Taught Itself to See Self-supervised learning is fascinating! How can AI learn from images only without labels? In this video, we’ll build the method from first principles and uncover the key ideas behind CLIP, MAE, SimCLR, and DINO (v1–v3). Video link: youtu.be/oGTasd3cliMyoutu.beHow AI Taught Itself to See [DINOv3]YouTube video by Jia-Bin Huang 0123
Jia-Bin Huang @jbhuang0604.bsky.social · 15/08/2025New video! A quick dive into the recent Hierarchical Reasoning Model (HRM) through the lens of algorithm synthesis. Check it out: youtu.be/RK7lysjz_G0youtu.beThe Weirdly Small AI That Cracks Reasoning Puzzles [HRM]YouTube video by Jia-Bin Huang 0151
Jia-Bin Huang @jbhuang0604.bsky.social · 08/08/2025Diffusion LLMs are promising ways to overcome the limitations of autoregressive LLMs. Less error propagation, easier to control, and faster to sample! But how do Diffusion LLMs actually work? 🤔 In this video, let's explore some ideas on this fascinating topic! youtu.be/8BTOoc0yDVA 0111
Jia-Bin Huang @jbhuang0604.bsky.social · 16/07/2025In an era of billion-parameter models everywhere, it's incredibly refreshing to see how a fundamental question can be formulated and solved with simple, beautiful math. - How should we orient a solar panel ☀️🔋? - Zero AI! If you enjoy math, you'll love this! Video: www.youtube.com/watch?v=ZKzL... 182
Jia-Bin Huang @jbhuang0604.bsky.social · 08/07/2025Why is the "Title and Content" slide layout BAD? Most people prepare their presentation from this default layout. I used it for years without questioning it. BUT, this essentially guides you toward developing poor presentation. Why? 🤔 5222
Jia-Bin Huang @jbhuang0604.bsky.social · 01/07/2025Kids’ summer camp just kicked off, and that means... I finally have time to make new videos! What topics are you most interested in right now? 150
Jia-Bin Huang @jbhuang0604.bsky.social · 24/06/2025Why More Researchers Should be Content Creators Just trying something new! I recorded one of my recent talks, sharing what I learned from starting as a small content creator. youtu.be/0W_7tJtGcMI We all benefit when there are more content creators! 182
Reposted by Jia-Bin HuangKosta Derpanis @csprofkgd.bsky.social · 19/06/2025Fresh out of the oven! 🍞 @jbhuang0604.bsky.social breaks down Mean Flow from Kaiming’s group in his latest video. Video: youtu.be/swKdn-qT47Q?... 0182
Jia-Bin Huang @jbhuang0604.bsky.social · 21/06/2025Policy gradient methods rock! These are the core techniques for making your transformer "chat" and "reason", a robot that manipulates objects, and a drone that maneuvers in a complex environment. BUT, how do we learn all the developments in the past 30+ years? 131
Jia-Bin Huang @jbhuang0604.bsky.social · 20/06/2025Awesome! 🤩 So glad to hear the authors enjoyed the video, totally made my day! 1130
Jia-Bin Huang @jbhuang0604.bsky.social · 17/06/2025We had a blast at CVPR2025! There was so much to learn! I am particularly excited to meet many new friends and reconnect with old ones. I feel energized. Already looking forward to the next one! 060
Jia-Bin Huang @jbhuang0604.bsky.social · 04/06/2025Kullback–Leibler (KL) divergence is a cornerstone of machine learning. We use it everywhere, from training classifiers and distilling knowledge from models, to learning generative models and aligning LLMs. BUT, what does it mean, and how do we (actually) compute it? Video: youtu.be/tXE23653JrU 2315
Jia-Bin Huang @jbhuang0604.bsky.social · 03/06/2025My X/Twitter account has been hacked... Please don't believe what they said! Trying to get it back in the meantime. Sorry for the inconvenience! 050
Jia-Bin Huang @jbhuang0604.bsky.social · 21/05/2025RL is so back! Reinforcement learning is a key driver in aligning LLMs and enhancing their reasoning capabilities. BUT, it’s a tricky topic to wrap your head around (at least for myself 😵💫). So, I put up a video breaking down the basics in a way that clicked for me. I hope it helps you, too! 170
Jia-Bin Huang @jbhuang0604.bsky.social · 19/05/2025I find TRPO's idea of learning from others' experiences fascinating. So, I started running TRPO for my group, making all (previously individual) feedback on experiments, writing, rebuttals, and presentations public. Now everyone gets to learn from each other’s trajectories! 020
Jia-Bin Huang @jbhuang0604.bsky.social · 14/05/2025Exploration is key for robots to generalize, especially in open-ended environments with vague goals and sparse rewards. BUT, how do we go beyond random poking? Wouldn't it be great to have a robot that explores an environment just like a kid? Introducing Imagine, Verify, Execute (IVE)! 2102
Jia-Bin Huang @jbhuang0604.bsky.social · 26/04/2025Solving high-impact real-world problems with multimodal foundation models 020
Jia-Bin Huang @jbhuang0604.bsky.social · 15/03/2025Check out UrbanIR - Inverse rendering of unbounded scenes from a single video! It’s a super cool project led by the amazing Chih-Hao! @chih-hao.bsky.social is a rising star in 3DV! Follow him! Learn more here👇 0102
Jia-Bin Huang @jbhuang0604.bsky.social · 10/03/2025Interesting! I didn't realize how important a video title/packaging is until now. It's the same video, but with a better packaging it gets much more attention. 2150
Jia-Bin Huang @jbhuang0604.bsky.social · 09/03/2025How a 40-Year-Old Trick Solves Seamless Image Blending Laplacian pyramid blending is a simple yet effective tool for many applications, including object composition, seamless panorama stitching, and exposure fusion. Let’s learn this classic method that still works so well today. 1436
Jia-Bin Huang @jbhuang0604.bsky.social · 04/03/2025Fifth year grad students to incoming ones at the prospective student visit day: 0130
Jia-Bin Huang @jbhuang0604.bsky.social · 03/03/2025How to schedule your thesis defense? So you think publishing top-tier papers is hard? Wait until you need to schedule your prelim/defense! Some common mistakes and tips: 1152
Jia-Bin Huang @jbhuang0604.bsky.social · 28/02/2025Amid all the super-duper research advances, I am excited to share my video on a super-duper basic topic: How to sample your signals? Here we discuss sampling, why aliasing occurs, and how to prevent it. It's fascinating to view this through the lens of frequency analysis. youtu.be/fTJjPGaPsq4 13510
Jia-Bin Huang @jbhuang0604.bsky.social · 26/02/2025computer vision researchers’ conversations with friends be like 0282
Jia-Bin Huang @jbhuang0604.bsky.social · 22/02/2025The real challenge of a PhD defense is to get five academics to agree that they will show up in the same room at the same time. 0290
Jia-Bin Huang @jbhuang0604.bsky.social · 20/02/2025How does Template Matching Work? Template matching is a crucial technique for detecting and locating objects in images. It is widely applied in object recognition, medical imaging, and quality inspection. Let's build the algorithm and use it to find Waldo! Full video: youtu.be/EO1-MCWfXCUyoutu.beHow does Template Matching Work?YouTube video by Jia-Bin Huang 1152
Jia-Bin Huang @jbhuang0604.bsky.social · 14/02/2025Yay! Expressive Text-to-Image Generation with Rich Text now has a journal version (in IJCV 2025)! rich-text-to-image.github.io 1121
Jia-Bin Huang @jbhuang0604.bsky.social · 06/02/2025How Blur Filters Work? Blur filters reduce noise, extract image structures across scales, and create beautiful shallow depth-of-field effects. But HOW and WHY do they work? Let's find out together! Full video: youtu.be/xvDeFABiwrg 0141