Reposted by Aakash Kumar NainAndreas Kirsch @blackhc.bsky.social · 17/05/2025I want to share my latest (very short) blog post: "Active Learning vs. Data Filtering: Selection vs. Rejection." What is the fundamental difference between active learning and data filtering? Well, obviously, the difference is that: 1/11 14011
Aakash Kumar Nain @ak-nain.bsky.social · 10/03/2025What if you want to control the length of CoT sequences? Can you put a budget constraint at test time for the reasoner models while maintaining performance? This latest paper from CMU addresses these two questions via RL. Here is a summary of LCPO in case you are interested: 100
Aakash Kumar Nain @ak-nain.bsky.social · 14/02/2025Matryoshka Quantization: Another fantastic paper from GDM! MatQuant came out last week. It was a very refreshing read. Here is a summary in case you are interested: 110
Aakash Kumar Nain @ak-nain.bsky.social · 12/02/20251/3 Two years ago, we started a series on Diffusion Models that covered everything related to these models in-depth. We decided to write those tutorials covering intuition and the fundamentals because we could not find any high-quality diffusion tutorials then. 100
Aakash Kumar Nain @ak-nain.bsky.social · 28/01/2025JanusPro is here, the next generation of the Janus model, with a few surprises (even for me!). I liked JanusFlow a lot, but the JanusPro 1B is what caught my eye. Here is a summary of the paper in case you are interested: 130
Aakash Kumar Nain @ak-nain.bsky.social · 21/01/2025I read the R1 paper last night, and here is a summary cum highlights from the paper (technical report to be more precise) 130
Aakash Kumar Nain @ak-nain.bsky.social · 20/01/2025Everyone has heard enough about the scaling inference-time compute for LLMs in the past month. Diffusion models, on the other hand, have an innate flexibility for allocating varied compute at inference time. Here is a summary of how researchers at GDM exploit this property: 👇 120
Aakash Kumar Nain @ak-nain.bsky.social · 27/12/2024I just finished reading the DeepSeekv3 paper. Here is everything you need to know about it: 👇 x.com/A_K_Nain/sta... 061
Aakash Kumar Nain @ak-nain.bsky.social · 20/12/2024I just finished reading one of the latest papers from Meta Research, MetaMorph. Except for two things (both not good), it is an okay paper, simple, concise, and to the point. Here is a quick summary in case you are interested: x.com/A_K_Nain/sta... 020
Reposted by Aakash Kumar NainHernan Moraldo @hmoraldo.bsky.social · 16/12/2024Proud to see the release of Veo V2! deepmind.google/technologies... "Veo has achieved state of the art results in head-to-head comparisons of outputs by human raters over top video generation models"deepmind.googleVeo 2Veo is our state-of-the-art video generation model. It creates high quality video clips that match the style and content of a user's prompts, in resolutions up to 4K resolution. 2114
Aakash Kumar Nain @ak-nain.bsky.social · 16/12/2024What if I tell you you can train a SOTA Gaze estimation model in 1 hour on an RTX4090 GPU? Too good to be true? I was also skeptical of that claim made in the Gaze-LLE paper, but it is true. DINOv2 FTW! I finished reading the paper, and here is a summary : x.com/A_K_Nain/sta...x.comx.com 070
Aakash Kumar Nain @ak-nain.bsky.social · 13/12/2024Can you pre-train and fine-tune your VLMs in FP8? Can you get more than 2x efficiency with some simple tricks? Nvidia presents NVILA, an efficient frontier VLM that achieves all of the above. I finished reading the paper, and here is a summary in case you are interested: 110
Aakash Kumar Nain @ak-nain.bsky.social · 11/12/2024I am back to writing math-heavy yet intuitive blog posts. Almost two years ago, I wrote the diffusion tutorials with a similar intention. This time, I am targeting the fundamental concepts of LLMs and MLLMs. And here is the first post in that direction: Rotary Position Encodings. Enjoy reading! 🍻 2130
Aakash Kumar Nain @ak-nain.bsky.social · 09/12/20241/2 Google DeepMind announced PaliGemma 2 last week. It is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models. What does this generation of PaliGemma bring to the table? I finished reading the technical report, and here is a summary: 110
Aakash Kumar Nain @ak-nain.bsky.social · 08/12/2024Gemini 2.0 (if we are calling it 2.0 now) will be an interesting development. Why? It will be a good indicator of "Do we need test-time compute for now, or is there more left to juice out the transformers with some neat tricks?" 100
Aakash Kumar Nain @ak-nain.bsky.social · 07/12/2024Though the TTT used by the winners in ARC Prize 2024 definitely gave a huge boost to the performance and is a promising direction, I personally feel that in a few years we would have a solid model that will do all the tricks in a single forward pass. And it won't be a LLM 010
Aakash Kumar Nain @ak-nain.bsky.social · 04/12/2024Launch day! 💥💥 venturebeat.com/ai/emergence...venturebeat.comEmergence’s AI orchestrator launches to do what big tech offerings can’t: play well with othersIt aims to sit above the fray and work well with any application and vendor that the enterprise uses, uniting them all with its orchestrator. 010
Aakash Kumar Nain @ak-nain.bsky.social · 02/12/2024Nvidia presents Star Attention to improve LLM inference efficiency over long sequences. I was skeptical when I read the abstract the day it was published, but now that I have read the full paper, I think this is another good research x.com/A_K_Nain/sta...x.comx.com 0191
Aakash Kumar Nain @ak-nain.bsky.social · 27/11/2024Okay I like the idea of this app, but TBH this platform needs to step up to become what we need it to be. As of now: 1. Laggy 2. Half the time the tabs don't work 3. Feed is still broken 4. No bookmarks yet 5. Hyperlinks work randomly 210
Aakash Kumar Nain @ak-nain.bsky.social · 27/11/2024The multimodality space is now evolving in a much better way. The focus has shifted to finding the bottlenecks and fixing things on the fundamental level. This paper from **Apple** introduces **AIMv2**, and effort is in a similar direction, except that they only do it for the autoregressive models. 152
Aakash Kumar Nain @ak-nain.bsky.social · 26/11/2024Shameless plug but this is all you need to understand the fundamental of diffusion models: magic-with-latents.github.io/latent/posts...magic-with-latents.github.ioThe Latent: Code the Maths - A deep dive into DDPMs 071
Aakash Kumar Nain @ak-nain.bsky.social · 26/11/2024Generative World Explorer from John Hopkins University: an egocentric world exploration framework that allows an agent to mentally explore a large-scale 3D world arxiv.org/abs/2411.11844arxiv.orgGenerative World ExplorerPlanning with partial observation is a central challenge in embodied AI. A majority of prior works have tackled this challenge by developing agents that physically explore their environment to update ... 020
Aakash Kumar Nain @ak-nain.bsky.social · 26/11/2024Nvidia presents Hymba, another hybrid of attention and SSMs but for small family models: arxiv.org/abs/2411.13676arxiv.orgHymba: A Hybrid-head Architecture for Small Language ModelsWe propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates transformer attention mechanisms with state space models (SSMs) for enhanced efficienc... 031
Aakash Kumar Nain @ak-nain.bsky.social · 26/11/2024Efficient long video tokenization arxiv.org/abs/2411.14762arxiv.orgEfficient Long Video Tokenization via Coordinated-based Patch ReconstructionEfficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop a tokenizer that can encode long video clips, as it w... 030
Aakash Kumar Nain @ak-nain.bsky.social · 25/11/2024Ignoring the many missing features, the one thing that is a true delight on this app is to see your feed full of papers and ML discussions. 👌👌 050
Aakash Kumar Nain @ak-nain.bsky.social · 25/11/20241/2 We all have been impressed by the quality of models produced by Deepseek. I thought Qwen was good, but the main highlight is JanusFlow. Apart from the MM1 paper from Apple, I believe JanusFlow is one of the best papers on modern MLLMs. 170
Aakash Kumar Nain @ak-nain.bsky.social · 24/11/2024Except for the documentation, I think manim is freaking awesome! Spent half a day, and I am in delight playing with it. This is so cool! 010
Reposted by Aakash Kumar NainOpinion Editor @ Bluesky @dly.bsky.social · 24/11/2024the most problematic thing to come out of all these new bsky users is that a lot of y'all seem to think anything on here matters 892074228
Aakash Kumar Nain @ak-nain.bsky.social · 24/11/2024Well it feels stupid to ask to be added to a starter pack (FOMO much? 😂), but please add me to a few good ones please 110
Aakash Kumar Nain @ak-nain.bsky.social · 24/11/2024The drama over Anthropic's new post was not about stats vs research. It was always about the click-bait "New Anthropic Research" which is neutral but tragic in so many ways 010
Aakash Kumar Nain @ak-nain.bsky.social · 23/11/2024Got manim working finally. It is time to write a kickass blogpost now 020
Aakash Kumar Nain @ak-nain.bsky.social · 22/11/2024One of the goals is to create a small SOTA VLM, no matter what it takes 040
Reposted by Aakash Kumar Nainsimjeg.bsky.social @simjeg.bsky.social · 19/11/2024🚀 Excited to announce KVPress — our open-source library for efficient LLM KV cache compression! 👉 Check it out (and drop a ⭐): github.com/NVIDIA/kvpress 🔗 Full details in the thread 🧵 (1/4) 2496
Aakash Kumar Nain @ak-nain.bsky.social · 21/11/2024I think this space is actually picking up this time. Though this reminds me of good old twitter, there are certain good features on X that would be good to have on this platform 130
Reposted by Aakash Kumar NainGowthami Somepalli @gowthami.bsky.social · 21/11/2024Started a list of some researchers working on image/video generation. (Not comprehensive at all) Reply with a paper link and TLDR to get added to the list! I request all grad students to not feel imposter-y and just reply if you work in this field! #computervision #diffusion go.bsky.app/SP1uWoE 13348
Aakash Kumar Nain @ak-nain.bsky.social · 20/11/2024We all want to save GPU memory and increase its utilization, right? This latest paper from Apple is exactly what we all need (apart from attention, of course!). As we scale models, the vocabulary for these models will also grow (and we want to be in an ideal scenario). 120
Aakash Kumar Nain @ak-nain.bsky.social · 19/11/2024The first blog post of this season is already in progress. Wasted a full day to work with Manim, but matplotlib the saviour. Going to make such a solid post (like the diffusion series) that you would never have to search more on that topic from other sources 🍻 000
Aakash Kumar Nain @ak-nain.bsky.social · 19/11/2024I think the Gemini team needs to really address a few things. The quality of outputs of Claude is far superior to Gemini in all of my coding tests. In fact, I didn't expect Gemini to write the wrong code for cross-entropy. Also, I didn't expect it to interpret transpose wrongly 000
Aakash Kumar Nain @ak-nain.bsky.social · 19/11/2024The AI community is growing. We should either make people aware about this niche community or we should make lists! 🙌 000
Aakash Kumar Nain @ak-nain.bsky.social · 19/11/2024Why are people skipping Qwen again and again from their benchmarks? If your model doesn't outperform it, so what? You can always keep it for future research 🤷 000
Reposted by Aakash Kumar NainPatrick Kidger @patrickkidger.bsky.social · 18/11/2024Good software is an enabler for good science! 💥🧪 Inspired by the below post, I like to point people at libraries like github.com/patrick-kidg... as a template for what a modern Python library looks like: `pre-commit`, ruff, pyright, pyproject.toml, an open-source license, etc. 🤓github.comGitHub - patrick-kidger/equinox: Elegant easy-to-use neural networks + scientific computing in JAX. https://docs.kidger.site/equinox/Elegant easy-to-use neural networks + scientific computing in JAX. https://docs.kidger.site/equinox/ - patrick-kidger/equinox 68612
Aakash Kumar Nain @ak-nain.bsky.social · 18/11/2024The only thing that is responsible for Delhi choking because of the smog is the lack of accountability. Guess what, when you have a huge population, the value of lives decreases exponentially 110
Aakash Kumar Nain @ak-nain.bsky.social · 18/11/2024JAX is not about comparison with other frameworks. Core JAX is the only non-bloated and elegant framework in every aspect. 000
Aakash Kumar Nain @ak-nain.bsky.social · 17/11/2024Finding all my followers/following here now....help! 😂😂 000
Aakash Kumar Nain @ak-nain.bsky.social · 17/11/2024Will this grow or will we witness another "threads" moment here? Time will tell! PS: Why does this website take over a minute to load? I have never seen such lag in a log time 000
Aakash Kumar Nain @ak-nain.bsky.social · 16/11/2024I don't have I have ever cursed any python package as I did today while trying to install and use manin(community edition) for an upcoming blog post. Coolest visualizations, worst software 000
Aakash Kumar Nain @ak-nain.bsky.social · 15/11/2024Pretty sure it will be 2x hard to get the same followers/following here compared to the bird 😂 000
Aakash Kumar Nain @ak-nain.bsky.social · 09/11/2024Nothing is perfect. I bought a YT subscription, and this is one of the best things I did in terms of spending money on subscriptions, but man can we please remove the "shorts" tab? Please? 😭 100