Sign in

Aakash Kumar Nain

@ak-nain.bsky.social
940 followers 133 following 141 posts

Sr. ML Engineer | Keras 3 Collaborator | @GoogleDevExpert in Machine Learning | @TensorFlow addons maintainer l ML is all I do | Views are my own!

PostsRepliesMedia
Reposted by Aakash Kumar Nain
Andreas Kirsch @blackhc.bsky.social · 17/05/2025
I want to share my latest (very short) blog post: "Active Learning vs. Data Filtering: Selection vs. Rejection." What is the fundamental difference between active learning and data filtering? Well, obviously, the difference is that: 1/11
14011
Aakash Kumar Nain @ak-nain.bsky.social · 10/03/2025
What if you want to control the length of CoT sequences? Can you put a budget constraint at test time for the reasoner models while maintaining performance? This latest paper from CMU addresses these two questions via RL. Here is a summary of LCPO in case you are interested:
100
Aakash Kumar Nain @ak-nain.bsky.social · 14/02/2025
Matryoshka Quantization: Another fantastic paper from GDM! MatQuant came out last week. It was a very refreshing read. Here is a summary in case you are interested:
https://x.com/A_K_Nain/status/1890226873332092997
110
Aakash Kumar Nain @ak-nain.bsky.social · 12/02/2025
1/3 Two years ago, we started a series on Diffusion Models that covered everything related to these models in-depth. We decided to write those tutorials covering intuition and the fundamentals because we could not find any high-quality diffusion tutorials then.
100
Aakash Kumar Nain @ak-nain.bsky.social · 28/01/2025
JanusPro is here, the next generation of the Janus model, with a few surprises (even for me!). I liked JanusFlow a lot, but the JanusPro 1B is what caught my eye. Here is a summary of the paper in case you are interested:
130
Aakash Kumar Nain @ak-nain.bsky.social · 21/01/2025
I read the R1 paper last night, and here is a summary cum highlights from the paper (technical report to be more precise)
130
Aakash Kumar Nain @ak-nain.bsky.social · 20/01/2025
Everyone has heard enough about the scaling inference-time compute for LLMs in the past month. Diffusion models, on the other hand, have an innate flexibility for allocating varied compute at inference time. Here is a summary of how researchers at GDM exploit this property: 👇
120
Aakash Kumar Nain @ak-nain.bsky.social · 27/12/2024
I just finished reading the DeepSeekv3 paper. Here is everything you need to know about it: 👇 x.com/A_K_Nain/sta...
061
Aakash Kumar Nain @ak-nain.bsky.social · 20/12/2024
I just finished reading one of the latest papers from Meta Research, MetaMorph. Except for two things (both not good), it is an okay paper, simple, concise, and to the point. Here is a quick summary in case you are interested: x.com/A_K_Nain/sta...
https://x.com/A_K_Nain/status/1870068712709173645
020
Reposted by Aakash Kumar Nain
Hernan Moraldo @hmoraldo.bsky.social · 16/12/2024
Proud to see the release of Veo V2! deepmind.google/technologies... "Veo has achieved state of the art results in head-to-head comparisons of outputs by human raters over top video generation models"
deepmind.google
Veo 2
Veo is our state-of-the-art video generation model. It creates high quality video clips that match the style and content of a user's prompts, in resolutions up to 4K resolution.
2114
Aakash Kumar Nain @ak-nain.bsky.social · 16/12/2024
What if I tell you you can train a SOTA Gaze estimation model in 1 hour on an RTX4090 GPU? Too good to be true? I was also skeptical of that claim made in the Gaze-LLE paper, but it is true. DINOv2 FTW! I finished reading the paper, and here is a summary : x.com/A_K_Nain/sta...
x.com
x.com
070
Aakash Kumar Nain @ak-nain.bsky.social · 13/12/2024
Can you pre-train and fine-tune your VLMs in FP8? Can you get more than 2x efficiency with some simple tricks? Nvidia presents NVILA, an efficient frontier VLM that achieves all of the above. I finished reading the paper, and here is a summary in case you are interested:
110
Aakash Kumar Nain @ak-nain.bsky.social · 11/12/2024
I am back to writing math-heavy yet intuitive blog posts. Almost two years ago, I wrote the diffusion tutorials with a similar intention. This time, I am targeting the fundamental concepts of LLMs and MLLMs. And here is the first post in that direction: Rotary Position Encodings. Enjoy reading! 🍻
https://aakashkumarnain.github.io/posts/ml_dl_concepts/rope.html
2130
Aakash Kumar Nain @ak-nain.bsky.social · 09/12/2024
1/2 Google DeepMind announced PaliGemma 2 last week. It is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models. What does this generation of PaliGemma bring to the table? I finished reading the technical report, and here is a summary:
110
Aakash Kumar Nain @ak-nain.bsky.social · 08/12/2024
Gemini 2.0 (if we are calling it 2.0 now) will be an interesting development. Why? It will be a good indicator of "Do we need test-time compute for now, or is there more left to juice out the transformers with some neat tricks?"
100
Aakash Kumar Nain @ak-nain.bsky.social · 07/12/2024
Though the TTT used by the winners in ARC Prize 2024 definitely gave a huge boost to the performance and is a promising direction, I personally feel that in a few years we would have a solid model that will do all the tricks in a single forward pass. And it won't be a LLM
010
Aakash Kumar Nain @ak-nain.bsky.social · 04/12/2024
Launch day! 💥💥 venturebeat.com/ai/emergence...
venturebeat.com
Emergence’s AI orchestrator launches to do what big tech offerings can’t: play well with others
It aims to sit above the fray and work well with any application and vendor that the enterprise uses, uniting them all with its orchestrator.
010
Aakash Kumar Nain @ak-nain.bsky.social · 02/12/2024
Nvidia presents Star Attention to improve LLM inference efficiency over long sequences. I was skeptical when I read the abstract the day it was published, but now that I have read the full paper, I think this is another good research x.com/A_K_Nain/sta...
x.com
x.com
0191
Aakash Kumar Nain @ak-nain.bsky.social · 27/11/2024
Okay I like the idea of this app, but TBH this platform needs to step up to become what we need it to be. As of now: 1. Laggy 2. Half the time the tabs don't work 3. Feed is still broken 4. No bookmarks yet 5. Hyperlinks work randomly
210
Aakash Kumar Nain @ak-nain.bsky.social · 27/11/2024
The multimodality space is now evolving in a much better way. The focus has shifted to finding the bottlenecks and fixing things on the fundamental level. This paper from **Apple** introduces **AIMv2**, and effort is in a similar direction, except that they only do it for the autoregressive models.
Summary: https://x.com/A_K_Nain/status/1861598387059167248
152
Aakash Kumar Nain @ak-nain.bsky.social · 26/11/2024
Shameless plug but this is all you need to understand the fundamental of diffusion models: magic-with-latents.github.io/latent/posts...
magic-with-latents.github.io
The Latent: Code the Maths - A deep dive into DDPMs
071
Aakash Kumar Nain @ak-nain.bsky.social · 26/11/2024
Generative World Explorer from John Hopkins University: an egocentric world exploration framework that allows an agent to mentally explore a large-scale 3D world arxiv.org/abs/2411.11844
arxiv.org
Generative World Explorer
Planning with partial observation is a central challenge in embodied AI. A majority of prior works have tackled this challenge by developing agents that physically explore their environment to update ...
020
Aakash Kumar Nain @ak-nain.bsky.social · 26/11/2024
Nvidia presents Hymba, another hybrid of attention and SSMs but for small family models: arxiv.org/abs/2411.13676
arxiv.org
Hymba: A Hybrid-head Architecture for Small Language Models
We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates transformer attention mechanisms with state space models (SSMs) for enhanced efficienc...
031
Aakash Kumar Nain @ak-nain.bsky.social · 26/11/2024
Efficient long video tokenization arxiv.org/abs/2411.14762
arxiv.org
Efficient Long Video Tokenization via Coordinated-based Patch Reconstruction
Efficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop a tokenizer that can encode long video clips, as it w...
030
Aakash Kumar Nain @ak-nain.bsky.social · 25/11/2024
Ignoring the many missing features, the one thing that is a true delight on this app is to see your feed full of papers and ML discussions. 👌👌
050
Aakash Kumar Nain @ak-nain.bsky.social · 25/11/2024
This is 🔥
020
Aakash Kumar Nain @ak-nain.bsky.social · 25/11/2024
1/2 We all have been impressed by the quality of models produced by Deepseek. I thought Qwen was good, but the main highlight is JanusFlow. Apart from the MM1 paper from Apple, I believe JanusFlow is one of the best papers on modern MLLMs.
170
Aakash Kumar Nain @ak-nain.bsky.social · 24/11/2024
Except for the documentation, I think manim is freaking awesome! Spent half a day, and I am in delight playing with it. This is so cool!
010
Reposted by Aakash Kumar Nain
Opinion Editor @ Bluesky @dly.bsky.social · 24/11/2024
the most problematic thing to come out of all these new bsky users is that a lot of y'all seem to think anything on here matters
892074228
Aakash Kumar Nain @ak-nain.bsky.social · 24/11/2024
Well it feels stupid to ask to be added to a starter pack (FOMO much? 😂), but please add me to a few good ones please
110
Aakash Kumar Nain @ak-nain.bsky.social · 24/11/2024
The drama over Anthropic's new post was not about stats vs research. It was always about the click-bait "New Anthropic Research" which is neutral but tragic in so many ways
010
Aakash Kumar Nain @ak-nain.bsky.social · 23/11/2024
Got manim working finally. It is time to write a kickass blogpost now
020
Aakash Kumar Nain @ak-nain.bsky.social · 23/11/2024
The OSS space is picking up!
030
Aakash Kumar Nain @ak-nain.bsky.social · 22/11/2024
One of the goals is to create a small SOTA VLM, no matter what it takes
040
Reposted by Aakash Kumar Nain
simjeg.bsky.social @simjeg.bsky.social · 19/11/2024
🚀 Excited to announce KVPress — our open-source library for efficient LLM KV cache compression! 👉 Check it out (and drop a ⭐): github.com/NVIDIA/kvpress 🔗 Full details in the thread 🧵 (1/4)
2496
Aakash Kumar Nain @ak-nain.bsky.social · 21/11/2024
I think this space is actually picking up this time. Though this reminds me of good old twitter, there are certain good features on X that would be good to have on this platform
130
Reposted by Aakash Kumar Nain
Gowthami Somepalli @gowthami.bsky.social · 21/11/2024
Started a list of some researchers working on image/video generation. (Not comprehensive at all) Reply with a paper link and TLDR to get added to the list! I request all grad students to not feel imposter-y and just reply if you work in this field! #computervision #diffusion go.bsky.app/SP1uWoE
13348
Aakash Kumar Nain @ak-nain.bsky.social · 20/11/2024
We all want to save GPU memory and increase its utilization, right? This latest paper from Apple is exactly what we all need (apart from attention, of course!). As we scale models, the vocabulary for these models will also grow (and we want to be in an ideal scenario).
120
Aakash Kumar Nain @ak-nain.bsky.social · 19/11/2024
The first blog post of this season is already in progress. Wasted a full day to work with Manim, but matplotlib the saviour. Going to make such a solid post (like the diffusion series) that you would never have to search more on that topic from other sources 🍻
000
Aakash Kumar Nain @ak-nain.bsky.social · 19/11/2024
I think the Gemini team needs to really address a few things. The quality of outputs of Claude is far superior to Gemini in all of my coding tests. In fact, I didn't expect Gemini to write the wrong code for cross-entropy. Also, I didn't expect it to interpret transpose wrongly
000
Aakash Kumar Nain @ak-nain.bsky.social · 19/11/2024
The AI community is growing. We should either make people aware about this niche community or we should make lists! 🙌
000
Aakash Kumar Nain @ak-nain.bsky.social · 19/11/2024
Why are people skipping Qwen again and again from their benchmarks? If your model doesn't outperform it, so what? You can always keep it for future research 🤷
000
Reposted by Aakash Kumar Nain
Patrick Kidger @patrickkidger.bsky.social · 18/11/2024
Good software is an enabler for good science! 💥🧪 Inspired by the below post, I like to point people at libraries like github.com/patrick-kidg... as a template for what a modern Python library looks like: `pre-commit`, ruff, pyright, pyproject.toml, an open-source license, etc. 🤓
github.com
GitHub - patrick-kidger/equinox: Elegant easy-to-use neural networks + scientific computing in JAX. https://docs.kidger.site/equinox/
Elegant easy-to-use neural networks + scientific computing in JAX. https://docs.kidger.site/equinox/ - patrick-kidger/equinox
68612
Aakash Kumar Nain @ak-nain.bsky.social · 18/11/2024
The only thing that is responsible for Delhi choking because of the smog is the lack of accountability. Guess what, when you have a huge population, the value of lives decreases exponentially
110
Aakash Kumar Nain @ak-nain.bsky.social · 18/11/2024
JAX is not about comparison with other frameworks. Core JAX is the only non-bloated and elegant framework in every aspect.
000
Aakash Kumar Nain @ak-nain.bsky.social · 17/11/2024
Finding all my followers/following here now....help! 😂😂
000
Aakash Kumar Nain @ak-nain.bsky.social · 17/11/2024
Will this grow or will we witness another "threads" moment here? Time will tell! PS: Why does this website take over a minute to load? I have never seen such lag in a log time
000
Aakash Kumar Nain @ak-nain.bsky.social · 16/11/2024
I don't have I have ever cursed any python package as I did today while trying to install and use manin(community edition) for an upcoming blog post. Coolest visualizations, worst software
000
Aakash Kumar Nain @ak-nain.bsky.social · 15/11/2024
Pretty sure it will be 2x hard to get the same followers/following here compared to the bird 😂
000
Aakash Kumar Nain @ak-nain.bsky.social · 09/11/2024
Nothing is perfect. I bought a YT subscription, and this is one of the best things I did in terms of spending money on subscriptions, but man can we please remove the "shorts" tab? Please? 😭
100