Sign in

Paul Chang

@mummitrollet.bsky.social
109 followers 210 following 70 posts

ML + stuff @Datacrunch

PostsRepliesMedia
Reposted by Paul Chang
Magnus Ross @magnusar.bsky.social · 29/08/2025
I wrote something about building systems people actually want as an academic in ML. It's pretty much an open letter to 6-months-ago me. magnusross.github.io/posts/moms-m...
magnusross.github.io
Moms, Models and Medicine | Magnus Ross
A good friend of mine is deep in the world of startups and spends a lot of his time doing idea validation—that is, trying to understand if there is a market for a given idea or product. Despite the fact that in the startup world success is eventually judged by sales or profits, whereas in ML4H it is more likely to be adoption by clinicians and, hopefully, an associated improvement in clinical outcomes, both rely on designing something that people actually want and will use. Therefore, I think many of the tools that help entrepreneurs validate ideas can be repurposed to help researchers undertake projects with real impact.
122
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 25/06/2025
❗️ We just expanded our capacity of B200 SXM6 180GB servers – available in the DataCrunch Cloud Platform. The best thing is… You can deploy the Blackwell platform without approvals. Just sign in, select the instance type, and start your deployment: cloud.datacrunch.io?utm_source=b...
001
Paul Chang @mummitrollet.bsky.social · 30/05/2025
Also pretty cool to see open source community building on top of each other!
010
Paul Chang @mummitrollet.bsky.social · 30/05/2025
The paper also suggests Group Tied Attention (GTA), which works in the opposite direction and draws inspiration from MLA, incorporating those techniques into GQA.
100
Paul Chang @mummitrollet.bsky.social · 30/05/2025
The technique called Grouped Latent Attention (GLA) can now be split across devices according to the group, providing higher throughput without a drop in performance by maintaining high arithmetic intensity and achieving better parallelism.
100
Paul Chang @mummitrollet.bsky.social · 30/05/2025
Well, the paper suggests a hybrid method. What about using MLA and adding groups?
100
Paul Chang @mummitrollet.bsky.social · 30/05/2025
Instead, one must make a copy of the latent component across GPUs, which feels wasteful.
100
Paul Chang @mummitrollet.bsky.social · 30/05/2025
This is where MLA is somewhat awkward, and GQA scores some points back. MLA uses a single large latent head that must be replicated across all tensor-parallel GPUs, which means that sharding the attention computations across GPUs cannot be done.
100
Paul Chang @mummitrollet.bsky.social · 30/05/2025
First of all, a confession! In the blog titled 'Multi-Head Latent Attention: Benefits in Memory and Computation', we didn't tell the whole story—the benchmarking on a single GPU. In reality, for DeepSeek V3-style models, parallelization is needed.
100
Paul Chang @mummitrollet.bsky.social · 30/05/2025
The paper focuses on designing more effective decoding attention for inference in light of Multi-head Latent Attention (MLA) and Group Query Attention (GQA).
100
Paul Chang @mummitrollet.bsky.social · 30/05/2025
A new paper just dropped from Tri Dao(🐐)'s lab! arxiv.org/abs/2505.21487 Here is my hot take!
arxiv.org
Hardware-Efficient Attention for Fast Decoding
LLM decoding is bottlenecked for large batches and long contexts by loading the key-value (KV) cache from high-bandwidth memory, which inflates per-token latency, while the sequential nature of decodi...
111
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 29/05/2025
🆕 Inference API for FLUX.1 Kontext [max] & [pro] are now available on DataCrunch! We are an infrastructure partner of Black Forest Labs for Kontext, a suite of generative flow matching models for text-to-image and image-to-image editing. Learn more: datacrunch.io/managed-endp...
101
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 26/05/2025
🚨 Summer Inference by Symposium AI is happening next Wednesday, June 4, at 16:00-22:00. 🇫🇮 This event will bring together 250 AI engineers, researchers, and founders under one roof in Helsinki. 🔗 You can still grab one of the last remaining seats: lu.ma/x5hhj79x
lu.ma
Symposium AI - Summer Inference · Luma
Join 250 leading AI builders for an epic night in Helsinki! Symposium AI events bring together top AI talent, researchers, and engineers who are actively…
011
Paul Chang @mummitrollet.bsky.social · 09/05/2025
Some links that helped me to understand the roofline model. jax-ml.github.io/scaling-book... kipp.ly/transformer-...
jax-ml.github.io
All About Rooflines | How To Scale Your Model
When we run algorithms on hardware, we're bounded by three things: how fast it can do math (OPs/second), the bandwidth available for moving data around (bytes/second), and the total memory available t...
000
Paul Chang @mummitrollet.bsky.social · 09/05/2025
datacrunch.io/blog/multi-h... The blog post explains these terms and how they relate to algorithm intensity. Let us know if you have any questions or spot errors. #MLSky
datacrunch.io
Multi-Head Latent Attention: Benefits in Memory and Computation
Multi-Head Latent Attention (MLA) vs. Group Query Attention (GQA): Transformer inference optimization in DeepSeek V3 with lower KV cache and higher FLOPs/s.
100
Paul Chang @mummitrollet.bsky.social · 09/05/2025
However, more is at play; revisiting Kipply's infamous Transformer Inference Arithmetic article shows that the MLA mechanism used during inference is now compute-bound 🖥️ and not memory-bound 💾.
100
Paul Chang @mummitrollet.bsky.social · 09/05/2025
Looking at the projections involved in DeepSeeek's attention (MLA) of the KV cache automatically makes one think it means less memory needed in HBM, preventing dreaded out-of-memory errors 👿 .
100
Paul Chang @mummitrollet.bsky.social · 09/05/2025
Algorithm hardware co-design was a big reason the whale 🐋(DeepSeek) made such a splash 💦 with its V3 and R1 releases.
120
Reposted by Paul Chang
Ayush Bharti @ayushbharti.bsky.social · 02/05/2025
"Cost-aware simulation-based inference" is accepted at AISTATS 2025. Check out our poster #205 on Sunday May 4th in Hall A-E if you are in Phuket. Finland's rising star @huangdaolang.bsky.social will be there to assist you :D arxiv.org/abs/2410.07930 @fxbriol.bsky.social @samikaski.bsky.social
arxiv.org
Cost-aware simulation-based inference
Simulation-based inference (SBI) is the preferred framework for estimating parameters of intractable models in science and engineering. A significant challenge in this context is the large computation...
2185
Paul Chang @mummitrollet.bsky.social · 27/04/2025
This is very true! Go and speak to people in more old-school businesses and you quickly realize that with current models you could already do so much.
040
Reposted by Paul Chang
Ethan Mollick @emollick.bsky.social · 26/04/2025
I don’t mean to be a broken record but AI development could stop at the o3/Gemini 2.5 level and we would have a decade of major changes across entire professions & industries (medicine, law, education, coding…) as we figure out how to actually use it & adapt our systems. AI disruption is baked in.
1322621
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 27/04/2025
1/ If you are at ICLR / AABI / AISTATS, check out work from our lab and collaborators on *inference everywhere anytime all at once*! Go talk to my incredible PhD students @huangdaolang.bsky.social & @chengkunli.bsky.social + amazing collaborator Severi Rissanen. @univhelsinkics.bsky.social FCAI
1225
Reposted by Paul Chang
Naomi Saphra @nsaphra.bsky.social · 26/04/2025
I wrote something up for AI people who want to get into bluesky and either couldn't assemble an exciting feed or gave up doomscrolling when their Following feed switched to talking politics 24/7.
nsaphra.net
The AI Researcher's Guide to a Non-Boring Bluesky Feed | Naomi Saphra
How to migrate to bsky without a boring feed.
2336094
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 22/04/2025
1/10🔥 New paper alert in #AABI2025 Proceedings! Normalizing Flow Regression (NFR) — an offline Bayesian inference method. What if you could get a full posterior using *only* the evaluations you *already* have, maybe from optimization runs?
1236
Paul Chang @mummitrollet.bsky.social · 21/04/2025
@aidanscannell.bsky.social
010
Reposted by Paul Chang
Frank Schneider @fsschneider.bsky.social · 16/04/2025
Tired of your open-source ML work not getting the academic recognition it deserves? 🤔 Submit to the first-ever CodeML workshop at #ICML2025! It focuses on new libraries, improvements to established ones, best practices, retrospectives, and more. codeml-workshop.github.io/codeml2025/
codeml-workshop.github.io
CODEML Workshop
Championing Open-source Development in Machine Learning.
0346
Paul Chang @mummitrollet.bsky.social · 16/04/2025
Average cost for a student is 86,000$ a year just saying 😜
010
Paul Chang @mummitrollet.bsky.social · 16/04/2025
Congrats Pierre!
110
Paul Chang @mummitrollet.bsky.social · 16/04/2025
This is so true!
000
Paul Chang @mummitrollet.bsky.social · 16/04/2025
Sounds fun! I want to hear about it when you are back!
020
Paul Chang @mummitrollet.bsky.social · 15/04/2025
Wow very dystopian where are you heading in Indonesia?
110
Paul Chang @mummitrollet.bsky.social · 14/04/2025
I feel it's worse in some ways big labs can take ideas from academia and not put anything back. At least when it comes to core revenue models.
010
Paul Chang @mummitrollet.bsky.social · 14/04/2025
Yeah i agree papers end up being written to get past reviewers which screws up readability.
110
Paul Chang @mummitrollet.bsky.social · 14/04/2025
I have been doing more webdev stuff but cool it works with research code too.
010
Paul Chang @mummitrollet.bsky.social · 14/04/2025
Yeah its pretty cool the latest capabilities. I have found first getting Claude or Gemini to build a plan step in an md or text file. Then executing on the components of the plan in single steps or giving it to cursor or claude code (if you want to burn some api $).
110
Reposted by Paul Chang
Sarah Drasner @sarahedo.bsky.social · 13/04/2025
This is a great list, things that “the best engineers I know” do, stuff like: - understanding things deeply, reading the actual source - being willing to help other people - status doesn’t matter, good ideas come from anywhere endler.dev/2025/best-pr...
endler.dev
The Best Programmers I Know | Matthias Endler
I have met a lot of developers in my life. Late…
726040
Paul Chang @mummitrollet.bsky.social · 11/04/2025
Great story! We're also working to enable EU's AI sovereignty. We're are an AI cloud based in Finland. We're soon exceeding LUMI's capacity without relying on state funding. Eager to connect about your future stories on EU’s other contributors to AI sovereignty.
010
Paul Chang @mummitrollet.bsky.social · 10/04/2025
They also released these: blog.google/products/goo...
blog.google
Ironwood: The first Google TPU for the age of inference
We’re introducing Ironwood, our seventh-generation Tensor Processing Unit (TPU) designed to power the age of generative AI inference.
000
Paul Chang @mummitrollet.bsky.social · 09/04/2025
B200 go brrrr! It seems by doubling the TFLOPs you get double the speed. Cool stuff by Antonio and WavespeedAI team to get FLUX-dev inference (SOTA diffusion) in under a second on a B200. datacrunch.io/blog/flux-on...
datacrunch.io
FLUX on B200: Real-Time Image Inference with WaveSpeedAI + DataCrunch Collaboration
How WaveSpeedAI and DataCrunch achieved an up to 6x faster image inference by optimizing FLUX-dev's latency and efficiency: NVIDIA B200 vs. H100 benchmark.
021
Paul Chang @mummitrollet.bsky.social · 09/04/2025
Grandmas that troll! Many levels to my misspelt Finnish,
100
Paul Chang @mummitrollet.bsky.social · 09/04/2025
🤣
120
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 08/04/2025
1/ We asked GPT-4.5 -- allegedly the model with the best sense of humor, according to the site we do not talk about here -- to write a comic about our recent AISTATS paper on the Amortized Conditioning Engine (ACE). Then gpt-4o drew it. You judge the result... (text continues 👇)
A comic "Bayes explains everything!"
2175
Reposted by Paul Chang
Ethan Mollick @emollick.bsky.social · 08/04/2025
The Llama 4 model that won in LM Arena is different than the released version. I have been comparing the answers from Arena to the released model. They aren't close. The data is worth a look also as it shows how LM Arena results can be manipulated to be more pleasing to humans. t.co/rqAey9SMwh
05111
Paul Chang @mummitrollet.bsky.social · 08/04/2025
This is really not looking good PR at the moment! Such a contrast with previous releases. I wonder what they do next? www.reddit.com/r/LocalLLaMA...
reddit.com
“Serious issues in Llama 4 training. I Have Submitted My Resignation to GenAI“
000
Paul Chang @mummitrollet.bsky.social · 08/04/2025
I wanted to change the color scheme on a blog for some plots so I decided to test Claude code. github.com/datacrunch-r... I didnt relaize it inserts "Co-Authored-By: Claude <noreply@anthropic.com>" I would have got away with it if it wasn't for that pesky Claude code.
github.com
Update plot-script.py to use 2025 brand colors · datacrunch-research/blogs@6955de0
- Added brand colors 2025 palette - Updated all plots to use the new color scheme - Regenerated all plot images with the new colors - Set consistent style across all plots 🤖 Generated with [Claude...
030
Paul Chang @mummitrollet.bsky.social · 08/04/2025
It reminds me of a story about the Sky cycling team when they were on top. They would take their own mattresses on tour to ensure that every athlete got optimal sleep. They were concerned with squeezing out every percentage point.
010
Paul Chang @mummitrollet.bsky.social · 08/04/2025
In the blog, we look at the impact of each optimization, such as data parallelism versus tensor parallelism or speculative decoding. Essentially, many smaller optimizations can make a big difference in throughput and latency.
100
Paul Chang @mummitrollet.bsky.social · 08/04/2025
This is a new blog looking at the individual optimizations that went into serving the DeepSeek model class in SGLang. I have been observing the SGLang repo for a few months now, and it's crazy how quickly they integrate new optimized features. It's a very cool open-source project!
121
Paul Chang @mummitrollet.bsky.social · 06/04/2025
Hi Nathan - Enjoying your substack! Thanks for putting our great info.
010
Paul Chang @mummitrollet.bsky.social · 06/04/2025
Llama 4 uses both interleaved chunked attention and global (NoPE) attention mechanisms, similar to a recent Cohere paper. It's cool to see innovation in attention layer architectures for the large models, and it showed to the world. arxiv.org/abs/2501.18795.
arxiv.org
Rope to Nope and Back Again: A New Hybrid Attention Strategy
Long-context large language models (LLMs) have achieved remarkable advancements, driven by techniques like Rotary Position Embedding (RoPE) (Su et al., 2023) and its extensions (Chen et al., 2023; Liu...
180