Sign in

Paul Chang

@mummitrollet.bsky.social
109 followers 210 following 70 posts

ML + stuff @Datacrunch

PostsRepliesMedia
Reposted by Paul Chang
Magnus Ross @magnusar.bsky.social · 29/08/2025
I wrote something about building systems people actually want as an academic in ML. It's pretty much an open letter to 6-months-ago me. magnusross.github.io/posts/moms-m...
magnusross.github.io
Moms, Models and Medicine | Magnus Ross
A good friend of mine is deep in the world of startups and spends a lot of his time doing idea validation—that is, trying to understand if there is a market for a given idea or product. Despite the fact that in the startup world success is eventually judged by sales or profits, whereas in ML4H it is more likely to be adoption by clinicians and, hopefully, an associated improvement in clinical outcomes, both rely on designing something that people actually want and will use. Therefore, I think many of the tools that help entrepreneurs validate ideas can be repurposed to help researchers undertake projects with real impact.
122
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 25/06/2025
❗️ We just expanded our capacity of B200 SXM6 180GB servers – available in the DataCrunch Cloud Platform. The best thing is… You can deploy the Blackwell platform without approvals. Just sign in, select the instance type, and start your deployment: cloud.datacrunch.io?utm_source=b...
001
Paul Chang @mummitrollet.bsky.social · 30/05/2025
A new paper just dropped from Tri Dao(🐐)'s lab! arxiv.org/abs/2505.21487 Here is my hot take!
arxiv.org
Hardware-Efficient Attention for Fast Decoding
LLM decoding is bottlenecked for large batches and long contexts by loading the key-value (KV) cache from high-bandwidth memory, which inflates per-token latency, while the sequential nature of decodi...
111
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 29/05/2025
🆕 Inference API for FLUX.1 Kontext [max] & [pro] are now available on DataCrunch! We are an infrastructure partner of Black Forest Labs for Kontext, a suite of generative flow matching models for text-to-image and image-to-image editing. Learn more: datacrunch.io/managed-endp...
101
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 26/05/2025
🚨 Summer Inference by Symposium AI is happening next Wednesday, June 4, at 16:00-22:00. 🇫🇮 This event will bring together 250 AI engineers, researchers, and founders under one roof in Helsinki. 🔗 You can still grab one of the last remaining seats: lu.ma/x5hhj79x
lu.ma
Symposium AI - Summer Inference · Luma
Join 250 leading AI builders for an epic night in Helsinki! Symposium AI events bring together top AI talent, researchers, and engineers who are actively…
011
Paul Chang @mummitrollet.bsky.social · 09/05/2025
Algorithm hardware co-design was a big reason the whale 🐋(DeepSeek) made such a splash 💦 with its V3 and R1 releases.
120
Reposted by Paul Chang
Ayush Bharti @ayushbharti.bsky.social · 02/05/2025
"Cost-aware simulation-based inference" is accepted at AISTATS 2025. Check out our poster #205 on Sunday May 4th in Hall A-E if you are in Phuket. Finland's rising star @huangdaolang.bsky.social will be there to assist you :D arxiv.org/abs/2410.07930 @fxbriol.bsky.social @samikaski.bsky.social
arxiv.org
Cost-aware simulation-based inference
Simulation-based inference (SBI) is the preferred framework for estimating parameters of intractable models in science and engineering. A significant challenge in this context is the large computation...
2185
Reposted by Paul Chang
Ethan Mollick @emollick.bsky.social · 26/04/2025
I don’t mean to be a broken record but AI development could stop at the o3/Gemini 2.5 level and we would have a decade of major changes across entire professions & industries (medicine, law, education, coding…) as we figure out how to actually use it & adapt our systems. AI disruption is baked in.
1322621
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 27/04/2025
1/ If you are at ICLR / AABI / AISTATS, check out work from our lab and collaborators on *inference everywhere anytime all at once*! Go talk to my incredible PhD students @huangdaolang.bsky.social & @chengkunli.bsky.social + amazing collaborator Severi Rissanen. @univhelsinkics.bsky.social FCAI
1225
Reposted by Paul Chang
Naomi Saphra @nsaphra.bsky.social · 26/04/2025
I wrote something up for AI people who want to get into bluesky and either couldn't assemble an exciting feed or gave up doomscrolling when their Following feed switched to talking politics 24/7.
nsaphra.net
The AI Researcher's Guide to a Non-Boring Bluesky Feed | Naomi Saphra
How to migrate to bsky without a boring feed.
2335994
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 22/04/2025
1/10🔥 New paper alert in #AABI2025 Proceedings! Normalizing Flow Regression (NFR) — an offline Bayesian inference method. What if you could get a full posterior using *only* the evaluations you *already* have, maybe from optimization runs?
1236
Reposted by Paul Chang
Frank Schneider @fsschneider.bsky.social · 16/04/2025
Tired of your open-source ML work not getting the academic recognition it deserves? 🤔 Submit to the first-ever CodeML workshop at #ICML2025! It focuses on new libraries, improvements to established ones, best practices, retrospectives, and more. codeml-workshop.github.io/codeml2025/
codeml-workshop.github.io
CODEML Workshop
Championing Open-source Development in Machine Learning.
0346
Reposted by Paul Chang
Sarah Drasner @sarahedo.bsky.social · 13/04/2025
This is a great list, things that “the best engineers I know” do, stuff like: - understanding things deeply, reading the actual source - being willing to help other people - status doesn’t matter, good ideas come from anywhere endler.dev/2025/best-pr...
endler.dev
The Best Programmers I Know | Matthias Endler
I have met a lot of developers in my life. Late…
726040
Paul Chang @mummitrollet.bsky.social · 09/04/2025
B200 go brrrr! It seems by doubling the TFLOPs you get double the speed. Cool stuff by Antonio and WavespeedAI team to get FLUX-dev inference (SOTA diffusion) in under a second on a B200. datacrunch.io/blog/flux-on...
datacrunch.io
FLUX on B200: Real-Time Image Inference with WaveSpeedAI + DataCrunch Collaboration
How WaveSpeedAI and DataCrunch achieved an up to 6x faster image inference by optimizing FLUX-dev's latency and efficiency: NVIDIA B200 vs. H100 benchmark.
021
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 08/04/2025
1/ We asked GPT-4.5 -- allegedly the model with the best sense of humor, according to the site we do not talk about here -- to write a comic about our recent AISTATS paper on the Amortized Conditioning Engine (ACE). Then gpt-4o drew it. You judge the result... (text continues 👇)
A comic "Bayes explains everything!"
2175
Reposted by Paul Chang
Ethan Mollick @emollick.bsky.social · 08/04/2025
The Llama 4 model that won in LM Arena is different than the released version. I have been comparing the answers from Arena to the released model. They aren't close. The data is worth a look also as it shows how LM Arena results can be manipulated to be more pleasing to humans. t.co/rqAey9SMwh
05111
Paul Chang @mummitrollet.bsky.social · 08/04/2025
I wanted to change the color scheme on a blog for some plots so I decided to test Claude code. github.com/datacrunch-r... I didnt relaize it inserts "Co-Authored-By: Claude <noreply@anthropic.com>" I would have got away with it if it wasn't for that pesky Claude code.
github.com
Update plot-script.py to use 2025 brand colors · datacrunch-research/blogs@6955de0
- Added brand colors 2025 palette - Updated all plots to use the new color scheme - Regenerated all plot images with the new colors - Set consistent style across all plots 🤖 Generated with [Claude...
030
Paul Chang @mummitrollet.bsky.social · 08/04/2025
This is a new blog looking at the individual optimizations that went into serving the DeepSeek model class in SGLang. I have been observing the SGLang repo for a few months now, and it's crazy how quickly they integrate new optimized features. It's a very cool open-source project!
121
Paul Chang @mummitrollet.bsky.social · 06/04/2025
Llama 4 uses both interleaved chunked attention and global (NoPE) attention mechanisms, similar to a recent Cohere paper. It's cool to see innovation in attention layer architectures for the large models, and it showed to the world. arxiv.org/abs/2501.18795.
arxiv.org
Rope to Nope and Back Again: A New Hybrid Attention Strategy
Long-context large language models (LLMs) have achieved remarkable advancements, driven by techniques like Rotary Position Embedding (RoPE) (Su et al., 2023) and its extensions (Chen et al., 2023; Liu...
180
Reposted by Paul Chang
Samuel Vaiter @samuelvaiter.com · 04/04/2025
The Johnson–Lindenstrauss Lemma states that a set of high-dimensional points can be mapped into a much lower-dimensional space while approximately preserving pairwise distances. This is useful for dimensionality reduction, clustering, etc. stanford.edu/class/cs114/...
1225
Paul Chang @mummitrollet.bsky.social · 04/04/2025
Free Ghibli images through the WaveSpeed API pretty cool! wavespeed.ai/blog/posts/2...
wavespeed.ai
WaveSpeedAI Docs
Ultimate APl for Accelerating Al Image and Video Generation
100
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 31/03/2025
🚨 NVIDIA HGX B200: available NOW on DataCrunch! Be among the first to gain instant access to 1x, 2x, 4x, and 8x B200 GPUs with our high-performance VMs. Sign up and enjoy expert support with secure service where performance meets sustainability. 🔗 cloud.datacrunch.io
012
Reposted by Paul Chang
Magnus Ross @magnusar.bsky.social · 31/03/2025
I just came across this paper from ICLR 2024 which proposes an intricate combination of transformers and diffusion models to generate forecasts with these uncertainty bounds (red box), which are clearly inappropriate and could likely be outperformed by modelling the data as a random walk...
A grid of plots of time series forecasts, the proposed model's forecasts are highlighted by a red box. The uncertainty estimates for the proposed model are constant over the forecast horizon and are poorly calibrated.
162
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 28/03/2025
🥉 SemiAnalysis awarded DataCrunch with bronze on the GPU Cloud ClusterMAX™ Rating! We thank their team for this independent evaluation, validating our approach to pushing the boundary of resource-efficient AI infrastructure ⬇️
101
Paul Chang @mummitrollet.bsky.social · 28/03/2025
Pretty exciting idea about enforcing sparsity by changing the activation function used. arxiv.org/abs/2503.16672
arxiv.org
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
In this paper, we demonstrate how to leverage 2:4 sparsity, a popular hardware-accelerated GPU sparsity pattern, to activations to accelerate large language model training and inference. Crucially we ...
120
Reposted by Paul Chang
Jacob Springer @jacobspringer.bsky.social · 26/03/2025
Training with more data = better LLMs, right? 🚨 False! Scaling language models by adding more pre-training data can decrease your performance after post-training! Introducing "catastrophic overtraining." 🥁🧵👇 arxiv.org/abs/2503.19206 1/10
13314
Paul Chang @mummitrollet.bsky.social · 26/03/2025
This looks a cool workshop and great lineup! The structured foundation models, feels like part 2 of time series foundation models lots of cool problems to solve.
020
Reposted by Paul Chang
Nick Erickson @nickerickson.bsky.social · 25/03/2025
We are excited to announce #FMSD: "1st Workshop on Foundation Models for Structured Data" has been accepted to #ICML 2025! Call for Papers: icml-structured-fm-workshop.github.io
01510
Paul Chang @mummitrollet.bsky.social · 26/03/2025
I've been diving into the "black magic" world of CUDA recently. More posts may follow, but I think we're at an interesting point. Perhaps the "CUDA moat" is under pressure and perhaps changing how we interact with GPU programming. 🧵
152
Paul Chang @mummitrollet.bsky.social · 26/03/2025
Even though I know this is who I am talking to, I still can't help saying thank you when they help me out 😆.
030
Paul Chang @mummitrollet.bsky.social · 25/03/2025
@lacerbi.bsky.social wrote a very nice summary post of our paper here if anyone missed it: bsky.app/profile/lace... I can give some more behind-the-scenes information. 🧵
172
Paul Chang @mummitrollet.bsky.social · 24/03/2025
medium.com/@NoamShazeer... I came across this post from one of the attention GOATs (thnx @conorhassan.bsky.social for term). I know there are more complicated solutions to this but I find this quite a simple solution to implement to keep track of shapes.
medium.com
Shape Suffixes — Good Coding Style
If you code neural networks, I believe that this convention can make your life more pleasant. We keep this pretty religiously at…
011
Paul Chang @mummitrollet.bsky.social · 24/03/2025
Speaking with @trappmartin.bsky.social, I realised there is no good space online to discuss ML-related stuff. It feels like a lot of people are on Bsky, but not much discussion I am myself guilty. So I am going to try and be more active here.
4170
Reposted by Paul Chang
Marcus Klasson @marcusklasson.bsky.social · 17/03/2025
Submission deadline is extended to March 20 for submitting your paper to our #CVPR2025 workshop on Uncertainty Quantification for Computer Vision. Looking forward to see your submissions on recognizing failure scenarios and enabling robust vision systems! More info: uncertainty-cv.github.io/2025/
uncertainty-cv.github.io
UNCV Workshop @ CVPR 2025
CVPR 2025 Workshop on Uncertainty Quantification for Computer Vision.
0105
Paul Chang @mummitrollet.bsky.social · 12/03/2025
Multi-Head Latent Attention vs Group Query Attention: We break down why MLA is a more expressive memory compression technique AND why naive implementations can backfire. Check it out!
02710
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 06/03/2025
1/ Introducing ACE (Amortized Conditioning Engine)! Our new AISTATS 2025 paper presents a transformer framework that unifies tasks from image completion to BayesOpt & simulator-based inference under *one* probabilistic conditioning approach. It's Bayes all the way down!
13514
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 05/03/2025
👋 We're attending NVIDIA GTC and we'd love to meet you in the SF Bay Area. 📨 We invite you to our series of meetups for AI developers throughout the week: March 16-22. ⬇️ We'll announce all of them over the coming days. Here are the first two:
111
Reposted by Paul Chang
datacrunch.io @datacrunch.io · 07/02/2025
(1/5) We'd like to congratulate Prime Intellect on the release of SYNTHETIC-1, the largest open-source dataset of verified reasoning traces for math, coding, and science. More on the release of SYNTHETIC-1: www.primeintellect.ai/blog/synthet...
primeintellect.ai
SYNTHETIC-1: Scaling Distributed Synthetic Data Generation for Verified Reasoning
Today, we are excited to introduce SYNTHETIC-1, a collaborative effort to create the largest open-source dataset of verified reasoning traces for math, coding and science, leveraging DeepSeek-R1. Our ...
111
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 14/01/2025
1/📯I am hiring! Postdoc position in Probabilistic Machine Learning and Amortized Inference. @univhelsinkics.bsky.social with strong links to the Finnish Center for AI (FCAI) Please see blurb and link in thread below. Applications evaluated on a rolling basis. Please reshare!
12714
Reposted by Paul Chang
The Data Therapist in the Blue Sky @datatherapist.bsky.social · 22/12/2024
If you are a #PhD student or a #postdoc in #AI / #NLP #nlproc you might want to read this thoughtful blog post
1112
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 05/12/2024
1/ Did they fool you into thinking that variational inference is just "the lesser way" of doing Bayesian inference? Please let me argue otherwise. lacerbi.github.io/blog/2024/vi... Or just come to play with the interactive variational inference widget that took me way too long to code up.
44910
Reposted by Paul Chang
Daolang Huang @huangdaolang.bsky.social · 05/12/2024
Optimizing decision utility in Bayesian experimental design is key to improving downstream decision-making. Excited to share our #NeurIPS2024 paper on Amortized Decision-Aware Bayesian Experimental Design: arxiv.org/abs/2411.02064 @lacerbi.bsky.social @samikaski.bsky.social Details below.
13912
Reposted by Paul Chang
Luigi Acerbi @lacerbi.bsky.social · 04/12/2024
1/ Excuse me, can I interest you in eliciting your beliefs as flexible probability distributions? No worries, we only need pairwise comparisons or rankings, no personal details. Led by **Petrus Mikkola** and joint with **Arto Klami**, to be presented soon at @neuripsconf.bsky.social #NeurIPS2024
26914