Sign in

Nicolas Beltran-Velez

@velezbeltran.bsky.social
1.8K followers 1K following 84 posts

Machine Learning PhD Student @ Blei Lab & Columbia University. Working on probabilistic ML | uncertainty quantification | LLM interpretability. Excited about everything ML, AI and engineering!

PostsRepliesMedia
Reposted by Nicolas Beltran-Velez
Irving Institute for Cancer Dynamics @cancerdynamics.bsky.social · 21/05/2025
🎓 Hats off to the 2025 IICD graduates: Yining Ma Junze Huang Yichi Yang Ruilin Dai Boan Zhu Cameron Park @jlfan.bsky.social & Achille Nazaret! Wishing you all the best in your next chapter — we’re proud of you! 💙 #Columbia2025 @bleilab.bsky.social @khanhndinh.bsky.social @elhamazizi.bsky.social
071
Reposted by Nicolas Beltran-Velez
kyunghyuncho.bsky.social @kyunghyuncho.bsky.social · 20/05/2025
this is probably not the complete picture of KD, but i can definitely sleep better after writing down and confirming this minimal working explanation. arXiv: arxiv.org/abs/2505.13111 (3/4)
arxiv.org
Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
Knowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits are well documented-...
272
Reposted by Nicolas Beltran-Velez
Freda Shi @fredashi.bsky.social · 28/03/2025
I received a review like this five years ago. It’s probably the right time now to share it with everyone who wrote or got random discouraging reviews from ICML/ACL.
1635
Reposted by Nicolas Beltran-Velez
Nathan Lambert @natolambert.bsky.social · 26/02/2025
First 11 chapters of RLHF Book have v0 draft done. Should be useful now. Next: * Crafting more blog content into future topics, * DPO+ chapter, * Meeting with publishers to get wheels turning on physical copies, * Cleaning & cohesiveness rlhfbook.com
0489
Reposted by Nicolas Beltran-Velez
briantrippe.bsky.social @briantrippe.bsky.social · 19/02/2025
🔥 Benchmark Alert! MotifBench sets a new standard for evaluating protein design methods in motif scaffolding. Why does this matter? Reproducibility & fair comparison have been lacking—until now. Paper: arxiv.org/abs/2502.12479 | Repo: github.com/blt2114/Moti... A thread ⬇️
14117
Reposted by Nicolas Beltran-Velez
Alexander Doria @dorialexander.bsky.social · 19/02/2025
The HuggingFace/Nanotron team just shipped an entire pretraining textbook in interactive format. huggingface.co/spaces/nanot... It’s not just a great pedagogic support, but many unprecedented data and experiments presented for the first time in a systematic way.
0409
Nicolas Beltran-Velez @velezbeltran.bsky.social · 19/02/2025
I just wanted to see what it looked like 😭
030
Nicolas Beltran-Velez @velezbeltran.bsky.social · 17/02/2025
Good God, please. I just want some gradients that don't vanish 😭
140
Reposted by Nicolas Beltran-Velez
Juan Diego Rodriguez @juand-r.bsky.social · 05/02/2025
I was hoping that recent events would lead to a mass exodus from X. Many have left, but most of the ML and LLM people have not. I have lost a lot of respect for the ML community.
9714
Reposted by Nicolas Beltran-Velez
lebellig @lebellig.bsky.social · 05/02/2025
Now that bluesky has gifs (it didn't work?), I can share (again) my educational notebook on discrete flow matching (by Itai Gat et al.). Also please check the original article and official implementation by Meta! 🐍 github.com/gle-bellier/... 🐍 github.com/facebookrese... 📄 arxiv.org/abs/2407.15595
1152
Reposted by Nicolas Beltran-Velez
Christian A. Naesseth @canaesseth.bsky.social · 05/02/2025
Really excited about this! We note a connection between diffusion/flow models and neural/latent SDEs. We show how to use this for simulation-free learning of fully flexible SDEs. We refer to this as SDE Matching and show speed improvements of several orders of magnitude. arxiv.org/abs/2502.02472
arxiv.org
SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations
The Latent Stochastic Differential Equation (SDE) is a powerful tool for time series and sequence modeling. However, training Latent SDEs typically relies on adjoint sensitivity methods, which depend ...
05010
Reposted by Nicolas Beltran-Velez
Ted Underwood @tedunderwood.com · 03/02/2025
I have a sinking feeling that by 2029 I'm going to be faking a British accent so no one will think I was one of the *Americans* working on AI during the regime.
This is a scatterplot with the following key features:

Axes:
The x-axis represents "Interest in AI," with values ranging approximately from -2 to 2.
The y-axis represents "Willingness to Tolerate Closed, Autocratic Systems," also ranging from about -2 to 2.
Data Points:
Black dots dominate the plot, distributed across all four quadrants, indicating diverse positions on both variables.
A few red dots labeled "my peeps" are clustered in the bottom-right quadrant, signifying high interest in AI but low tolerance for closed, autocratic systems.
Blue Lines:
The plot includes horizontal and vertical blue lines at zero, dividing it into four quadrants for visual reference.
This visualization highlights a subset of individuals ("my peeps") who stand out from the majority based on their distinct combination of interest and values.
910910
Nicolas Beltran-Velez @velezbeltran.bsky.social · 03/02/2025
NGL, it's kind of surprising that more people haven't migrated here, especially given what Musk has been doing these days. I don't get it.
010
Reposted by Nicolas Beltran-Velez
Nathan Lambert @natolambert.bsky.social · 01/02/2025
Since everyone wants to learn RL for language models now post DeepSeek, reminder that I've been working on this book quietly in the background for months. Policy gradient chapter is coming together. Plugging away at the book every day now. rlhfbook.com/c/11-policy-...
215519
Reposted by Nicolas Beltran-Velez
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 29/01/2025
Please stop anthropomorphizing language models, it makes them feel really bad
3702
Reposted by Nicolas Beltran-Velez
Ben Collins @bencollins.bsky.social · 29/01/2025
This comments section is the first time I've felt even a shred of hope in eight days.
reddit.com
From the fednews community on Reddit
Explore this post and more from the fednews community
567202933829
Reposted by Nicolas Beltran-Velez
Jeff Dean @jeffdean.bsky.social · 26/01/2025
Nazi salutes and speaking at neo-Nazi rallies seems bad. There's history that we should learn from.
31155
Nicolas Beltran-Velez @velezbeltran.bsky.social · 25/01/2025
Something I really like about NLP research is that it makes everything super intuitive. This week I have been thinking about variational inference in NLP and a lot of the things that seemed to require mathematical intuition just become trivial when thinking about language. So cool:)
020
Reposted by Nicolas Beltran-Velez
Ethan Mollick @emollick.bsky.social · 15/01/2025
New randomized, controlled trial by the World Bank of students using GPT-4 as a tutor in Nigeria. Six weeks of after-school AI tutoring = 2 years of typical learning gains, outperforming 80% of other educational interventions. And it helped all students, especially girls who were initially behind.
1535388
Nicolas Beltran-Velez @velezbeltran.bsky.social · 28/12/2024
Does anyone have any good resources to learn about quantization? Any essential papers to read and resources about how to use/quantize models in practice are greatly appreciated!
030
Reposted by Nicolas Beltran-Velez
Mark Riedl @markriedl.bsky.social · 20/12/2024
1-> 2 -> 3 -> 3.5 -> 4 -> 4o -> o1 -> o3 I guess we need AGI just to figure out how to name things
7716
Reposted by Nicolas Beltran-Velez
Csaba Szepesvari @skiandsolve.bsky.social · 19/12/2024
If you are into ML theory (RL or not) with a proven track record, and you are interested in an industry research position, PM me. Feel free to spread the word.
27531
Reposted by Nicolas Beltran-Velez
Joy Fan @jlfan.bsky.social · 18/12/2024
🧵 Excited to share #Echidna, a Bayesian framework for quantifying the impact of gene dosage on phenotypic plasticity: tinyurl.com/296kf7hf! With @elhamazizi.bsky.social and @mingxz.bsky.social, we integrate scRNA-seq & WGS to uncover how CNAs drive tumor evolution and transcriptional variability.
biorxiv.org
2156
Reposted by Nicolas Beltran-Velez
Elham Azizi @elhamazizi.bsky.social · 18/12/2024
Proud of this work spearheaded by the phenomenal @jlfan.bsky.social and @mingxz.bsky.social in collaboration w/ Ben Izar! The past 3 years we've worked hard to unravel how #CNVs shape #tumor phenotypic plasticity seen in #singlecell #RNAseq data ➡️ #Echidna 🦔
1176
Nicolas Beltran-Velez @velezbeltran.bsky.social · 12/12/2024
Hello! We will be presenting Estimating the Hallucination Rate of Generative AI at NeurIPS. Come if you'd like to chat about epistemic uncertainty for In-Context Learning, or uncertainty more generally. :) Location: East Exhibit Hall A-C #2703 Time: Friday @ 4:30 Paper: arxiv.org/abs/2406.07457
0234
Reposted by Nicolas Beltran-Velez
Yuli Slavutsky @yulislavutsky.bsky.social · 10/12/2024
I'm on my way to #NeurIPS2024. On Friday I'm going to present my latest paper with Yuval Benjamini. The gist is in the comments, and come chat with me to hear more!
174
Reposted by Nicolas Beltran-Velez
claudia shi @claudiashi.bsky.social · 10/12/2024
The circuit hypothesis proposes that LLM capabilities emerge from small subnetworks within the model. But how can we actually test this? 🤔 joint work with @velezbeltran.bsky.social @maggiemakar.bsky.social @anndvision.bsky.social @bleilab.bsky.social Adria @far.ai Achille and Caro
2156
Reposted by Nicolas Beltran-Velez
David Cox @neurobongo.bsky.social · 08/12/2024
I have been a little bit scarce on social media in the last few months. Some of that has just been from being busy at work, but some of it has had to do with my father's passing. He hated the very idea of social media, but he religiously followed my Twitter, and then Bluesky, feeds. /n
5451
Reposted by Nicolas Beltran-Velez
Elham Azizi @elhamazizi.bsky.social · 16/11/2023
We are thrilled to share #Decipher 🔍! Extremely proud of Achille and Joy for leading the development of this creative #ML tool combining VAEs with deep exponential models for comparing disease & healthy trajectories ➕ applying it to study leukemic onset! biorxiv.org/conte…
163
Reposted by Nicolas Beltran-Velez
Andrew Gordon Wilson @andrewgwils.bsky.social · 05/12/2024
I wanted to make my first post about a project close to my heart. Linear algebra is an underappreciated foundation for machine learning. Our new framework CoLA (Compositional Linear Algebra) exploits algebraic structure arising from modelling assumptions for significant computational savings! 1/4
313821
Reposted by Nicolas Beltran-Velez
Keyon Vafa @keyonv.bsky.social · 04/12/2024
Thank you Nature and @anilananth.bsky.social for this great feature on LLMs and AGI (and for highlighting our work arxiv.org/abs/2406.03689)
0104
Nicolas Beltran-Velez @velezbeltran.bsky.social · 04/12/2024
From our lab finally I convinced Yuli Slavutsky (generalization/uncertainty/robustness) @yulislavutsky.bsky.social and Claudia Shi (science of LLMs and AI safety) @claudiashi.bsky.social to use bluesky. Go give them a follow! 😁
080
Reposted by Nicolas Beltran-Velez
Sweta Karlekar @swetakar.bsky.social · 02/12/2024
Very happy to share some recent work by my colleagues @velezbeltran.bsky.social, @aagrande.bsky.social and @anazaret.bsky.social! Check out their work on tree-based diffusion models (especially the website—it’s quite superb 😊)!
1151
Nicolas Beltran-Velez @velezbeltran.bsky.social · 02/12/2024
I am very excited to share our new Neurips 2024 paper + package, Treeffuser! 🌳 We combine gradient-boosted trees with diffusion models for fast, flexible probabilistic predictions and well-calibrated uncertainty. paper: arxiv.org/abs/2406.07658 repo: github.com/blei-lab/tre... 🧵(1/8)
Samples y | x from Treeffuser vs. true densities, for multiple values of x under three different scenarios. Treeffuser captures arbitrarily complex conditional distributions that vary with x.
415223
Reposted by Nicolas Beltran-Velez
ruiqigao.bsky.social @ruiqigao.bsky.social · 02/12/2024
A common question nowadays: Which is better, diffusion or flow matching? 🤔 Our answer: They’re two sides of the same coin. We wrote a blog post to show how diffusion models and Gaussian flow matching are equivalent. That’s great: It means you can use them interchangeably.
625459
Reposted by Nicolas Beltran-Velez
Jia-Bin Huang @jbhuang0604.bsky.social · 01/12/2024
How to drive your research forward? “I tested the idea we discussed last time. Here are some results. It does not work. (… awkward silence)” Such conversations happen so many times when meetings with students. How do we move forward? You need …
19118
Reposted by Nicolas Beltran-Velez
Gabriel Peyré @gabrielpeyre.bsky.social · 30/11/2024
I wrote a summary of the main ingredients of the neat proof by Hugo Lavenant that diffusion models do not generally define optimal transport. github.com/mathematical...
523845
Reposted by Nicolas Beltran-Velez
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 29/11/2024
Look if ClosedForProfitAI wants to train ChatPilotLLM on bluesky, they're not gonna use 1 or 2 million posts from huggingface they'll just download the entire firehose, have extremely skilled people update the model as new posts get downloaded, and not tell anybody what data they're using to train.
2658
Reposted by Nicolas Beltran-Velez
merve @merve.bsky.social · 27/11/2024
It's pretty sad to see the negative sentiment towards Hugging Face on this platform due to a dataset put by one of the employees. I want to write a small piece. 🧵 Hugging Face empowers everyone to use AI to create value and is against monopolization of AI it's a hosting platform above all.
2945570
Reposted by Nicolas Beltran-Velez
Jeremy Howard @howard.fm · 28/11/2024
FYI, here's the entire code to create a dataset of every single bsky message in real time: ``` from atproto import * def f(m): print(m.header, parse_subscribe_repos_message()) FirehoseSubscribeReposClient().start(f) ```
1944262
Reposted by Nicolas Beltran-Velez
Jeremy Howard @howard.fm · 28/11/2024
Rather than restricting data to only the richest and most powerful (as reddit, facebook, and twitter do), bsky makes it available to everyone. Personally, I think that's a good thing.
118915
Reposted by Nicolas Beltran-Velez
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 28/11/2024
I've never understood someone less than the journalists who snarkily write about how no one could conceivably ever want any AI tool. I can't steelman their position. I couldn't pass an ideological turing test. I cannot put myself in their shoes. It's very frustrating.
15995
Reposted by Nicolas Beltran-Velez
Omar Sanseviero @osanseviero.bsky.social · 27/11/2024
I'm disheartened by how toxic and violent some responses were here. There was a mistake, a quick follow up to mitigate and an apology. I worked with Daniel for years and is one of the persons most preoccupied with ethical implications of AI. Some replies are Reddit-toxic level. We need empathy.
2933237
Reposted by Nicolas Beltran-Velez
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 26/11/2024
Excited to share OLMo 2! 🐟 7B and 13B weights, trained up to 4-5T tokens, fully open data, code, etc 🐠 better architecture and recipe for training stability 🐡 staged training, with new data mix Dolmino🍕 added during annealing 🦈 state-of-the-art OLMo 2 Instruct models #nlp #mlsky links below👇
A scatter plot comparing language models by performance (y-axis, measured in average performance on 10 benchmarks) versus training computational cost (x-axis, in approximate FLOPs). The plot shows OLMo 2 models (marked with stars) achieving Pareto-optimal efficiency among open models, with OLMo-2-13B and OLMo-2-7B sitting at the performance frontier relative to other open models like DCLM, Llama 3.1, StableLM 2, and Qwen 2.5. The x-axis ranges from 4x10^22 to 2x10^24 FLOPs, while the y-axis ranges from 35 to 70 benchmark points.
16812
Reposted by Nicolas Beltran-Velez
Luca Soldaini 🎀 @soldaini.net · 26/11/2024
"i can fix her 🥹" meme but it's me and tokenizers
2632
Reposted by Nicolas Beltran-Velez
Clément Canonne @ccanonne.github.io · 24/11/2024
On the TCS job market: Juspreet Singh Sandhu! Optimization, complexity & probabilistic analysis of random CSPs, and models in classical & quantum statistical physics. Techniques: random matrices, free probability, concentration inequalities, stochastic analysis. 1/2 #TCSSky #AcademicJobMarket
1111