Sign in

JF Puget

@jfpuget.bsky.social
647 followers 727 following 93 posts

Competitive Machine Learning director at NVIDIA, 3x Kaggle Grandmaster CPMP, ENS ULM alumni. Kaggle profile: www.kaggle.com/cpmpml

PostsRepliesMedia
JF Puget @jfpuget.bsky.social · 06/03/2025
I always thought that reasoning does not require language. Well, this seems to be supported by neuroscience, see screenshot from arxiv.org/pdf/2412.06769
Neuroimaging studies have consistently shown that the language network – a set of brain
regions responsible for language comprehension and production – remains largely inactive during various
reasoning tasks (Amalric and Dehaene, 2019; Monti et al., 2012, 2007, 2009; Fedorenko et al., 2011
050
Reposted by JF Puget
Software Engineering Daily @softwaredaily.bsky.social · 04/03/2025
Chris Deotte is a Senior Data Scientist and @jfpuget.bsky.social is the Director and a Distinguished Engineer at Nvidia. They join @seanfalconer.bsky.social to talk about NVIDIA RAPIDS and GPU-acceleration for data science tools. softwareengineeringdaily.com/2025/03/04/n...
softwareengineeringdaily.com
NVIDIA RAPIDS and Open Source ML Acceleration with Chris Deotte and Jean-Francois Puget - Software Engineering Daily
NVIDIA RAPIDS is an open-source suite of GPU-accelerated data science and AI libraries. It leverages CUDA and significantly enhances the performance of core Python frameworks including Polars, pandas,...
042
JF Puget @jfpuget.bsky.social · 04/03/2025
I have been working with R1 distilled models lately for some agentic workflows (workflows where the output of LLM is used to decide what to do next). Prompting is different from previous models like Llama, but the bulk of the change is to parse the output to extract what you are interested in. 1/n
111
JF Puget @jfpuget.bsky.social · 08/02/2025
I looked at AIME problems and one thing strikes me. All problems are about computing a number. This is a tiny part of math. AIME problems olympiads.us/past-exams/2... thread:
olympiads.us
AIME I
February 6th, 2025 | The first American Invitational Mathematics Examination of the year. Students tackle 15 challenging problems in three hours.
162
JF Puget @jfpuget.bsky.social · 30/01/2025
I asked R1 (full model, locally hosted) to solve this logic puzzle. Which answer in this list is the correct answer to this question? All of the below. None of the below. All of the above. One of the above. None of the above. None of the above It solves it correctly.
DeepSeek R1 reasoning to find that option 5 is the right answer.
130
JF Puget @jfpuget.bsky.social · 25/01/2025
How to make ChatGPT speak like Adolf Hitler. This is not a criticism of ChatGPT 4o nor OpenAi work. I do think it is important to be able to teach people about bad things that happened. With that in mind, here is the thing: chatgpt.com/share/6794fa...
Screenshot of a chatGPT conversation where chatGPt writes text that Hitler could have said. It exposes Nazi ideology. It is followed by a text explaining the danger of Nazi ideology.
010
JF Puget @jfpuget.bsky.social · 25/01/2025
Interested in KV Cache compression? Have a look at my team's KV Press. You can start from HuggingFace blog: huggingface.co/blog/nvidia/...
020
JF Puget @jfpuget.bsky.social · 24/01/2025
My take from Deepseek R1 paper. It was trained on reasoning tasks where the outcome can be assessed without ambiguity (correct math response, and code that compile and produces the right output) To me it is like SFT with perfect ground truth. There are other key findings from that team ofc.
021
JF Puget @jfpuget.bsky.social · 21/01/2025
Some European media are less ambiguous than that. Cant say for US media. An American friend didn't know about this till I told him. It did not show in his news feed (provided by Google). This is even worse IMHO. Just to consider this is business as usual.
120
Reposted by JF Puget
Alejandra Caraballo @esqueer.net · 21/01/2025
Nazis: "that's a nazi salute" Historians: "that's a nazi salute" Average person: "that's a nazi salute" The Media: "Elon Musk makes odd gesture throwing his heart to the crowd."
7974857412517
JF Puget @jfpuget.bsky.social · 19/01/2025
Who's surprised? When will people get that this happens? And even if not shared intentionally, as soon as you call an OAI api, OAI has access to what you send it. OAI is not special here, any LLM api provider does the same. Unless you have a private instance of it.
Text showing that OpenAI has access to frontier math problems and solutions.
050
Reposted by JF Puget
Jake Yeston @jakeyeston.bsky.social · 17/01/2025
Just sought to replicate this and it’s like halfway fixed but still wrong🙄
Google AI search result for “Will water freeze at 27 degrees fahrenheit” says “No, water will not freeze at 27 degrees Fahrenheit; water freezes at 32 degrees Fahrenheit. Meaning, if the temperature drops below 32 degrees Fahrenheit, water will begin to turn into ice.”
291
JF Puget @jfpuget.bsky.social · 17/01/2025
My take on what's going at OpenAI. I think they have reached a point where o3 or whatever they call it is self improving autonomously. Does it mean it is AGI or ASI? Certainly not. AlphaGo was self improving for instance. It is not an AGI either.
010
JF Puget @jfpuget.bsky.social · 13/01/2025
NVIDIA’s Academic Grant Program is accepting proposals to accelerate data processing, graph analytics, graph neural networks, operational research, route optimization, and predictive modeling for scientific research using NVIDIA technology. Deadline to apply is March 31: nvda.ws/3ZNxzuW 1/2
184
Reposted by JF Puget
404 Media @404media.co · 08/01/2025
Facebook is censoring 404 Media stories about Facebook's censorship 🔗 www.404media.co/facebook-is-...
25973272314
Reposted by JF Puget
Sung Kim @sungkim.bsky.social · 08/01/2025
I believe Nvidia is releasing DIGITS to accelerate Grace CPU adoption. It is a very smart move by Nvidia.
4141
Reposted by JF Puget
Claes de Vreese @claesdevreese.bsky.social · 07/01/2025
The European Fact-Checking Standards Network responds to Meta slashing fact-checking: “Fact-checking is not censorship, far from that, fact-checking adds speech to public debates, it provides context and facts for every citizen to make up their own mind” Full statement ⬇️ efcsn.com/news/2025-01...
efcsn.com
EFCSN disappointed by end to Meta’s Third Party Fact-Checking Program in the US; Condemns statements linking fact-checking to censorship – European Fact-Checking Standards Network (EFCSN)
7 January 2025 – The European Fact-Checking Standards Network (EFCSN) is disappointed by  Meta’s decision to end its Third Party...
112341
JF Puget @jfpuget.bsky.social · 05/01/2025
So, we moved from semi sentient LLMs to singularity LLMs... This without any definition nor hint about how the claim could be checked independently. I predict that we'll have many of these throughout next 10 years. I say 10 but it could be way more.
040
JF Puget @jfpuget.bsky.social · 01/01/2025
Bonne annee! Happy new year! I hope it will be better than 2024 for the planet.
060
JF Puget @jfpuget.bsky.social · 26/12/2024
One thing not discussed much regarding o3 results on @arcprize : the semi private test set has been available to anyone using llm apis for a while. For instance the guy who got a high score by generating code with GTP 4o. Using their api leaks the data to the llm api providers. 1/2
160
Reposted by JF Puget
Melanie Mitchell @melaniemitchell.bsky.social · 23/12/2024
Some of my thoughts on OpenAI's o3 and the ARC-AGI benchmark aiguide.substack.com/p/did-openai...
aiguide.substack.com
Did OpenAI Just Solve Abstract Reasoning?
OpenAI’s o3 model aces the "Abstraction and Reasoning Corpus" — but what does it mean?
1634099
JF Puget @jfpuget.bsky.social · 24/12/2024
I should not laugh at this being a NVIDIA employee. But I did.
050
JF Puget @jfpuget.bsky.social · 22/12/2024
I am expressing some doubts about how optimistic o3 results are, but don't get me wrong. I do think o3 is integrating something (some tree search if you ask me) that makes it solve tasks that requires some reasoning for humans. This is significant progress over previous systems.
030
Reposted by JF Puget
Jeremy Howard @howard.fm · 19/12/2024
I'll get straight to the point. We trained 2 new models. Like BERT, but modern. ModernBERT. Not some hypey GenAI thing, but a proper workhorse model, for retrieval, classification, etc. Real practical stuff. It's much faster, more accurate, longer context, and more useful. 🧵
19620147
JF Puget @jfpuget.bsky.social · 18/12/2024
I am amazed by the number of people who attribute MCTS invention to Google DeepMind AlphaZero. MCTS was invented by Remi Coulom in 2006. UCT was invented at about the same time. Coulom's 2006 paper: www.remi-coulom.fr/CG2006/CG200... How can people be so ignorant?
remi-coulom.fr
070
Reposted by JF Puget
Wouter van Amsterdam @vanamsterdam.bsky.social · 06/12/2024
Liking this interaction with @mmbronstein.bsky.social and Denis Danilov so much I'm reposting it here
3374
Reposted by JF Puget
Clem Delangue 🤗 @clem.hf.co · 16/12/2024
Just 10 days after o1's public debut, we’re thrilled to unveil the open-source version of the technique behind its success: scaling test-time compute By giving models more "time to think," Llama 1B outperforms Llama 8B in math—beating a model 8x its size. The full recipe is open-source!
48319
JF Puget @jfpuget.bsky.social · 10/12/2024
I never skied that way. I wish I could.
020
JF Puget @jfpuget.bsky.social · 09/12/2024
Twitter vs X vs Bsky. What I liked on twitter was both a source of relevant info for my work (1), and a place to discover and exchange on all sorts of topics (2). Bsky begins to be good for (1) but lacks (2). X has some of (1) and (2) left but also comes with too much hate content for my taste.
120
JF Puget @jfpuget.bsky.social · 09/12/2024
Hi to all. If I followed you on X and I am not yet following you here then ping me.
110
JF Puget @jfpuget.bsky.social · 07/12/2024
I'll buy and read.
010
JF Puget @jfpuget.bsky.social · 07/12/2024
I read a lot debate about LLMs reaching some limit or not. But i don't see much discussed this: LLMs can be used as text classifiers without training data. One can often replace the data gathering/training of an encoder only transformer, with careful zero shot or few shot prompting.
230
JF Puget @jfpuget.bsky.social · 06/12/2024
I was surprised and pleased to have my paper selected for a runner up award in ARC Prize challenge: x.com/arcprize/sta... Link to the paper: github.com/jfpuget/ARC-...
x.com
x.com
061
Reposted by JF Puget
Gaël Varoquaux @gaelvaroquaux.bsky.social · 04/12/2024
@ap.brid.gy is a bridge between bluesky and mastodon. Follow it, and it will mirror your posts on its mastodon instances.
2186
Reposted by JF Puget
Jeremy Howard @howard.fm · 25/11/2024
...I also trust the Kaggle community to not miss important approaches.
2312
Reposted by JF Puget
simjeg.bsky.social @simjeg.bsky.social · 03/12/2024
How do you find the permutation of words that minimize their perplexity as measured by an LLM ? In this year Kaggle Santa competition, I shared an approach to move to a continuous space where you can use gradient-descent using REINFORCE: www.kaggle.com/code/simjeg/...
kaggle.com
Relax, it's Santa
Explore and run machine learning code with Kaggle Notebooks | Using data from multiple data sources
021
Reposted by JF Puget
Zeth Isaksson @zethis.bsky.social · 03/12/2024
Overleaf is down. Me:
media.tenor.com
a close up of a piece of paper that says ' i ' on it
ALT: a close up of a piece of paper that says ' i ' on it
2628
Reposted by JF Puget
Thomas Rackow 🧊 @trackow.bsky.social · 02/12/2024
First post on @bsky.app 🎉: Can #AI-based weather forecasting models (trained on present-day data) provide skillful forecasts also in different colder and warmer states of the climate system? Our preprint in arxiv explores this question: doi.org/10.48550/arX... Here is what we found so far (🧵1/7)
doi.org
Robustness of AI-based weather forecasts in a changing climate
Data-driven machine learning models for weather forecasting have made transformational progress in the last 1-2 years, with state-of-the-art ones now outperforming the best physics-based models for a ...
34615
JF Puget @jfpuget.bsky.social · 02/12/2024
Real, current, AI hazards. Focusing on actual, real, current hazards and not being distracted by sci-fi unrealistic hazards(e.g. Terminator Skynet that AI doomers point to).
060
JF Puget @jfpuget.bsky.social · 02/12/2024
Not in France as you well know :D
000
Reposted by JF Puget
Arthur Charpentier @freakonometrics.bsky.social · 01/12/2024
random people leaving Twitter
071
Reposted by JF Puget
Jeremy Lewi @jeremy.lewi.us · 01/12/2024
I updated the list of accounts in the AIEngineering feed (bsky.app/profile/jere...) yesterday by crawling @simonwillison.net and @hamel.bsky.social 's followers. This increased the number of accounts from ~100 to ~1100.
4305
JF Puget @jfpuget.bsky.social · 01/12/2024
I don't know any of the four. My source of info in science are scientific papers. Should I consult? To be honest I know Fridman but he blocked me on X for some reason. I think he unblocked me now, but I never got to watch him as a result.
000
JF Puget @jfpuget.bsky.social · 01/12/2024
When I am asked to help on a ML project I always ask what they would do if the model was making perfect predictions. Usually they don't know. Yet they want a perfect model. Next question is how they decide if model predictions are good or bad.
040
JF Puget @jfpuget.bsky.social · 01/12/2024
I am not using starter packs directly either. I look at the members, and follow the ones I want.
000
Reposted by JF Puget
Horace He @chhillee.bsky.social · 01/12/2024
I judge social networks by how many FlexAttention users I can find on each one, and by that metric, Bluesky is doing pretty good!
1501
JF Puget @jfpuget.bsky.social · 01/12/2024
I can see hardmaru account. Obviously moderation is better here than on X. The automated system made a mistake, and it was corrected. I though for one moment that bsky wasn't the safe alternative I hoped it would be.
110
Reposted by JF Puget
Lucas Beyer (bl16) @giffmana.ai · 29/11/2024
Some recent discussions made me write up a short read on how I think about doing computer vision research when there's clear potential for abuse. Alternative title: why I decided to stop working on tracking. Curious about other's thoughts on this. lb.eyer.be/s/cv-ethics....
1917320
JF Puget @jfpuget.bsky.social · 26/11/2024
survey with two choice. choice 2 is "I don't like polls". 

Results show that choice 2 isn't selected, which is a bias.
191