Sign in

Philip Bontrager

@pbontrager.bsky.social
794 followers 963 following 181 posts

AI researcher & engineer @Meta working on @PyTorch torchtune in NYC; interests in generative models, RL, and evolutionary strategies 💻 github.com/pbontrager 📝 tinyurl.com/philips-papers

PostsRepliesMedia
Philip Bontrager @pbontrager.bsky.social · 08/06/2025
What goes into saving checkpoints is not something that many people think about, but as models get bigger this becomes a challenge. The biggest open models now have checkpoints over 700gb that can take tens of minutes every time you want to consolidate into a checkpoint. pytorch.org/blog/hugging...
162
Reposted by Philip Bontrager
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 06/02/2025
We've built a simulated driving agent that we trained on 1.6 billion km of driving with no human data. It is SOTA on every planning benchmark we tried. In self-play, it goes 20 years between collisions.
2229855
Philip Bontrager @pbontrager.bsky.social · 25/01/2025
In the Alice In Wonderland (github.com/LAION-AI/AIW) reasoning and generalization benchmark, DeepSeek R1 appears to perform much more like o1 mini than o1 -preview. (Plot from laion-ai)
240
Philip Bontrager @pbontrager.bsky.social · 14/01/2025
Can we just study LLM activations/behavior because it’s interesting and it can tell us things about language and AI without imbuing artificial importance or meaning on top of it?
020
Philip Bontrager @pbontrager.bsky.social · 01/01/2025
Plagiarize other people’s research
010
Reposted by Philip Bontrager
mr. TIM @timkellogg.me · 23/12/2024
this seems like a decent LLM test. 3 sequential game states of Qwirkle. Sonnet 3.5 gets the first play but not the second o1 is very bad at this. first play it takes 59 seconds and it’s answer isn’t even a play that’s on the board. same with second play, but only 36 seconds
3122
Philip Bontrager @pbontrager.bsky.social · 24/12/2024
Contrary to what I see in a lot of online discussions, AI benchmarks are not meant to show how capable an AI system is, but instead they show what they can’t do.
100
Philip Bontrager @pbontrager.bsky.social · 22/12/2024
If you have a lot of experience training and fine-tuning ML models and want to help bring that expertise to the community, we’re looking to hire a new member for the torchtune team! www.metacareers.com/jobs/5121890...
metacareers.com
Software Engineer - PyTorch Domains
Meta's mission is to build the future of human connection and the technology that makes it possible.
130
Philip Bontrager @pbontrager.bsky.social · 20/12/2024
New release of torchtune right before Christmas! We have new recipes, better integration with vLLM and HF, support for Gemma2, and more. We've also now added support for Kaggle notebooks! www.kaggle.com/code/felipem...
kaggle.com
torchtune in kaggle
Explore and run machine learning code with Kaggle Notebooks | Using data from multiple data sources
050
Philip Bontrager @pbontrager.bsky.social · 19/12/2024
New encoder using all the latest training tricks! One thing I’m wondering is how this compares to something like SmolLM (similar size). I know encoder models should provide better embeddings but I wonder what this looks like in practice.
000
Reposted by Philip Bontrager
Clem Delangue 🤗 @clem.hf.co · 16/12/2024
Just 10 days after o1's public debut, we’re thrilled to unveil the open-source version of the technique behind its success: scaling test-time compute By giving models more "time to think," Llama 1B outperforms Llama 8B in math—beating a model 8x its size. The full recipe is open-source!
48319
Reposted by Philip Bontrager
Blake Richards @tyrellturing.bsky.social · 16/12/2024
1/ Okay, one thing that has been revealed to me from the replies to this is that many people don't know (or refuse to recognize) the following fact: The unts in ANN are actually not a terrible approximation of how real neurons work! A tiny 🧵. 🧠📈 #NeuroAI #MLSky
2115238
Philip Bontrager @pbontrager.bsky.social · 15/12/2024
Excited to see diffusion language models getting scaled up to sizes where we can start to compare them to auto-regressive approaches (though this model is a bit of a hybrid)
020
Philip Bontrager @pbontrager.bsky.social · 14/12/2024
I’m at the age where I have to go on LinkedIn if I want to see what my old high school friends are up too.
010
Reposted by Philip Bontrager
Yoav Goldberg @yoavgo.bsky.social · 14/12/2024
"pre training as we know it will end (because we will run out of data)" is, in other words, "learning to complete partial observations is not sufficient to get to intelligence". i think this was kinda obvious to many, but maybe noteworthy that a true scale-believer said it.
1323
Philip Bontrager @pbontrager.bsky.social · 14/12/2024
The way you can tell if an image is AI generated or not is by looking at the hands. If the hands look weird they’re probably human drawn.
000
Philip Bontrager @pbontrager.bsky.social · 11/12/2024
As a counterpoint, when I was applying for grad schools, a professor where I was applying told me that ML was just linear algebra and my PhD would just be that. Almost made me reconsider
020
Reposted by Philip Bontrager
Christian A. Naesseth @canaesseth.bsky.social · 09/12/2024
#NeurIPS2024 #ML
2754
Reposted by Philip Bontrager
Petar Veličković @petar-v.bsky.social · 08/12/2024
A very nice blog from Przemek Pietrzkiewicz, offering thoughts on our recent result in AI for competitive programming 🏆 Przemek co-led the Hash Code contest, which we used as the main test-bed to evaluate our approach 🚀 Worth a read if you want to understand implications of our work! Link below ⬇️
1223
Philip Bontrager @pbontrager.bsky.social · 07/12/2024
If the internet gets filled up with AI generated text, presumably it’s the good text that humans decided to keep from the models. Does that mean over time all model training becomes RLHF? 🤔
100
Philip Bontrager @pbontrager.bsky.social · 06/12/2024
Llama 3.3 70B is out getting very close benchmarks results to the 405B model. If you want to fine-tune it on a bit more than 48GB of VRAM, checkout this torchtune config gist.github.com/pbontrager/b...
gist.github.com
Ultra Low Memory Llama 3.3 Finetuning Config
Ultra Low Memory Llama 3.3 Finetuning Config. GitHub Gist: instantly share code, notes, and snippets.
050
Philip Bontrager @pbontrager.bsky.social · 04/12/2024
Really cool new work out of Deep Mind for video game world generation using latent diffusion! Soon you'll be able to speed run a game just by tricking a model to morph you from one location to another. deepmind.google/discover/blo...
deepmind.google
Genie 2: A large-scale foundation world model
Generating unlimited diverse training environments for future general agents
13910
Reposted by Philip Bontrager
Jia-Bin Huang @jbhuang0604.bsky.social · 01/12/2024
How to drive your research forward? “I tested the idea we discussed last time. Here are some results. It does not work. (… awkward silence)” Such conversations happen so many times when meetings with students. How do we move forward? You need …
19118
Philip Bontrager @pbontrager.bsky.social · 27/11/2024
When building torchtune we’ve had lots of discussions on where to put code. All in the top level recipe? In utilities? Build a trainer? The goal is always to make experimentation and hacking with the recipes easy. I’m curious what your opinions are on using trainers vs recipes style scripts.
120
Philip Bontrager @pbontrager.bsky.social · 24/11/2024
Looking at CMA-ES again and noticed that the next generation sampling rule is “x = m + σ * N(0,C)” looks very similar to an online version of the diffusion algorithm, the main difference is that diffusion usually assumes feature independence.
110
Reposted by Philip Bontrager
vmoens @vmoens.bsky.social · 20/11/2024
Just got started with the PyTorch Starter Pack Who did I forget?
032
Philip Bontrager @pbontrager.bsky.social · 23/11/2024
Thought I’d reshare a talk here that I gave last year. Using an experimental library I helped build, I show how you can build your own Dalle style diffusion model from scratch. t.co/oheTSbpkkI
t.co
https://www.datacamp.com/code-along/building-a-diffuser-model-from-scratch-with-pytorch
010
Philip Bontrager @pbontrager.bsky.social · 21/11/2024
First post on here, let me introduce myself: I started in ML when LSTMs and GANs were cool. I’ve worked on RL for game design, generative biometrics, fashion search, and a little bit more.
100