Sign in

Michael Kirchhof

@mkirchhof.bsky.social
368 followers 187 following 101 posts

RS on uncertainty quantification for agents at Apple

PostsRepliesMedia
Michael Kirchhof @mkirchhof.bsky.social · 30/09/2026
We let an agent self-improve by exploring and internalizing ”on this sort of task, keep this sort of thing in mind“ bits of self-feedback. It learned to solve tasks where the original policy failed 128 attempts in a row and GRPO flatlined. Paper: arxiv.org/abs/2609.37633 🧵1/8
210
Reposted by Michael Kirchhof
lucazapp.bsky.social @lucazapp.bsky.social · 29/06/2026
🚀 We're hiring a Research Scientist, Machine Learning Research (MLR), Barcelona ☀️ 👉 Apply here: jobs.apple.com/en-us/detail... #MachineLearning #MLResearch #AI #Hiring #Apple #Barcelona #ResearchScientist #DeepLearning
jobs.apple.com
AIML - Machine Learning Researcher - MLR - Jobs - Careers at Apple
Apply for a AIML - Machine Learning Researcher - MLR job at Apple. Read about the role and find out if it’s right for you.
283
Michael Kirchhof @mkirchhof.bsky.social · 23/04/2026
Chat with me about uncertainty + LLM (agents) at #ICLR2026, in either of these 6 sessions: 1. Thu 10:30 AM - 1 PM, Pavilion 3 P3-#309, Memory LLM paper 2. Thu 3:15 PM – 5:45 PM, Pavilion 3 P3-#312, SelfReflect 3. Fri, 9:30 AM - 11:30 AM, Apple booth
110
Michael Kirchhof @mkirchhof.bsky.social · 23/04/2026
Advice for ICLR (or any other conference): Don't collect posters. Collect understanding. 🧵
100
Michael Kirchhof @mkirchhof.bsky.social · 06/03/2026
New paper 🥳 RL relies a lot on an agent’s capability to explore. Our strategy-guided exploration makes the agent find new solutions more efficiently. It learns faster, and in some environments its Pass@1 surpasses the base model’s Pass@128. 🧵 1/6 📄 arxiv.org/abs/2603.02045
142
Michael Kirchhof @mkirchhof.bsky.social · 13/02/2026
Deciding which tokens are learnable and which not is not just a question of loss. Check out our new foundational research paper on training small language models that remain factual 👇
050
Michael Kirchhof @mkirchhof.bsky.social · 27/01/2026
4 ICLR papers 🥳 There’s an insightful story between them: If you sample LLMs multiple times, they are calibrated, even on higher levels [1], but they cannot talk about this uncertainty in a single prompt [2], so you have to help them out to gather information Bayes-optimally [3]
291
Reposted by Michael Kirchhof
Bruno Mlodozeniec @brunokm.bsky.social · 06/01/2026
In our new work — Complete(d)P — we try to answer 3 questions about hyperparameter (HP) scaling: ● How to transfer across model size, tokens&batch-size?→ Complete(d)P ● Do per-module HPs matter? ✔️2x speed-ups possible ● Do they transfer to larger scale? ✔️ With the right parameterisation
184
Michael Kirchhof @mkirchhof.bsky.social · 19/12/2025
If you want some holiday reflections: This is not just a blogpost, but an insight into the philosophy of one of the best scientific minds (and best humans, really) I had the honor to share a bit of my life with.
000
Michael Kirchhof @mkirchhof.bsky.social · 14/11/2025
Our research team is hiring PhD interns 🍏 Spend your next summer in Paris and explore the next frontiers of LLMs for uncertainty quantification, calibration, RL and post-training, and Bayesian experimental design. Details & Application ➡️ jobs.apple.com/en-my/detail...
jobs.apple.com
Internship - Machine Learning Research on Uncertainty - Jobs at Apple (MY)
Apply for a Internship - Machine Learning Research on Uncertainty job at Apple. Read about the role and find out if it’s right for you.
132
Reposted by Michael Kirchhof
sineadwilliamson.bsky.social @sineadwilliamson.bsky.social · 07/11/2025
📢 We’re looking for a researcher in in cogsci, neuroscience, linguistics, or related disciplines to work with us at Apple Machine Learning Research! We're hiring for a one-year interdisciplinary AIML Resident to work on understanding reasoning and decision making in LLMs. 🧵
1105
Reposted by Michael Kirchhof
Marco Cuturi @marcocuturi.bsky.social · 05/11/2025
We have been working with Michal Klein on pushing a module to train *flow matching* models using JAX. This is shipped as part of our new release of the OTT-JAX toolbox (github.com/ott-jax/ott) The tutorial to do so is here: ott-jax.readthedocs.io/tutorials/ne...
1147
Reposted by Michael Kirchhof
Marco Cuturi @marcocuturi.bsky.social · 17/10/2025
It's that time of the year! 🎁 The Apple Machine Learning Research (MLR) team in Paris is hiring a few interns, to do cool research for ±6 months 🚀🚀 & work towards publications/OSS. Check requirements and apply: ➡️ jobs.apple.com/en-us/detail... More❓→ ✉️ mlr_paris_internships@group.apple.com
074
Michael Kirchhof @mkirchhof.bsky.social · 06/10/2025
LLMs are currently this one big parameter block that stores all sort of facts. In our new preprint, we add context-specific memory parameters to the model, and pretrain the model along with a big bank of memories. 📑 arxiv.org/abs/2510.02375 [1/10]🧵
1134
Reposted by Michael Kirchhof
Marco Cuturi @marcocuturi.bsky.social · 03/10/2025
Our two phenomenal interns, Alireza Mousavi-Hosseini and Stephen Zhang @syz.bsky.social have been cooking some really cool work with Michal Klein and me over the summer. Relying on optimal transport couplings (to pick noise and data pairs) should, in principle, be helpful to guide flow matching 🧵
2307
Michael Kirchhof @mkirchhof.bsky.social · 01/10/2025
Many treat uncertainty = a number. At Apple, we're rethinking this: LLMs should output strings that reveal all information of their internal distributions. We find that Reasoning, SFT, CoT can't do it - yet. To get there, we introduce the SelfReflect benchmark. arxiv.org/pdf/2505.20295
3336
Reposted by Michael Kirchhof
Shubhendu Trivedi @shubhendu.bsky.social · 01/09/2025
Natural idea. Looks like a nice paper too. arxiv.org/abs/2508.21184
arxiv.org
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
We propose a general-purpose approach for improving the ability of Large Language Models (LLMs) to intelligently and adaptively gather information from a user or other external source using the framew...
1256
Michael Kirchhof @mkirchhof.bsky.social · 14/07/2025
I'll present my view on the future of uncertainties in LLMs and vision models at @icmlconf.bsky.social, in penal discussions, posters, and workshops. Reach out if you wanna chat :) Here's everything from me and other folks at Apple: machinelearning.apple.com/updates/appl...
051
Michael Kirchhof @mkirchhof.bsky.social · 03/07/2025
Can LLMs access and describe their own internal distributions? With my colleagues at Apple, I invite you to take a leap forward and make LLM uncertainty quantification what it can be. 📄 arxiv.org/abs/2505.20295 💻 github.com/apple/ml-sel... 🧵1/9
1236
Reposted by Michael Kirchhof
silingao.bsky.social @silingao.bsky.social · 23/06/2025
NEW PAPER ALERT: Recent studies have shown that LLMs often lack robustness to distribution shifts in their reasoning. Our paper proposes a new method, AbstRaL, to augment LLMs’ reasoning robustness, by promoting their abstract thinking with granular reinforcement learning.
163
Michael Kirchhof @mkirchhof.bsky.social · 11/06/2025
I‘ll talk today about our latest research on uncertainty quantification at Apple (papers are 2 weeks old) and what I see as the future for UQ in vision and LLMs. See you at 102B, 4:30pm! PS: Lmk if you wanna chat :)
030
Michael Kirchhof @mkirchhof.bsky.social · 29/05/2025
At the end of my PhD, I reflected on uncertainty quantification research, and what might change with chatbots and LLM agents. This was now accepted as position paper at @icmlconf.bsky.social. Some of those future topics are already picking up pace, so have an evening read ☕ arxiv.org/abs/2505.22655
0184
Reposted by Michael Kirchhof
Maureen de Seyssel @maureendeseyssel.bsky.social · 27/05/2025
Now that @interspeech.bsky.social registration is open, time for some shameless promo! Sign-up and join our Interspeech tutorial: Speech Technology Meets Early Language Acquisition: How Interdisciplinary Efforts Benefit Both Fields. 🗣️👶 www.interspeech2025.org/tutorials ⬇️ (1/2)
interspeech2025.org
https://www.interspeech2025.org/tutorials
Your cookies are disabled, please enable them.
195
Michael Kirchhof @mkirchhof.bsky.social · 08/05/2025
Aleatoric and epistemic uncertainty are clear-cut concepts, right? ... right? 😵‍💫 In our new ICLR blogpost we let different schools of thought speak and contradict each other, and revisit chatbots where “the character of aleatory ‘transforms’ into epistemic” iclr-blogposts.github.io/2025/blog/re...
1309
Reposted by Michael Kirchhof
Cem Koç @cemkoch.bsky.social · 07/05/2025
Today we have released the code and a demo iOS application for FastVLM - our extremely efficient and fast vision language model which runs on your device using MLX! You can check out the code and the app here: github.com/apple/ml-fas...
143
Reposted by Michael Kirchhof
Preetum Nakkiran @preetumnakkiran.bsky.social · 11/02/2025
Paper🧵 (cross-posted at X): When does composition of diffusion models "work"? Intuitively, the reason dog+hat works and dog+horse doesn’t has something to do with independence between the concepts being composed. The tricky part is to formalize exactly what this means. 1/
Left Image: A shaggy dog-horse hybrid standing in a rural landscape.
Right Image: A golden dog wearing a red beret against a blurred outdoor background.
23915
Reposted by Michael Kirchhof
Vimal Thilak @aggieinca.bsky.social · 07/02/2025
🚨 Apple Machine Learning Research Internship opportunity! My colleagues in Apple MLR are looking for a PhD research intern with a strong interest in reinforcement learning/post-training for LLMs. If interested, apply by sending an email to Etai Littwin (elittwin at apple dot com)
031
Michael Kirchhof @mkirchhof.bsky.social · 24/01/2025
Wow, OpenAI's o1 has a whopping 93% ECE on Humanity's Last Exam. So if you just prompt o1 to tell you how sure it is about its answer, it will basically produce gibberish. And that's how most users will ask for uncertainties. We have work to do!
Results of state-of-the-art LLMs on Humanity's Last Exam are surprisingly bad, especially their uncertainties.
0161
Reposted by Michael Kirchhof
Marco Cuturi @marcocuturi.bsky.social · 22/01/2025
Today is a great day for optimal transport 🎉! Lots of gratitude 🙏 for all folks who contributed to ott-jax.readthedocs.io and pushed for the MOSCOT (now @ nature!) paper, from visionaries @dominik1klein.bsky.social, G. Palla, Z. Piran to the magician, Michal Klein! ❤️ www.nature.com/articles/s41...
nature.com
Mapping cells through time and space with moscot - Nature
Moscot is an optimal transport approach that overcomes current limitations of similar methods to enable multimodal, scalable and consistent single-cell analyses of datasets across spatial and temporal...
0227
Michael Kirchhof @mkirchhof.bsky.social · 13/12/2024
Many LLM uncertainty estimators perform similarly, but does that mean they do the same? No! We find that they use different cues, and combining them gives even better performance. 🧵1/5 📄 openreview.net/forum?id=QKR... NeurIPS: Sunday, East Exhibition Hall A, Safe Gen AI workshop
1114
Reposted by Michael Kirchhof
Andrea Santilli @asantilli.bsky.social · 12/12/2024
Interested in learning how to evaluate uncertainty in LLMs? Check out our work at NeurIPS! Feel free to reach out for a chat!
131
Reposted by Michael Kirchhof
Bálint Mucsányi @bmucsanyi.bsky.social · 12/12/2024
Excited to present our spotlight paper on uncertainty disentanglement at #NeurIPS! Drop by today between 11 am and 2 pm PST at West Ballroom A-D #5509 and let's chat!
0101
Michael Kirchhof @mkirchhof.bsky.social · 12/12/2024
Evaluating your LLM uncertainties with Rougle-L will show clear winners... except that they aren't actually good. We find that Rouge-L spuriously favors some methods over others. 🧵1/4 📄 openreview.net/forum?id=jGt... NeurIPS: Sunday, East Exhibition Hall A, Safe Gen AI workshop
173
Reposted by Michael Kirchhof
Alexander Kolesnikov @handle.invalid · 04/12/2024
Ok, it is yesterdays news already, but good night sleep is important. After 7 amazing years at Google Brain/DM, I am joining OpenAI. Together with @xzhai.bsky.social and @giffmana.ai, we will establish OpenAI Zurich office. Proud of our past work and looking forward to the future.
811611
Reposted by Michael Kirchhof
Bálint Mucsányi @bmucsanyi.bsky.social · 03/12/2024
Thrilled to share our NeurIPS spotlight on uncertainty disentanglement! ✨ We study how well existing methods disentangle different sources of uncertainty, like epistemic and aleatoric. While all tested methods fail at this task, there are promising avenues ahead. 🧵 👇 1/7 📖: arxiv.org/abs/2402.19460
4566
Michael Kirchhof @mkirchhof.bsky.social · 03/12/2024
Proud to announce our NeurIPS spotlight, which was in the works for over a year now :) We dig into why decomposing aleatoric and epistemic uncertainty is hard, and what this means for the future of uncertainty quantification. 📖 arxiv.org/abs/2402.19460 🧵1/10
37412
Michael Kirchhof @mkirchhof.bsky.social · 22/11/2024
My last week as an Apple intern was insane. 3 paper deadlines? Sure, I can do it. "Wanna interview this week?" Sure, I can do it! "Wanna present your side project to the senior VP?" Sure, I... wait 🤯 I had such a flow! It was so fun! I want more. I'm joining Apple as a Research Scientist 🍎
The apple park as a Lego miniature building.
6240
Michael Kirchhof @mkirchhof.bsky.social · 19/11/2024
Hello there 🦋 I'll continue my uncertainty quantification research here :) Expect a NeurIPS spotlight and some work-in-progress papers here in the next days. You'll hear it here first!
080