Sign in

Marco Cuturi

@marcocuturi.bsky.social
823 followers 58 following 30 posts

machine learning researcher @ Apple machine learning research

PostsRepliesMedia
Reposted by Marco Cuturi
szilviaujvary.bsky.social @szilviaujvary.bsky.social · 13/02/2026
Small Language Models (SLMs) don’t have the capacity to remember everything in their training data. Which tokens should they learn to predict, and when should they ask for help? We tackle this question in our new preprint. You can check it out on arxiv: arxiv.org/abs/2602.12005 🧵1/7
1467
Marco Cuturi @marcocuturi.bsky.social · 06/01/2026
With other folks at 🍏, @brunokm.bsky.social has worked on a complete(d) parameterisation for NNs that can *transfer* locally tuned hyperparameters: tune optimizers' parameters (e.g. LR) *per module/depth* using an evolutionary search on small models → they transfer perf. gains to much larger models
051
Reposted by Marco Cuturi
davidgrangier.bsky.social @davidgrangier.bsky.social · 02/12/2025
I am at #NeurIPS2025. We can chat about data mixing, efficient training, ML@Apple and more.
021
Marco Cuturi @marcocuturi.bsky.social · 01/12/2025
my 2 cents on the ICLR drama: It's been years that the system has been under attack. But it's also been years that we hear, year after year, that there is no way to enforce protection mechanisms (e.g. deny lists for dishonest authors or reviewers etc..) for legal reasons.
161
Reposted by Marco Cuturi
sineadwilliamson.bsky.social @sineadwilliamson.bsky.social · 07/11/2025
📢 We’re looking for a researcher in in cogsci, neuroscience, linguistics, or related disciplines to work with us at Apple Machine Learning Research! We're hiring for a one-year interdisciplinary AIML Resident to work on understanding reasoning and decision making in LLMs. 🧵
1105
Marco Cuturi @marcocuturi.bsky.social · 05/11/2025
We have been working with Michal Klein on pushing a module to train *flow matching* models using JAX. This is shipped as part of our new release of the OTT-JAX toolbox (github.com/ott-jax/ott) The tutorial to do so is here: ott-jax.readthedocs.io/tutorials/ne...
1147
Reposted by Marco Cuturi
Rocío Mercado Oropeza @rociomer.bsky.social · 29/10/2025
Afternoon talks by: @marcocuturi.bsky.social Elena Agliari Jan Gerken Thanks all for the great talks, conversations, and engagement! Fingers crossed we get to host this event a 4th time next year and see many of you back in Gothenburg 🤞🇸🇪
051
Reposted by Marco Cuturi
Pau Rodriguez @paurodriguez.bsky.social · 21/10/2025
🚀 Excited to share LinEAS, our new activation steering method accepted at NeurIPS 2025! It approximates optimal transport maps e2e to precisely guide 🧭 activations achieving finer control 🎚️ with ✨ less than 32 ✨ prompts! 💻https://github.com/apple/ml-lineas 📄https://arxiv.org/abs/2503.10679
121
Marco Cuturi @marcocuturi.bsky.social · 17/10/2025
It's that time of the year! 🎁 The Apple Machine Learning Research (MLR) team in Paris is hiring a few interns, to do cool research for ±6 months 🚀🚀 & work towards publications/OSS. Check requirements and apply: ➡️ jobs.apple.com/en-us/detail... More❓→ ✉️ mlr_paris_internships@group.apple.com
074
Marco Cuturi @marcocuturi.bsky.social · 09/10/2025
While working on semidiscrete flow matching this summer (➡️ arxiv.org/abs/2509.25519), I kept looking for a video illustrating that the velocity field solving the Benamou-Brenier OT problem is NOT constant w.r.t. time ⏳... so I did it myself, take a look! ott-jax.readthedocs.io/tutorials/th...
0111
Reposted by Marco Cuturi
Michael Kirchhof @mkirchhof.bsky.social · 06/10/2025
LLMs are currently this one big parameter block that stores all sort of facts. In our new preprint, we add context-specific memory parameters to the model, and pretrain the model along with a big bank of memories. 📑 arxiv.org/abs/2510.02375 [1/10]🧵
1134
Reposted by Marco Cuturi
David Picard @davidpicard.eurosky.social · 04/10/2025
Wow! Finally OT done on the entire training set to train a diffusion model!
0123
Marco Cuturi @marcocuturi.bsky.social · 03/10/2025
Our two phenomenal interns, Alireza Mousavi-Hosseini and Stephen Zhang @syz.bsky.social have been cooking some really cool work with Michal Klein and me over the summer. Relying on optimal transport couplings (to pick noise and data pairs) should, in principle, be helpful to guide flow matching 🧵
2307
Reposted by Marco Cuturi
Peter Gray @peteryugray.bsky.social · 21/08/2025
New Apple #ML Research Highlight: The "Super Weight:" How Even a Single Parameter can Determine an #LLM's Behavior machinelearning.apple.com/research/the...
machinelearning.apple.com
The
A recent paper from Apple researchers,
142
Marco Cuturi @marcocuturi.bsky.social · 01/08/2025
scaling up the computation of optimal transport couplings to hundreds of thousands of 3k dimensional vectors made easy using sharding and OTT-JAX! check this notebook, it only takes a few lines of code thanks to JAX's native sharding abilities ott-jax.readthedocs.io/en/latest/tu...
ott-jax.readthedocs.io
Sharded Sinkhorn — ott 0.5.1.dev34+g3462f28 documentation
0142
Reposted by Marco Cuturi
Peter Gray @peteryugray.bsky.social · 23/07/2025
New Apple #ML Research Highlight: "FastVLM: Efficient Vision Encoding for Vision Language Models" machinelearning.apple.com/research/fas...
machinelearning.apple.com
FastVLM: Efficient Vision Encoding for Vision Language Models
Vision Language Models (VLMs) enable visual understanding alongside textual inputs. They are typically built by passing visual tokens from a…
161
Reposted by Marco Cuturi
Pauline Luc @paulineluc.bsky.social · 08/07/2025
So pleased and proud to share with you what our team has been up to, on an ambitious journey to build a video foundation model for scientific domains ! ✨ 🚀 🎞️ 🧪 #ICCV2025 #AI4Science
0112
Reposted by Marco Cuturi
Ben Recht @beenwrekt.bsky.social · 03/07/2025
The NeurIPS paper checklist corroborates the bureaucratic theory of statistics.
argmin.net
Standard error of what now?
The NeurIPS checklist corroborates the bureaucratic theory of statistics.
46413
Reposted by Marco Cuturi
Michael Kirchhof @mkirchhof.bsky.social · 03/07/2025
Can LLMs access and describe their own internal distributions? With my colleagues at Apple, I invite you to take a leap forward and make LLM uncertainty quantification what it can be. 📄 arxiv.org/abs/2505.20295 💻 github.com/apple/ml-sel... 🧵1/9
1236
Reposted by Marco Cuturi
λ³🎲 @cubiclogic.bsky.social · 29/06/2025
www.kyotoprize.org/en/laureates...
kyotoprize.org
Shun-ichi Amari | Kyoto Prize
Shun-ichi Amari
021
Reposted by Marco Cuturi
λ³🎲 @cubiclogic.bsky.social · 20/06/2025
Shunichi Amari has been awarded the 40th (2025) Kyoto Prize in recognition of his pioneering research in the fields of artificial neural networks, machine learning, and information geometry www.riken.jp/pr/news/2025...
riken.jp
甘利 俊一 栄誉研究員が「京都賞」を受賞
甘利 俊一栄誉研究員(本務:帝京大学 先端総合研究機構 特任教授)は、人工ニューラルネットワーク、機械学習、情報幾何学分野での先駆的な研究が評価され、第40回(2025)京都賞(先端技術部門 受賞対象分野:情報科学)を受賞しました。
23512
Reposted by Marco Cuturi
silingao.bsky.social @silingao.bsky.social · 23/06/2025
NEW PAPER ALERT: Recent studies have shown that LLMs often lack robustness to distribution shifts in their reasoning. Our paper proposes a new method, AbstRaL, to augment LLMs’ reasoning robustness, by promoting their abstract thinking with granular reinforcement learning.
163
Reposted by Marco Cuturi
Maureen de Seyssel @maureendeseyssel.bsky.social · 27/05/2025
Now that @interspeech.bsky.social registration is open, time for some shameless promo! Sign-up and join our Interspeech tutorial: Speech Technology Meets Early Language Acquisition: How Interdisciplinary Efforts Benefit Both Fields. 🗣️👶 www.interspeech2025.org/tutorials ⬇️ (1/2)
interspeech2025.org
https://www.interspeech2025.org/tutorials
Your cookies are disabled, please enable them.
195
Reposted by Marco Cuturi
Cem Koç @cemkoch.bsky.social · 07/05/2025
Today we have released the code and a demo iOS application for FastVLM - our extremely efficient and fast vision language model which runs on your device using MLX! You can check out the code and the app here: github.com/apple/ml-fas...
143
Reposted by Marco Cuturi
davidgrangier.bsky.social @davidgrangier.bsky.social · 22/04/2025
#ICLR #TrainBetterLM I am at ICLR, come to our posters for improved language model training! Recycle gradients for faster neural net training with AdEMAmix iclr.cc/virtual/2025... (Fri Apr 25, 10 am). 1/3
123
Reposted by Marco Cuturi
Peter Gray @peteryugray.bsky.social · 21/04/2025
New post: "Apple Machine Learning Research at #ICLR 2025" - highlighting a selection of the many Apple #ML research papers to be presented at @iclr-conf.bsky.social this week: machinelearning.apple.com/research/icl...
machinelearning.apple.com
Apple Machine Learning Research at ICLR 2025
Apple researchers are advancing machine learning (ML) and AI through fundamental research that improves the world’s understanding of this…
063
Reposted by Marco Cuturi
Pau Rodriguez @paurodriguez.bsky.social · 10/12/2024
Thrilled to share the latest work from our team at @Apple where we achieve interpretable and fine-grained control of LLMs and Diffusion models via Activation Transport 🔥 📄 arxiv.org/abs/2410.23054 🛠️ github.com/apple/ml-act 0/9 🧵
34715
Reposted by Marco Cuturi
Unusual Whales @unusualwhales.bsky.social · 10/04/2025
The trader made millions. Why was this unusual? For a few reasons: - Firstly new opening volume on chain, minutes after market open - IVR of +80 on $QQQ, with iv percentile of medium. - happened at once, at ask -otm - market was bearish, across the board Unusual. Come learn: unusualwhales.com
1125448
Reposted by Marco Cuturi
Elizabeth Warren @warren.senate.gov · 09/04/2025
I'm calling for an investigation into whether President Trump manipulated the market to benefit his Wall Street donors—all while working people and small businesses paid the price. Did Trump help insiders cash in on his tariff flip-flopping? It sure looks like corruption.
youtu.be
Did Trump help insiders cash in on his tariff flip-flopping?
YouTube video by Senator Elizabeth Warren
812154803924
Reposted by Marco Cuturi
Peter Gray @peteryugray.bsky.social · 10/04/2025
Accepted as a Spotlight at @iclr-conf.bsky.social the work shares a new method for fine-grained control over #genAI output - without the computational overhead, complexity, and volume of data needed by #RLHF or fine-tuning, and with more reliable results than prompt engineering.
041
Reposted by Marco Cuturi
Preetum Nakkiran @preetumnakkiran.bsky.social · 11/02/2025
Paper🧵 (cross-posted at X): When does composition of diffusion models "work"? Intuitively, the reason dog+hat works and dog+horse doesn’t has something to do with independence between the concepts being composed. The tricky part is to formalize exactly what this means. 1/
Left Image: A shaggy dog-horse hybrid standing in a rural landscape.
Right Image: A golden dog wearing a red beret against a blurred outdoor background.
23915
Reposted by Marco Cuturi
kyunghyuncho.bsky.social @kyunghyuncho.bsky.social · 28/01/2025
because "attention is all you need"
071
Reposted by Marco Cuturi
Samira @samiraabnar.bsky.social · 28/01/2025
🚨 One question that has always intrigued me is the role of different ways to increase a model's capacity: parameters, parallelizable compute, or sequential compute? We explored this through the lens of MoEs:
1188
Reposted by Marco Cuturi
Pierre Ablin @pierreablin.bsky.social · 24/01/2025
Excited to see Sigmoid Attention accepted at ICLR 2025 !! Make attention ~18% faster with a drop-in replacement 🚀 Code: github.com/apple/ml-sig... Paper arxiv.org/abs/2409.04431
arxiv.org
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are typically obtained as...
1285
Reposted by Marco Cuturi
Josh Susskind @kindsuss.bsky.social · 23/01/2025
Here's a really cool cross-institution study leveraging optimal transport techniques developed by my Apple ML Research colleagues! It's great to see basic research in machine learning translate into scientific tools like this. Cuts into the AI hype a bit ;)
021
Reposted by Marco Cuturi
dominik1klein.bsky.social @dominik1klein.bsky.social · 23/01/2025
Missing the deep learning part? go check out the follow up work @neuripsconf.bsky.social (tinyurl.com/yvf72kzf) and @iclr-conf.bsky.social (tinyurl.com/4vh8vuzk)
tinyurl.com
GENOT: Entropic (Gromov) Wasserstein Flow Matching with...
Single-cell genomics has significantly advanced our understanding of cellular behavior, catalyzing innovations in treatments and precision medicine. However, single-cell sequencing technologies are...
0113
Reposted by Marco Cuturi
Heiko Lickert @heikolickert.bsky.social · 23/01/2025
Exciting new tool to make the most out of multimodal single-cell data from our friends and long-term collaborators of the @fabian_theis lab - check it out, it works beautifully and really makes sense out of your complex data! 👇👍😉
02710
Reposted by Marco Cuturi
Fabian Theis @fabiantheis.bsky.social · 22/01/2025
Excited to see Moscot (moscot-tools.org) published in @Nature! We scaled Optimal Transport (OT) in single-cell genomics & added multimodality together with spatiotemporal trajectory inference, finding exciting new biology in the pancreas! 🚀 Read at www.nature.com/articles/s41...
nature.com
Mapping cells through time and space with moscot - Nature
Moscot is an optimal transport approach that overcomes current limitations of similar methods to enable multimodal, scalable and consistent single-cell analyses of datasets across spatial and temporal...
212541
Reposted by Marco Cuturi
dominik1klein.bsky.social @dominik1klein.bsky.social · 23/01/2025
@AimeeBastidas, @pacotael, @MartaTarquis, @ShreyParikh07, Ilan Gold, @heikolickert.bsky.social , @mostafabakhti.bsky.social @marcocuturi.bsky.social , @fabiantheis.bsky.social
041
Reposted by Marco Cuturi
Nature @nature.com · 23/01/2025
Nature research paper: Mapping cells through time and space with moscot go.nature.com/40pqpvw
go.nature.com
Mapping cells through time and space with moscot - Nature
Moscot is an optimal transport approach that overcomes current limitations of similar methods to enable multimodal, scalable and consistent single-cell analyses of datasets across spatial and temporal dimensions.
0478
Marco Cuturi @marcocuturi.bsky.social · 22/01/2025
Today is a great day for optimal transport 🎉! Lots of gratitude 🙏 for all folks who contributed to ott-jax.readthedocs.io and pushed for the MOSCOT (now @ nature!) paper, from visionaries @dominik1klein.bsky.social, G. Palla, Z. Piran to the magician, Michal Klein! ❤️ www.nature.com/articles/s41...
nature.com
Mapping cells through time and space with moscot - Nature
Moscot is an optimal transport approach that overcomes current limitations of similar methods to enable multimodal, scalable and consistent single-cell analyses of datasets across spatial and temporal...
0227
Reposted by Marco Cuturi
Nomad Matt @nomadmatt.bsky.social · 21/01/2025
For those saying it wasn't a Nazi salute. Keep posting it. Keep sharing it.
612115328195
Reposted by Marco Cuturi
Helmholtz Munich @helmholtzmunich.bsky.social · 22/01/2025
AI in Cell Research: Unveiling Organ Development #Moscot, a groundbreaking #AI tool developed by an international team led by #HelmholtzMunich, tracks cells in real time, revealing how organs like the #pancreas form. t1p.de/viil5 @fabiantheis.bsky.social @nature.com
AI in Cell Research: Moscot Reveals Cell Dynamics in Unprecedented Detail
064
Marco Cuturi @marcocuturi.bsky.social · 18/12/2024
The Apple Machine Learning Research (MLR) team in Paris has openings for both FTE roles and a short-term post-doc position to contribute to our team's research agenda. Researchers at Apple's MLR (led by Samy Bengio) target impactful publications in top-tier ML venues and OSS.
1133
Reposted by Marco Cuturi
UniReps @unireps.bsky.social · 12/12/2024
The UniReps Workshop is happening THIS SATURDAY at #NeurIPS 🤖🧠 Join us for a day of insightful talks and discussions with @sueyeonchung.bsky.social, @eringrant.bsky.social, @leavittron.bsky.social, @itsneuronal.bsky.social, @marcocuturi.bsky.social, Philip Isola, Neel Nanda and Stefanie Jegelka! 🎤
1209
Marco Cuturi @marcocuturi.bsky.social · 10/12/2024
Check this really nice thread by @paurodriguez.bsky.social !!
010
Marco Cuturi @marcocuturi.bsky.social · 08/12/2024
I will be at hashtag#neurips2024 for the entire week from Tuesday! Please feel free to reach out if interested in Apple MLR, either in the Apple booth, one of the posters I am involved in (neurips.cc/virtual/2024..., neurips.cc/virtual/2024... neurips.cc/virtual/2024...),
neurips.cc
NeurIPS Poster Progressive Entropic Optimal Transport SolversNeurIPS 2024
181
Reposted by Marco Cuturi
Felix Petersen @petersen.ai · 28/11/2024
I'm excited to share our NeurIPS 2024 paper "Newton Losses: Using Curvature Information for Learning with Differentiable Algorithms" 🤖. Paper link 📜: arxiv.org/abs/2410.19055
arxiv.org
Newton Losses: Using Curvature Information for Learning with Differentiable Algorithms
When training neural networks with custom objectives, such as ranking losses and shortest-path losses, a common problem is that they are, per se, non-differentiable. A popular approach is to continuou...
1184
Marco Cuturi @marcocuturi.bsky.social · 22/11/2024
this is starting to feel like transfer to mastodon a couple of years ago except this time it might work? 😊
061
Reposted by Marco Cuturi
Alaa El-Nouby @alaaelnouby.bsky.social · 22/11/2024
𝗗𝗼𝗲𝘀 𝗮𝘂𝘁𝗼𝗿𝗲𝗴𝗿𝗲𝘀𝘀𝗶𝘃𝗲 𝗽𝗿𝗲-𝘁𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝘄𝗼𝗿𝗸 𝗳𝗼𝗿 𝘃𝗶𝘀𝗶𝗼𝗻? 🤔 Delighted to share AIMv2, a family of strong, scalable, and open vision encoders that excel at multimodal understanding, recognition, and grounding 🧵 paper: arxiv.org/abs/2411.14402 code: github.com/apple/ml-aim HF: huggingface.co/collections/...
35819