Reposted by Louis BethuneBruno Mlodozeniec @brunokm.bsky.social · 06/01/2026In our new work — Complete(d)P — we try to answer 3 questions about hyperparameter (HP) scaling: ● How to transfer across model size, tokens&batch-size?→ Complete(d)P ● Do per-module HPs matter? ✔️2x speed-ups possible ● Do they transfer to larger scale? ✔️ With the right parameterisation 184
Reposted by Louis BethuneMarco Cuturi @marcocuturi.bsky.social · 18/12/2024Both jobs are onsite in our Paris office. Please apply to this ad for the FTE position: jobs.apple.com/en-us/detail... and/or reach out to MLR_Paris_FTE@group.apple.com for questions. We are looking for candidates with a track record of papers in ML conferences and familarity with JAX / pytorchjobs.apple.comAIML - Machine Learning Researcher, MLR - Careers at AppleApply for a AIML - Machine Learning Researcher, MLR job at Apple. Read about the role and find out if it’s right for you. 121
Reposted by Louis BethuneAlaa El-Nouby @alaaelnouby.bsky.social · 22/11/2024𝗗𝗼𝗲𝘀 𝗮𝘂𝘁𝗼𝗿𝗲𝗴𝗿𝗲𝘀𝘀𝗶𝘃𝗲 𝗽𝗿𝗲-𝘁𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝘄𝗼𝗿𝗸 𝗳𝗼𝗿 𝘃𝗶𝘀𝗶𝗼𝗻? 🤔 Delighted to share AIMv2, a family of strong, scalable, and open vision encoders that excel at multimodal understanding, recognition, and grounding 🧵 paper: arxiv.org/abs/2411.14402 code: github.com/apple/ml-aim HF: huggingface.co/collections/... 35819