Reposted by Mark IbrahimKaren Ullrich (s/h) @karen-ullrich.bsky.social · 30/06/2026We are launching a new blog; Reliable-AI.Review. First post is up: On the Impossibility of Mitigating AI Jailbreaks. 195
Mark Ibrahim @markibrahim.bsky.social · 10/12/2025Want to teach AI agents to use apps like humans? Get started with digital agents research using OpenApps, our new Python-based environment. 133
Mark Ibrahim @markibrahim.bsky.social · 07/11/2025We introduce, Common-O, a new multimodal benchmark for hallucination when reasoning across scenes. We find leading multimodal LLMs can reliably identify objects, yet hallucinate when reasoning across scenes. 🧵1/3 100
Mark Ibrahim @markibrahim.bsky.social · 16/10/2025If you’re an NYU student, come learn about this wonderful opportunity to collaborate with us at FAIR events.atmeta.com/metanyuaimen... Panel is tomorrow 10am at NYU Center for Data Science.events.atmeta.com 100
Mark Ibrahim @markibrahim.bsky.social · 09/10/2025One can manipulate LLM rankings to put any model in the lead—only by modifying the single character separating demonstration examples. Learn more in our new paper arxiv.org/abs/2510.05152 w/ Jingtong Su, Jianyu Zhang, @karen-ullrich.bsky.social , and Léon Bottou. 🧵 1313
Mark Ibrahim @markibrahim.bsky.social · 21/07/2025Open-weights for our Llip multimodal vision-language model led by @lavoiems.bsky.social are public! LLIP proposes new pre-training objective to capture the many ways to describe an image leading to strong performance across a suite of 22-zero shot benchmarks. bsky.app/profile/lavo... 010
Mark Ibrahim @markibrahim.bsky.social · 17/06/2025A good language model should say “I don’t know” by reasoning about the limits of its knowledge. Our new work AbstentionBench carefully measures this overlooked skill in an open-codebase others can build on! We find frontier reasoning degrades models’ ability to know when NOT to answer. 🧵1/2 110
Mark Ibrahim @markibrahim.bsky.social · 02/05/2025Join us as a PhD research intern at FAIR w/ @polkirichenko.bsky.social and Kamalika Chaudhuri to start this summer or fall with a focus on open science into multimodal models, agents and beyond! Email polkirichenko@meta.com with the title [Prospective Intern 2025] and attach your CV if interested! 000
Mark Ibrahim @markibrahim.bsky.social · 11/12/2024Can we boost transformers’ ability to retrieve knowledge and plan in maze navigation by only tweaking the learning objective? We emphatically say YES in our #NeurIPS 2024 study! 🧵 w/ Ouail Kitouni, Niklas Nolte, Diane Bouchacourt, Adina Williams, and Mike Rabbat Paper arxiv.org/abs/2406.05183 240