Sign in

Oscar Mañas

@oscmansan.bsky.social
180 followers 172 following 9 posts

Research scientist at Meta, PhD from Mila and Université de Montréal. Working on multimodal vision+language generation. Català a Zúric.

PostsRepliesMedia
Oscar Mañas @oscmansan.bsky.social · 10/06/2025
Headed to @cvprconference.bsky.social in Nashville! I'll be presenting our work on Multimodal Reward-guided Decoding. Let's connect if you're around!
020
Oscar Mañas @oscmansan.bsky.social · 15/05/2025
TFW you find a memory leak in your code two days before the rebuttal's deadline
010
Oscar Mañas @oscmansan.bsky.social · 22/04/2025
Heading to Singapore for the next 1.5 weeks for @iclr-conf.bsky.social. If you're around and want to meet up, hit me up!
010
Reposted by Oscar Mañas
Simons Institute for the Theory of Computing @simonsinstitute.bsky.social · 01/04/2025
"Tokenize Everything!" Luke Zettlemoyer of @uofwa.bsky.social on using GPT-like autoregressive techniques for training multimodal models (text, images, audio etc.) at the Simons Institute workshop on The Future of Language Models and Transformers simons.berkeley.edu/workshops/fu...
031
Oscar Mañas @oscmansan.bsky.social · 22/12/2024
I quite like this analogy by Oriol Vinyals: * LLM ~= core electric brain * Agent ~= LLM with a digital body youtu.be/78mEYaztGaw
youtu.be
Gemini 2.0 and the evolution of agentic AI with Oriol Vinyals
YouTube video by Google DeepMind
030
Oscar Mañas @oscmansan.bsky.social · 14/12/2024
Curious about how to effectively steer the behavior of multimodal LLMs during inference to improve their visual grounding? Join me today at 4:30pm at the AFM workshop at @NeurIPSConf, where I'll be presenting a poster on my work. Come by to learn more! openreview.net/forum?id=VWJ...
031
Oscar Mañas @oscmansan.bsky.social · 10/12/2024
Tomorrow at 3:15pm I'll be presenting my work at @mila-quebec.bsky.social's booth (#104) at @neuripsconf.bsky.social. Come to learn more about controlling multimodal LLMs via reward-guided decoding! 🔗 openreview.net/forum?id=VWJ...
openreview.net
Controlling Multimodal LLMs via Reward-guided Decoding
As Multimodal Large Language Models (MLLMs) gain widespread applicability, it is becoming increasingly desirable to adapt them for diverse user needs. In this paper, we study the adaptation of...
0113
Reposted by Oscar Mañas
François Fleuret @francois.fleuret.org · 05/12/2024
All this being said, Meta/FAIR remains the only place where you can do open AI research with a group of stellar colleagues ten times larger than any university + big-tech computational capabilities level.
2391