Reposted by Jiaang LiDongyan Lin @dongyanl1n.bsky.social · 26/06/2026(1/n) Thrilled to share my first paper at Meta FAIR! "EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data" 👶 Human infants learn language from sparse, noisy multimodal input. Today's VLMs can't. We built a benchmark + challenge to close that gap. 🧵 13916
Jiaang Li @jiaangli.bsky.social · 17/04/2026🎉 Excited to share our work "𝗥𝗔𝗩𝗘𝗡𝗘𝗔: 𝗔 𝗕𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸 𝗳𝗼𝗿 𝗠𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹 𝗥𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹-𝗔𝘂𝗴𝗺𝗲𝗻𝘁𝗲𝗱 𝗩𝗶𝘀𝘂𝗮𝗹 𝗖𝘂𝗹𝘁𝘂𝗿𝗲 𝗨𝗻𝗱𝗲𝗿𝘀𝘁𝗮𝗻𝗱𝗶𝗻𝗴", accepted at #ICLR2026! 🇧🇷 I'll be attending ICLR in person — would love to connect and chat there! 🤝 🗓️ Sat, Apr 25, 2026, 10:30 AM – 1:00 PM GMT-03 📍 Pavilion 4 P4-# 3618 131
Reposted by Jiaang LiSerge Belongie @serge.belongie.com · 20/03/2026Feeling overwhelmed by all the recent developments in video understanding? What used to require dozens of modular computational workflows involving SLAM, feature tracking, optical flow, camera calibration, multiview geometric constraints, and resnet backbones is now... (1/3) 1508
Jiaang Li @jiaangli.bsky.social · 14/07/2025Feel free to reach out and chat with Xinyi on July 18th in Vancouver at the #ICML 000
Reposted by Jiaang LiSerge Belongie @serge.belongie.com · 30/03/2025Would you present your next NeurIPS paper in Europe instead of traveling to San Diego (US) if this was an option? Søren Hauberg (DTU) and I would love to hear the answer through this poll: (1/6)docs.google.comNeurIPS participation in EuropeWe seek to understand if there is interest in being able to attend NeurIPS in Europe, i.e. without travelling to San Diego, US. In the following, assume that it is possible to present accepted papers ... 6279161
Reposted by Jiaang LiSebastian Loeschcke @sloeschcke.bsky.social · 03/06/2025Check out our new preprint 𝐓𝐞𝐧𝐬𝐨𝐫𝐆𝐑𝐚𝐃. We use a robust decomposition of the gradient tensors into low-rank + sparse parts to reduce optimizer memory for Neural Operators by up to 𝟕𝟓%, while matching the performance of Adam, even on turbulent Navier–Stokes (Re 10e5). 2317
Reposted by Jiaang LiPioneer Centre for AI @aicentre.dk · 02/06/2025PhD student, Jiaang Li and his collaborators, with insights into cultural understanding of vision-language models 👇 011
Reposted by Jiaang LiSrishti @srishtiy.bsky.social · 02/06/2025I am excited to announce our latest work 🎉 "Cultural Evaluations of Vision-Language Models Have a Lot to Learn from Cultural Theory". We review recent works on culture in VLMs and argue for deeper grounding in cultural theory to enable more inclusive evaluations. Paper 🔗: arxiv.org/pdf/2505.22793 35818
Jiaang Li @jiaangli.bsky.social · 23/05/2025🚀New Preprint🚀 Can Multimodal Retrieval Enhance Cultural Awareness in Vision-Language Models? Excited to introduce RAVENEA, a new benchmark aimed at evaluating cultural understanding in VLMs through RAG. arxiv.org/abs/2505.14462 More details:👇 1177
Reposted by Jiaang LiYifei Yuan @yfyuan01.bsky.social · 21/04/2025I won’t be attending #ICLR in person this year😢. But feel free to check our paper ‘Revisiting the Othello World Model Hypothesis’ with Anders Søgaard, accepted at ICLR world models workshop! Paper link arxiv.org/abs/2503.04421arxiv.orgRevisiting the Othello World Model HypothesisLi et al. (2023) used the Othello board game as a test case for the ability of GPT-2 to induce world models, and were followed up by Nanda et al. (2023b). We briefly discuss the original experiments, ... 012
Reposted by Jiaang LiZhaochong An @zhaochongan.bsky.social · 11/02/2025Thrilled to announce "Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation" is accepted as a Spotlight (5%) at #ICLR2025! Our model MM-FSS leverages 3D, 2D, & text modalities for robust few-shot 3D segmentation—all without extra labeling cost. 🤩 arxiv.org/pdf/2410.22489 More details👇 1257
Reposted by Jiaang LiChengzu @chengzu-li.bsky.social · 14/01/2025Forget just thinking in words. 🔔Our New Preprint: 🚀 New Era of Multimodal Reasoning🚨 🔍 Imagine While Reasoning in Space with MVoT Multimodal Visualization-of-Thought (MVoT) revolutionizes reasoning by generating visual "thoughts" that transform how AI thinks, reasons, and explains itself. 161
Reposted by Jiaang LiNico Lang @nicolang.bsky.social · 09/01/2025FGVC12 Workshop is coming to #CVPR 2025 in Nashville! Are you working on fine-grained visual problems? This year we have two peer-reviewed paper tracks: i) 8-page CVPR Workshop proceedings ii) 4-page non-archival extended abstracts CALL FOR PAPERS: sites.google.com/view/fgvc12/...sites.google.comFGVC12 Workshop - SubmissionCall for Papers Workshop - Date TBC (either June 11th or 12th 2025) FGVC12 will have two paper tracks and a nectar track: Proceedings track: 8-page papers that will appear in the official CVPR worksho... 0103
Reposted by Jiaang LiSerge Belongie @serge.belongie.com · 30/12/2024Here’s a short film produced by the Danish Royal Academy of Sciences, showcasing the WineSensed 🍷 project of Þóranna Bender et al. thoranna.github.io/learning_to_...youtu.beVidenSkaber | Min AI forstår mig ikke - professor Serge BelongieYouTube video by Videnskabernes Selskab 0173
Reposted by Jiaang LiBelongie Lab @belongielab.org · 21/12/2024From San Diego to New York to Copenhagen, wishing you Happy Holidays!🎄 0404
Reposted by Jiaang LiBelongie Lab @belongielab.org · 03/12/2024With @neuripsconf.bsky.social right around the corner, we’re excited to be presenting our work soon! Here’s an overview (1/5) 1166
Reposted by Jiaang LiBelongie Lab @belongielab.org · 25/11/2024Here’s a starter pack with members of our lab that have joined Blueskygo.bsky.appBelongie LabJoin the conversation 0134
Reposted by Jiaang LiChristoph Molnar @christophmolnar.bsky.social · 24/11/2024No one can explain stochastic gradient descent better than this panda.media.tenor.coma panda bear is rolling around in the grass in a zoo enclosure .Alt: a panda bear is rolling around in the grass in a zoo enclosure . 1021632
Jiaang Li @jiaangli.bsky.social · 19/11/2024🤔Do Vision and Language Models Share Concepts? 🚀 We present an empirical evaluation and find that language models partially converge towards representations isomorphic to those of vision models. #EMNLP 📃 direct.mit.edu/tacl/article... 2267
Reposted by Jiaang LiMaria Antoniak @mariaa.bsky.social · 19/11/2024I'm recruiting 1-2 PhD students to work with me at the University of Colorado Boulder! Looking for creative students with interests in #NLP and #CulturalAnalytics. Boulder is a lovely college town 30 minutes from Denver and 1 hour from Rocky Mountain National Park 😎 Apply by December 15th! 9302136
Reposted by Jiaang LiBelongie Lab @belongielab.org · 17/11/2024Logging on! 🧑💻🦋 We're the Belongie Lab led by @sergebelongie.bsky.social. We study Computer Vision and Machine Learning, located at the University of Copenhagen and Pioneer Centre for AI. Follow along to hear about our research past and present! www.belongielab.orgbelongielab.orgBelongie Lab - HomeBelongie Lab -- Home. 0289
Reposted by Jiaang LiSerge Belongie @serge.belongie.com · 17/11/2024A new approach to training models in memory-constrained settings, LoQT allows for the pre-training of a 13B LLM on a 24GB GPU without model parallelism, checkpointing, or offloading strategies during training Code: github.com/sebulo/LoQTgithub.comGitHub - sebulo/LoQTContribute to sebulo/LoQT development by creating an account on GitHub. 2405