Sign in

Ameya P.

@bayesiankitten.bsky.social
600 followers 129 following 15 posts

Postdoctoral Researcher @ Bethgelab, University of Tübingen Benchmarking | LLM Agents | Data-Centric ML | Continual Learning | Unlearning drimpossible.github.io

PostsRepliesMedia
Reposted by Ameya P.
ELLIOT Project @elliot-eu.bsky.social · 24/06/2025
🚀 A new era in European #AIresearch begins! ELLIOT is a €25M #HorizonEurope project launching July 2025 to build open, trustworthy Multimodal Generalist Foundation Models. 30 partners, 12 countries, EU values. 🔗 Press release: apigateway.agilitypr.com/distribution...
0133
Reposted by Ameya P.
Andreas Geiger @andreasgeiger.bsky.social · 14/04/2025
🚀 Never miss a beat in science again! 📬 Scholar Inbox is your personal assistant for staying up to date with your literature. It includes: visual summaries, collections, search and a conference planner. Check out our white paper: arxiv.org/abs/2504.08385 #OpenScience #AI #RecommenderSystems
19420
Reposted by Ameya P.
Andreas Hochlehnert @ahochlehnert.bsky.social · 10/04/2025
🧵1/ 🚨 New paper: A Sober Look at Progress in Language Model Reasoning We re-evaluate recent SFT and RL models for mathematical reasoning and find most gains vanish under rigorous, multi-seed, standardized evaluation. 📊 bethgelab.github.io/sober-reason... 📄 arxiv.org/abs/2504.07086
1145
Reposted by Ameya P.
arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 10/04/2025
Hochlehnert, Bhatnagar, Udandarao, Albanie, Prabhu, Bethge: A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility arxiv.org/abs/2504.07086 arxiv.org/pdf/2504.07086 arxiv.org/html/2504.07086
112
Ameya P. @bayesiankitten.bsky.social · 10/04/2025
Great work! A much-needed upgrade for continual learning datasets—excited to see progress on long-timespan tasks beyond classification. Deets below👇
010
Ameya P. @bayesiankitten.bsky.social · 12/03/2025
Deadline extended to March 19 for the EVAL-FoMo workshop @cvprconference.bsky.social! We welcome submissions (incl. published papers) analyzing emerging capabilities & limits in visual foundation models. Details: sites.google.com/view/eval-fo... #CVPR2025
sites.google.com
EVAL-FoMo 2
A Vision workshop on Evaluations and Analysis
031
Ameya P. @bayesiankitten.bsky.social · 28/02/2025
LMs excel at solving problems (~48% success) but falter at debunking them (<9% counterexample rate)! Could form an AI Brandolini's Law: "Capability needed to refute bullshit is far larger than that needed to generate it"
020
Reposted by Ameya P.
shiven-s.bsky.social @shiven-s.bsky.social · 28/02/2025
AI can generate correct-seeming hypotheses (and papers!). Brandolini's law states BS is harder to refute than generate. Can LMs falsify incorrect solutions? o3-mini (high) scores just 9% on our new benchmark REFUTE. Verification is not necessarily easier than generation 🧵
142
Reposted by Ameya P.
Hilde Kuehne @hildekuehne.bsky.social · 19/02/2025
🚀 Call for Papers – CVPR 3rd Workshop on Multi-Modal Foundation Models (MMFM) @cvprconference.bsky.social ! 🚀 🔍 Topics: Multi-modal learning, vision-language, audio-visual, and more! 📅 Deadline: March 14, 2025 📝 Submission: cmt3.research.microsoft.com/MMFM2025 🌐 sites.google.com/view/mmfm3rd...
cmt3.research.microsoft.com
Conference Management Toolkit - Login
Microsoft's Conference Management Toolkit is a hosted academic conference management system. Modern interface, high scalability, extensive features and outstanding support are the signatures of Micros...
162
Reposted by Ameya P.
Prasanna Mayilvahanan @prasannamayil.bsky.social · 18/02/2025
New preprint out! 🎉 How does LLM training loss translate to downstream performance? We show that pretraining data and tokenizer shape loss-to-loss scaling, while architecture and other factors play a surprisingly minor role! brendel-group.github.io/llm-line/ 🧵1/8
1188
Ameya P. @bayesiankitten.bsky.social · 17/02/2025
CuratedThoughts: Data curation focus for RL post-training! (Update 1) 🚀 25% of Openthoughts-114k-math filtered — issues included proofs, missing figures, and multiple questions with one answer. Check out work by @ahochlehnert.bsky.social & @hrdkbhatnagar.bsky.social below 👇
020
Reposted by Ameya P.
A. Sophia Koepke @askoepke.bsky.social · 13/02/2025
Our 2nd Workshop on Emergent Visual Abilities and Limits of Foundation Models (EVAL-FoMo) is accepting submissions. We are looking forward to talks by our amazing speakers that include @saining.bsky.social, @aidanematzadeh.bsky.social, @lisadunlap.bsky.social, and @yukimasano.bsky.social. #CVPR2025
073
Ameya P. @bayesiankitten.bsky.social · 12/02/2025
🔥 #CVPR2025 Submit your cool papers to Workshop on Emergent Visual Abilities and Limits of Foundation Models 📷📷🧠🚀✨ sites.google.com/view/eval-fo... Submission Deadline: March 12th!
sites.google.com
EVAL-FoMo 2
A Vision workshop on Evaluations and Analysis
032
Ameya P. @bayesiankitten.bsky.social · 07/02/2025
LMs are used for annotation, evaluation and distillation! We identify critical issues! LMs of a similar capability class (not model family tho!) behave similarly and this skews oversight far more than I expected. Check the 4-in-1 mega paper below to 👀 how 👇
020
Ameya P. @bayesiankitten.bsky.social · 13/12/2024
New Work: RanDumb!🚀 Poster @NeurIPS, East Hall #1910- come say hi👋 Core claim: Random representations Outperform Online Continual Learning Methods! How: We replace the deep network by a *random projection* and linear clf, yet outperform all OCL methods by huge margins [1/n]
110
Reposted by Ameya P.
Alfredo Canziani @alfcnz.bsky.social · 12/12/2024
The Practitioner's Guide to Continual Multimodal Pretraining @dziadzio.bsky.social @confusezius.bsky.social @vishaalurao.bsky.social @bayesiankitten.bsky.social
0244
Ameya P. @bayesiankitten.bsky.social · 11/12/2024
Breaking the 8-model merge limit was tough, but we scaled to merging 200+ models! The secret? Iterative finetuning + merging *over time*. The time axis unlocks scalable mergeability. Merging has surprising scaling gains across size & compute budgets. All the gory details ⬇️
010
Ameya P. @bayesiankitten.bsky.social · 10/12/2024
How do we benchmark the vast capabilities of foundation models? Introducing ONEBench – a unifying benchmark to test them all, led by @adhirajghosh.bsky.social and @dziadzio.bsky.social!⬇️ Sample-level benchmarks could be the new generation- reusable, recombinable & evaluate lots of capabilities!
021
Ameya P. @bayesiankitten.bsky.social · 10/12/2024
Come chat with us @ NeurIPS for hot takes on the future of continual learning with foundation models!
010
Reposted by Ameya P.
Nezihe Merve Gürel @nmervegurel.bsky.social · 03/12/2024
The list of accepted workshops for ICLR 2025 is available at openreview.net/group?id=ICL... @iclr-conf.bsky.social We received 120 wonderful proposals, with 40 selected as workshops.
openreview.net
ICLR 2025 Workshop Proposals
Welcome to the OpenReview homepage for ICLR 2025 Workshop Proposals
15715