Sign in

André Cruz

@andcrz.bsky.social
160 followers 481 following 3 posts

🎓 PhD student at the Max Planck Institute for Intelligent Systems 🔬 Safe and robust AI, algorithms and society 🔗 andrefcruz.github.io 📍 researcher in 🇩🇪, from 🇵🇹

PostsRepliesMedia
Reposted by André Cruz
Yatong Chen @yatongchen.bsky.social · 03/07/2026
LLM leaderboards aren't passive measurements — they're mechanisms that create incentives! Why does the current evaluation protocol reward "benchmaxxing": post-training for the benchmark without improving model capability? Our #ICML2026 paper gives a theoretical answer! 🧵
2144
André Cruz @andcrz.bsky.social · 26/06/2026
Excited to share that I've started my summer internship at @msftresearch.bsky.social NYC, working with Solon Barocas in the FATE team! If you're in the city and want to grab a coffee or talk research, reach out 👋
011
Reposted by André Cruz
Florian Dorner @flodorner.bsky.social · 23/04/2026
At ICLR and interested in theory for LLMs? Join us at our poster to learn more about the (im)possibility of scaling laws for test-time scaling methods like Best-of-N when verification is imperfect!
132
Reposted by André Cruz
Shubhendu Trivedi @shubhendu.bsky.social · 25/03/2026
There is now a whole sub-industry around LLM routing, gateways, even compute arbitrageurs (e.g. inference dot net). This is a basic but nice study on arbitrage in such settings and implications for the ecosystem (e.g. price drops, market entry, revenue cannibalization etc.) arxiv.org/abs/2603.22404
arxiv.org
Computational Arbitrage in AI Model Markets
Consider a market of competing model providers selling query access to models with varying costs and capabilities. Customers submit problem instances and are willing to pay up to a budget for a verifi...
021
Reposted by André Cruz
Florian Dorner @flodorner.bsky.social · 05/12/2025
Meet me at the Benchmarking workshop (sites.google.com/view/benchma...) at EurIPS on Saturday: We’ll present two works on errors in LLM-as-Judge and their impacts on benchmarking and test-time-scaling:
173
Reposted by André Cruz
Yatong Chen @yatongchen.bsky.social · 22/09/2025
We (w/ Moritz Hardt, Olawale Salaudeen and @joavanschoren.bsky.social) are organizing the Workshop on the Science of Benchmarking & Evaluating AI @euripsconf.bsky.social 2025 in Copenhagen! 📢 Call for Posters: rb.gy/kyid4f 📅 Deadline: Oct 10, 2025 (AoE) 🔗 More info: rebrand.ly/bg931sf
1217
Reposted by André Cruz
Stand Up for Science! @standupforscience.net · 12/02/2025
Welcome to the Bluesky account for Stand Up for Science 2025! Keep an eye on this space for updates, event information, and ways to get involved. We can't wait to see everyone #standupforscience2025 on March 7th, both in DC and locations nationwide! #scienceforall #sciencenotsilence
289114575402
André Cruz @andcrz.bsky.social · 06/02/2025
Tomorrow at 1:30pm ET at the Harvard EconCS seminar, I'm presenting our paper on LLMs as risk scorers: We build benchmarks using US Census data & show how miscalibrated LLMs are on real-world tabular data distributions. 📍Harvard SEC LL2.221-open to the public econcs.seas.harvard.edu/event/spring...
econcs.seas.harvard.edu
Spring EconCS 2025 Seminars | EconCS Group
121