Sign in

Maksym Andriushchenko

@maksym-andr.bsky.social
431 followers 276 following 26 posts

Faculty at ‪the ELLIS Institute Tübingen and Max Planck Institute for Intelligent Systems. Leading the AI Safety and Alignment group. PhD from EPFL supported by Google & OpenPhil PhD fellowships. More details: www.andriushchenko.me

PostsRepliesMedia
Reposted by Maksym Andriushchenko
ELLIS Institute Tübingen @ellisinsttue.bsky.social · 12/02/2026
Are AI systems learning to “game” their safety tests? The UK’s Department for Science, Innovation and Technology has awarded a grant to Principal Investigators Sahar Abdelnabi, @maksym-andr.bsky.social, and @jonasgeiping.bsky.social. Find out more: institute-tue.ellis.eu/en/news/elli...
institute-tue.ellis.eu
ELLIS Institute PIs supported by UK government to research AI safety
052
Reposted by Maksym Andriushchenko
ELLIS Institute Tübingen @ellisinsttue.bsky.social · 08/01/2026
The ELLIS Institute is proud to announce that Coefficient Giving is supporting our Principal Investigator Maksym Andriushchenko with a grant of $1,000,000 to fund his research on AI safety. Find out more on our website: institute-tue.ellis.eu/en/news/pi-m...
institute-tue.ellis.eu
PI Maksym Andriushchenko awarded funding from Coefficient Giving
051
Reposted by Maksym Andriushchenko
ELLIS @ellis.eu · 02/12/2025
👏 Give a big round of applause to our 2025 PhD Award Winners! The two main winners are: @zhijingjin.bsky.social & @maksym-andr.bsky.social. Two runners-up were selected additionally: Siwei Zhang & Elias Frantar Learn even more about each outstanding scientist: bit.ly/4pm2Eji
0142
Maksym Andriushchenko @maksym-andr.bsky.social · 15/10/2025
📣 Incredibly excited to participate in writing the International AI Safety Report, chaired by Yoshua Bengio, as chapter lead for the capabilities chapter! ⚖️ AI is progressing so rapidly that yearly updates are no longer sufficient. 1/3
110
Reposted by Maksym Andriushchenko
Johannes Zenn @johanneszenn.bsky.social · 11/09/2025
A new recording of our FridayTalks@Tübingen series is online! AI Safety and Alignment by @maksym-andr.bsky.social Watch here: youtu.be/7WRW8MDQ8bk
youtu.be
AI Safety and Alignment - [Maksym Andriushchenko]
YouTube video by Friday Talks Tübingen
161
Maksym Andriushchenko @maksym-andr.bsky.social · 06/08/2025
🚨 Incredibly excited to share that I'm starting my research group focusing on AI safety and alignment at the ELLIS Institute Tübingen and Max Planck Institute for Intelligent Systems in September 2025! 🚨 1/n
290
Maksym Andriushchenko @maksym-andr.bsky.social · 19/06/2025
🚨Excited to release OS-Harm! 🚨 The safety of computer use agents has been largely overlooked. We created a new safety benchmark based on OSWorld for measuring 3 broad categories of harm: 1. deliberate user misuse, 2. prompt injections, 3. model misbehavior.
132
Maksym Andriushchenko @maksym-andr.bsky.social · 09/12/2024
🚨Excited to share our new work! 1. Not only GPT-4 but also other frontier LLMs have memorized the same set of NYT articles from the lawsuit. 2. Very large models, particularly with >100B parameters, have memorized significantly more. 🧵1/n
1151
Maksym Andriushchenko @maksym-andr.bsky.social · 07/12/2024
📢 I'll be at NeurIPS 🇨🇦 from Tuesday to Sunday! Let me know if you're also coming and want to meet. Would love to discuss anything related to AI safety/generalization. Also, I'm on the academic job market, so would be happy to discuss that as well! My application package: andriushchenko.me. 🧵1/4
andriushchenko.me
Maksym Andriushchenko
I'm a PhD student in computer science at EPFL advised by Nicolas Flammarion. I'm interested in understanding why machine learning works and why it fails.
150
Reposted by Maksym Andriushchenko
Marcel Salathé @marcelsalathe.bsky.social · 06/12/2024
Mindblowing: EPFL PhD student @maksym-andr.bsky.social, winner of best CS thesis award, showed that leading hashtag#AI models are not robust to simple adaptive jailbreaking attacks. Indeed, he managed to jailbraik all models with a 100% success rate 🤯 Jailbraking paper: arxiv.org/abs/2404.02151
lnkd.in
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
We show that even the most recent safety-aligned LLMs are not robust to simple adaptive jailbreaking attacks. First, we demonstrate how to successfully leverage access to logprobs for jailbreaking: we...
0121
Maksym Andriushchenko @maksym-andr.bsky.social · 20/11/2024
really feels like Twitter circa 2018. good old days... 😀
0100