Sign in

Simon Geisler

@geislersi.bsky.social
31 followers 9 following 9 posts

Interests: machine learning, algorithms, coding Current: Machine Learning PhD Student at TU Munich with Prof. Stephan Günnemann Past: @deep-mind.bsky.social, Bosch Center for Artificial Intelligence, Bosch Connected Services

PostsRepliesMedia
Reposted by Simon Geisler
Marten Lienen @martenlienen.bsky.social · 14/07/2025
Real data is noisy but HiPPO assumes it's clean. Our UnHiPPO initialization resists noise with implicit Kalman filtering and makes SSMs robust without architecture changes. Learn more at our #ICML poster: Thu 11am E-2409 Paper: openreview.net/forum?id=U8G... Code: github.com/martenlienen...
033
Reposted by Simon Geisler
Nicholas Gao @n-gao.bsky.social · 09/04/2025
I am truly excited to share our latest work with @mscherbela.bsky.social, Philipp Grohs, and @guennemann on "Accurate Ab-initio Neural-network Solutions to Large-Scale Electronic Structure Problems"! arxiv.org/abs/2504.06087
arxiv.org
Accurate Ab-initio Neural-network Solutions to Large-Scale Electronic Structure Problems
We present finite-range embeddings (FiRE), a novel wave function ansatz for accurate large-scale ab-initio electronic structure calculations. Compared to contemporary neural-network wave functions, Fi...
1185
Simon Geisler @geislersi.bsky.social · 27/02/2025
This work was possible due to my awesome coauthors @wollschlager.bsky.social , M. H. I. Abdalla, Vincent Cohen-Addad, Johannes Gasteiger, Stephan Günnemann, and the support by Center for AI Safety (CAIS) as well as Google Reserarch. 9/
030
Simon Geisler @geislersi.bsky.social · 27/02/2025
Want to dive deeper and better understand LLM robustness? Read our paper for details on our REINFORCE objective, its implementation, and experimental evaluation. Let us know what you think! 💬👇 arxiv.org/abs/2502.17254 #AI #MachineLearning #AISafety 8/
arxiv.org
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
To circumvent the alignment of large language models (LLMs), current optimization-based adversarial attacks usually craft adversarial prompts by maximizing the likelihood of a so-called affirmative re...
140
Simon Geisler @geislersi.bsky.social · 27/02/2025
The results speak for themselves! 📈 On HarmBench standard, Our REINFORCE-GCG doubles the attack success rate on Llama3 & raises it from 2% to 50% on the circuit breaker defense! This shows that our method gets closer to the estimation of the LLMs’ true robustness. 🔬 7/
140
Simon Geisler @geislersi.bsky.social · 27/02/2025
Our objective can be used with various attack algorithms. In the context of jailbreaks and with appropriate reward, maximizing the reward turns out to be equivalent to maximizing the probability of the model responding in a harmful manner. 6/
130
Simon Geisler @geislersi.bsky.social · 27/02/2025
What makes our REINFORCE objective better? ① Adaptive: Tailored to the specific LLM being attacked. ② Distributional: Considers the model’s distribution of responses. ③ Semantic: Focuses on genuinely harmful behavior (LLM-as-a-judge), not just a static starting phrase. 5/
130
Simon Geisler @geislersi.bsky.social · 27/02/2025
Our solution? REINFORCE adversarial attacks! 💪 We use reinforcement learning principles to guide the attack towards semantically harmful outcomes, considering the distribution of possible LLM responses. It dynamically adapts to the model and targets its actual generation. 🎯 4/
140
Simon Geisler @geislersi.bsky.social · 27/02/2025
Considering that the attack objective navigates the LLM during the attack, in terms of real-world navigation systems, the affirmative objective is like the instruction “exit the driveway” but lacks instructions about getting to your actual destination. 🚗💥3/
130
Simon Geisler @geislersi.bsky.social · 27/02/2025
What's the problem with existing #LLM adversarial attacks? 🤔 They often just try to make the LLM start its response inappropriately (affirmative objective). The LLM might neither complete the response in a harmful way nor does the target adapt to model-specific preferences. 2/
130
Simon Geisler @geislersi.bsky.social · 27/02/2025
Do you think your LLM is robust?⚠️With current adversarial attacks it is hard to find out since they optimize the wrong thing! We fix this with our adaptive, semantic, and distributional objective. By Günnemann's lab @ TU Munich's lab & Google Research, w/ CAIS support Here's how we did it. 🧵
152
Reposted by Simon Geisler
Lukas Gosch @lukasgosch.bsky.social · 14/12/2024
Super happy & honored that our work on certifying NNs against poisoning won the Best Paper Award at AdvML-Frontiers@ #NeurIPS2024. Come by our poster 10:40am-12&4-5pm (or talk) today :) Joint work w/ Mahalakshmi Sabanayagam, Debarghya Ghoshdastidar & Stephan Günnemann L: arxiv.org/pdf/2407.10867
043
Reposted by Simon Geisler
Nicholas Gao @n-gao.bsky.social · 09/12/2024
Excited to present our work on Neural Pfaffians at #NeurIPS. 🗣️ Oral: Friday 3:30pm, East Ballroom A, B 📊 Post: Friday 4:30pm - 7:30pm, East Exhibit Hall A-C #3600 📝 Paper: openreview.net/forum?id=HRk... Happy to chat!
0114