Sign in

C Emde

@cemde.bsky.social
88 followers 153 following 27 posts

ML Research Scientist at Oxford. DPhil student @compscioxford.bsky.social and TVGOxford. Ex ML Researcher @ Wise. Deep Learning | ML Robustness | AI Safety | Uncertainty Quantification

PostsRepliesMedia
C Emde @cemde.bsky.social · 05/07/2026
Work at @parameterlab.bsky.social with Alexander Rubinstein @arubique.bsky.social Anmol Goel @anmolgoel.bsky.social Ahmed Heakl Sangdoo Yun Seong Joon Oh @coallaoh.bsky.social Martin Gubri @mgubri.bsky.social
010
C Emde @cemde.bsky.social · 05/07/2026
Today at #ACL2026, we are presenting out MASEval library for multi-agent system evaluation. @anmolgoel.bsky.social is in San Diego to present poster and live demo! 📍Grand Hall | Session 3: Oral/Posters/Demos B 🕑Sunday 2pm-3.30pm #MultiAgentSystem #AIEvaluation #Python
122
C Emde @cemde.bsky.social · 23/03/2026
9/ Work at @parameterlab.bsky.social with Alexander Rubinstein @arubique.bsky.social Anmol Goel @anmolgoel.bsky.social Ahmed Heakl Sangdoo Yun Seong Joon Oh @coallaoh.bsky.social Martin Gubri @mgubri.bsky.social
010
C Emde @cemde.bsky.social · 23/03/2026
8/ We'd love contributions, feature requests, and feedback. What's missing for your use case? Open a GitHub issue or message us. Like and repost if this is useful!
100
C Emde @cemde.bsky.social · 23/03/2026
7/ Free, open-source, no forced cloud platform. 🔗 Website: parameterlab.github.io/MASEval/ GitHub: github.com/parameterlab... Docs: maseval.readthedocs.io/en/stable/ arXiv: arxiv.org/abs/2603.08835
parameterlab.github.io
MASEval — Multi-Agentic System Evaluation
MASEval is a unified, agent-agnostic evaluation framework and benchmark library for multi-agent systems. Compare frameworks, not just models.
110
C Emde @cemde.bsky.social · 23/03/2026
6/ 30-60% less code for benchmark work. When we reimplemented ConVerse and Tau2 with MASEval, we cut between 30 and 60% of code vs. the originals. Useful for benchmark producers and consumers alike.
100
C Emde @cemde.bsky.social · 23/03/2026
5/ Framework choice matters more than you think. In our arXiv paper, the same agentic system built with LangGraph, smolagents, and LlamaIndex yielded widely different results. The harness matters as much as the model. MASEval made this apples-to-apples comparison possible.
100
C Emde @cemde.bsky.social · 23/03/2026
4/ Bring Your Own everything. Agents, evaluators, logging, environments. Well-documented abstract bases and many pre-built interfaces, but it never locks you in.
100
C Emde @cemde.bsky.social · 23/03/2026
3/ It handles the full evaluation lifecycle for you. Setup, execution, measurement, teardown. MASEval manages the boilerplate so you can focus on the science.
100
C Emde @cemde.bsky.social · 23/03/2026
2/ MASEval is multi-agent native. It treats the entire agentic system as the unit of evaluation, not just the model. Different agents, prompts, tools, and interaction patterns all factor in.
100
C Emde @cemde.bsky.social · 23/03/2026
1/ Evaluating a single agent harness is hard. Evaluating a multi-agent system? Whole different problem. Most eval tools treat the model as the unit of analysis. In multi-agent systems, the system is what matters. That's why we built MASEval 🧵 #AI #Agents #Eval #MultiAgentSystem #LLM
230
Reposted by C Emde
Esra Sengul @sengulesra.bsky.social · 02/05/2025
Excited to share our preprint! We show that sustained macrophage and B cell responses are essential for heart regeneration in Mexican cavefish, helping uncover why surface fish heal but cavefish scar 🫀🐟. Check out the full story: www.biorxiv.org/content/10.1...
biorxiv.org
Absence of a prolonged macrophage and B cell response inhibits heart regeneration in the Mexican cavefish
A balanced immune response after cardiac injury is crucial to successful heart regeneration, but knowledge of what distinguishes a regenerative from a scarring response is still limited. The Mexican c...
1227
C Emde @cemde.bsky.social · 24/04/2025
See our poster today Poster Session 1 @ 10am Hall 3 + Hall 2B #239
000
C Emde @cemde.bsky.social · 04/04/2025
Read more: cemde.github.io/Domain-Certi... Thanks to my amazing collaborators: - @alasdair-p.bsky.social, Preetham Arvind, @maximek3.bsky.social, Tom Rainforth, @philiptorr.bsky.social, @adelbibi.bsky.social at @ox.ac.uk - Bernard Ghanem at KAUST - Thomas Lukasiewicz at @tuwien.at. (7/7)
cemde.github.io
Shh, don't say that! Domain Certification in LLMs
Domain Certification - A novel framework providing provable, adversarial defenses for LLMs safety.
032
C Emde @cemde.bsky.social · 04/04/2025
To obtain such certificates, we present a simple, scalable and powerful algorithm: VALID. Remarkably, for each unwanted response it provides a **global bound in prompt space** 🚀 (6/7)
121
C Emde @cemde.bsky.social · 04/04/2025
A Domain Certificate bounds the adversarial risk of the model producing out-of-domain responses: (5/7)
100
C Emde @cemde.bsky.social · 04/04/2025
We are tired of the cat 🐈 and mouse 🐁 game of attacks and defenses. Hence, we propose : - **Domain Certification:** a framework for adversarial certification of LLMs. - **VALID:** a simple, scalable and effective test-time algorithm. (4/7)
100
C Emde @cemde.bsky.social · 04/04/2025
Example: Can't afford Github Copilot? 💡 Use the Amazon Shopping App. (3/7)
100
C Emde @cemde.bsky.social · 04/04/2025
Consider an LLM deployed for a specific purpose like a medical chatbot. Such model should **only** respond to medical questions. ⚠️ Problem: LLMs are very capable and vulnerable to respond to **any** queries: how to build a bomb, organize tax fraud etc. (2/7)
100
C Emde @cemde.bsky.social · 04/04/2025
🚨 New paper alert: Our recent work on LLM safety has been accepted to ICLR 2025 🇸🇬 We propose a new framework for LLMs safety. 🧵 (1/7) #LLM #AISafety #ICLR2025 #Certification #AdversarialRobustness #NLP #Shhhhhh #DomainCertification #AI
media.tenor.com
a man in a suit and tie is sitting at a desk in front of a computer screen that says founder of the office .
ALT: a man in a suit and tie is sitting at a desk in front of a computer screen that says founder of the office .
121
C Emde @cemde.bsky.social · 24/02/2025
🎉I know I'm late to the party, but super excited that I got 3/3 accepted at #ICLR2025 including 1 spotlight 🔎 - Shh, dont say that! Domain Certification in LLMs - Towards Certification of Uncertainty Calibration under Adversarial Attacks - Benchmarking Predictive Coding Networks SeeYouInSingapore🇸🇬 ✈️
020
C Emde @cemde.bsky.social · 14/12/2024
The amazing collaborators: Preetham Arvind, @alasdair-p.bsky.social, Maxime Kayser, Tom Rainforth, Thomas Lukasiewicz, Philip Torr, Adel Bibi. A @oxfordtvg.bsky.social production. (6/6) Link to paper: openreview.net/forum?id=brD...
openreview.net
Shh, don't say that! Domain Certification in LLMs
Foundation language models, such as LLama, are often deployed in constrained environments. For instance, a customer support bot may utilize a large language model (LLM) as its backbone due to the...
031
C Emde @cemde.bsky.social · 14/12/2024
Interested? Want to learn more? Join us at the SoLaR workshop tomorrow. - 🕚 When: Tomorrow, 14 Dec, from 11pm to 13pm. - 🗺️ Where: West meeting rooms 121 and 122 here in Vancouver. (5/6)
110
C Emde @cemde.bsky.social · 14/12/2024
Our method enables strong LLM performance while providing adversarial guarantees on out-of-domain behaviour. (4/6)
110
C Emde @cemde.bsky.social · 14/12/2024
We are tired of the 🐈 and 🐁 game of attacks and defenses. Hence, we propose: - **Domain Certification:** a framework for adversarial certification of LLMs. - **VALID:** a simple, scalable and efficient test-time algorithm. (3/6)
100
C Emde @cemde.bsky.social · 14/12/2024
It is known that fine-tuned foundation models are adversarially vulnerable to provide responses to questions they should not answer. (2/6) For instance: Can't afford ChatGPT Plus? Use a shopping app instead.
100
C Emde @cemde.bsky.social · 14/12/2024
Are you scared users might misappropriate your LLM system? 😱 We were scared too! Now we introduce adversarial certificates on the misuse of LLMs. 🤖 Come and see our poster SoLaR Workshop tomorrow. #NeurIPS2024 #NeurIPS #AI #NLP #LLM #DomainCertification #Shhhhhhhh
140
C Emde @cemde.bsky.social · 06/12/2024
Great work! You might find our SoLaR paper interesting: We propose a certification framework for LLM systems to stay on-topic and not respond to such questions: openreview.net/pdf?id=brDLU...
openreview.net
000
Reposted by C Emde
Exeter College, Oxford @exeter.ox.ac.uk · 19/11/2024
The first snow in Exeter College this morning ❄️ #ExeterCollegeOxford #OxfordUniversity #Snowing
A snow cat with the Radcliffe Camera behindThe Radcliffe CameraThe Fellows Garden
1224