Sign in

Parameter Lab

@parameterlab.bsky.social
35 followers 8 following 37 posts

Empowering individuals and organisations to safely use foundational AI models. parameterlab.de

PostsRepliesMedia
Parameter Lab @parameterlab.bsky.social · 23/03/2026
If you care about rigorous evaluation of agentic systems, give it a look at MASEval! The harness is an important element of agents. MASEval makes it straightforward to change its components and evaluate their impact. MASEval is our first software! parameterlab.github.io/MASEval/ ⬇️
parameterlab.github.io
MASEval — Multi-Agentic System Evaluation
MASEval is a unified, agent-agnostic evaluation framework and benchmark library for multi-agent systems. Compare frameworks, not just models.
020
Parameter Lab @parameterlab.bsky.social · 03/02/2026
‼️New paper from Parameter Lab! ⛓️‍💥 We identify privacy collapse, a silent failure mode of LLMs: LLMs fine-tuned on seemingly benign data can lose their ability to respect contextual privacy norms. Done by @anmolgoel.bsky.social during his internship! Check-out 👇
051
Parameter Lab @parameterlab.bsky.social · 28/01/2026
👏 Proud to share that the paper that Ahmed Heakl authored during his internship at Parameter Lab was accepted at #ICLR2026! See how 🩺Dr.LLM increases accuracy and decreases inference computations of frozen LLMs: www.linkedin.com/posts/ahmed-...
linkedin.com
#llm #ai #efficientai #nlp #mlresearch #reasoning #adaptivecompute | Ahmed Heakl | 19 comments
Super excited to share the last work from my internship in Germany 🇩🇪! 🚀 Dr.LLM: Dynamic Layer Routing for LLMs > What if we can reduce computation AND increase accuracy? 🤯 Most prompts don’t need ev...
040
Reposted by Parameter Lab
Martin Gubri @mgubri.bsky.social · 04/11/2025
Our #EMNLP2025 paper Leaky Thoughts 🫗 shows that Large Reasoning Models (LRMs) can unintentionally leak sensitive information hidden in their internal thoughts. 📍 Come chat with Tommaso at our poster on Friday 7th, 10:30–12:00 in Hall C3 📄 aclanthology.org/2025.emnlp-m...
aclanthology.org
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
Tommaso Green, Martin Gubri, Haritz Puerto, Sangdoo Yun, Seong Joon Oh. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
021
Parameter Lab @parameterlab.bsky.social · 21/08/2025
We challenge the view that reasoning traces are a safe internal part of a model’s process. Our work shows they can leak information, through both deliberate attacks and accidental leakage. RTAI: researchtrend.ai/papers/2506.... ArXiv: arxiv.org/abs/2506.15674 Code: github.com/parameterlab... 2/2
researchtrend.ai
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
We study privacy leakage in the reasoning traces of large reasoning models used as personal agents. Unlike final outputs, reasoning traces are often assume...
010
Parameter Lab @parameterlab.bsky.social · 21/08/2025
🫗 An LLM's "private" reasoning may leak your sensitive data! 🎉 Excited to share our paper "Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers" was accepted at #EMNLP main! 1/2
Overall diagram about contextual privacy & LRMs
151
Parameter Lab @parameterlab.bsky.social · 23/06/2025
Work done with: Haritz Puerto, Martin Gubri ‪‪@mgubri.bsky.social‬ , Tommaso Green, Sangdoo Yun and Seong Joon Oh @coallaoh.bsky.social‬ #SEO #AI #LLM #GenerativeAI #Marketing #DigitalMarketing #Perplexity #NLProc
010
Parameter Lab @parameterlab.bsky.social · 23/06/2025
Key takeaways: ❌ C-SEO doesn’t help improve visibility in AI answers. 🔎 Traditional SEO is your tool for online visibility. 🚀 Our benchmark sets the stage to develop C-SEO methods that might work in the future.
100
Parameter Lab @parameterlab.bsky.social · 23/06/2025
🔎 The results are clear: current C-SEO strategies don’t work. This challenges the recent hype and suggests that creators don’t need to game LLMs and create even more clickbaits. Just focus on producing genuinely good content and let traditional SEO do its work.
100
Parameter Lab @parameterlab.bsky.social · 23/06/2025
C-SEO Bench evaluates Conversational Search Engine Optimization (C-SEO) techniques on two key tasks: 🔍 Product Recommendation ❓ Question Answering Spanning multiple domains, it tests both domain-specific performance and the generalization of C-SEO methods.
100
Parameter Lab @parameterlab.bsky.social · 23/06/2025
💥 With the rise of conversational search, a new technique of "Conversational SEO" (C-SEO) emerged, claiming it can boost content inclusion in AI-generated answers. We put these claims to the test by building C-SEO Bench, the first comprehensive benchmark to rigorously evaluate these new strategies.
 Illustration of a conversational search engine for product recommendation. After applying a C-SEO method on the third document, its ranking gets boosted by +2 positions.
100
Parameter Lab @parameterlab.bsky.social · 23/06/2025
🔎Does Conversational SEO actually work? Our new benchmark has an answer! Excited to announce our new paper: C-SEO Bench: Does Conversational SEO Work? 🌐 RTAI: researchtrend.ai/papers/2506.... 📄 Paper: arxiv.org/abs/2506.11097 💻 Code: github.com/parameterlab... 📊 Data: huggingface.co/datasets/par...
Paper thumbnail.
121
Parameter Lab @parameterlab.bsky.social · 26/04/2025
Excited to share that our paper "Scaling Up Membership Inference: When and How Attacks Succeed on LLMs" will be presented next week at #NAACL2025! 🖼️ Catch us at Poster Session 8 - APP: NLP Applications 🗓️ May 2, 11:00 AM - 12:30 PM 🗺️ Hall 3 Hope to see you there!
021
Parameter Lab @parameterlab.bsky.social · 14/02/2025
Ready to Join? Send your resume + a short note on why you’re a great fit to recruit@parameterlab.de. Be part of a team that’s redefining research with AI! #Hiring #DataEngineer #AI #RemoteJobs
000
Parameter Lab @parameterlab.bsky.social · 14/02/2025
Why Join Us? 🚀 Make a Difference – Your work directly enhances how research is shared and discovered. 🌍 Flexibility – Choose full-time or part-time, work remotely or locally. ⚡ Innovative Environment – AI, research, and data-driven solutions all in one place. 🤝 Great Team
100
Parameter Lab @parameterlab.bsky.social · 14/02/2025
What You Bring: ✅ Proficiency in Airflow & PostgreSQL – Complex workflows and databases. ✅ Strong Python Skills – Clean, efficient, and maintainable code is your thing. ✅ (Bonus) Experience with LLMs – A huge plus as we integrate AI-driven solutions. ✅ Problem-Solving Mindset ✅ Team Spirit
100
Parameter Lab @parameterlab.bsky.social · 14/02/2025
What You’ll Do: ✔ Build Scalable Data Pipelines – Design and optimize workflows using tools like Airflow. ✔ Work Closely with AI Experts & Engineers – Collaborate to solve real-world data challenges. ✔ Optimize and Maintain Systems – Keep our data infrastructure fast, secure, and adaptable.
100
Parameter Lab @parameterlab.bsky.social · 14/02/2025
Our LLM-powered ecosystem also bridges the gap between cutting-edge research and industry leaders. If you're passionate about data, AI, and making an impact, we’d love to have you on board!
100
Parameter Lab @parameterlab.bsky.social · 14/02/2025
👥 We're Hiring: Senior/Junior Data Engineer! 📍 Remote or Local | Full-Time or Part-Time At ResearchTrend.AI, we’re building a platform that connects researchers and AI engineers worldwide—helping them stay ahead with daily digests, insightful summaries, and interactive events.
researchtrend.ai
ResearchTrend.AI
Explore the most trending research topics in AI
120
Parameter Lab @parameterlab.bsky.social · 06/02/2025
🔎 Wonder how to prove an LLM was trained on a specific text? The camera ready of our Findings of #NAACL 2025 paper is available! 📌 TLDR: longs texts are needed to gather enough evidence to determine whether specific data points were included in training of LLMs: arxiv.org/abs/2411.00154
arxiv.org
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
Membership inference attacks (MIA) attempt to verify the membership of a given data sample in the training set for a model. MIA has become relevant in recent years, following the rapid development of ...
051
Parameter Lab @parameterlab.bsky.social · 23/01/2025
We are delighted to announce that our research paper on the scale of LLM membership inference has been accepted for publication in the Findings of #NAACL2025! 🎉
040
Reposted by Parameter Lab
Seong Joon Oh @coallaoh.bsky.social · 22/11/2024
There's an internship opening at @parameterlab.bsky.social : parameterlab.de/careers The research outputs have been quite successful so far: researchtrend.ai/organization...
parameterlab.de
Careers | Parameter Lab
Join us at Parameter Lab to shape the future of safe AI. In our dynamic and inclusive environment, we focus not only on our mission but also on fostering your personal growth through rewarding work ex...
262
Parameter Lab @parameterlab.bsky.social · 20/11/2024
🎉We’re pleased to share the release of the models from our Apricot🍑 paper, accepted at ACL 2024! At Parameter Lab, we believe openness and reproducibility are essential for advancing science, and we've put in our best effort to ensure it. 🤗 huggingface.co/collections/... 🧵 bsky.app/profile/dnns...
huggingface.co
🍑 Apricot Models - a parameterlab Collection
Fine-tuned models for black-box LLM calibration, trained for "Apricot: Calibrating Large Language Models Using Their Generations Only" (ACL 2024)
093
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🔗 Links: Code and results github.com/parameterlab/mia-scaling Project Website: haritzpuerto.github.io/scaling-mia Paper: arxiv.org/pdf/2411.00154
github.com
GitHub - parameterlab/mia-scaling: Source code of "Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models"
Source code of "Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models" - parameterlab/mia-scaling
000
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🙌 Team Credits: This research was conducted by Haritz Puerto @mgubri.bsky.social @oodgnas.bsky.social and @coallaoh.bsky.social with support from NAVER AI Lab. Stay tuned for more updates! 🚀
110
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🤓 Want More? Check out the community page of MIA for LLMs in ReserachTrend.AI researchtrend.ai/communities/MIALM You can see related works, the evolution of the community, and top authors!
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
💬 What Do You Think? Could MIA reach a level where data owners use it as legal evidence? How might this affect LLM deployment? Let us know! #AI #LLM #NLProc
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🌐 Implications for Data Privacy: Our findings have real-world relevance for data owners worried about unauthorized use of their content in model training. It can also be used to support accountability of LLM evaluation in end-tasks.
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🔎 Better Results in Fine-Tuning: Fine-tuned models show even stronger MIA results. The table shows the performance at sentence level and for collections of 20 sentences, evaluated on Phi-2 fine-tuned for QA (huggingface.co/haritzpuerto/phi-2-d… ).
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🔬 Our Testing Setup: We ran experiments using Pythia models (2.8B and 6.9B parameters) with training samples from The Pile dataset, comparing them to validation and test sets. This setup avoids data leakage to ensure a reliable evaluation of MIA.
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🚀 The Key? Number of tokens & Aggregation: MIA’s accuracy improves as we aggregate MIA scores across multiple paragraphs. Longer documents or larger document collections significantly boost MIA effectiveness.
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🛠️ First Success on Pre-Trained LLMs: By adapting recent work on Dataset Inference (@pratyushmaini.bsky.social ), we successfully applied MIA on pre-trained LLMs. Check out our figure below: MIA achieves an AUROC of 0.75 on documents of up to 20k tokens!
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🔍 New Benchmark, New Insights: We developed a new benchmark to assess MIA effectiveness across data scales, from single sentences to document collections. This lets us identify precisely when and how MIA succeeds on LLMs.
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🚨 What’s MIA? It’s a method to detect if a specific data sample was used in model training. We show that MIA works effectively on long documents (~20k tokens) and in collections of documents (>100 docs) —just the scale relevant for legal applications! 👮
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
It is assumed that Membership Inference Attacks (MIA) do not work on LLMs, but our new paper shows it can work at the right scale! MIA is effective if the number of input tokens is large enough, such as in long documents and collections of them.
100
Parameter Lab @parameterlab.bsky.social · 19/11/2024
🚨📄 Exciting new research! Discover when and at what scale we can detect if specific data was used in training LLMs — a method known as Membership Inference (MIA)! Our findings open new doors for using MIA as potential legal evidence in AI. 🧵 arxiv.org/abs/2411.00154
151
Parameter Lab @parameterlab.bsky.social · 18/11/2024
📄 We’ve been working on the calibration of confidence score for black-box LLMs. See the thread below for an overview of the 🍑 Apricot paper, proudly accepted at #ACL24! #uncertainty
031
Parameter Lab @parameterlab.bsky.social · 18/11/2024
Check out one of our latest papers about LLM fingerprinting!
000
Parameter Lab @parameterlab.bsky.social · 18/11/2024
We are excited to join Bluesky! At Parameter Lab, we're committed to enhancing AI safety and trustworthiness. Our research addresses privacy, copyright, and security challenges in foundational models. Follow us for insights and updates on trustworthy AI research!
000