Sign in

Zara Siddique

@zarasiddique.bsky.social
153 followers 666 following 33 posts

Working on ethics and bias in NLP @CardiffNLP #NLP #NLProc

PostsRepliesMedia
Zara Siddique @zarasiddique.bsky.social · 23/10/2025
Did it.. work?!
110
Zara Siddique @zarasiddique.bsky.social · 26/07/2025
Loved giving my second tutorial on steering vectors at #CardiffNLPWorkshop Lots of enthusiastic participants! @cardiffnlp.bsky.social
042
Zara Siddique @zarasiddique.bsky.social · 14/07/2025
#CardiffNLPWorkshop off to a flying start with talks from Jennifer Foster and Marianna Apidianaki @cardiffnlp.bsky.social
042
Zara Siddique @zarasiddique.bsky.social · 26/05/2025
Pleased to say this has been accepted to ACL System Demos :)
0101
Zara Siddique @zarasiddique.bsky.social · 21/05/2025
Come to my hackathon! Last one was super fun I promise
032
Zara Siddique @zarasiddique.bsky.social · 14/05/2025
Shoutout to supervisors Liam Turner and Luis Espinosa-Anke and @cardiffnlp.bsky.social. I'm also interested in future collaborations on the topic so please message if you are interested :)
000
Zara Siddique @zarasiddique.bsky.social · 14/05/2025
I highly encourage people to play around, you can get started in just a few lines. Here's a Colab notebook: tinyurl.com/yysmb45c Note that the results from this Colab won't be the best because it's using a smaller model to reduce loading times. I would recommend using at least a 7B.
drive.google.com
Dialz Tutorial - Zara Siddique - KnitTogether 2025.ipynb
Colab notebook
100
Zara Siddique @zarasiddique.bsky.social · 14/05/2025
As part of our validation, we see if we can reduce stereotypicality in outputs from Mistral 7B, using GPT-4o as a judge. There is a notable reduction compared to baselines and prompting, which is cool.
100
Zara Siddique @zarasiddique.bsky.social · 14/05/2025
For those that are new to the topic, steering vectors are constructed using a set of paired sentences, where one elicits a 'positive' activation of neurons and the other elicits a 'negative' activation of neurons - by taking the difference, we isolate activations responsible for a certain 'concept'.
110
Zara Siddique @zarasiddique.bsky.social · 14/05/2025
🚨 NEW PAPER ALERT 🚨 Dialz: A Python Toolkit for Steering Vectors ArXiv: arxiv.org/abs/2505.06262 Docs: cardiffnlp.github.io/dialz/ Repo: github.com/cardiffnlp/d... A Python package to help you create, apply and visualise steering vectors for anything you want - from sycophancy to bias.
2268
Zara Siddique @zarasiddique.bsky.social · 13/05/2025
New friends! Old friends! Please register if you’d like 2 whole days packed with NLP fun
022
Zara Siddique @zarasiddique.bsky.social · 03/04/2025
Super interesting!
010
Zara Siddique @zarasiddique.bsky.social · 27/03/2025
Love this take: "Society appears far more willing to critically examine and address bias in AI systems than confront human bias directly"
theguardian.com
Could AI help us build a more racially just society? | Sanmi Koyejo
We have an opportunity to build systems that don’t just replicate our current inequities. Will we take them?
010
Zara Siddique @zarasiddique.bsky.social · 25/03/2025
I’d hire you
110
Reposted by Zara Siddique
Dustin Wright @dustinbwright.com · 25/03/2025
I am still in need of emergency reviewers for ARR this cycle for the computational social science track, please DM me if you have capacity 🙏
026
Zara Siddique @zarasiddique.bsky.social · 25/03/2025
Do it! When interviewers ask me about them it’s usually a good sign that it’s a nice workplace.
110
Zara Siddique @zarasiddique.bsky.social · 13/03/2025
The work presents the first systematic investigation of steering vectors for bias mitigation, and we demonstrate that SVE is a powerful and computationally efficient strategy for reducing bias in LLMs, with broader implications for enhancing AI safety.
010
Zara Siddique @zarasiddique.bsky.social · 13/03/2025
Building on these promising results, we introduce Steering Vector Ensembles (SVE), a method that averages multiple individually optimized steering vectors, each targeting a specific bias axis such as age, race, or gender.
110
Zara Siddique @zarasiddique.bsky.social · 13/03/2025
When optimized on the BBQ dataset, our individually tuned steering vectors achieve average improvements of 12.2%, 4.7%, and 3.2% over the baseline for Mistral, Llama, and Qwen, respectively.
110
Zara Siddique @zarasiddique.bsky.social · 13/03/2025
We present a novel approach to bias mitigation in large language models (LLMs) by applying steering vectors to modify model activations in forward passes. We employ Bayesian optimization to systematically identify effective contrastive pair datasets across nine bias axes.
130
Zara Siddique @zarasiddique.bsky.social · 13/03/2025
NEW PAPER 📜 Shifting Perspectives: Steering Vector Ensembles for Robust Bias Mitigation in LLMs ArXiv: arxiv.org/abs/2503.05371 GitHub: github.com/groovychoons... Extremely Unofficial Blog Post: zarasiddique.com/blog/shiftin...
arxiv.org
Shifting Perspectives: Steering Vector Ensembles for Robust Bias Mitigation in LLMs
We present a novel approach to bias mitigation in large language models (LLMs) by applying steering vectors to modify model activations in forward passes. We employ Bayesian optimization to systematic...
1175
Zara Siddique @zarasiddique.bsky.social · 10/03/2025
Strongly encourage you to register for our free NLP workshop, previously had speakers from DeepMind, Microsoft, Amazon and top university NLP labs etc. and it’s looking like it’s going to be a great line up this year too. If you can’t make it, please share with others who may be interested!
011
Reposted by Zara Siddique
Cardiff NLP @cardiffnlp.bsky.social · 05/03/2025
We've created a Cardiff NLP Starter Pack to make it easy to follow #NLP researchers at Cardiff Uni.
064
Zara Siddique @zarasiddique.bsky.social · 03/03/2025
Super interesting work!
000
Reposted by Zara Siddique
404 Media @404media.co · 29/01/2025
OpenAI furious DeepSeek might have stolen all the data OpenAI stole from us 🔗 www.404media.co/openai-furio...
1073681929
Zara Siddique @zarasiddique.bsky.social · 26/01/2025
Severance episode 2, Traitors final AND this It's a weekend of watching for me 🍿
youtube.com
Joy Buolamwini and Sam Altman: Unmasking the Future of AI
YouTube video by Commonwealth Club World Affairs (CCWA)
010
Zara Siddique @zarasiddique.bsky.social · 26/01/2025
Need to be spending less time on deepseek and more time on deep sleep 😴
020
Zara Siddique @zarasiddique.bsky.social · 23/01/2025
Any thoughts to whether this would extend well to more traditional CompSci courses?
100
Zara Siddique @zarasiddique.bsky.social · 23/01/2025
Welcome Wikipedian! And totally agree.
010
Zara Siddique @zarasiddique.bsky.social · 23/01/2025
+1 I would also like to see this.
100
Reposted by Zara Siddique
Joanne Boisson @joanneboisson.bsky.social · 20/01/2025
Our paper on extraction of metaphoric analogies from literary texts will be presented in COLING ( Wed 11:00 , Atrium, poster) by @camachocollados.bsky.social and Luis Espinosa-Anke. Done with @zarasiddique.bsky.social , @hsuvas.bsky.social and @antypasd.bsky.social
164
Zara Siddique @zarasiddique.bsky.social · 21/01/2025
If you ever find yourself getting worried about catastrophic AI risk, remember you can always just turn off the AI at the plug
000
Reposted by Zara Siddique
The Guardian @theguardian.com · 16/01/2025
I knew one day I’d have to watch powerful men burn the world down – I just didn’t expect them to be such losers | Rebecca Shaw
theguardian.com
I knew one day I’d have to watch powerful men burn the world down – I just didn’t expect them to be such losers | Rebecca Shaw
Elon Musk and Mark Zuckerberg’s desperation to be cool as they suck up to Donald Trump is so cringe it makes my skin crawl I don’t know if anyone else has noticed this but everything seems to be going down the tubes quite fast. And not fun tubes, like at…
992941891
Zara Siddique @zarasiddique.bsky.social · 16/01/2025
arxiv.org/abs/2304.13734 This is an interesting one to add
arxiv.org
The Internal State of an LLM Knows When It's Lying
While Large Language Models (LLMs) have shown exceptional performance in various tasks, one of their most prominent drawbacks is generating inaccurate or false information with a confident tone. In th...
110
Zara Siddique @zarasiddique.bsky.social · 31/12/2024
Wrote a little blog post on my favourite EMNLP papers :) zarasiddique.com/blog/my-favo...
zarasiddique.com
Zara Siddique
020
Zara Siddique @zarasiddique.bsky.social · 30/12/2024
What is “Latin”?
110
Zara Siddique @zarasiddique.bsky.social · 03/12/2024
Had fun presenting my favourite #EMNLP2024 papers today at our secret reading group + bonus raccoon pics from Miami 🦝 Will follow up with favourite papers in blog post form soon!
020
Zara Siddique @zarasiddique.bsky.social · 20/11/2024
Had an amazing time at #EMNLP2024 and excited to connect with other researchers on here :)
050