Sign in

Mimansa Jaiswal

@mimansaj.bsky.social
1.8K followers 1.6K following 142 posts

Robustness, Data & Annotations, Evaluation & Interpretability in LLMs mimansajaiswal.github.io

PostsRepliesMedia
Reposted by Mimansa Jaiswal
Steve Klabnik @steveklabnik.com · 28/05/2025
I am disappointed in the AI discourse steveklabnik.com/writing/i-am...
steveklabnik.com
I am disappointed in the AI discourse
209911178
Reposted by Mimansa Jaiswal
Jeremy Morrell @jeremymorrell.dev · 05/04/2025
Meta introduced Llama 4 models and added this section near the very bottom of the announcement 😬 “[LLMs] historically have leaned left when it comes to debated political and social topics.” ai.meta.com/blog/llama-4...
Meta
Addressing bias in LLMs

It's well-known that all leading LLMs have had issues with bias-specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet.

Our goal is to remove bias from our Al models and to make sure that Llama can understand and articulate both sides of a contentious issue. As part of this work, we're continuing to make Llama more responsive so that it answers questions, can respond to a variety of different viewpoints without passing judgment, and doesn't favor some views over others.

We have made improvements on these efforts with this release—Llama 4 performs significantly better than Llama 3 and is comparable to Grok:• Llama 4 refuses less on debated political and social topics overall (from 7% in Lama 3.3 to below 2%).
• Llama 4 is dramatically more balanced with which prompts it refuses to respond to (the proportion of unequal response refusals is now less than 1% on a set of debated topical questions).
• Our testing shows that Llama 4 responds with strong political lean at a rate comparable to Grok (and at half of the rate of Llama 3.3) on a contentious set of political or social topics. While we are making progress, we know we have more work to do and will continue to drive this rate further down.
We're proud of this progress to date and remain committed to our goal of eliminating overall bias in our models.
513438
Reposted by Mimansa Jaiswal
Joe Littrell @gentlemanjoe.bsky.social · 06/04/2025
"We train our LLMs on art and literature and educational materials, and for some reason they keep turning out progressive."
011914
Reposted by Mimansa Jaiswal
Sarah Wiegreffe @sarah-nlp.bsky.social · 03/04/2025
Have work on the actionable impact of interpretability findings? Consider submitting to our Actionable Interpretability workshop at ICML! See below for more info. Website: actionable-interpretability.github.io Deadline: May 9
02010
Reposted by Mimansa Jaiswal
Somnath Basu Roy Chowdhury @somnathbrc.bsky.social · 02/04/2025
𝐇𝐨𝐰 𝐜𝐚𝐧 𝐰𝐞 𝐩𝐞𝐫𝐟𝐞𝐜𝐭𝐥𝐲 𝐞𝐫𝐚𝐬𝐞 𝐜𝐨𝐧𝐜𝐞𝐩𝐭𝐬 𝐟𝐫𝐨𝐦 𝐋𝐋𝐌𝐬? Our method, Perfect Erasure Functions (PEF), erases concepts perfectly from LLM representations. We analytically derive PEF w/o parameter estimation. PEFs achieve pareto optimal erasure-utility tradeoff backed w/ theoretical guarantees. #AISTATS2025 🧵
2378
Reposted by Mimansa Jaiswal
Sian Gooding @siangooding.bsky.social · 02/04/2025
New paper from our team @GoogleDeepMind! 🚨 We've put LLMs to the test as writing co-pilots – how good are they really at helping us write? LLMs are increasingly used for open-ended tasks like writing assistance, but how do we assess their effectiveness? 🤔 arxiv.org/pdf/2503.19711
arxiv.org
1208
Reposted by Mimansa Jaiswal
SE Gyges @segyges.bsky.social · 23/03/2025
pre aca you would specifically avoid being diagnosed or seeking treatment if you didn't have health insurance to prevent it from making it impossible for you to get health insurance. when you bought health insurance after doing this you committed fraud. i did this.
67810
Reposted by Mimansa Jaiswal
Jay Rosen @jayrosen.bsky.social · 07/03/2025
Some of his readers have asked Mike Masnick @mmasnick.bsky.social why his technology news site, Tech Dirt, has been covering politics so intensely lately. www.techdirt.com/2025/03/04/w... I cannot recommend Mike's reply enough. It's exactly what readers need to hear, what journalists need to do.
8645421810
Reposted by Mimansa Jaiswal
Martin Wattenberg @wattenberg.bsky.social · 25/02/2025
Neat visualization that came up in the ARBOR project: this shows DeepSeek "thinking" about a question, and color is the probability that, if it exited thinking, it would give the right answer. (Here yellow means correct.)
68116
Reposted by Mimansa Jaiswal
Nathan Lambert @natolambert.bsky.social · 25/02/2025
Come work with me! We are looking to bring on more top talent to our language modeling workstream at @ai2.bsky.social building the open ecosystem. We are hiring: * Research scientists * Senior research engineers * Post docs (Young investigators) * Pre docs job-boards.greenhouse.io/thealleninst...
job-boards.greenhouse.io
The Allen Institute for AI
45415
Mimansa Jaiswal @mimansaj.bsky.social · 24/02/2025
I interviewed for LLM/ML research scientist/engineering positions last Fall. Over 200 applications, 100 interviews, many rejections & some offers later, I decided to write the process down, along with the resources I used. Links to the process & resources in the following tweets
OCR'ed text from screenshot of top of post: LLM (ML) Job Interviews (Fall 2024) - Process A retelling of my experience interviewing for ML/LLM research science/engineering focused roles in Fall 2024.  This post has two parts:  Job Search Mechanics (including context, applying, and industry information), which you can continue reading below, and, Preparation Material and Overview of Questions, which you can read at LLM (ML) Job Interviews - Resources  Disclaimer Last Updated:  Dec 24, 2024  This is the process I used, which may work differently for you depending on your circumstances. I am writing this in December 2024, and the process occurred during Fall 2024. Given how rapidly the field of LLMs evolves, this information might become outdated quickly, but the general principles should remain relevant. (more...)  Read at: https://mimansajaiswal.github.io/posts/llm-ml-job-interviews-fall-2024-process/
35511
Reposted by Mimansa Jaiswal
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 13/02/2025
Obsessed with the work coming out of Finale Doshi-Velez's group; they don't just take the limits of the real world for ML deployment seriously but instead turn it into new algorithmic ideas arxiv.org/abs/2406.08636
arxiv.org
Towards Integrating Personal Knowledge into Test-Time Predictions
Machine learning (ML) models can make decisions based on large amounts of data, but they can be missing personal knowledge available to human users about whom predictions are made. For example, a mode...
0619
Reposted by Mimansa Jaiswal
Alt CDC (they/them) @altcdc.altgov.info · 01/02/2025
The entire archive of CDC datasets can be found here. HUGE shoutout to data archivists- this work is important 👏🙌🏻 archive.org/details/2025...
224117784629
Reposted by Mimansa Jaiswal
Ai2 @ai2.bsky.social · 21/01/2025
Can AI really help with literature reviews? 🧐 Meet Ai2 ScholarQA, an experimental solution that allows you to ask questions that require multiple scientific papers to answer. It gives more in-depth and contextual answers with table comparisons and expandable sections 💡 Try it now: scholarqa.allen.ai
Ai2 ScholarQA logo
13412
Reposted by Mimansa Jaiswal
Ryan Moulton @moultano.bsky.social · 21/01/2025
It is such a slap in the face to the Indian American community to delay their green cards for decades and then declare that because of that delay their American children aren't citizens.
1293
Mimansa Jaiswal @mimansaj.bsky.social · 21/01/2025
Is it time for a social media break again? It has not been a great day. 😅
010
Mimansa Jaiswal @mimansaj.bsky.social · 06/01/2025
This is pretty cool! Learn more: developer.chrome.com/docs/devtool... (seems to cover CSS and network requests, which might be fun to lay around with)
AI innovations tab in developer tools settings in Chrome
040
Reposted by Mimansa Jaiswal
Yoav Goldberg @yoavgo.bsky.social · 05/01/2025
i was annoyed at having many chrome tabs with PDF papers having uninformative titles, so i created a small chrome extension to fix it. i'm using it for a while now, works well. today i put it on github. enjoy. github.com/yoavg/pdf-ta...
59822
Reposted by Mimansa Jaiswal
Maria Antoniak @mariaa.bsky.social · 31/12/2024
It's ready! 💫 A new blog post in which I list of all the tools and apps I've been using for work, plus all my opinions about them. maria-antoniak.github.io/2024/12/30/o... Featuring @kagi.com, @warp.dev, @paperpile.bsky.social, @are.na, Fantastical, @obsidian.md, Claude, and more.
3621525
Reposted by Mimansa Jaiswal
Maggie Appleton @maggieappleton.com · 31/12/2024
I've always wanted to build things with D3, but the learning curve was too high. At least for the bespoke stuff I wanted to make (not just simple bar charts). I can finally make things like this thanks to Cursor! I just art directed this, and it made everything work beautifully. Even on mobile 🎉
817414
Mimansa Jaiswal @mimansaj.bsky.social · 29/12/2024
If you like working with a canvas, Muse (museapp.com) currently has a 30% off. 2 major things that make it different than freeform -- ink sticks to sticky notes (so it feels more like writing in the real world), and you can snippet out sections from pdf that link back to source.
museapp.com
Inspired & focused thinking with Muse
Muse is a canvas for thinking that helps you get clarity on things that matter. Think in private or collaborate with others. Available for iPad and Mac.
121
Mimansa Jaiswal @mimansaj.bsky.social · 26/12/2024
I don't usually discuss software, PKM, or tools here, but I found a valuable tip today. Not only can you use this to create beautiful video tutorials similar to what Screenstudio creates automatically, but you can also use this feature during live sharing & streaming, making it incredibly useful!
020
Mimansa Jaiswal @mimansaj.bsky.social · 26/12/2024
This probably won’t reach many people, but if someone who feels the same way ends up reading this, I hope it helps them realize they’re not alone. Lately, my mood has been heavily influenced by the ‘tech(adjacent) twitter' vibes, & the past few months have been really rough. ⏎
270
Reposted by Mimansa Jaiswal
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 25/12/2024
The other major Chinese AI lab, DeepSeek, just dropped their own last-minute entry into the 2024 model race: DeepSeek v3 is a HUGE model (685B parameters) which showed up, mostly undocumented, on Hugging Face this morning. My notes so far: simonwillison.net/2024/Dec/25/deeps…
simonwillison.net
deepseek-ai/DeepSeek-V3-Base
No model card or announcement yet, but this new model release from Chinese AI lab DeepSeek (an arm of Chinese hedge fund [High-Flyer](https://en.wikipedia.org/wiki/High-Flyer_(company))) looks very significant. It's a huge model …
1116
Reposted by Mimansa Jaiswal
Surya Ganguli @suryaganguli.bsky.social · 23/12/2024
Our new paper "Fooling LLM graders into giving better grades through neural activity guided adversarial prompting" lead expertly by Atsushi Yamamura shows how even in closed models like Gemini, we can add a small adversarial suffix to an essay to get unreasonably high scores arxiv.org/abs/2412.15275
arxiv.org
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
The deployment of artificial intelligence (AI) in critical decision-making and evaluation processes raises concerns about inherent biases that malicious actors could exploit to distort decision outcom...
5395
Mimansa Jaiswal @mimansaj.bsky.social · 23/12/2024
A question for those working on pre-training models (specifically regarding Olmo's logs): at what training step can you begin comparing two models' evaluation performance to get an estimated sense of their relative quality (for example, Olmo7B versus Olmo-2-7B)?
130
Reposted by Mimansa Jaiswal
Ahmad Beirami @abeirami.bsky.social · 21/12/2024
Echoing Theertha Suresh We are hiring! Our team at Google Research NY is seeking a Research Scientist! Our recent research efforts include developing algorithms for improving inference efficiency and alignment of LLMs. If you are interested, please consider applying! www.google.com/about/career...
google.com
Research Scientist, Speech and Language Algorithms, Research — Google Careers
0258
Mimansa Jaiswal @mimansaj.bsky.social · 23/12/2024
ImageFX is awesome at creating character sheets, and here are some examples. The text is gibberish though compared to Recraft. I like to use these to add consistent, interesting visuals to presentations with pretty minimal effort. (Prompt expansion by GPT4o in alt)
Prompt:
A character sheet featuring a cute chibi-style girl inspired by kawaii aesthetics. She has big expressive eyes , soft pastel colors , and is designed to convey a wide range of emotions and actions. She has her hair styled in two pigtails tied with large ribbons, and she is wearing a skirt tunic over a white shirt, paired with simple shoes. The sheet should include the following 6 poses or concepts:
1. Thinking : She is holding her chin with one hand, a question mark above her head, and a curious expression.
2. Pointing : She is confidently pointing at something, with a cheerful and proud expression.
3. Shocked : Her eyes wide open, hands on her cheeks, and a dramatic exclamation mark above her head.
4. Saying No : She has her arms crossed in an “X” shape with a pouty, determined expression.
5. Saying Yes : She is giving a thumbs-up with a big, cheerful smile and sparkling eyes.
6. Holding a Sign : She is holding a blank sign in front of her.

Alt text: An image showing a character with various poses and expressions, labeled with descriptions like “thinking,” “shocked,” “saying yes,” and “holding a sign.”  It has 8 such poses, but the text and the actual character poses aren't really related.Prompt:
A character sheet featuring a quirky, chibi-style scientist character designed for a research presentation. The character has oversized glasses, a lab coat slightly too big, and a quirky hairstyle resembling a “ mad scientist ” but in an endearing way. Their colors are vibrant , and they should radiate a humorous yet knowledgeable vibe. The sheet should include the following six poses or concepts:
	1. Eureka! : Holding a glowing lightbulb above their head with an excited grin and sparkles around them.
	2. Oops! : Covered in soot with a surprised expression , holding a beaker that’s overflowing with foam.
	3. Questioning : Stroking their chin thoughtfully, a small cloud of question marks surrounding their head.
	4. Explaining : Pointing to a chalkboard with formulas and diagrams, wearing a confident expression .
	5. Facepalm : Slapping their forehead with a comically exasperated look , accompanied by a small “oops” bubble.
	6. Celebrating : Jumping in the air with both hands up, confetti falling, and a triumphant expression.

Alt text: Pretty good representation of the prompt above but has 8 poses, and has gibberish on the blackboard.Prompt: A character sheet featuring a mischievous chibi-style AI assistant character representing the quirks and issues of large language models. The character is a floating holographic figure with glitchy edges , pixelated features , and an “error” motif woven into their design (e.g., a bow tie made of 404 signs ). Their design is sleek but cheeky , and their actions highlight the humorous yet problematic quirks of LLMs . The sheet should include the following six poses or concepts :
	1.	Context Collapse: Holding a giant scroll that unrolls and crushes them, their eyes spinning in confusion with text fragments spilling everywhere.
	2.	Hallucination: Dramatically presenting a floating , glowing object labeled “100% Wrong” with a smug expression and a sparkly , “I’m totally sure!” vibe.
	3.	Repetition Loop: Spinning around like a broken record , saying “Did you mean…?” repeatedly with a dizzy expression.
	4.	Overconfident Answering: Standing on a soapbox labeled “Trust Me,” confidently pointing while a nearby thought bubble shows a totally incorrect statement .
	5.	Context Cutoff: Hitting their head against a giant wall with “Token Limit Reached” written on it, looking frustrated and glitchy .
	6.	Overwhelm: Flailing under a downpour of input texts , holding a tiny umbrella , with an “I can’t handle this!” expression.

Alt: Pretty good representation of the original prompt above, but the gibberish in text stands out. It has 8 characters instead of 6 characters that I asked for.
030
Reposted by Mimansa Jaiswal
kyunghyuncho.bsky.social @kyunghyuncho.bsky.social · 21/12/2024
feeling a but under the weather this week … thus an increased level of activity on social media and blog: kyunghyuncho.me/i-sensed-anx...
kyunghyuncho.me
i sensed anxiety and frustration at NeurIPS’24 – Kyunghyun Cho
1917836
Mimansa Jaiswal @mimansaj.bsky.social · 16/12/2024
Honestly whisk is pretty good and really fun to play around with. Probably one of the best implementations of auto-prompting I have seen. And I like that the default presets are so cute. Here is a preview using my display picture. Check out more at: labs.google/fx/tools/whisk
This app is Google’s “Whisk” experiment, a generative AI tool designed to create image+text → image using style+subject+scene descriptors. Here’s a description of its core functionality:
	1.	Input a Subject: Users upload a photo of themselves (or someone else).
	2.	Style Customization: The tool generates an image of the character with a chosen aesthetic or theme, like the plush-toy dinosaur theme seen here.
	3.	Scene and Actions: Users can modify what the character does (e.g., eating pizza, waving, sitting) or place it in specific scenarios.
	4.	Output: The tool provides different AI-generated variations of the character based on the photo and selected prompts.

The prompts can be modified or refined based on what is generated.
010
Reposted by Mimansa Jaiswal
Chise @sailorrooscout.bsky.social · 13/12/2024
Hey, guess what prevents Polio? Psst. The answer is a VACCINE. Guess what happens when enough people DON’T get vaccinated for it? It RETURNS.
12650021244
Reposted by Mimansa Jaiswal
Ben Schmidt @bschmidt.bsky.social · 13/12/2024
I will never stop being amazed at the ability of Google infoboxes to display information that is not just untrue but completely senseless. What kind of structured data graph can even produce "San Jose, CA, United Kingdom"? Are they just running LLMs with non-zero temperature even on this stuff now?
Google info box with text. Highlighted portion says "San Jose, CA, United Kingdom"

About
Cisco Systems, Inc. is an American multinational digital communications technology conglomerate corporation headquartered in San Jose, California. Cisco develops, manufactures, and sells networking hardware, software, telecommunications equipment and other high-technology services and products. Wikipedia
CEO: Chuck Robbins (Jul 26, 2015–)
Founded: December 10, 1984, San Francisco, CA
Founders: Sandy Lerner, Leonard Bosack
Headquarters: San Jose, CA, United Kingdom
Number of employees: 90,400 (2024)
Parent organization: Cisco
Revenue: 56.99 billion USD (2023)
Disclaimer
4403
Reposted by Mimansa Jaiswal
Aparna Nair @disabilitystor1.bsky.social · 13/12/2024
As an Indian, my contempt for (rich white) anti-vaxxers in the west is so deep I often cannot quite find the words to express my contempt. I grew up seeing polio-disabled children and adults crawling over rubble and roads in India. My father worked for the Pulse Polio program. What is happening?
17146511325
Reposted by Mimansa Jaiswal
TechCrunch @techcrunch.com · 13/12/2024
Microsoft debuts Phi-4, a new generative AI model, in research preview
tcrn.ch
Microsoft debuts Phi-4, a new generative AI model, in research preview
Microsoft has announced the newest addition to its Phi family of generative AI models. Called Phi-4, the model is improved in several areas over its predecessors, Microsoft claims — in particular math problem solving. That’s partly the result of improved…
0274
Reposted by Mimansa Jaiswal
Greg Bodwin @gbodwin.bsky.social · 12/12/2024
I might be the last horse to cross the finish line here, but I learned today that the length "1em" in tex doesn't stand for anything, rather, it's called that because it's exactly the length of one uppercase M 🤯
4678
Mimansa Jaiswal @mimansaj.bsky.social · 11/12/2024
Trying out image and video generation models is interesting and different for me -- because I do not have any internal desire to "break" the model. I have no easy or hard prompts pre-defined for those models. I try to use them in the flow, i.e., oh I want this video/image, can I use AI for it.
010
Reposted by Mimansa Jaiswal
Simon Willison @simonwillison.net · 11/12/2024
Wrote up my initial impressions of the new Google Gemini 2.0 Flash model - it's really good, and the streaming mode (where you can stream video and audio to it and get audio streamed right back) is pure science-fiction simonwillison.net/2024/Dec/11/...
simonwillison.net
Gemini 2.0 Flash: An outstanding multi-modal LLM with a sci-fi streaming mode
Huge announcment from Google this morning: Introducing Gemini 2.0: our new AI model for the agentic era. There’s a ton of stuff in there (including updates on Project Astra and …
916936
Reposted by Mimansa Jaiswal
Jennifer Hu @jennhu.bsky.social · 11/12/2024
Thanks everyone for your interest!! The slides are now posted here: neurips.cc/media/neurip...
neurips.cc
0133
Reposted by Mimansa Jaiswal
Adina Williams @adinawilliams.bsky.social · 11/12/2024
Our paper PRISM alignment won a best paper award at #neurips2024! All credits to @hannahrosekirk.bsky.social A.Whitefield, P.Röttger, A.M.Bean, K.Margatina, R.Mosquera-Gomez, J.Ciro, @maxbartolo.bsky.social H.He, B.Vidgen, S.Hale Catch Hannah tomorrow at neurips.cc/virtual/2024/poster/97804
blog.neurips
2679
Reposted by Mimansa Jaiswal
Stella Biderman @stellaathena.bsky.social · 09/12/2024
Neural networks are not black boxes. They're the opposite of black boxes: we have extensive access to their internals. I think people have accepted this framing so innately that they've forgotten it's not true and it even warps how they do experiments.
8619
Reposted by Mimansa Jaiswal
Besmira Nushi @besmiranushi.bsky.social · 09/12/2024
@vidhishab.bsky.social and I will be presenting Eureka during NeurIPS Expo, on Wednesday, December 11 at 4:30 PM (West Meeting Room 301). Join us to get a glimpse of a demo, recent results, and an overall in-depth comparison of 12 frontier foundation models.
092
Reposted by Mimansa Jaiswal
Jennifer Hu @jennhu.bsky.social · 07/12/2024
Stop by our #NeurIPS tutorial on Experimental Design & Analysis for AI Researchers! 📊 neurips.cc/virtual/2024/tutorial/99528 Are you an AI researcher interested in comparing models/methods? Then your conclusions rely on well-designed experiments. We'll cover best practices + case studies. 👇
neurips.cc
NeurIPS Tutorial Experimental Design and Analysis for AI ResearchersNeurIPS 2024
68614
Mimansa Jaiswal @mimansaj.bsky.social · 07/12/2024
Google slides still doesn't have morph transitions 😠
020
Reposted by Mimansa Jaiswal
Vicki @vickiboykis.com · 06/12/2024
Great post that captures the tension between classic ML approaches and modern deep learning while acknowledging the nuances of both. “Working with LLMs doesn’t feel the same. It’s like fitting pieces into a pre-defined puzzle instead of building the puzzle itself.” www.reddit.com/r/MachineLea...
reddit.com
From the MachineLearning community on Reddit
Explore this post and more from the MachineLearning community
813418
Reposted by Mimansa Jaiswal
Blake Watson @blakewatson.com · 04/12/2024
I guess it is time to introduce Bluesky to my latest work, HTML for People. Anyone can make a website with HTML. No previous coding experience required. I cover everything you need to know to get started in an approachable and friendly way. And it’s free for all. 🚀 htmlforpeople.com
htmlforpeople.com
HTML for People
HTML isn't only for people working in the tech field. It's for everyone. Learn how to make a website from scratch in this beginner friendly web book.
17742662095
Reposted by Mimansa Jaiswal
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 05/12/2024
Increasingly thinking of my PhD work as directing “sorry, what did you means” at a field as in “sorry what do you mean your field doesn’t have benchmarks” “sorry what do you mean you didn’t tune your baselines” “sorry what do you mean your simulators run slow”
2161
Reposted by Mimansa Jaiswal
Jenn Wortman Vaughan @jennwv.bsky.social · 25/11/2024
The FATE group at @msftresearch.bsky.social NYC is accepting applications for 2025 interns. 🥳🎉 For full consideration, apply by 12/18. jobs.careers.microsoft.com/global/en/jo... Interested in AI evaluation? Apply for the STAC internship too! jobs.careers.microsoft.com/global/en/jo...
47335