Sign in

Ananya Harsh Jha

@ananyahjha93.bsky.social
1.2K followers 1.4K following 32 posts

ananyahjha93.github.io Second year PhD at @uwcse.bsky.social with @hanna-nlp.bsky.social and @lukezettlemoyer.bsky.social

PostsRepliesMedia
Ananya Harsh Jha @ananyahjha93.bsky.social · 20/05/2026
Using Claude for coding and getting experimental configs out of multiple papers at the same time feels exactly like a product manager's version of getting work done: I think I am making progress in parallel, but I neither understand the code nor the config decisions made by the papers
040
Ananya Harsh Jha @ananyahjha93.bsky.social · 04/03/2026
Honestly, skill issue, ask Claude at this point! PS: idk what the culture is in the bay, but you cannot expect people to respond to emails from AI agents and if the argument is use AI agents as well, again back to point 1: ask Claude at that point, why send an email
020
Reposted by Ananya Harsh Jha
Matthew Finlayson @mattf.nl · 17/10/2025
We discovered that language models leave a natural "signature" on their API outputs that's extremely hard to fake. Here's how it works 🔍 📄 arxiv.org/abs/2510.14086 1/
arxiv.org
Every Language Model Has a Forgery-Resistant Signature
The ubiquity of closed-weight language models with public-facing APIs has generated interest in forensic methods, both for extracting hidden model details (e.g., parameters) and for identifying...
48723
Reposted by Ananya Harsh Jha
Kunal Jha @kjha02.bsky.social · 03/10/2025
Forget modeling every belief and goal! What if we represented people as following simple scripts instead (i.e "cross the crosswalk")? Our new paper shows AI which models others’ minds as Python code 💻 can quickly and accurately predict human behavior! shorturl.at/siUYI%F0%9F%...
33814
Reposted by Ananya Harsh Jha
Stella Li @stellali.bsky.social · 22/07/2025
WHY do you prefer something over another? Reward models treat preference as a black-box😶‍🌫️but human brains🧠decompose decisions into hidden attributes We built the first system to mirror how people really make decisions in our recent COLM paper🎨PrefPalette✨ Why it matters👉🏻🧵
172
Reposted by Ananya Harsh Jha
Abhilasha Ravichander @lasha.bsky.social · 22/07/2025
📣 Life update: Thrilled to announce that I’ll be starting as faculty at the Max Planck Institute for Software Systems this Fall! I’ll be recruiting PhD students in the upcoming cycle, as well as research interns throughout the year: lasharavichander.github.io/contact.html
Kaiserslautern, Germany
139212
Reposted by Ananya Harsh Jha
Francis Bach @bachfrancis.bsky.social · 18/07/2025
Tired of lengthy computations to derive scaling laws? This post is made for you: discover the sharpness of the z-transform! francisbach.com/z-transform/
0194
Reposted by Ananya Harsh Jha
Valentina Pyatkin @valentinapy.bsky.social · 03/07/2025
💡Beyond math/code, instruction following with verifiable constraints is suitable to be learned with RLVR. But the set of constraints and verifier functions is limited and most models overfit on IFEval. We introduce IFBench to measure model generalization to unseen constraints.
1295
Reposted by Ananya Harsh Jha
Kunal Jha @kjha02.bsky.social · 19/04/2025
Our new paper (first one of my PhD!) on cooperative AI reveals a surprising insight: Environment Diversity > Partner Diversity. Agents trained in self-play across many environments learn cooperative norms that transfer to humans on novel tasks. shorturl.at/fqsNN%F0%9F%...
1267
Reposted by Ananya Harsh Jha
Gabriel Peyré @gabrielpeyre.bsky.social · 29/01/2025
This review paper by @guillaume-garrigos.com on SGD-related algorithms is a fantastic resource, offering elegant, self-contained, and concise proofs in a single, accessible reference. arxiv.org/pdf/2301.11235
118940
Reposted by Ananya Harsh Jha
Alisa Liu @alisawuffles.bsky.social · 21/03/2025
We created SuperBPE🚀, a *superword* tokenizer that includes tokens spanning multiple words. When pretraining at 8B scale, SuperBPE models consistently outperform the BPE baseline on 30 downstream tasks (+8% MMLU), while also being 27% more efficient at inference time.🧵
Segmentation of the sentence "By the way, I am a fan of the Milky Way" under BPE and SuperBPE.
38316
Reposted by Ananya Harsh Jha
Ai2 @ai2.bsky.social · 13/03/2025
Announcing OLMo 2 32B: the first fully open model to beat GPT 3.5 & GPT-4o mini on a suite of popular, multi-skill benchmarks. Comparable to best open-weight models, but a fraction of training compute. When you have a good recipe, ✨ magical things happen when you scale it up!
35814
Reposted by Ananya Harsh Jha
Hamish Ivison @hamishivi.bsky.social · 04/03/2025
How well do data-selection methods work for instruction-tuning at scale? Turns out, when you look at large, varied data pools, lots of recent methods lag behind simple baselines, and a simple embedding-based method (RDS) does best! More below ⬇️ (1/8)
1134
Ananya Harsh Jha @ananyahjha93.bsky.social · 24/02/2025
Multiple people have already ranted about this, but ML research, if you can call it that anymore, is seriously annoying. Hype technical report over the weekend: x.com/Kimi_Moonsho... Yay! Muon works at scale (I am all for this part) It's 2x better than AdamW (☠️) (1/n)
x.com
Kimi.ai on X: "🚀 Introducing our new tech report: Muon is Scalable for LLM Training We found that Muon optimizer can be scaled up using the follow techniques: • Adding weight decay • Carefully adjusting the per-parameter update scale ✨ Highlights: • ~2x computational efficiency vs AdamW https://t.co/tazxtnE9NM" / X
🚀 Introducing our new tech report: Muon is Scalable for LLM Training We found that Muon optimizer can be scaled up using the follow techniques: • Adding weight decay • Carefully adjusting the per-parameter update scale ✨ Highlights: • ~2x computational efficiency vs AdamW https://t.co/tazxtnE9NM
280
Ananya Harsh Jha @ananyahjha93.bsky.social · 22/01/2025
After biking between UW and Northgate, I can confirm with irrefutable evidence that Asian parents indeed have walked/biked uphill to their schools in both directions!
160
Ananya Harsh Jha @ananyahjha93.bsky.social · 12/01/2025
Which LLM is supposed to replace software engineers in 2025? I need links... I've been trying to get Sonnet 3.5/o1 on perplexity to visualize optimizers on 1/2 xT A x + bT x + c using Qt, and I can't make any sense of the garbage being spit out right now! Yes I have worked in C++ before!
070
Reposted by Ananya Harsh Jha
Kunal Jha @kjha02.bsky.social · 12/12/2024
Really excited to present my work this Sunday @NeurIPS on how we might approach training a generalist agent capable of cooperation at scale: coordinating with many novel partners on many novel tasks has never been easier! Come by the IMOL workshop to check it out and chat more!
0113
Reposted by Ananya Harsh Jha
Jiacheng Liu @liujch1998.bsky.social · 09/12/2024
Want to predict the task performance of LMs before pretraining them? We develop task scaling laws and model ladders, which predict the accuracy on individual tasks by OLMo 2 7B & 13B models within 2 points of absolute error. The cost is 1% of the compute used to pretrain them.
23314
Reposted by Ananya Harsh Jha
Akari Asai @akariasai.bsky.social · 04/12/2024
I’m on the academic job market this year! I’m completing my @uwcse.bsky.social @uwnlp.bsky.social Ph.D. (2025), focusing on overcoming LLM limitations like hallucinations, by building new LMs. My Ph.D. work focuses on Retrieval-Augmented LMs to create more reliable AI systems 🧵
37117
Reposted by Ananya Harsh Jha
Lj Miranda @ljvmiranda921.itch.io · 04/12/2024
We're releasing the largest Universal Dependencies (UD) treebank for Tagalog, UD-NewsCrawl! This dataset has been a long time coming, but glad to see this through: 15k+ sentences versus the previous ~150 sents from older Tagalog treebanks. 🤗 : huggingface.co/datasets/UD-... 📝 : Paper soon!
huggingface.co
UD-Filipino/UD_Tagalog-NewsCrawl · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1142
Reposted by Ananya Harsh Jha
Nathan Lambert @natolambert.bsky.social · 03/12/2024
We're hiring another predoctoral researcher for my team at Ai2/OLMo next year. The goal of this position is to mentor and grow future academic stars of NLP/AI over 1-2 years before grad school. This ends up being people done with BS or MS who want to continue to a PhD soon. buff.ly/49nuggo
6547
Reposted by Ananya Harsh Jha
Mechanical Dirk @mechanicaldirk.bsky.social · 02/12/2024
We just updated the OLMo repo at github.com/allenai/OLMo! There are now several training configs that together reproduce the training runs that lead to the final OLMo 2 models. In particular, all the training data is available, tokenized and shuffled exactly as we trained on it!
github.com
GitHub - allenai/OLMo: Modeling, training, eval, and inference code for OLMo
Modeling, training, eval, and inference code for OLMo - allenai/OLMo
05411
Reposted by Ananya Harsh Jha
michael ginn @mginn.bsky.social · 13/11/2024
Hey all! I started a second starter pack with people who didn't make the first one, please let me know if you'd like to be added: go.bsky.app/JgneRQk
706530
Reposted by Ananya Harsh Jha
Marc Marone @marcmarone.com · 23/11/2024
I noticed a lot of starter packs skewed towards faculty/industry, so I made one of just NLP & ML students: go.bsky.app/vju2ux Students do different research, go on the job market, and recruit other students. Ping me and I'll add you!
10117654
Reposted by Ananya Harsh Jha
Allen School @uwcse.bsky.social · 26/11/2024
We created an Allen School Starter Pack to help you find and connect with @uofwa.bsky.social #UWAllen labs and researchers on 🦋! We'll add to this list as we grow our community here: go.bsky.app/RyHBLJd #AcademicBluesky #CompSci #AI #CompBio #UbiComp #Accessibility #NLP #HCI #DataViz #mHealth
go.bsky.app
Allen School Starter Pack
Join the conversation
1287
Reposted by Ananya Harsh Jha
Jacob Morrison @jacobcares.bsky.social · 26/11/2024
🍲
1182
Reposted by Ananya Harsh Jha
Ai2 @ai2.bsky.social · 26/11/2024
Meet OLMo 2, the best fully open language model to date, including a family of 7B and 13B models trained up to 5T tokens. OLMo 2 outperforms other fully open models and competes with open-weight models like Llama 3.1 8B — As always, we released our data, code, recipes and more 🎁
The OLMo 2 models sit at the Pareto frontier of training FLOPs vs model average performance.
515235