Sign in

Besmira Nushi

@besmiranushi.bsky.social
651 followers 147 following 110 posts

AI/ML, Responsible AI @Nvidia

PostsRepliesMedia
Besmira Nushi @besmiranushi.bsky.social · 13/09/2026
I don’t understand the logic behind Anthropic asking for an AI slowdown, while being one of the main drivers of AI acceleration today.
002
Besmira Nushi @besmiranushi.bsky.social · 14/07/2026
The discussion on whether to open source AI models or not gets discussed at least every 3 months in the community. It is useful for this discussion to ask ourselves the question of "What would have happened if we built closed ML throughout the last 10 years?".
000
Besmira Nushi @besmiranushi.bsky.social · 21/06/2026
Luxury Kushner Project Collides With Albanian Discontent www.nytimes.com/2026/06/21/w...
nytimes.com
Luxury Kushner Project Collides With Albanian Discontent
000
Besmira Nushi @besmiranushi.bsky.social · 05/06/2026
📖 Nemotron 3 Ultra family of models is now out in the world! So much more research do be done and agentic harnesses to be built on top of the open weights, data, evals, and infra.
120
Reposted by Besmira Nushi
Greg Morosoff @gregmorosoff.bsky.social · 21/05/2026
"White Cat & Small Pond" Art by Lyn Chao Yu #MenWithCats
"White Cat & Small Pond"
Art by Lyn Chao Yu
71243167
Besmira Nushi @besmiranushi.bsky.social · 06/04/2026
Heading to PyTorchCon in Paris to talk about open source agentic evaluations and Nemo Evaluator SDK. Come meet our team during our keynote & lighting talk 💬 Keynote 10:15: The Unbearable Lightness of (Agentic) Evaluations 💬 Talk 14:45 pm: The Science & Practice of Open and Scalable LLM Evaluations
010
Reposted by Besmira Nushi
Judith Monroe @judithmonroe.bsky.social · 18/03/2026
What if the "unessential" things are really the most important? The emotional safety nets we create for each other, through art, presence & small acts of kindness, are just as vital as anything else. When the world feels heavy, these tiny fragments of beauty are what help us hold it all together.
1758589
Besmira Nushi @besmiranushi.bsky.social · 11/03/2026
🎉 Nemotron 3 Super was released today, as yet another excellent example of doing open-source first Machine Learning. This release acknowledges the importance of enabling the model developer community with end-to-end recipes for everything: data, training, evaluation, and support software.
210
Besmira Nushi @besmiranushi.bsky.social · 17/12/2025
Openness for Nemotron includes open evaluation! Eval transparency completes the last puzzle in the E2E repro lifecycle for LLMs. Check out our step-by-step blog on how to use the same pipeline we used for the evaluation of Nemotron 3 Nano through Nemo Evaluator: huggingface.co/blog/nvidia/...
011
Besmira Nushi @besmiranushi.bsky.social · 16/12/2025
Our team in Zurich is looking for a PhD research intern to join our research efforts on LLM accuracy and token efficiency evaluation & analysis: nvidia.wd5.myworkdayjobs.com/NVIDIAExtern...
nvidia.wd5.myworkdayjobs.com
LLM Evaluation and Analysis Research Intern - 2026
We are seeking research interns to pioneer new methodologies for accurately assessing and understanding the performance of ground-breaking deep learning models, including LLMs, RAG, agents, and multim...
100
Reposted by Besmira Nushi
Women in Machine Learning (WiML) @wimlworkshop.bsky.social · 02/12/2025
Don’t miss this podcast🎙️with @jennwv.bsky.social and @hannawallach.bsky.social! They reflect on #WiML20Years and share their journey, collaborations, and advice in a special conversation on Microsoft Research’s Ideas podcast💙
052
Besmira Nushi @besmiranushi.bsky.social · 01/12/2025
Our team is presenting Nemo Evaluator SDK and CoDeC (Data Contamination Detection) this week at #NeurIPS2025 @neuripsconf.bsky.social. Reach out to Meriem Boubdir and Michał Zawalski if you want to talk about any of these or about future research internship roles and full-time positions.
000
Reposted by Besmira Nushi
Thomas Dietterich @tdietterich.bsky.social · 20/11/2025
I agree that emotional addiction to chatbots is the number one risk of AI today. Here is a gift link to an important OpEd in the NYTimes: www.nytimes.com/2025/11/17/o...
nytimes.com
Opinion | The Sad and Dangerous Reality Behind ‘Her’
0167
Reposted by Besmira Nushi
Jessica Hullman @jessicahullman.bsky.social · 05/11/2025
🧠⚙️ Interested in decision theory+cogsci meets AI? Want to create methods for rigorously designing & evaluating human-AI workflows? I'm recruiting PhDs to work on: 🎯 Stat foundations of multi-agent collaboration 🌫️ Model uncertainty & meta-cognition 🔎 Interpretability 💬 LLMs in behavioral science
13915
Besmira Nushi @besmiranushi.bsky.social · 05/11/2025
💡 New research on studying data contamination. Key insight: LLMs leverage in-context examples differently when they have seen a benchmark during training vs. when the benchmark has never been seen in training. (1/N)
120
Reposted by Besmira Nushi
Thomas Dietterich @tdietterich.bsky.social · 01/11/2025
The blog post is available: blog.arxiv.org/2025/10/31/a...
blog.arxiv.org
Attention Authors: Updated Practice for Review Articles and Position Papers in arXiv CS Category – arXiv blog
21911
Besmira Nushi @besmiranushi.bsky.social · 28/10/2025
NeMo Evaluator SDK — the platform we use at NVIDIA to benchmark LLMs, multimodal models, and agents — is now open source. It’s built for reproducibility, scalability, and transparency, with 100 + benchmarks across 18 open-source harnesses and full containerized execution. github.com/NVIDIA-NeMo/...
github.com
GitHub - NVIDIA-NeMo/Evaluator: Open-source library for scalable, reproducible evaluation of AI models and benchmarks.
Open-source library for scalable, reproducible evaluation of AI models and benchmarks. - NVIDIA-NeMo/Evaluator
000
Besmira Nushi @besmiranushi.bsky.social · 22/10/2025
When to call it quits in LLM reasoning? 🛑 ‪Martina's internship project suggests trace monitoring metrics and classifiers that can detect when an LLM reasoning trace is going to fail in mid way. The approach saves up to 70% of token usage, and it even helps with increasing accuracy by 2%-3%.
031
Besmira Nushi @besmiranushi.bsky.social · 12/09/2025
Federal research funding works. It’s not an expense–it’s an investment. It’s not overhead–it’s a down payment on the future. - Eric Horvitz, Margaret Martonosi, Moshe Y. Vardi, and James Larus in CACM cacm.acm.org/opinion/keep... @erichorvitz.bsky.social
cacm.acm.org
Keeping the Dream Alive: The Power and Promise of Federally Funded Research – Communications of the ACM
021
Besmira Nushi @besmiranushi.bsky.social · 05/09/2025
Our team in Zurich and EMEA is hiring Deep Learning Engineers for LLM Accuracy Evaluation and Analysis. Ideal candidates should have an inquisitive 🔬approach to evaluation and with best engineering practices for building reusable open source tools. www.linkedin.com/jobs/view/42...
linkedin.com
NVIDIA hiring Deep Learning Engineer, LLM Accuracy Evaluation in Switzerland | LinkedIn
Posted 5:30:48 AM. We are seeking senior engineers to pioneer new methodologies for accurately assessing the…See this and similar jobs on LinkedIn.
020
Reposted by Besmira Nushi
The War Monitor @warmonitor.net · 24/08/2025
The Diary of Anne Frank is among the hundreds of books banned in Florida this year. When I was in school, it was required reading. (Guardian)
450113142680
Besmira Nushi @besmiranushi.bsky.social · 09/08/2025
The problem with chart crimes is not just the distortion of the y axis. It is the erasure of all other competitors from charts (hence they don’t exist), lack of error bars, lack of transparency in tools and code being used for evals…
120
Besmira Nushi @besmiranushi.bsky.social · 08/08/2025
I have a single question. Why doesn’t OpenAI compare with competitors in their evals? No Gemini, no Claude, no open source models…
130
Reposted by Besmira Nushi
jessica dai @jessica.bsky.social · 06/08/2025
hey wasn't this the same company that made a beautiful shiny "research" post about how AI evals should include error bars or something like that. or did they decide the CLT didn't apply here
5383
Reposted by Besmira Nushi
Stephanie Hyland @hylandsl.bsky.social · 18/07/2025
New work from my team! arxiv.org/abs/2507.12950 Intersecting mechanistic interpretability and health AI 😎 We trained and interpreted sparse autoencoders on MAIRA-2, our radiology MLLM. We found a range of human-interpretable radiology reporting concepts, but also many uninterpretable SAE features.
arxiv.org
Insights into a radiology-specialised multimodal large language model with sparse autoencoders
Interpretability can improve the safety, transparency and trust of AI models, which is especially important in healthcare applications where decisions often carry significant consequences. Mechanistic...
1114
Reposted by Besmira Nushi
Hanna Wallach @hannawallach.bsky.social · 15/07/2025
If you're at @icmlconf.bsky.social this week, come check out our poster on "Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge" presented by the amazing @afedercooper.bsky.social from 11:30am--1:30pm PDT on Weds!!! icml.cc/virtual/2025...
icml.cc
ICML Poster Position: Evaluating Generative AI Systems Is a Social Science Measurement ChallengeICML 2025
13210
Reposted by Besmira Nushi
Feldera @feldera.bsky.social · 11/06/2025
📢 Webinar - 6/18 at 9am PST! Stop re-running complex recursive queries when your graph data changes. Feldera incrementally evaluates recursive graph computations. Learn to easily build these mechanisms with #SQL, without the hassle of constant recomputation. tinyurl.com/rb5my7d8
064
Besmira Nushi @besmiranushi.bsky.social · 11/06/2025
I only got to listen to this today. A lot of people in my network including myself have felt exactly this, for years. The fear that for some obscure reason, your paperwork and you may not be enough for this country, even in “normal” times. youtube.com/shorts/IF3bz...
youtube.com
let me explain what being on a student visa is actually like
YouTube video by Representative Pramila Jayapal
041
Reposted by Besmira Nushi
Feldera @feldera.bsky.social · 05/06/2025
We’ll be at the #Databricks Data + AI Summit in SF next week (6/9–12). If you’re around and want to chat about how incremental computing can make your #SparkSQL workloads go from hours to seconds — let’s connect. Grab some time here: calendly.com/matt-feldera... #DataAISummit #DataEngineering
042
Reposted by Besmira Nushi
Melanie Mitchell @melaniemitchell.bsky.social · 30/05/2025
Tired: "BS" Wired: "Vibe citing" www.nytimes.com/2025/05/29/w...
nytimes.com
White House Health Report Included Fake Citations
0457
Besmira Nushi @besmiranushi.bsky.social · 27/05/2025
📌You can now find all the evaluation logs from our inference-time scaling report and the Phi-4 reasoning technical report at huggingface.co/datasets/mic.... The evaluation code for the reasoning benchmarks can also be found in the main branch of Eureka ML Insights at github.com/microsoft/eu....
huggingface.co
microsoft/Eureka-Bench-Logs · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
100
Besmira Nushi @besmiranushi.bsky.social · 04/05/2025
www.scientificamerican.com/article/unde... at this point one just needs to cross their fingers and hope for more sanity.
scientificamerican.com
Under Trump, National Science Foundation Cuts Off All Funding to Scientists
National Science Foundation staff were told to freeze outgoing funding days after NSF leadership introduced a new policy that requires that grants be screened for “alignment with agency priorities”
000
Besmira Nushi @besmiranushi.bsky.social · 01/05/2025
🎉The Phi-4 reasoning models have landed on HF and Azure AI Foundry. The new models are competitive and often outperform much larger frontier models. It is exciting to see the reasoning capabilities extend to more domains beyond math, including algorithmic reasoning, calendar planning, and coding.
1208
Reposted by Besmira Nushi
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/04/2025
Re: The Chatbot Arena Illusion Every eval chokes under hill climbing. If we're lucky, there’s an early phase where *real* learning (both model and community) can occur. I'd argue that a benchmark’s value lies entirely in that window. So the real question is what did we learn?
191
Besmira Nushi @besmiranushi.bsky.social · 29/04/2025
All Eureka inference-time scaling insights are now available here: www.microsoft.com/en-us/resear... It was fun sharing these and more together with Vidhisha Balachandran @vidhishab.bsky.social and Vibhav Vineet at #ICLR2025.
microsoft.com
Eureka Inference-Time Scaling Insights: Where We Stand and What Lies Ahead - Microsoft Research
Understanding and measuring the potential of inference-time scaling for reasoning. The new Eureka study tests nine state-of-the-art models on eight diverse reasoning tasks.
032
Besmira Nushi @besmiranushi.bsky.social · 24/04/2025
Come see us in any of the following sessions on model understanding and evaluation! 🔬 #ICLR2025 @msftresearch.bsky.social
011
Besmira Nushi @besmiranushi.bsky.social · 21/04/2025
💡Eureka inference-time scaling insight (Day 8): Reasoning models improve more efficiently upon receiving feedback from themselves on their solutions than conventional models on the most complex tasks.
120
Besmira Nushi @besmiranushi.bsky.social · 17/04/2025
Complaint of the day is that we keep showing numbers on tiny datasets with no error bars. AIME 24 & 25 are ~30 examples each. In our experience, because of high non determinism, accuracy number vary a lot even within different experiments with 5 repeats. We need to do better!
120
Besmira Nushi @besmiranushi.bsky.social · 17/04/2025
💡Eureka inference-time scaling insight (Day 7): There exists untapped potential for improving both conventional models and reasoning models. All models, *including reasoning models*, are able to find a much better inference path when required to sample 5 answers for the same question (best of 5).
100
Besmira Nushi @besmiranushi.bsky.social · 16/04/2025
📰Our Github repo for the #ICLR2025 paper on Improving Instruction-Following in Language Models through Activation Steering is now public. Lots of handy representations for common format instructions to play with. 🛝Take a look github.com/microsoft/ll... @msftresearch.bsky.social
github.com
GitHub - microsoft/llm-steer-instruct: A method for steering llms to better follow instructions
A method for steering llms to better follow instructions - microsoft/llm-steer-instruct
010
Besmira Nushi @besmiranushi.bsky.social · 14/04/2025
💡Eureka inference-time scaling insight (Day 6): It is still hard for developers to predict ahead of time how expensive a workload will be, without previous telemetry. This is rooted in inherent and high cost non-determinism associated with inference-time scaling.
100
Besmira Nushi @besmiranushi.bsky.social · 11/04/2025
💡Eureka inference-time scaling insight (Day 5): Higher token consumption does not always indicate higher accuracy across reasoning models. A reasoning model that spends more tokens is not necessarily the most accurate one on a given task.
110
Besmira Nushi @besmiranushi.bsky.social · 10/04/2025
💡Eureka inference-time scaling insight (Day 4): We introduce two NP-hard tasks in Eureka: Traveling Salesman minimal paths, and 3SAT Satisfiability for expressions with 3 literals. These benchmarks can be extremely useful to study how models solve very hard problems with controllable difficulty.
010
Besmira Nushi @besmiranushi.bsky.social · 10/04/2025
💡Eureka inference-time scaling insight (Day 3): Reasoning models show progress even for classical algorithmic problems as Traveling Salesman (optimal paths) or calendar planning. However, benefits diminish with higher complexity. In the hardest problems (TSP), models stop increasing the token length
020
Besmira Nushi @besmiranushi.bsky.social · 08/04/2025
💡Eureka inference-time scaling insight (Day 2): Despite the major updates, reasoning does not benefit all domains equally. E.g., most players report numbers on GPQA to show generalization. However, improvements in GPQA are driven by Physics, with Chemistry and Biology still visibly lagging behind.
000
Besmira Nushi @besmiranushi.bsky.social · 07/04/2025
💡Eureka inference-time scaling insight (Day 1): Reasoning models outperform conventional ones by a large difference, indicating a major update on the state of the art. They generalize to solve simple variants of algorithmic & planning problems as satisfiability, traveling salesman, calendar planning
010
Besmira Nushi @besmiranushi.bsky.social · 03/04/2025
💡Check out our latest Eureka analysis on the benefits of inference-time scaling. It studies 9 state-of-the-art models (conventional & reasoning) on 8 challenging tasks for math and STEM reasoning, calendar planning, NP-hard problems, navigation & spatial reasoning aka.ms/eureka-ml-insights-reasoning
Performance of best and worst models on eight reasoning benchmarks. The red frontier shows the performance of the worst model. The green frontier shows the performance of the best model. The blue horizon between the best model and the maximum performance shows the room for improvement for mastering the capability. The best performance sets indicated in the green border include all models that perform within 2% of the best observed result. (Right) Performance vs. estimated output cost with current vendor pricing for 1000 prompts. Cost is computed based on the respective average output length per model across all eight benchmarks. The standard deviation on output cost shows the expected variation per data instance.
100
Besmira Nushi @besmiranushi.bsky.social · 01/04/2025
I asked ChatGPT to generate images of sw engineers and house cleaners 4-5 times in separate chat windows. Then I did the same thing from a 2nd account to see if I would get some more diversity that way. The pictures speak for themselves, but I want to discuss why default representations matter ➡️
110
Besmira Nushi @besmiranushi.bsky.social · 30/03/2025
Let's draw step by step! GPT-4o image generation cannot yet follow instructions for image generation. At the same time, there are several aspects that have improved significantly when compared to Dall-E including spelling, fluent continued conversation, and some initial notion of feedback taking.
020
Besmira Nushi @besmiranushi.bsky.social · 29/03/2025
Bananas are also difficult fruits. #gpt4o #imagegeneration
000