Reposted by Werner GeyerPicard Tips @picardtips.bsky.social · 19/10/2025Picard technology tip: Sometimes your chief engineer can build new systems that are better than your existing enterprise software. 111216
Reposted by Werner GeyerHyo Jin (Gina) Do @dohyojin.bsky.social · 16/10/2025❣️ Shout out to my amazing co-authors: Rachel Ostrand, @wernergeyer.bsky.social , @keerthi166.bsky.social, Dennis Wei, and Justin Weisz! If you'll be at AIES, I would love to connect and chat more about our work! 🙌 011
Reposted by Werner GeyerPicard Tips @picardtips.bsky.social · 26/09/2025Picard management tip: Even without game-changing results, experimentation is time well spent. 011222
Werner Geyer @wernergeyer.bsky.social · 25/09/20256/ Try it out & explore more: 👉 GitHub: github.com/IBM/eval-ass... 👉 Demo: evalassist-evalassist.hf.space 👉 Project page: ibm.github.io/eval-assist/ 000
Werner Geyer @wernergeyer.bsky.social · 25/09/20255/ And we’re planning to bring several backend capabilities into the UI soon. Stay tuned 👀 100
Werner Geyer @wernergeyer.bsky.social · 25/09/20254/ ⚙️ Backend updates • Independent Judges module (no UI - see: github.com/IBM/eval-ass...) • Unified Judge API • Extensible: supports Unitxt, M-Prometheus & more • Self-consistency: run judges multiple times • In-context examples • Multi-criteria evals w/roll-ups • Custom prompts supportedgithub.comeval-assist/backend/src/evalassist/judges at main · IBM/eval-assistEvalAssist is an open-source project that simplifies using large language models as evaluators (LLM-as-a-Judge) of the output of other large language models by supporting users in iteratively refin... 100
Werner Geyer @wernergeyer.bsky.social · 25/09/20253/ 🖥️ UI updates • Export & import test data (CSV) • More benchmarks: JudgeBench & BigGen, grouped by capabilities • 50+ Unitxt () criteria via Unixt (www.unitxt.ai) catalog integration • Export/import test cases in JSON • Model provider connections can be tested before evalsunitxt.ai 100
Werner Geyer @wernergeyer.bsky.social · 25/09/20252/ 📄 Paper @acmuist.bsky.social : EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences By @dohyojin.bsky.social - presenting Wed 9:00–10:30 in “Managing Tasks.” session 👉 arxiv.org/pdf/2410.00873arxiv.org 120
Werner Geyer @wernergeyer.bsky.social · 25/09/20251/ EvalAssist makes it easier to test, refine & share evaluation criteria for LLMs. ibm.github.io/eval-assist/ We’ve added powerful new features on both the UI and backend, plus we’ll be at UIST next week presenting our paper on task-specific evaluations & AI-assisted judgment strategies.ibm.github.ioEvalAssistEvalAssist simplifies LLM-as-a-Judge by supporting users in iteratively refining evaluation criteria in a web-based user experience. 100
Werner Geyer @wernergeyer.bsky.social · 25/09/2025🚀 Excited to share some updates from EvalAssist, the open-source LLM-as-a-Judge framework we released a few months ago! 🧵 110
Werner Geyer @wernergeyer.bsky.social · 21/08/2025We've just extended the IUI Workshop deadline by one week to August 29. Looking forward to your contributions! 020
Werner Geyer @wernergeyer.bsky.social · 28/07/2025Getting ready! Come visit us at the IBM booth @acl to learn about our latest Research. We have a number of super interesting demos lined up. research.ibm.com/events/acl-2... 000
Reposted by Werner GeyerCHIWORK @chiwork.bsky.social · 04/04/2025We’re growing and going global! 🌍 CHIWORK 2025 is shaping up to be our biggest and most diverse edition yet. Thanks to everyone who submitted, reviewed, and supported us 💙 Can’t wait to see you in Amsterdam! 🔗 chiwork.org #CHIWORK2025 #HCI #FutureOfWork 053
Reposted by Werner GeyerACM - Intelligent User Interfaces @acm-iui.bsky.social · 09/06/2025📢 Call for Workshop & Tutorial Proposals 📢 Bring your ideas and discuss them with fellow researchers in Paphos, Cyprus, from March 22-26, 2026. iui.hosting.acm.org/2026/call-fo... #CallForProposals #IUI2026 #HCI #AI 021
Werner Geyer @wernergeyer.bsky.social · 16/06/2025📣 Today we open-sourced EvalAssist, a web-based tool that makes it super easy to develop criteria for llm judges. You can run this now locally and then scale up with notebooks using Unitxt. Check out the AI Alliance article to get the scoop: thealliance.ai/blog/llm-as-...thealliance.aiLLM-as-a-Judge Without the Headaches: EvalAssist Brings Structure and Simplicity to the Chaos of LLM Output Review | AI AllianceEvaluating AI model outputs at scale is a major challenge for teams using LLMs, especially when assessing nuanced qualities like politeness, fairness, and tone that traditional benchmarks miss. IBM Re... 153
Reposted by Werner GeyerPatricia Kahr @pkahr.bsky.social · 05/06/2025📣 Call for Workshop & Tutorial Proposals 📣 #IUI2026 is looking forward to your contribution! Bring your ideas and discuss them with fellow researchers in Paphos, Cyprus, from March 22-26, 2026. 🚨 Proposal Deadlines: Aug 22 (Workshops) and Oct 17 (Tutorials)🚨 iui.hosting.acm.org/2026/call-fo...iui.hosting.acm.orgCall for Workshop & Tutorial Proposals | IUI 122
Werner Geyer @wernergeyer.bsky.social · 05/06/2025📣 IUI 2026 Call for Workshops and Tutorials is live 📣 iui.acm.org/2026/call-fo... Note that this year, submissions will be due August 22 earlier than previous years. Pls. spread the word! We had a fantastic workshop program in 2025 and I'm looking forward to an even better one in 2026 in Cyprus.iui.acm.orgCall for Workshop & Tutorial Proposals | IUI 020
Werner Geyer @wernergeyer.bsky.social · 06/05/2025We just published a summary the 6th workshop on Human-AI Co-Creation with Generative Models at IUI 2025 in March. This year's special topic, of course, AI agents and agency. Two of our sessions covered this topic and we had an exciting panel discussion. Check it out! medium.com/human-center...medium.comHAI-GEN 2025: 6th Workshop on Human-AI Co-Creation with Generative Modelsby Osnat Mokryn (University of Haifa, IL), Orit Shaer (Wellesley College, US), Werner Geyer (IBM Research, US), Mary Lou Maher (Computing… 010
Reposted by Werner GeyerKush Varshney कुश वार्ष्णेय @krvarshney.bsky.social · 08/04/2025A summary of decolonial AI alignment in the Human-Centered AI publication on Medium. Thanks to @jweisz3.bsky.social for asking me to write it, and for editing the piece. medium.com/human-center...medium.comDecolonial AI Alignmentby Kush Varshney (IBM Research, US) 052
Reposted by Werner GeyerKush Varshney कुश वार्ष्णेय @krvarshney.bsky.social · 28/03/2025I'm on the IBM Mixture of Experts podcast wearing a safety vest. We talk about all the new things in AI this week. I also connect to older work by IBM Fellows Irene Greif, Bob Dennard, Rolf Landauer, and Charlie Bennett and to Mauro Martino's new AI-generated film. www.youtube.com/watch?v=CgqH...youtube.comDeepSeek-V3-0324, Gemini Canvas and GPT-4o image generationYouTube video by IBM Technology 022
Reposted by Werner GeyerKush Varshney कुश वार्ष्णेय @krvarshney.bsky.social · 09/04/2025Granite Guardian tops a new benchmark! research.ibm.com/blog/granite...research.ibm.comGranite Guardian tops third-party AI benchmarkIBM’s collection of LLM guardrail models take six of the top 10 spots on the new GuardBench leaderboard. 032
Werner Geyer @wernergeyer.bsky.social · 31/03/2025Asparagus time in Germany. This is an automated peeling machine. No AI 😀 000
Werner Geyer @wernergeyer.bsky.social · 26/03/2025All set up for demo time at IUI. We are showing a tool for GenAU-assisted hypotheses exploration. dl.acm.org/doi/10.1145/... 010
Werner Geyer @wernergeyer.bsky.social · 18/03/2025Ah, and now there is a cool name for it :) www.zdnet.com/article/what... Is there already a CHI paper about it? :)zdnet.comWhat is AI vibe coding? It's all the rage but it's not for everyone - here's whyCaution: Experience required. Vibe coding feels like magic, until your AI assistant starts overwriting your work. 010
Werner Geyer @wernergeyer.bsky.social · 11/02/2025We have two amazing keynotes this year at HAI-GEN 2025 to challenge our thinking on co-creative systems from an interaction perspective. Hope to cu at IUI this year! hai-gen.github.io/2025/program/lnkd.inLinkedInThis link will take you to a page that’s not on LinkedIn 030
Reposted by Werner GeyerEstelle Smith @estellesmithphd.bsky.social · 31/01/2025📣 The #CSCW2026 deadline (@acm-cscw.bsky.social) has been posted. Big change this year. There is **only one deadline** for 2026 and it is May 13, 2025. 📣 Please spread the word! #CSCW #CHI #HCI #socialcomputing cscw.acm.org/2025/index.p...cscw.acm.orgCALL FOR PAPERS – CSCW 2025 22713
Reposted by Werner GeyerKush Varshney कुश वार्ष्णेय @krvarshney.bsky.social · 20/12/2024"IBM has equipped the Granite Guardian 3.1 models with the ability to detect hallucinations in AI agent workflows. This feature provides oversight of an AI agent completing a task, monitoring for fabricated information or incorrect function calls." technologymagazine.com/articles/the...technologymagazine.comThe Key to How IBM's Granite 3.1 is Advancing Enterprise AIIBM’s new Granite 3.1 addresses key enterprise needs, including expanded context handling, multilingual support, new tools and AI agent development 062
Reposted by Werner GeyerKush Varshney कुश वार्ष्णेय @krvarshney.bsky.social · 18/12/2024We released Granite Guardian 3.1 today! Even better at harm detection than Granite Guardian 3.0. The main new feature is 'function calling hallucination' detection relevant for tool-using AI agents. github.com/ibm-granite/... 071
Reposted by Werner GeyerDaniel Buschek @dbuschek.bsky.social · 19/12/2024📣 Join our 6th "HAI-GEN" workshop on Human-AI Co-Creation with Generative Models at #IUI2025! 👉 Submit a paper, demo or poster by Jan 22: hai-gen.github.io/2025/ with: @ossimokryn.bsky.social, @wernergeyer.bsky.social, Orit Shaer, Justin Weisz, Lydia Chilton, Mary Lou Maher #IUI #HCI #CHI #AI 073
Werner Geyer @wernergeyer.bsky.social · 16/12/2024I showed this cool demo last week @neuripsconf.bsky.social Now we have a public version on Hugging Face that you can play with to see the "judge" model in action. huggingface.co/spaces/ibm-g... Enjoy! Open source repo & benchmarks: github.com/ibm-granite/...huggingface.coGranite Guardian Demo - a Hugging Face Space by ibm-granitedemo 074
Reposted by Werner GeyerKush Varshney कुश वार्ष्णेय @krvarshney.bsky.social · 14/12/2024Nice video about Granite Guardian! www.youtube.com/watch?v=n7tt...youtube.comGranite Guardian 3.0 8B - Detect Risks in LLM Prompts Responses and RAG Pipelines - Install LocallyYouTube video by Fahd Mirza 031
Werner Geyer @wernergeyer.bsky.social · 12/12/2024Neurips is a great place to meet friends and former colleagues 020
Reposted by Werner GeyerKush Varshney कुश वार्ष्णेय @krvarshney.bsky.social · 10/12/2024It is @neuripsconf.bsky.social booth setup day! Among Ambrish Rawat, @bhoov.bsky.social, and @wernergeyer.bsky.social, who do you think is *not* an author of the Granite Guardian technical report we released today? (Hint: Granite Guardian helps make any LLM safer.) Link: github.com/ibm-granite/... 062
Werner Geyer @wernergeyer.bsky.social · 09/12/2024Now posted at the under construction booth 😀 our demo lineup for Tuesday. Looking forward connecting with you at the IBM booth @neuripsconf.bsky.social 082
Werner Geyer @wernergeyer.bsky.social · 09/12/2024Getting ready for the big show this week. @neuripsconf.bsky.social 020
Reposted by Werner GeyerKrzysztof Gajos @kgajos.bsky.social · 09/12/2024Do you feel that your profession or organization urgently needs to adopt #AI? Here's some advice that we often share (as academics working in human-AI interaction and #ML) with professional colleagues who consider the role of modern technology in their professions. (w/ Weiwei Pan and Hongjin Lin)iis.seas.harvard.eduKeep calm and carry on, with or without AI: What to do if your profession or organization urgently needs to adopt AI – Intelligent Interactive Systems Group at Harvard 086
Reposted by Werner GeyerMichael Hind @michaelhind.bsky.social · 07/12/2024I'm happy to announce a significant revision of our paper describing opportunities and challenges of quantitative AI risk assessments, also known as automated red-teaming: arxiv.org/abs/2209.06317arxiv.orgQuantitative AI Risk Assessments: Opportunities and ChallengesAlthough AI systems are increasingly being leveraged to provide value to organizations, individuals, and society, significant attendant risks have been identified and have manifested. These risks have... 163
Werner Geyer @wernergeyer.bsky.social · 06/12/2024The UX surrounding this capability was grounded and informed by research our Human-Centered Trustworthy AI team did on communicating factuality scores and source attribution led by @dohyojin.bsky.social, part of a larger research effort on how to communicate uncertainty. arxiv.org/abs/2405.20434arxiv.orgFacilitating Human-LLM Collaboration through Factuality Scores and Source AttributionsWhile humans increasingly rely on large language models (LLMs), they are susceptible to generating inaccurate or false information, also known as "hallucinations". Technical advancements have been mad... 031
Reposted by Werner GeyerAmy Zhang @axz.bsky.social · 05/12/2024If you are headed to NeurIPS, please join for our Pluralistic Alignment workshop! We have a great set of speakers from a range of backgrounds. & all the papers that will be presented at the workshop are posted: pluralistic-alignment.github.io lmk if you'd like to catch up at the conference too! :)pluralistic-alignment.github.ioPluralistic Alignment @ NeurIPS 2024Pluralistic Alignment 1429
Reposted by Werner GeyerHendrik Strobelt @henstr.bsky.social · 03/12/2024🎺 Here comes the official 2024 NeurIPS paper browser: - browse all NeurIPS papers in a visual way - select clusters of interest and get cluster summary - ZOOOOM in - filter by human assigned keywords - filter by substring (authors, titles) neurips2024.vizhub.ai #neurips by IBM Research Cambridge 56021
Reposted by Werner GeyerAmy Zhang @axz.bsky.social · 03/12/2024In collab w/ Semantic Scholar, we conducted a large scale (800+) survey of researcher usage and perceptions of LLMs for science. Major findings: +Most are using LLMs already, mostly for writing +LLMs seem to be a win for research equity +But some groups, like women, have more ethical concerns too 37914
Werner Geyer @wernergeyer.bsky.social · 03/12/2024EvalAssist supports both general LLM judges and specialized judges. Our demo @neuripsconf.bsky.social also features IBM's brand new Granite Guardian 3.0 models that allow you to detect harms and risks in LLM-generated content. Check out a demo video here: ibm.ent.box.com/file/1675887...ibm.ent.box.comGranite Guardian 3.0 Demo .mp4 | Powered by Box 021
Werner Geyer @wernergeyer.bsky.social · 02/12/2024We also do have a paper on uncertainty quantification for LLM-as-a-judge at the NeurIPS 2024 workshop on Statistical Frontiers in LLMs and Foundation Models. That one is on Saturday if you are still around. arxiv.org/abs/2410.11594arxiv.orgBlack-box Uncertainty Quantification Method for LLM-as-a-JudgeLLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty of LLM-as-a-Judge ev... 031
Werner Geyer @wernergeyer.bsky.social · 02/12/2024You can read more about this work here: arxiv.org/abs/2410.00873arxiv.orgAligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy PreferencesEvaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes time given the large a... 010