Reposted by Spandana GellaGaurav Kamath @grvkamath.bsky.social · 04/03/2026 🚨New Paper!🚨 How do reasoning LLMs handle inferences that have no deterministic answer? We find that they diverge from humans in some significant ways, and fail to reflect human uncertainty… 🧵(1/10) 35820
Spandana Gella @spandanagella.bsky.social · 17/06/2025Our team is hiring an intern discrete diffusion of text and/or code. Please apply! 020
Reposted by Spandana GellaPatrice Bechard @patricebechard.bsky.social · 29/05/2025🚀 New paper from our team at @servicenowresearch.bsky.social! 💫𝐒𝐭𝐚𝐫𝐅𝐥𝐨𝐰: 𝐆𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐧𝐠 𝐒𝐭𝐫𝐮𝐜𝐭𝐮𝐫𝐞𝐝 𝐖𝐨𝐫𝐤𝐟𝐥𝐨𝐰 𝐎𝐮𝐭𝐩𝐮𝐭𝐬 𝐅𝐫𝐨𝐦 𝐒𝐤𝐞𝐭𝐜𝐡 𝐈𝐦𝐚𝐠𝐞𝐬 We use VLMs to turn 𝘩𝘢𝘯𝘥-𝘥𝘳𝘢𝘸𝘯 𝘴𝘬𝘦𝘵𝘤𝘩𝘦𝘴 and diagrams into executable workflows 🖍️→⚙️ 🔗 arxiv.org/abs/2503.218... 📝 tinyurl.com/3utdbn97%E2%... #Sketch2Flow #AI #VLM 101
Reposted by Spandana GellaXiangru (Edward) Jian @edwardjian.bsky.social · 15/05/2025🚀 Excited to share that UI-Vision has been accepted at ICML 2025! 🎉 We have also released the UI-Vision grounding datasets. Test your agents on it now! 🚀 🤗 Dataset: huggingface.co/datasets/Ser... #ICML2025 #AI #DatasetRelease #Agentshuggingface.coServiceNow/ui-vision · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 001
Spandana Gella @spandanagella.bsky.social · 24/03/2025Very excited to announce our GUI benchmarking dataset UI-Vision : uivision.github.io Our evals reveal current GUI-models struggle with grounding small elements, dense UIs and has limited domain/spatial/motion understanding. Watch out this space for more exciting stuff from us!uivision.github.ioUI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and InteractionUI-Vision 030
Spandana Gella @spandanagella.bsky.social · 10/03/2025Web agents powered by LLMs can solve complex tasks, but our analysis shows that they can also be easily misused to automate harmful tasks. See the thread below for more details on our new web agent safety benchmark: SafeArena and Agent Risk Assessment framework (ARIA). 052
Reposted by Spandana GellaKarolina Stańczak @karstanczak.bsky.social · 04/03/2025📢New Paper Alert!🚀 Human alignment balances social expectations, economic incentives, and legal frameworks. What if LLM alignment worked the same way?🤔 Our latest work explores how social, economic, and contractual alignment can address incomplete contracts in LLM alignment🧵 12713
Reposted by Spandana GellaAarash Feizi @aarashfeizi.bsky.social · 27/02/2025🚨 Excited to introduce PairBench! 🚨 💡 TL;DR: VLM-judges can fail at data comparison! ✅ PairBench helps you pick the right one by testing alignment, symmetry, smoothness & controllability—ensuring reliable auto-evaluation. 📄 Paper: arxiv.org/abs/2502.15210 🧵 Thread: 👇 112
Reposted by Spandana GellaAlexandre Lacoste @alex-lacoste.bsky.social · 12/12/2024We’re really excited to release this large collaborative work for unifying web agent benchmarks under the same roof. In this TMLR paper, we dive in-depth into #BrowserGym and #AgentLab. We also present some unexpected performances from Claude 3.5-Sonnet 12111
Spandana Gella @spandanagella.bsky.social · 12/12/2024If you want to know all about the exciting stuff we do with web agents @servicenowresearch.bsky.social register here and interact with our team including the amazing @alex-lacoste.bsky.social and @adrouinenv.bsky.social 020
Spandana Gella @spandanagella.bsky.social · 10/12/2024Thrilled to launch BigDocs—an open multimodal dataset set to transform document understanding! Our contribution to VLM community, supporting transparency in multimodal document reasoning. Proud to work with the most passionate and amazing team @servicenowresearch.bsky.social ! 140
Reposted by Spandana GellaAlexandre Lacoste @alex-lacoste.bsky.social · 03/12/2024🧵-1 We are thrilled to release #AgentLab, a new open-source package for developing and evaluating web agents. This builds on the new #BrowserGym package which supports 10 different benchmarks, including #WebArena. 21815