Sign in

Alex Gill

@agill32.bsky.social
399 followers 366 following 15 posts

NLP researcher at U of U

PostsRepliesMedia
Alex Gill @agill32.bsky.social · 02/10/2026
Work done with an amazing team! @farhanishmam.bsky.social , Xuyen Nguyen, Neha Bhat, Parker DeYoung, @fatemehc.bsky.social , Nathan Stringham, @kennethmarino.bsky.social and @anamarasovic.bsky.social. Github: github.com/alexgill321/KNOWS-benchmark Paper: arxiv.org/abs/2609.30604
github.com
GitHub - alexgill321/KNOWS-benchmark: KNOWS: a benchmark for web agents on realistic Google Workspace tasks (Docs/Sheets/Slides) with hybrid white-box evaluators and step-level failure categories
KNOWS: a benchmark for web agents on realistic Google Workspace tasks (Docs/Sheets/Slides) with hybrid white-box evaluators and step-level failure categories - alexgill321/KNOWS-benchmark
021
Alex Gill @agill32.bsky.social · 02/10/2026
Key Findings: • Long-horizon computer use tasks remain challenging. The best baseline, Comet, achieves only a 3% complete SR. • Harness matters as much as the model. Opus 4.7 with Comet outperforms the Browsergym harness by +27.5 ACF.
110
Alex Gill @agill32.bsky.social · 02/10/2026
Evaluation of agents is typically a tradeoff between richness, reliability and automation. We use a hybrid approach with our evaluators, combining the richness of LLM judgements with robust deterministic checks through Google Workspace APIs.
100
Alex Gill @agill32.bsky.social · 02/10/2026
KNOWS has 110 tasks across 20 domains, each of which conforms to a rubric of design requirements and multiple rounds of review. Each task took between 9-18 hours of expert effort to complete.
100
Alex Gill @agill32.bsky.social · 02/10/2026
We task agents with a detailed workflow, then grade the artifact it produces against fine-grained evaluation criteria.
100
Alex Gill @agill32.bsky.social · 02/10/2026
A useful computer-use agent doesn't just find information. It turns it into something you can share. Introducing KNOWS: agents research the live web, then build a real Google Doc, Sheet or Slide deck in the browser.
161
Reposted by Alex Gill
Nathan Kalman-Lamb @nkalamb.bsky.social · 21/11/2025
Folks, I don’t know how it’s possible, but it gets funnier.
16467106
Alex Gill @agill32.bsky.social · 03/11/2025
I'll be in Suzhou 🇨🇳 at #EMNLP this week presenting "What has been Lost with Synthetic Evaluation?" done with @anamarasovic.bsky.social & @lasha.bsky.social! 🎉 📍Findings Session 1 - Hall C 📅 Wed, November 5, 13:00 - 14:00 arxiv.org/abs/2505.22830
0122
Reposted by Alex Gill
Women in AI Research - WiAIR @wiair.bsky.social · 20/10/2025
🧠 Can large language models build the very benchmarks used to evaluate them? In “What Has Been Lost with Synthetic Evaluation”, Ana Marasović (@anamarasovic.bsky.social) and collaborators ask what happens when LLMs start generating the datasets used to test their reasoning. (1/6🧵)
2103
Alex Gill @agill32.bsky.social · 04/06/2025
More results and analysis can be found in the paper. We welcome any discussion, thanks for reading!!
010
Alex Gill @agill32.bsky.social · 04/06/2025
We hope that our work will inspire future research into: - Can further prompt review improve the difficulty of synthetic data? - What other axes (representativeness, diversity) are affected when using LLMs to generate benchmarks?
110
Alex Gill @agill32.bsky.social · 04/06/2025
Key takeways: - While LLM generated evals may be 𝑣𝑎𝑙𝑖𝑑, as a whole they lose crucial aspects in complexity. - LLMs are promising where complexity is less critical, but human annotators are vital for benchmarks assessing real-world generalization & nuanced scenarios.
100
Alex Gill @agill32.bsky.social · 04/06/2025
But are these instances similarly difficult? We explore the difficulty of synthetic benchmarks by comparing performance on synthetic & human-written data across a suite of models. We find that performance is consistently higher on generated versions of the datasets.
100
Alex Gill @agill32.bsky.social · 04/06/2025
We perform a human study and even find that LLM-generated data is preferred! We ask NLP researchers to act as dataset creators and gather preferences between synthetic and human-authored data.
100
Alex Gill @agill32.bsky.social · 04/06/2025
We examine both the 𝑣𝑎𝑙𝑖𝑑𝑖𝑡𝑦 and 𝑑𝑖𝑓𝑓𝑖𝑐𝑢𝑙𝑡𝑦 of LLM-generated versions of two high-quality reading comprehension datasets: CondaQA & DROP. We find that validity is not an issue. We are able to get LLMs to generate instances that are highly valid according to our dataset specs.
100
Alex Gill @agill32.bsky.social · 04/06/2025
We are increasingly seeing LLMs being used to create challenging benchmarks that are then used for evaluating LLMs. Is this a valid approach to evaluation construction? Do we lose anything in this process?
100
Alex Gill @agill32.bsky.social · 04/06/2025
𝐖𝐡𝐚𝐭 𝐇𝐚𝐬 𝐁𝐞𝐞𝐧 𝐋𝐨𝐬𝐭 𝐖𝐢𝐭𝐡 𝐒𝐲𝐧𝐭𝐡𝐞𝐭𝐢𝐜 𝐄𝐯𝐚𝐥𝐮𝐚𝐭𝐢𝐨𝐧? (arxiv.org/abs/2505.22830) I'm happy to announce that the preprint release of my first project is online! Developed with the amazing support of @lasha.bsky.social & @anamarasovic.bsky.social
arxiv.org
What Has Been Lost with Synthetic Evaluation?
Large language models (LLMs) are increasingly used for data generation. However, creating evaluation benchmarks raises the bar for this emerging paradigm. Benchmarks must target specific phenomena, pe...
1124