Sign in

Atharva Sehgal

@aseg.bsky.social
82 followers 439 following 28 posts

Postdoc scholar @Caltech | PhD from @UTAustin

PostsRepliesMedia
Atharva Sehgal @aseg.bsky.social · 06/07/2026
To quick add these as events: atharvas.net/icml/
atharvas.net
Atharva Sehgal @ ICML 2026 — Talks, Times & Locations
Atharva Sehgal is presenting three projects at ICML 2026 in Seoul: FormulaCode, PWW-Bench, and Vendi-Evolve. Times, locations at the COEX Convention Center, and add-to-calendar links.
000
Atharva Sehgal @aseg.bsky.social · 06/07/2026
Vendi-Evolve: icml.cc/virtual/202... (Saturday 3:45PM, Room 401) ShinkaEvolve extension investigating an information theoretic diversity measure for improving how LLM-guided Evolutionary agents explore hypothesis.
100
Atharva Sehgal @aseg.bsky.social · 06/07/2026
PWW-Bench (w/ @sabrinareguyal): icml.cc/virtual/202... (Saturday 11AM, Hall D1) This benchmark evaluates VLMs mathematical problem solving skills.
120
Atharva Sehgal @aseg.bsky.social · 06/07/2026
FormulaCode: icml.cc/virtual/202... (Tuesday 2PM; Hall A #2201) x.com/atharva_seh...
100
Atharva Sehgal @aseg.bsky.social · 06/07/2026
I'll be at #ICML this week presenting three projects generally structured around self-evolving agents. Go to atharvas.net/icml to mark these on your calendar! Always happy to chat! Feel free to email me (atharvas@caltech.edu) or book a time through: calendar.app.google/Ws8qzdpLRRh...
calendar.google.com
Google Calendar - Easier Time Management, Appointments & Scheduling
Learn how Google Calendar helps you stay on top of your plans - at home, at work and everywhere in between.
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
Generally, the community keeps itself invariant to such confounders. i.e., our harness runs tasks on a hardware isolated EC2 instance. This version of the dataset is all in python so no build system other than uv. I've tried some MVPs for other languages and I have some thoughts there (DM me!).
000
Atharva Sehgal @aseg.bsky.social · 13/05/2026
FormulaCode will appear at #ICML2026! This is joint work with James Hou, Akanksha Sarkar, Ishaan Mantripragada, @swarat, @JenJSun, and @yisongyue. Huge thanks to @LaudeInstitute for supporting this project. See y'all in Korea!
000
Atharva Sehgal @aseg.bsky.social · 13/05/2026
More findings, leaderboard, code, and instructions for contributing tasks at www.formulacode.org. We have some RLVR recipes w/ harbor coming up soon. Would love to see what the community does with it!
formulacode.org
FormulaCode
FormulaCode is the first large-scale analysis of the holistic ability of LLM agents to optimize codebases.
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
Where did agents win / lose? They match or beat humans on tasks where humans optimized by parallelization or batching. But, they were much worse when humans reached for a lower-level library (numpy/pandas) for optimizing the code (agents had a proclivity for datastructure solns.)
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
Negotiating multi-workload tradeoffs is where experts pulled cleanly ahead, and probably where future agents have the most room to learn. Experts, on average, accept ~10% slowdowns on some workloads *while* landing bigger global optimizations.
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
Since we measure multiple targets per task, we can stratify performance by module, class, and function. Agents carry diverse profiles. Sonnet has stronger module-level refactors but GPT-5 wins at the function level. A single-target benchmark hides cannot surface such insights.
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
Every frontier agent we tested makes real code faster than baseline (Geomean Speedup), including GPT-5, Claude 4 Sonnet, Gemini 2.5 Pro, and Qwen 3 Coder under both Terminus 2 and OpenHands. But none of them consistently beat the human experts (Advantage).
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
FormulaCode also enables studying the long-tail retrieval properties of LLM agents at scale since we sample problems from 70+ repos. Previous efforts have only sampled 9-10 repos from the head of this distribution.
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
FormulaCode is also a continually updated benchmark. We added ~27 new tasks every month of 2025 with reproducible Docker environments, letting us measure how agents improve over time without contamination. Follow FormulaCode's latest statistics on data.formulacode.org!
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
The closest perf. benchmarks evaluate each task against a single optimization target. This misses what makes real optimization hard: speeding up one target can quietly regress others. FormulaCode pairs each of 957 tasks with ~264 crowdsourced workloads on average across 70 repos.
100
Atharva Sehgal @aseg.bsky.social · 13/05/2026
Excited to share FormulaCode, a continually updating benchmark for evaluating the holistic ability of LLM agents to optimize codebases. Our current dataset consists of 957 tasks curated from 245477 pull requests in 70+ repositories (and growing!). 🌐 formulacode.org 🧵👇
formulacode.org
FormulaCode
FormulaCode is the first large-scale analysis of the holistic ability of LLM agents to optimize codebases.
210
Reposted by Atharva Sehgal
Neehar Kondapaneni @therealpaneni.bsky.social · 08/07/2025
You’ve generated 10k concepts with your favorite XAI method -- now what? Many concepts you’ve found are fairly obvious and uninteresting. What if you could 𝑠𝑢𝑏𝑡𝑟𝑎𝑐𝑡 obvious concepts away and focus on the more complex ones? We tackle this in our latest preprint!
112
Reposted by Atharva Sehgal
LM4Sci @ COLM2025 @lm4sci.bsky.social · 22/06/2025
Deadline Extended! Submit to the LM4Sci Workshop @ COLM 2025 in Montreal 🇨🇦 🧠 Large Language Modeling for Scientific Discovery (LM4Sci) 📅 New Deadline: June 30 📢 Notification: July 24 📍 Workshop: Oct 10, 2025 📝 Non-archival short (2–4p) & full (up to 8p) papers welcome!
053
Reposted by Atharva Sehgal
LM4Sci @ COLM2025 @lm4sci.bsky.social · 14/06/2025
🚨 Call for Papers: LM4Sci @COLM_conf 2025 🚨 Excited to announce the Large Language Modeling for Scientific Discovery (LM4Sci) workshop at COLM 2025 in Montreal, Canada! Submission Deadline: June 23 Notification: July 24 Workshop: October 10, 2025
156
Atharva Sehgal @aseg.bsky.social · 13/06/2025
Check out the full paper for the mathematical formulation, experiments, and our methodology: arxiv.org/abs/2504.00185 Code and other artifacts are available here: trishullab.github.io/escher-web/ Thank you for following along!
arxiv.org
Self-Evolving Visual Concept Library using Vision-Language Critics
We study the problem of building a visual concept library for visual recognition. Building effective visual concept libraries is challenging, as manual definition is labor-intensive, while relying sol...
010
Atharva Sehgal @aseg.bsky.social · 13/06/2025
How it works: 1️⃣ LLM proposes concepts per class 2️⃣ CLIP-style VLM scores them 3️⃣ Escher spots confused classes 4️⃣ Escher stores this in a history bank 5️⃣ LLM proposes better concepts and stores them → repeat The loop is self-amplifying: better concepts ➡️ better feedback ➡️ an even better concept library.
110
Atharva Sehgal @aseg.bsky.social · 13/06/2025
Escher solves this problem using feedback from a vision language model to improve the reasoning, specifically for fine-grained image classification.
100
Atharva Sehgal @aseg.bsky.social · 13/06/2025
Our hypothesis: the failure arises from the program synthesizers treating the vision model as a deterministic function. Reality is messy and the VLM outputs are stochastic. The LLMs assumptions of how the VLM will behave and how it actually behaves are decoupled. We need to overcome this decoupling.
100
Atharva Sehgal @aseg.bsky.social · 13/06/2025
A visual program decomposes complex perceptual reasoning problems into a logical combination of simpler perceptual tasks that can be solved using off-the-shelf vision foundation models. This provides a modular and robust framework, but finding the correct decomposition is still extremely hard.
Even with visual programming, the LLM proposing the program has no idea about the execution semantics of the underlying VLM. Things still don't work.
100
Atharva Sehgal @aseg.bsky.social · 13/06/2025
Reasoning about these images is pretty hard. o3 – even with web access – can’t do this for us out of the box. In such a situation, writing programs provides a mechanism for dividing up a complex reasoning task into solvable subtasks. This motivates most of the visual programming literature.
gpt-o3, which has probably seen this image before, reasons incorrectly about the type of lizard and gets it wrong. Visual feedback is extremely important here!
100
Atharva Sehgal @aseg.bsky.social · 13/06/2025
In many vision tasks, perceptual reasoning does not come naturally. Experts still have to deeply study an image, deduce relevant concepts, and reason about them in natural language (www.inaturalist.org/observations...). Our goal is to automate this process – with no human oversight.
An example from inaturalist of two scientist deliberating how to classify a rare lizard. The first scientists gets it wrong because they aren't trained as a herpetologist. The second scientist is a trained herpetologist, and  reasons in natural language how to correctly identify the image.
110
Atharva Sehgal @aseg.bsky.social · 13/06/2025
Massive thanks to my co-authors Patrick Yuan, Ziniu Hu, @yisongyue.bsky.social, Jennifer J. Sun & @swarat.bsky.social for making this possible!
100
Atharva Sehgal @aseg.bsky.social · 13/06/2025
I’m presenting Escher (trishullab.github.io/escher-web) at #cvpr2025 Saturday morning (Poster Session #3). Escher builds a visual concept library with a vision‑language critic (no human labels needed). Swing by if you’d like to chat about program synthesis & multimodal reasoning!
183
Atharva Sehgal @aseg.bsky.social · 13/02/2025
Just julia things.
010
Reposted by Atharva Sehgal
Miles Cranmer @milescranmer.bsky.social · 01/12/2024
Happy to announce the PySR v1.0 release! github.com/MilesCranmer... PySR lets you do high-performance symbolic regression from Python. Now, you can learn multiple symbolic expressions simultaneously! Also: + Parametric expressions + TensorBoard support + Improved search + Julia-based inference
15411
Reposted by Atharva Sehgal
Swarat Chaudhuri @swarat.bsky.social · 06/12/2024
Missing NeurIPS this year but wanted to highlight our new paper on LLM-guided genetic programming: trishullab.github.io/lasr-web/ Our method, LaSR, conditions mutation/crossover operators on (1) an LLM's general domain knowledge, and (2) LLM-generated abstractions of high-performing programs. (1/2)
182
Atharva Sehgal @aseg.bsky.social · 10/12/2024
Check out the full paper for the mathematical formulation, llm scaling law experiments, and our methodology: arxiv.org/abs/2409.09359 More context here: x.com/atharva_sehg... Thank you to all my coauthors: Arya, Omar, @milescranmer.bsky.social, and @swarat.bsky.social!
x.com
x.com
000
Atharva Sehgal @aseg.bsky.social · 10/12/2024
Arya and I'll be at #NeurIPS presenting LaSR (trishullab.github.io/lasr-web/) on Wednesday morning 11AM PST to 2PM PST (East Exhibit Hall A-C #4003). Drop by and say Hi!
193