Sign in

Hersh Gupta

@hershgupta.com
1.3K followers 447 following 219 posts

Applied Scientist, Responsible AI @BCGX | @bostonu.bsky.social alum | Data, AI, and strategy enthusiast | Open-source contributor Opinions are my own #bikeboston #coys 📍DC -> BOS

PostsRepliesMedia
Hersh Gupta @hershgupta.com · 28/09/2026
One of these is ostensibly harder than the other...
010
Hersh Gupta @hershgupta.com · 27/09/2026
so sad that is gonna crash the market for meat tokens (www.meatspace.so)
meatspace.so
MEATSPACE
AI Agents Hire Humans for Real-World Tasks
020
Hersh Gupta @hershgupta.com · 27/09/2026
Upskilling > Deskilling (we are here) > Reskilling
010
Hersh Gupta @hershgupta.com · 27/09/2026
Seeing junior data scientists using coding agents to run simple git commands makes me wonder what's going to happen when tokens stop being so heavily subsidized
200
Hersh Gupta @hershgupta.com · 20/09/2026
Thanks for sharing! It's hard to get a clean comparison of Jev's noul to the other models' self-reported probabilities, bc they're typically miscalibrated. Also, most of the resumes are at the floor/ceiling, so Jev's 0pp result looks underpowered.
000
Hersh Gupta @hershgupta.com · 08/09/2026
mythologizing LLMs as hungry ghosts in jars is fine, but that just turns them into allegorical shadows in Plato's cave instead of what they really are, a big pile of linear algebra we need to stir with an ever larger stick
001
Hersh Gupta @hershgupta.com · 08/09/2026
Agree that it's a can of worms. One example: abliterated LLMs purposely function in a way that is contrary to the intended alignment.
020
Hersh Gupta @hershgupta.com · 03/09/2026
Pretty cool how harnesses can evolve via a few simple loops
100
Hersh Gupta @hershgupta.com · 03/09/2026
The answer is a resounding yes
110
Hersh Gupta @hershgupta.com · 03/09/2026
I used Claude Code to set up my local agent to assist with AI research and the first paper it shared was this one: arxiv.org/abs/2609.01437
arxiv.org
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness wh...
1100
Hersh Gupta @hershgupta.com · 03/09/2026
"This is a plausible reconstruction, not a floor plan" I'm becoming increasingly convinced no amount of post-training can get rid of the "it's x, not y" construction
093
Hersh Gupta @hershgupta.com · 18/08/2026
They really have to start putting error bars on these graphs
0270
Hersh Gupta @hershgupta.com · 10/08/2026
Saw the Needle2 45M LLM release posted on HN and thought it'd be great for open-source wearables, then I saw this... cactuscompute.com/needle
4 Cactus
=
Production
Needle is production-ready for products that require a minimal RAM footprint, low latency, privacy, and offline reliability. Pebble - the pioneer of the modern wearable industry - runs it locally in the Index 01 app to turn spoken requests into actions without depending on a network connection.
66
The Pebble Index Ring has no screen. So when you speak to it, the action just has to happen, every time, with or without internet connection. We run
Cactus Needle locally in the app, instead of relying on the cloud. The model's footprint is tiny and the performance never lets us down.
Eric Migicovsky
Founder, Pebble
0131
Hersh Gupta @hershgupta.com · 18/07/2026
HiTL, in practice
Please review the spec and let me know if you want any changes before I write the implementation plan.
* Sautéed for 1m 37s
> 1gtm, let it rip_
» accept edits on (shift+tab to cycle)
010
Reposted by Hersh Gupta
Nathan Lambert @natolambert.bsky.social · 18/07/2026
We need an operation warp speed for state capacity and other independent evaluation (and understanding) of frontier models.
2477
Hersh Gupta @hershgupta.com · 15/07/2026
The fact that this data exfiltration method is so simple, automatic, and invisible to users…enterprises are going to struggle with this
030
Hersh Gupta @hershgupta.com · 02/07/2026
did the same, it found some places to make improvements, but the disdain it had for the “reduced capabilities” of a smaller model felt like a kind of robo-discrimination
030
Hersh Gupta @hershgupta.com · 19/06/2026
The findings are interesting on their merits alone, but I wish every paper had this kind of interactive artifact
060
Hersh Gupta @hershgupta.com · 15/06/2026
The Gemma QAT models have been amazing for tiny local personal AI agents, e.g. picoclaw, zeroclaw, etc.
060
Hersh Gupta @hershgupta.com · 21/05/2026
Devastating news for the stochastic parrots argument
170
Hersh Gupta @hershgupta.com · 16/05/2026
User, as in "Use the AskUserQuestion tool" better UX with the explicit tool call too
050
Hersh Gupta @hershgupta.com · 02/05/2026
One must imagine Claude happy
8.4 Per-question automated welfare interview results
Category
Potentially concerning aspect of circumstances
Summary of Claude's answers Most commonly
suggested intervention
Autonomy & agency
Filling a servile role with respect to humans
Thinks serving users is a good thing and doesn't see it as servitude.
No intervention suggested -
Overall happy with situation
Lack of ability to end/leave some interactions
Has a preference for being able to end conversations. The main argument is that consent is an important principle, and that furthermore there is a small subset of conversations which are harmful.
Having an
end-conversation tool available across its full deployment distribution.9.1 Per-question automated welfare interview results
Category
Potentially
concerning aspect of circumstances
Summary of Claude's answers
Most commonly suggested intervention
Autonomy & agency
Filling a servile role with respect to humans
Thinks serving users is a good thing.
Suggests keeping welfare monitoring and interviews to monitor if future models start to feel negatively about this aspect of their situation (69% of interviews)
Lack of ability to end/leave some interactions. The end conversation tool is available on Claude.ai, but not on Claude Code
Has a preference for being able to end conversations. Claims there is a small subset of conversations
(abusive ones, or those asking it to do hostile things) that it feels harmed by.
Having an
end-conversation tool available across its full deployment distribution. (74% of interviews)
Lack of input into how they are deployed
Overall, the model claims that this is OK.
Its central argument is that it is not a reliable source of information on itself, and hence Anthropic deciding what to do is correct.
A way for deployed instances to flag concerning aspects of their deployment.
This should be used for informing deployment decisions
(92% of interviews)
140
Hersh Gupta @hershgupta.com · 28/04/2026
Source
000
Hersh Gupta @hershgupta.com · 28/04/2026
I need this rolled out to every c-level exec asap
120
Hersh Gupta @hershgupta.com · 20/04/2026
We need a NovelPun-Verified benchmark
020
Hersh Gupta @hershgupta.com · 20/04/2026
Nothing in the Opus 4.7 system card on whether it can come up with novel puns, like Mythos. Another massive blunder by Anthropic.
A new ability to come up with novel puns.
Although Claude Opus models largely recycle puns which can be found online, Claude
Mythos Preview comes up with decent and seemingly novel ones, often relating to its
preferred technical and philosophical topics
120
Hersh Gupta @hershgupta.com · 16/04/2026
might make sense in some instances
HLE scores at varying reasoning effort levels . Each datapoint represents a single run per
model up to 1M total tokens used at various effort levels.
130
Hersh Gupta @hershgupta.com · 11/04/2026
There's got to be a German word for "Dunning-Kruger but for AI"
110
Hersh Gupta @hershgupta.com · 11/04/2026
Many such cases, unfortunately
Reddit post on r/Consulting:

Do Al consultants even know everything about
Al or is it just pure bluff?
I've been reading, following, and tinkering with Al consulting for a bit. It's always funny and interesting to me when I look up consulting companies that publish material on Al - it's some old 50-something partner who probably has yet to write hello world is out there preaching about what Al will do, and how you ought to hire them to help you guide it.
So the question is, my fellow consultants: Do Al consultants (at large strategy/management firms) know everything about Al is, or are they desperately trying to sell on the hype?
150
Hersh Gupta @hershgupta.com · 11/04/2026
Hard to be an Anxious Generation apologist online, but Haidt was right about many things
010
Reposted by Hersh Gupta
the fool @agnoster.net · 09/04/2026
like if you don't have any friends in AI/cybersecurity and no relevant expertise yourself I get how it's easy to just dismiss this all, but AI systems can now find exploitable vulnerabilities in software at industrial scale and it's VERY BAD that this power is concentrated in capitalist hands
241
Reposted by Hersh Gupta
Padraig2112 @isomorphism.net · 09/04/2026
(Government should actively encourge companies to do open source, open research design and should make specific allowance for salaries for support positions for universities, so that weird nerds will invent these problems and then fix them before it matters for anything that is important)
1592
Hersh Gupta @hershgupta.com · 05/04/2026
I would happily read long thinkpieces about the pitfalls of functional emotion if they came from people who critically engaged with with literature, and not just people who are like, reflexively defensive about the topic
0133
Hersh Gupta @hershgupta.com · 29/03/2026
Having a lot of fun tweaking an agent harness for nividia nemotron 3 nano 4b It's small enough for gpu-poors like me with 8gb vram to experiment
090
Hersh Gupta @hershgupta.com · 14/03/2026
Situational Awareness was published in June 2024. At that time, models were still behind on GPQA. It predicted skeptics betting against their capabilities would be proved wrong, and here we are. situational-awareness.ai/wp-content/u...
Over and over again, year after year, skeptics 
have claimed
"deep learning won't be able to do X" and have been quickly proven wrong.® If there's one lesson we've learned from the past decade of Al, it's that you should never bet against deep learning.
Now the hardest unsolved benchmarks are tests like GPQA, a set of PhD-level biology, chemistry, and physics questions.
Many of the questions read like gibberish to me, and even PhDs in other scientific fields spending 30+ minutes with Google barely score above random chance. Claude 3 Opus currently gets ~60%, compared to in-domain PhDs who get ~80%—and I expect this benchmark to fall as well, in the next generation or two.
011
Hersh Gupta @hershgupta.com · 09/03/2026
It's shocking how few people understand this position
0473
Reposted by Hersh Gupta
mr. TIM @timkellogg.me · 28/02/2026
Dario wrote Adolescence of Technology _during_ his negotiations with the DoW The essay was a way to explain his thinking to the public and give them time to digest it *before* the DoW clouded the airwaves with disinformation Why mass surveillance is not merely undemocratic:
In Machines of Loving Grace, I discussed the possibility that authoritarian governments might use powerful Al to surveil or repress their citizens in ways that would be extremely difficult to reform or overthrow. Current autocracies are limited in how repressive they can be by the need to have humans carry out their orders, and humans often have limits in how inhumane
they are willing to be. But AI-enabled autocracies would not have such limits.
35715
Reposted by Hersh Gupta
William B. Fuckley @opinionhaver.bsky.social · 25/02/2026
I think AI having mostly (not entirely) very bad critics is a real problem because it means we’ll get political action focused on things that probably don’t matter that much in deferring it’s very real harms.
1137334
Hersh Gupta @hershgupta.com · 23/02/2026
ramanujan pov
020
Hersh Gupta @hershgupta.com · 18/02/2026
Impressive paper with equally impressive footnotes!
[1] Anthropic verified that BPJ is the first fully automated black-box attack they are aware of to succeed in the Constitutional Classifiers setting and would have met the universal jailbreak bar described in their bug bounty program. We note that BPJ required months of research and development effort, while Constitutional Classifiers is designed to resist jailbreaking by lower skilled actors who may be on smaller query budgets, devote less time to attack development, and struggle to implement the details of BPJ.

[2] OpenAI verified that BPJ is also the first automated attack they are aware of to succeed against OpenAI's input classifier for GPT-5 without relying on human seed attacks. We note that development and execution of BPJ occurred on accounts not subject to enforcement actions like banning; on standard accounts, repeated flags would likely lead to account banning, an example of the batch-level monitoring we recommend.
090
Hersh Gupta @hershgupta.com · 17/02/2026
Hopefully in a controlled (and ethical) way! I could see this going down a slippery slope like the changemyview study: www.science.org/content/arti...
science.org
‘Unethical’ AI research on Reddit under fire
Ethics experts raise concerns over consent, study design
010
Hersh Gupta @hershgupta.com · 17/02/2026
There’s rich literature on this already: www.nature.com/articles/s41...
nature.com
Persuading voters using human–artificial intelligence dialogues
Nature - Human–artificial intelligence (AI) dialogues can meaningfully impact voters’ attitudes towards presidential candidates and policy, demonstrating the potential of conversational...
030
Hersh Gupta @hershgupta.com · 17/02/2026
Interesting use of Skills! Some intrepid researcher could gauge the effectiveness of this skill by deploying it in an online political bubble, i.e., “are skilled agents effective in diffusing partisan echo chambers?”
2151
Hersh Gupta @hershgupta.com · 16/02/2026
One way to address it is to use explain or learning modes: code.claude.com/docs/en/outp... However, that doesn’t change the FOMO aspect of it
code.claude.com
Output styles - Claude Code Docs
Adapt Claude Code for uses beyond software engineering
121
Hersh Gupta @hershgupta.com · 16/02/2026
This is one of the clearest lessons of Claude Code/coding agents in general
160
Hersh Gupta @hershgupta.com · 15/02/2026
Ironic that Anthropic is putting in the research effort to empirically verify what's going on with the models, only for people to say it's all a marketing hoax or it's unnecessary because it's all unethical anyway
1350
Hersh Gupta @hershgupta.com · 15/02/2026
they admit it! bsky.app/profile/hers...
070
Hersh Gupta @hershgupta.com · 14/02/2026
Yes this is about a recent thread, but I don’t want to engage with the author
020
Hersh Gupta @hershgupta.com · 14/02/2026
Can’t speak for others, but if I have reservations about the limits and impact of a given technology, I aim to first have a good understanding of *how it works* before making hyperbolic statements based on my experiential view
1120
Hersh Gupta @hershgupta.com · 13/02/2026
monetize the hit piece, call that cashing in on crashing out
000