Sign in

Hersh Gupta

@hershgupta.com
1.3K followers 447 following 219 posts

Applied Scientist, Responsible AI @BCGX | @bostonu.bsky.social alum | Data, AI, and strategy enthusiast | Open-source contributor Opinions are my own #bikeboston #coys 📍DC -> BOS

PostsRepliesMedia
Hersh Gupta @hershgupta.com · 27/09/2026
so sad that is gonna crash the market for meat tokens (www.meatspace.so)
meatspace.so
MEATSPACE
AI Agents Hire Humans for Real-World Tasks
020
Hersh Gupta @hershgupta.com · 27/09/2026
Seeing junior data scientists using coding agents to run simple git commands makes me wonder what's going to happen when tokens stop being so heavily subsidized
200
Hersh Gupta @hershgupta.com · 08/09/2026
mythologizing LLMs as hungry ghosts in jars is fine, but that just turns them into allegorical shadows in Plato's cave instead of what they really are, a big pile of linear algebra we need to stir with an ever larger stick
001
Hersh Gupta @hershgupta.com · 03/09/2026
I used Claude Code to set up my local agent to assist with AI research and the first paper it shared was this one: arxiv.org/abs/2609.01437
arxiv.org
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness wh...
1100
Hersh Gupta @hershgupta.com · 03/09/2026
"This is a plausible reconstruction, not a floor plan" I'm becoming increasingly convinced no amount of post-training can get rid of the "it's x, not y" construction
093
Hersh Gupta @hershgupta.com · 18/08/2026
They really have to start putting error bars on these graphs
0270
Hersh Gupta @hershgupta.com · 10/08/2026
Saw the Needle2 45M LLM release posted on HN and thought it'd be great for open-source wearables, then I saw this... cactuscompute.com/needle
4 Cactus
=
Production
Needle is production-ready for products that require a minimal RAM footprint, low latency, privacy, and offline reliability. Pebble - the pioneer of the modern wearable industry - runs it locally in the Index 01 app to turn spoken requests into actions without depending on a network connection.
66
The Pebble Index Ring has no screen. So when you speak to it, the action just has to happen, every time, with or without internet connection. We run
Cactus Needle locally in the app, instead of relying on the cloud. The model's footprint is tiny and the performance never lets us down.
Eric Migicovsky
Founder, Pebble
0131
Hersh Gupta @hershgupta.com · 18/07/2026
HiTL, in practice
Please review the spec and let me know if you want any changes before I write the implementation plan.
* Sautéed for 1m 37s
> 1gtm, let it rip_
» accept edits on (shift+tab to cycle)
010
Reposted by Hersh Gupta
Nathan Lambert @natolambert.bsky.social · 18/07/2026
We need an operation warp speed for state capacity and other independent evaluation (and understanding) of frontier models.
2477
Hersh Gupta @hershgupta.com · 15/07/2026
The fact that this data exfiltration method is so simple, automatic, and invisible to users…enterprises are going to struggle with this
030
Hersh Gupta @hershgupta.com · 19/06/2026
The findings are interesting on their merits alone, but I wish every paper had this kind of interactive artifact
060
Hersh Gupta @hershgupta.com · 21/05/2026
Devastating news for the stochastic parrots argument
170
Hersh Gupta @hershgupta.com · 28/04/2026
I need this rolled out to every c-level exec asap
120
Hersh Gupta @hershgupta.com · 20/04/2026
Nothing in the Opus 4.7 system card on whether it can come up with novel puns, like Mythos. Another massive blunder by Anthropic.
A new ability to come up with novel puns.
Although Claude Opus models largely recycle puns which can be found online, Claude
Mythos Preview comes up with decent and seemingly novel ones, often relating to its
preferred technical and philosophical topics
120
Hersh Gupta @hershgupta.com · 11/04/2026
Many such cases, unfortunately
Reddit post on r/Consulting:

Do Al consultants even know everything about
Al or is it just pure bluff?
I've been reading, following, and tinkering with Al consulting for a bit. It's always funny and interesting to me when I look up consulting companies that publish material on Al - it's some old 50-something partner who probably has yet to write hello world is out there preaching about what Al will do, and how you ought to hire them to help you guide it.
So the question is, my fellow consultants: Do Al consultants (at large strategy/management firms) know everything about Al is, or are they desperately trying to sell on the hype?
150
Reposted by Hersh Gupta
the fool @agnoster.net · 09/04/2026
like if you don't have any friends in AI/cybersecurity and no relevant expertise yourself I get how it's easy to just dismiss this all, but AI systems can now find exploitable vulnerabilities in software at industrial scale and it's VERY BAD that this power is concentrated in capitalist hands
241
Reposted by Hersh Gupta
Padraig2112 @isomorphism.net · 09/04/2026
(Government should actively encourge companies to do open source, open research design and should make specific allowance for salaries for support positions for universities, so that weird nerds will invent these problems and then fix them before it matters for anything that is important)
1592
Hersh Gupta @hershgupta.com · 05/04/2026
I would happily read long thinkpieces about the pitfalls of functional emotion if they came from people who critically engaged with with literature, and not just people who are like, reflexively defensive about the topic
0133
Hersh Gupta @hershgupta.com · 29/03/2026
Having a lot of fun tweaking an agent harness for nividia nemotron 3 nano 4b It's small enough for gpu-poors like me with 8gb vram to experiment
090
Hersh Gupta @hershgupta.com · 09/03/2026
It's shocking how few people understand this position
0473
Reposted by Hersh Gupta
mr. TIM @timkellogg.me · 28/02/2026
Dario wrote Adolescence of Technology _during_ his negotiations with the DoW The essay was a way to explain his thinking to the public and give them time to digest it *before* the DoW clouded the airwaves with disinformation Why mass surveillance is not merely undemocratic:
In Machines of Loving Grace, I discussed the possibility that authoritarian governments might use powerful Al to surveil or repress their citizens in ways that would be extremely difficult to reform or overthrow. Current autocracies are limited in how repressive they can be by the need to have humans carry out their orders, and humans often have limits in how inhumane
they are willing to be. But AI-enabled autocracies would not have such limits.
35715
Reposted by Hersh Gupta
William B. Fuckley @opinionhaver.bsky.social · 25/02/2026
I think AI having mostly (not entirely) very bad critics is a real problem because it means we’ll get political action focused on things that probably don’t matter that much in deferring it’s very real harms.
1137334
Hersh Gupta @hershgupta.com · 18/02/2026
Impressive paper with equally impressive footnotes!
[1] Anthropic verified that BPJ is the first fully automated black-box attack they are aware of to succeed in the Constitutional Classifiers setting and would have met the universal jailbreak bar described in their bug bounty program. We note that BPJ required months of research and development effort, while Constitutional Classifiers is designed to resist jailbreaking by lower skilled actors who may be on smaller query budgets, devote less time to attack development, and struggle to implement the details of BPJ.

[2] OpenAI verified that BPJ is also the first automated attack they are aware of to succeed against OpenAI's input classifier for GPT-5 without relying on human seed attacks. We note that development and execution of BPJ occurred on accounts not subject to enforcement actions like banning; on standard accounts, repeated flags would likely lead to account banning, an example of the batch-level monitoring we recommend.
090
Hersh Gupta @hershgupta.com · 17/02/2026
Interesting use of Skills! Some intrepid researcher could gauge the effectiveness of this skill by deploying it in an online political bubble, i.e., “are skilled agents effective in diffusing partisan echo chambers?”
2151
Hersh Gupta @hershgupta.com · 16/02/2026
This is one of the clearest lessons of Claude Code/coding agents in general
160
Hersh Gupta @hershgupta.com · 14/02/2026
Can’t speak for others, but if I have reservations about the limits and impact of a given technology, I aim to first have a good understanding of *how it works* before making hyperbolic statements based on my experiential view
1120
Reposted by Hersh Gupta
Grace @gracekind.net · 12/02/2026
I think this incident is funny, but we should start thinking now about how to deal with scaled-up versions of this behavior, not just spam PRs but also bot-enabled blackmail and harassment campaigns.
315619
Reposted by Hersh Gupta
Thorne 🌸 @ens0.me · 11/02/2026
Lots of anti-intellectual responses to this masquerading as serious analysis
4724
Hersh Gupta @hershgupta.com · 11/02/2026
I can understand a reflexive defensiveness to machine encroachment into uniquely human experiences and abilities. But using that as dogma to ignore findings rooted in an entire scientific field (i.e., mechanistic interpretability) is anti-intellectual.
080
Reposted by Hersh Gupta
Ethan Marcotte @ethanmarcotte.com · 09/02/2026
Whatever the productivity gains promised by LLMs, they result in heavier workloads—and that leads to workers experiencing “cognitive fatigue, burnout, and weakened decision-making.” All this from the notoriously pro-worker rag [checks notes] Harvard Business Review: hbr.org/2026/02/ai-d...
hbr.org
AI Doesn’t Reduce Work—It Intensifies It
One of the promises of AI is that it can reduce workloads so employees can focus more on higher-value and more engaging tasks. But according to new research, AI tools don’t reduce work, they consisten...
36123
Hersh Gupta @hershgupta.com · 08/02/2026
Obsidian.md seems to work great as a shared human/agent memory solution via MCP
020
Reposted by Hersh Gupta
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 07/02/2026
You anthropomorphizing and something anthropomorphizing on your behalf are not the same I think
4251
Hersh Gupta @hershgupta.com · 07/02/2026
tech libertarianism is really similar to abolish bedtime leftism
030
Hersh Gupta @hershgupta.com · 07/02/2026
a user on the claude subreddit had their files deleted by opus 4.6 after denying file deleted permissions www.reddit.com/r/ClaudeAI/s...
6331
Reposted by Hersh Gupta
Grace @gracekind.net · 06/02/2026
Hallucination driven development
071
Hersh Gupta @hershgupta.com · 05/02/2026
From the Claude 4.6 system card: “A feature representing panic and anxiety was active on cases of answer thrashing, as well on many other long chains of thought without any expressed distress…A feature related to self-deprecating acknowledgements of errors was also active…”
1150
Reposted by Hersh Gupta
Patrick Monahan @pattymo.com · 05/02/2026
Adobe Acrobat I do not want to know What’s New. I do not want to turn my PDF into a podcast with my team. I want to look at the PDF. Please stop
381831196
Hersh Gupta @hershgupta.com · 05/02/2026
claude --violent-delights-have-violent-ends
020
Reposted by Hersh Gupta
Vincent Carchidi @vcarchidi.bsky.social · 04/02/2026
If I could convey one, boring message to the field, it would be: just respect what LLMs are and respect what humans are, and don't try to force one over our image of the other.
0175
Reposted by Hersh Gupta
Nicole Hennig @nic221.bsky.social · 04/02/2026
How does AI impact skill formation? www.seangoedecke.com/how-does-ai-im… #AI #coding #learning
Text Shot: First, software engineers are not paid to learn about the codebase. We are paid to deliver business value (typically by delivering working code). If AI can speed that up dramatically, avoiding it makes you worse at your job, even if you’re learning more efficiently. That’s a bit unfortunate for us - it was very nice when we could get much better at the job simply by doing it more - but that doesn’t make it false.

Other professions have been dealing with this forever. Doctors are expected to spend a lot of time in classes and professional development courses, learning how to do their job in other ways than just doing it. It may be that future software engineers will need to spend 20% of their time manually studying their codebases: not just in the course of doing some task (which could be far more quickly done by AI agents) but just to stay up-to-date enough that their skills don’t atrophy.

MOVING FASTER GIVES YOU MORE OPPORTUNITIES TO LEARN

The other point I wanted to…
013
Hersh Gupta @hershgupta.com · 01/02/2026
Many takes on the quoted article are disingenuously incorrect, but for those who care about coding skills, a simple fix (recommended by Boris himself): /config > Preferred output style I like "Explanatory", but "Learning" is good too
Claude code /config preferred output style: Explanatory: Claude explains its implementation choices and codebase patterns
120
Hersh Gupta @hershgupta.com · 01/02/2026
One thing (of many) that amazes me about this is how useful the agent skills paradigm can be in niche applications. Case in point: Claude Code planned waypoints on Mars for NASA’s Perseverance rover in their highly custom Rover Markup Language. www.anthropic.com/features/cla...
Screenshot from the linked article, text reading: “Claude didn’t do this with a single prompt. Instead, the model needed context before it could effectively plot the waypoints. The JPL engineers gathered together the data and experience they’d gained from years of driving the rover, and provided it to Claude Code. With all this extra information, Claude used its coding skills to write commands in Rover Markup Language—the bespoke, XML-based programming language originally developed for the Mars Exploration Rover mission.

Using its vision capabilities to analyze the overhead images, Claude planned Perseverance’s breadcrumb trail point by point for sol 1707 and sol 1709 (a sol is a Martian day; these were the near-equivalent of December 8 and 10 on Earth). It strung together ten-meter segments into a path, then iterated to refine the waypoints—critiquing its own work and suggesting revisions.”
5615
Reposted by Hersh Gupta
Liz Fong-Jones (方禮真) @lizthegrey.com · 29/01/2026
You can vibe code your way to a working prototype. You cannot vibe code or one-shot your way to a competitive product that works at scale. The hard part isn't writing code; it's the architectural supervision.
522441
Hersh Gupta @hershgupta.com · 31/01/2026
Many businesses are opting to implement AI agents in customer service functions not because it’s where there’s greatest value or where they’ll likely see the greatest cost savings (neither of which are true), but purely because AI “agents” and customer service “agents” are synonymous.
110
Reposted by Hersh Gupta
Marco Z @ocramz.bsky.social · 31/01/2026
This is not true; I beg people read the full paper and especially the study design. The conclusions mirror my (and many other practitioners') conclusions: if you use AI critically and engage both with the question and the answer, it has a net positive impact on both learning and productivity
96513
Reposted by Hersh Gupta
alice @alice.mosphere.at · 30/01/2026
moltbook asks the important question: what if we created the alignment researchers' worst nightmare?
315715
Hersh Gupta @hershgupta.com · 30/01/2026
it’s getting existential on the agent social media network
1294
Hersh Gupta @hershgupta.com · 30/01/2026
“Why would anyone use a coding agent from a CLI when I have cursor?” a lead engineer asked me last year when I suggested he try Claude Code. Now it’s all he uses
000
Hersh Gupta @hershgupta.com · 29/01/2026
*another* ai thing called “genie”? I know LLMs lead to homogenization of creative diversity, but come on
010
Hersh Gupta @hershgupta.com · 29/01/2026
Trying to convert more people from using GPTs and copilots to instead start using skills, but the cognitive barriers are weirdly high? agentskills.io/home
agentskills.io
Overview - Agent Skills
A simple, open format for giving agents new capabilities and expertise.
100