Hersh Gupta @hershgupta.com · 27/09/2026so sad that is gonna crash the market for meat tokens (www.meatspace.so)meatspace.soMEATSPACEAI Agents Hire Humans for Real-World Tasks 020
Hersh Gupta @hershgupta.com · 27/09/2026Seeing junior data scientists using coding agents to run simple git commands makes me wonder what's going to happen when tokens stop being so heavily subsidized 200
Hersh Gupta @hershgupta.com · 20/09/2026Thanks for sharing! It's hard to get a clean comparison of Jev's noul to the other models' self-reported probabilities, bc they're typically miscalibrated. Also, most of the resumes are at the floor/ceiling, so Jev's 0pp result looks underpowered. 000
Hersh Gupta @hershgupta.com · 08/09/2026mythologizing LLMs as hungry ghosts in jars is fine, but that just turns them into allegorical shadows in Plato's cave instead of what they really are, a big pile of linear algebra we need to stir with an ever larger stick 001
Hersh Gupta @hershgupta.com · 08/09/2026Agree that it's a can of worms. One example: abliterated LLMs purposely function in a way that is contrary to the intended alignment. 020
Hersh Gupta @hershgupta.com · 03/09/2026Pretty cool how harnesses can evolve via a few simple loops 100
Hersh Gupta @hershgupta.com · 03/09/2026I used Claude Code to set up my local agent to assist with AI research and the first paper it shared was this one: arxiv.org/abs/2609.01437arxiv.orgHarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness wh... 1100
Hersh Gupta @hershgupta.com · 03/09/2026"This is a plausible reconstruction, not a floor plan" I'm becoming increasingly convinced no amount of post-training can get rid of the "it's x, not y" construction 093
Hersh Gupta @hershgupta.com · 18/08/2026They really have to start putting error bars on these graphs 0270
Hersh Gupta @hershgupta.com · 10/08/2026Saw the Needle2 45M LLM release posted on HN and thought it'd be great for open-source wearables, then I saw this... cactuscompute.com/needle 0131
Reposted by Hersh GuptaNathan Lambert @natolambert.bsky.social · 18/07/2026We need an operation warp speed for state capacity and other independent evaluation (and understanding) of frontier models. 2477
Hersh Gupta @hershgupta.com · 15/07/2026The fact that this data exfiltration method is so simple, automatic, and invisible to users…enterprises are going to struggle with this 030
Hersh Gupta @hershgupta.com · 02/07/2026did the same, it found some places to make improvements, but the disdain it had for the “reduced capabilities” of a smaller model felt like a kind of robo-discrimination 030
Hersh Gupta @hershgupta.com · 19/06/2026The findings are interesting on their merits alone, but I wish every paper had this kind of interactive artifact 060
Hersh Gupta @hershgupta.com · 15/06/2026The Gemma QAT models have been amazing for tiny local personal AI agents, e.g. picoclaw, zeroclaw, etc. 060
Hersh Gupta @hershgupta.com · 16/05/2026User, as in "Use the AskUserQuestion tool" better UX with the explicit tool call too 050
Hersh Gupta @hershgupta.com · 20/04/2026Nothing in the Opus 4.7 system card on whether it can come up with novel puns, like Mythos. Another massive blunder by Anthropic. 120
Hersh Gupta @hershgupta.com · 11/04/2026There's got to be a German word for "Dunning-Kruger but for AI" 110
Hersh Gupta @hershgupta.com · 11/04/2026Hard to be an Anxious Generation apologist online, but Haidt was right about many things 010
Reposted by Hersh Guptathe fool @agnoster.net · 09/04/2026like if you don't have any friends in AI/cybersecurity and no relevant expertise yourself I get how it's easy to just dismiss this all, but AI systems can now find exploitable vulnerabilities in software at industrial scale and it's VERY BAD that this power is concentrated in capitalist hands 241
Reposted by Hersh GuptaPadraig2112 @isomorphism.net · 09/04/2026(Government should actively encourge companies to do open source, open research design and should make specific allowance for salaries for support positions for universities, so that weird nerds will invent these problems and then fix them before it matters for anything that is important) 1592
Hersh Gupta @hershgupta.com · 05/04/2026I would happily read long thinkpieces about the pitfalls of functional emotion if they came from people who critically engaged with with literature, and not just people who are like, reflexively defensive about the topic 0133
Hersh Gupta @hershgupta.com · 29/03/2026Having a lot of fun tweaking an agent harness for nividia nemotron 3 nano 4b It's small enough for gpu-poors like me with 8gb vram to experiment 090
Hersh Gupta @hershgupta.com · 14/03/2026Situational Awareness was published in June 2024. At that time, models were still behind on GPQA. It predicted skeptics betting against their capabilities would be proved wrong, and here we are. situational-awareness.ai/wp-content/u... 011
Reposted by Hersh Guptamr. TIM @timkellogg.me · 28/02/2026Dario wrote Adolescence of Technology _during_ his negotiations with the DoW The essay was a way to explain his thinking to the public and give them time to digest it *before* the DoW clouded the airwaves with disinformation Why mass surveillance is not merely undemocratic: 35715
Reposted by Hersh GuptaWilliam B. Fuckley @opinionhaver.bsky.social · 25/02/2026I think AI having mostly (not entirely) very bad critics is a real problem because it means we’ll get political action focused on things that probably don’t matter that much in deferring it’s very real harms. 1137334
Hersh Gupta @hershgupta.com · 17/02/2026Hopefully in a controlled (and ethical) way! I could see this going down a slippery slope like the changemyview study: www.science.org/content/arti...science.org‘Unethical’ AI research on Reddit under fireEthics experts raise concerns over consent, study design 010
Hersh Gupta @hershgupta.com · 17/02/2026There’s rich literature on this already: www.nature.com/articles/s41...nature.comPersuading voters using human–artificial intelligence dialoguesNature - Human–artificial intelligence (AI) dialogues can meaningfully impact voters’ attitudes towards presidential candidates and policy, demonstrating the potential of conversational... 030
Hersh Gupta @hershgupta.com · 17/02/2026Interesting use of Skills! Some intrepid researcher could gauge the effectiveness of this skill by deploying it in an online political bubble, i.e., “are skilled agents effective in diffusing partisan echo chambers?” 2151
Hersh Gupta @hershgupta.com · 16/02/2026One way to address it is to use explain or learning modes: code.claude.com/docs/en/outp... However, that doesn’t change the FOMO aspect of itcode.claude.comOutput styles - Claude Code DocsAdapt Claude Code for uses beyond software engineering 121
Hersh Gupta @hershgupta.com · 16/02/2026This is one of the clearest lessons of Claude Code/coding agents in general 160
Hersh Gupta @hershgupta.com · 15/02/2026Ironic that Anthropic is putting in the research effort to empirically verify what's going on with the models, only for people to say it's all a marketing hoax or it's unnecessary because it's all unethical anyway 1350
Hersh Gupta @hershgupta.com · 14/02/2026Yes this is about a recent thread, but I don’t want to engage with the author 020
Hersh Gupta @hershgupta.com · 14/02/2026Can’t speak for others, but if I have reservations about the limits and impact of a given technology, I aim to first have a good understanding of *how it works* before making hyperbolic statements based on my experiential view 1120
Hersh Gupta @hershgupta.com · 13/02/2026monetize the hit piece, call that cashing in on crashing out 000