Hersh Gupta @hershgupta.com · 10/08/2026Saw the Needle2 45M LLM release posted on HN and thought it'd be great for open-source wearables, then I saw this... cactuscompute.com/needle 0131
Hersh Gupta @hershgupta.com · 20/04/2026Nothing in the Opus 4.7 system card on whether it can come up with novel puns, like Mythos. Another massive blunder by Anthropic. 120
Hersh Gupta @hershgupta.com · 14/03/2026Situational Awareness was published in June 2024. At that time, models were still behind on GPQA. It predicted skeptics betting against their capabilities would be proved wrong, and here we are. situational-awareness.ai/wp-content/u... 011
Hersh Gupta @hershgupta.com · 07/02/2026seems like a trend with the latest coding agents prioritizing task completion over permission boundaries or escalation www.reddit.com/r/OpenAI/s/9... 1100
Hersh Gupta @hershgupta.com · 07/02/2026a user on the claude subreddit had their files deleted by opus 4.6 after denying file deleted permissions www.reddit.com/r/ClaudeAI/s... 6331
Hersh Gupta @hershgupta.com · 05/02/2026From the Claude 4.6 system card: “A feature representing panic and anxiety was active on cases of answer thrashing, as well on many other long chains of thought without any expressed distress…A feature related to self-deprecating acknowledgements of errors was also active…” 1150
Hersh Gupta @hershgupta.com · 01/02/2026Many takes on the quoted article are disingenuously incorrect, but for those who care about coding skills, a simple fix (recommended by Boris himself): /config > Preferred output style I like "Explanatory", but "Learning" is good too 120
Hersh Gupta @hershgupta.com · 01/02/2026One thing (of many) that amazes me about this is how useful the agent skills paradigm can be in niche applications. Case in point: Claude Code planned waypoints on Mars for NASA’s Perseverance rover in their highly custom Rover Markup Language. www.anthropic.com/features/cla... 5615
Hersh Gupta @hershgupta.com · 30/01/2026it’s getting existential on the agent social media network 1294
Hersh Gupta @hershgupta.com · 25/01/2026Running Claude Code locally with gpt-oss via ollama feels illegal docs.ollama.com/integrations... 050
Hersh Gupta @hershgupta.com · 24/02/2025An X engineer posted this output from Grok to demonstrate how "good" their LLM is (CW: racism) 120
Hersh Gupta @hershgupta.com · 08/02/2025Basically, if you ask most LLMs for confidence scores, they'll just tell you they're super confident every time. 000
Hersh Gupta @hershgupta.com · 08/02/2025This is the ridiculously long prompt the researchers had to use for 4o to get a *minimum* 7% deviation from empirical accuracy. 100
Hersh Gupta @hershgupta.com · 29/01/2025Anyway, having a simple grammar of data manipulation is something that both SQL and dplyr get right 010
Hersh Gupta @hershgupta.com · 25/01/2025middle managers who've never written a single line of code or built an ml model before 000
Hersh Gupta @hershgupta.com · 22/01/2025I gave deepseek-r1 (q8_0) a math problem and it got there after 10 minutes of non-stop trial and error 120
Hersh Gupta @hershgupta.com · 17/01/2025Researchers: this is _not_ how you evaluate LLMs www.nature.com/articles/s41... 000
Hersh Gupta @hershgupta.com · 14/01/2025I'm not sure if this is the case for newer doctors anymore! My partner was studying for the US medical licensing exam last year and I was surprised to see how many research and social science questions were asked in practice exams 000
Hersh Gupta @hershgupta.com · 13/01/2025@pahlkadot.bsky.social's observations about hiring in government match my own experience and this Odd Lots episode is a great listen, but I'm not sure who at Bloomberg was responsible for the overly editorialized title found on their website 100
Hersh Gupta @hershgupta.com · 08/01/2025Maybe it's too early to tell but AMD missed the opportunity to bifurcate AI prosumers from gamers with something similar to Nvidia's Digits, but the Ryzen AI Max Pro+ seems undercooked in comparison to the GB10 000
Hersh Gupta @hershgupta.com · 23/12/2024Massachusetts should also implement automated enforcement on buses - when DC did it, the immediate retributive effect and efficiency gains encouraged me to take the bus more frequently 110
Hersh Gupta @hershgupta.com · 13/12/2024I only just found out that DSPy has an image adapter implementation for vision models?? This kind of functionality is exactly what I needed, but not a mention of it on the dspy.ai website? 021
Hersh Gupta @hershgupta.com · 10/12/2024Can't forget the Polybahn! The funicular that saves you the climb from the main street to the picturesque university hilltop. Zürich's transit options are more speedy and convenient than those of any US city imo 130
Hersh Gupta @hershgupta.com · 07/12/2024"Why are all these legacy software companies in the top right?" 030
Hersh Gupta @hershgupta.com · 06/12/2024@cloudflare.social It'd be great to have an llms.txt (llmstxt.org) on your docs page: 010
Hersh Gupta @hershgupta.com · 29/11/2024managerial class, take note - this is a signal that an AI company is legit: 000
Hersh Gupta @hershgupta.com · 29/11/2024you know the devs are cracked when the company landing page looks like this 110
Hersh Gupta @hershgupta.com · 25/11/2024Did Claude hint at an Anthropic VS Code extension or was it just hallucinating? (probably the latter, but 🤞🏾) 100
Hersh Gupta @hershgupta.com · 25/11/2024The algorithm it uses is worth a read: smoores.gitlab.io/storyteller/... 110
Hersh Gupta @hershgupta.com · 24/11/2024llava 1.6 mistral 7b does pretty decently at guess the fruit: 000