Sign in

atharva

@atharvaraykar.com
117 followers 198 following 81 posts

i write at atharvaraykar.com i work @nilenso.com yes-anding the world.

PostsRepliesMedia
atharva @atharvaraykar.com · 23/04/2026
went to lobste.rs after a while. I think they are still not fans of AI.
72 upvotes. Title: AI as a Fascist Artifact | 47 comments
010
atharva @atharvaraykar.com · 17/04/2026
tbf gemini flash has also just gone up over each of the last two versions
220
atharva @atharvaraykar.com · 02/04/2026
analogies are like wet balloons, they are not very effective at explaining things
010
atharva @atharvaraykar.com · 16/02/2026
Hacker-types wanting to build random software side projects for fun, without needing to justify their utility or ability to generate capital is one of the oldest programmer stereotypes. Many such builders still exist!
000
atharva @atharvaraykar.com · 16/02/2026
I've been calling it Lee Sedol'd but Deep Blue is perhaps a better term for it. I thought the AlphaGo documentary is perhaps the greatest depiction of this "Deep Blue" feeling, especially knowing where we are now.
020
atharva @atharvaraykar.com · 14/02/2026
My colleague analysed the system prompts for codex and claude and realised the reason they feel different is because of deliberate product decisions in the prompt! blog.nilenso.com/blog/2026/02...
blog.nilenso.com
Codex CLI vs Claude Code on autonomy
060
atharva @atharvaraykar.com · 12/02/2026
It's also OpenAI-led and the reference schema is more-or-less identical to the current Responses API. It doesn't seem like Anthropic or Google have bought into this—they have competing formats. The vendors that are bought in don't have fully compliant implementations yet.
000
atharva @atharvaraykar.com · 12/02/2026
I was hoping the OpenResponses API would be a meaningful step forward deal with the LLM API standardisation headaches, but right now the spec is really undercooked. There are lots of inconsistencies/contradictions between the reference schemas and what the specification says! www.openresponses.org
openresponses.org
Open Responses
Open Responses documentation overview.
120
atharva @atharvaraykar.com · 11/02/2026
The problem it is solving makes sense (cross-vendor agent communication), but I don't understand why there's such a massive and detailed spec for an *anticipated* use case that hasn't properly materialised yet. Castles in the sky energy. a2a-protocol.org/latest/speci...
a2a-protocol.org
Overview - A2A Protocol
The official documentation for the Agent2Agent (A2A) protocol. The A2A protocol is an open standard that allows different AI agents to securely communicate, collaborate, and solve complex problems tog...
010
atharva @atharvaraykar.com · 11/02/2026
Is A2A protocol completely useless? I don't know of anyone building enterprise multi-agent communication. Why design such a thick protocol for a use case that does not yet exist in practice? a2a-protocol.org/latest/ Feels like another SOAP/CORBA etc
a2a-protocol.org
A2A Protocol
The official documentation for the Agent2Agent (A2A) protocol. The A2A protocol is an open standard that allows different AI agents to securely communicate, collaborate, and solve complex problems tog...
110
atharva @atharvaraykar.com · 09/02/2026
ese bsky.app/profile/grac...
010
atharva @atharvaraykar.com · 08/02/2026
thoughts on ese. atharvaraykar.com/ese/
atharvaraykar.com
Ese
Large language models are better thinkers than writers. Well okay, they don't think as humans do, but I've been letting it write the vast chunk of my computer programs over the last year or so, which ...
131
atharva @atharvaraykar.com · 06/02/2026
The other reason is that infrastructure noise and variations can affect benchmarks a lot, I wonder if that's the case with the SWE Bench Pro runs. www.anthropic.com/engineering/...
anthropic.com
Quantifying infrastructure noise in agentic coding evals
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
010
atharva @atharvaraykar.com · 06/02/2026
The harness matters a lot. SWE Bench Pro uses SWE-Agent-Mini by default. OpenAI likely reports a result on their own harness. Codex is a weird model that performs much worse in generic, minimal harnesses. That's likely why the tool shapes in Codex CLI are strange like "apply_patch".
110
atharva @atharvaraykar.com · 06/02/2026
While SWE Bench Pro is a pretty good benchmark (especially compared to Verified) the leaderboard rankings clearly "look wrong". No one will agree that Claude 4 Sonnet is better than 5.2 Codex. Really shows how insufficient public benchmarks are getting at conveying model capabilities.
010
atharva @atharvaraykar.com · 06/02/2026
It's particularly strange that Anthropic won't report SWE-Bench Pro in their announcements. Their models have always done better on it than OpenAI (at least on the public dataset): scale.com/leaderboard/... I think it might just be that ~80% solved looks more impressive than ~50% solved.
scale.com
SWE-Bench Pro (Public Dataset)
Explore the SEAL leaderboard with expert-driven LLM benchmarks and updated AI model leaderboards, ranking top models across coding, reasoning and more.
210
atharva @atharvaraykar.com · 06/02/2026
the screenshot is from my post: blog.nilenso.com/blog/2025/09...
020
atharva @atharvaraykar.com · 06/02/2026
The METR tasks are narrow (ie, "not messy") and not very numerous, so it's hard to generalise the automatability of software engineering from that alone. It looks like 100% replacement for software engineering can happen, but perhaps not in the next 2 years at least.


    Our tasks typically use environments that do not significantly change unless directly acted upon by the agent. In contrast, real tasks often occur in the context of a changing environment.

    […]

    Similarly, very few of our tasks are punishing of single mistakes. This is in part to reduce the expected cost of collecting human baselines.

This is not at all like the tasks I am doing.

METR acknowledges the messiness of the real world. They have come up with a “messiness rating” for their tasks, and the “mean messiness” of their tasks is 3.2/16.

By METR’s definitions, the kind of software engineering work that I’m mostly exposed to would score at least around 7-8, given that software engineering projects are path-dependent, dynamic and without clear counterfactuals. I have worked on problems that get to around 13/16 levels of messiness.

    An increase in task messiness by 1 point reduces mean success rates by roughly 8.1%

Extrapolating from METR’s measured effect of messiness, GPT-5 would go from 70% to around 40% success rate for 2-hour tasks. This maps to my experienced reality.
140
atharva @atharvaraykar.com · 03/02/2026
"Taking Jaggedness Seriously" by Helen Toner talks about this in some depth. helentoner.substack.com/p/taking-jag...
helentoner.substack.com
Taking Jaggedness Seriously
Why we should expect AI capabilities to keep being extremely uneven, and why that matters
051
atharva @atharvaraykar.com · 03/02/2026
my guess is due to this initiative by the bluesky team. bsky.social/about/blog/0...
screenshot:

PART I: 2025 KEY INITIATIVES Toxicity Filtering a Toxicity is a persistent challenge for all large-scale social apps. As communities grow, maintaining space for both friendly conversation and fierce disagreement requires intentional design choices. Our community doubled in size over the past year, and with that growth came tension: how to preserve healthy discourse while respecting genuine debate and diverse user preferences. Toxic and inflammatory discourse appears across all forms of social media; and almost universally, it's the case that a small percentage of people contribute disproportionately to causing this problem. A tiny number of users can have an outsize impact on conversation quality and on people's willingness to participate. In 2023-2024, anti-social behavior, such as harassment, trolling, and intolerance, consistently ranked among our top complaints reported by users. This content drives people away from forming connections, posting, or engaging, for fear of attacks and pile-ons.screenshot:

In October, we began experimenting with improving conversation quality, starting with replies. Rather than only reacting after users report abusive or toxic interactions, we launched an experiment to identify replies that are toxic, spammy, off-topic, or posted in bad faith, and reduce their visibility in the Bluesky app. This approach adds friction most viewers casually scanning a conversation won't encounter the toxic or potentially harmful replies while preserving content access in case we get it wrong. These replies remain accessible in the thread for those who want to see them. We also made sure this feature is aware of who you follow: Replies from accounts you follow appear above the fold, while toxic replies from people you don't follow require an additional click to view. After implementing this detection, daily reports of anti-social behavior dropped by approximately 79%. This reduction demonstrates measurable improvement in user experience: People are encountering substantially less toxicity in their day-to-day interactions on Bluesky.
2262
atharva @atharvaraykar.com · 30/01/2026
The actual issue to solve, at least for large projects, is getting DOS'd by a flood of low quality slop patches. It's a similar problem to the old Hacktoberfest spam issues, but perhaps worse in scale and scope.
130
atharva @atharvaraykar.com · 29/01/2026
I also like to think of forking to potentially be like a function call stack allocations, whose memory/"context" gets dumped out after the work is done and substituted with a return value, which for agents would be a summary of sorts.
100
atharva @atharvaraykar.com · 29/01/2026
nice to see someone else who is fork-pilled. @mariozechner.at's pi coding agent handles this pattern very well (still manual, like what you described with your claude code workflow, but with much smoother UX and first class support)
110
atharva @atharvaraykar.com · 28/01/2026
I've already had some aggressive muting and "not interested" preference spamming in place, and it isn't working quite as well anymore. It's just enough friction to make me try newer platforms for the time being!
010
atharva @atharvaraykar.com · 27/01/2026
do you have any idea what caused the inflection point?
clawdbot star history showing hockey stick growth, inflection point on Jan 20-ish
200
atharva @atharvaraykar.com · 27/01/2026
Some people in some normie-ish group chats I'm in thought this is a product by Anthropic. Bet they'd have got support requests for this already. Perhaps they don't want to have their name attached to this, which is horribly insecure for people who don't know what they are doing.
010
atharva @atharvaraykar.com · 27/01/2026
things I have dumped on the internet this month. How the lobsters algorithm works: atharvaraykar.com/lobsters/ 11:59 PM: atharvaraykar.com/reinforce/
atharvaraykar.com
How the Lobsters front page works
Lobsters is a computing-focused community centered around link aggregation and discussion. The code is open source, so I had a look at how the front page algorithm works. This is it: $$\textbf{hotn...
010
atharva @atharvaraykar.com · 27/01/2026
fwiw, I'm trying to use this site (and substack) more ever since the new X algorithm completely trashed my feed, it's surfacing only toxic sludge and slop the vibe here has improved quite a bit in the meantime. but I can also imagine a timeline where I stop microblogging altogether and touch grass
150
atharva @atharvaraykar.com · 27/12/2025
Exploring the weirdness of this would fall under the goals of AI village. At least as I understand it. But they definitely should not be unleashing these agents "outside the lab", hence my mention of this needing to be opt-in/consented or sandboxed in some way.
030
atharva @atharvaraykar.com · 26/12/2025
yeah they messed up with today's goal, these kind of things need to be opt-in.
130
atharva @atharvaraykar.com · 28/11/2025
The "excellence" still depends on whether the language is in the training distribution. It's pretty competent at Python and JS. Less so in Clojure. Or HashiCorp Language. Even so I agree that it's still a productivity boost across most languages.
050
atharva @atharvaraykar.com · 29/09/2025
I wrote a post looking into multiple SWE/coding benchmarks. Many of them measure something narrower than what their names suggests. blog.nilenso.com/blog/2025/09...
SWE-bench Verified and SWE-bench Pro
What it measures

How well a coding agent can submit a patch for a real-world GitHub issue that passes the unit tests for that issue.
The specifics

There are many variants: Full, Verified, Lite, Bash-only, Multimodal. Most labs in their chart report on SWE-bench Verified, which is a cleaned and human-reviewed subset.

Notes and quirks of SWE-bench Verified:

    It has 500 problems, all in Python. Over 40% are issues from the Django source repository; the rest are libraries. Web applications are entirely missing. The repositories that the agents have to operate are real, hefty open source projects.
    Solutions to these issues are small—think surgical edits or small function additions. The mean lines of code per solution are 11, and median lines of code are 4. Amazon found that over 77.6% of the solutions touch only one function.
    All the issues are from 2023 and earlier. This data was almost certainly in the training sets. Thus it’s hard to tell how much of the improvements are due to memorisation.
011
atharva @atharvaraykar.com · 19/09/2025
Wrote about units of work being a useful lever for getting good results from AI-assisted coding. blog.nilenso.com/blog/2025/09...
011
atharva @atharvaraykar.com · 12/09/2025
I've been poking Srihari, our most experienced engineer @nilenso.com to share his hard-earned knowledge for the benefit of others. Even if you're not an engineering leader like me, this checklist gives a lot of insight into what makes a great engineering org. blog.nilenso.com/blog/2025/09...
My Quarterly System Health Check-in

It is essential to periodically take a few steps back from the day to day and reflect on where we are against our strategic goals. If you’re an engineering leader, a head of engineering, a director, or a VP, you likely have a recurring meeting to this effect.

In this post, I propose a structure for this operational exercise (complementing a business review) that lasts 2-4 hours, every month or quarter. I see quality as solving for the Pareto front with the tangible dimensions of reliability, performance, cost, delivery and security, and the more intangible dimensions of simplicity and social structures. For each dimension, go through the list of questions below and try to answer them together.
022
atharva @atharvaraykar.com · 21/08/2025
I think this is the correct take if one is short-horizon AGI-pilled. But since I don't expect the shape of agent intelligence to fully be like human intelligence anytime soon (say ~10y), it helps to have a giant doorknob for agents, geared towards their style of processing information.
120
atharva @atharvaraykar.com · 15/08/2025
grateful for my low-quality education where teachers could not correctly explain the basics of anything right it pushed me to cultivate good epistemics and build truth-seeking habits early on it also forced me to never trust authoritative figures, many are very incompetent
000
atharva @atharvaraykar.com · 14/08/2025
This is a good example of the weirdness of AI.
It’s a recursive paradigm shifting paradigm shift

Code is data is code. Lispers get this, its turtles all the way down. Let’s walk through it.

1.    Software can use AI. We write pieces of text in between code that uses AI’s thinking or knowledge access capability. Like with autocompletion, or chatbots.
2.    AI can use software. It can access the filesystem and run programs on your OS, if you let it. It can access the web through search engines, or web browsers, if you let it.
3.    AI can build software. It writes and executes small python scripts to analyse data when thinking. The agency and autonomy needed to build full fledged meaningful software isn’t there with AI yet, but it’s advancing quickly. AI assisted coding is pretty big, you know this.
4.    AI IS software. AI is a trained neural network, and with sufficiently advanced capabilities, it can build itself.
030
atharva @atharvaraykar.com · 14/08/2025
Why Does AI Feel So Different? An enjoyable read from my colleague, Srihari. We've been talking about why this disruption feels different from other recent technological disruptions and he captured a lot of that really well in this post. Link: blog.nilenso.com/blog/2025/08...
Image description
While Kuhn doesn’t go into it, the technological diffusion, and the economic impact of scientific revolutions are better studied through the GPT (general purpose technology) paper from Bresnahan & Trajtenberg in 1995.

> “General Purpose Technologies (GPTs) are technologies that can affect an entire economy (usually at a national or global level). They have the potential for pervasive use in a wide range of sectors and, as they improve, they contribute to overall productivity growth.”

And Calvino et al in June 2025, finds that AI meets the key criteria of a General Purpose Technology. It’s pervasive, rapidly improving, enables new products, services and research methodologies, and enhances other sectors’ R&D and productivity.
101
atharva @atharvaraykar.com · 28/07/2025
Question for the "launch many parallel claude agents to go faster" crowd: Why isn't that the same problem as the Brooks's law all over again?
wikipedia screenshot of brooks's law page.
000
atharva @atharvaraykar.com · 25/07/2025
despite all the AI hype there's only been two kinds of tools with direct improvements to my life: - chat (like chatgpt, claude ai) - terminal coding agents (like claude code) surprising how nothing better has stuck yet
110
atharva @atharvaraykar.com · 04/07/2025
Several people ask me about how I'm keeping up with all the AI things and finding signal in this noisy landscape. I wrote a guide explaining this. blog.nilenso.com/blog/2025/06...
Table of Contents

    General guidelines
    Starting Points
    Official announcements, blogs and papers from those building AI
    High signal people to follow
    News and Media
    Esoterica
    Do I chug water from a firehose?
031
atharva @atharvaraykar.com · 24/06/2025
<3
000
atharva @atharvaraykar.com · 19/06/2025
I was seeing @karpathy.bsky.social's Software 3.0 talk, as one should. Surreal to see him recommend my writing to make a point on AI-assisted coding "Is that me on TV??" moment
Andrej Karpathy on the YCombinator startup school stage talking about my AI-assisted coding blog post
491
Reposted by atharva
Simon Willison @simonwillison.net · 11/06/2025
I loved it! simonwillison.net/2025/Jun/10/...
simonwillison.net
AI-assisted coding for teams that can’t get away with vibes
This excellent piece by Atharva Raykar offers a bunch of astute observations on AI-assisted development that I haven't seen written down elsewhere. Building with AI is fast. The gains in …
073
atharva @atharvaraykar.com · 11/06/2025
indeed, prompting is still extremely sensitive and brittle even today. Some of my colleagues @nilenso.com managed to get better code by just saying "write it like Rich Hickey would" which seemed to have activated the right neurons
040
atharva @atharvaraykar.com · 11/06/2025
in my case, I did not get into the weeds of the workflow, because my own workflows are constantly changing as these tools and models evolve I tried to focus on what consistently helped throughout all the changes, and the best rule of thumb I could come up with is "what helps the human helps the ai"
130
atharva @atharvaraykar.com · 11/06/2025
glad you found it useful!
1110
atharva @atharvaraykar.com · 01/06/2025
Quality, and the cutlets of enthusiasm atharvaraykar.com/quality-and-...
atharvaraykar.com
Quality, and the cutlets of enthusiasm
Zen and the art of motorcycle maintenance was on my reading wishlist for a while. From the title, I assumed the book would be something like a tasteful self-help book, a quirky blend of kōanic mystici...
132
atharva @atharvaraykar.com · 01/03/2025
switched to ghost
010
atharva @atharvaraykar.com · 26/01/2025
I have revamped my website of essays. it's now more searchable, supports commenting and also doubles as a newsletter. now back to writing again.
buff.ly
atharva's internet place
tldraw is canvas software that runs on your browser (like MS Paint, excalidraw etc). This project caught my attention when I saw demos of fun AI experiments with the canvas interface*. Even the plain…
100