Sign in

atharva

@atharvaraykar.com
117 followers 198 following 81 posts

i write at atharvaraykar.com i work @nilenso.com yes-anding the world.

PostsRepliesMedia
atharva @atharvaraykar.com · 23/04/2026
went to lobste.rs after a while. I think they are still not fans of AI.
72 upvotes. Title: AI as a Fascist Artifact | 47 comments
010
atharva @atharvaraykar.com · 06/02/2026
The METR tasks are narrow (ie, "not messy") and not very numerous, so it's hard to generalise the automatability of software engineering from that alone. It looks like 100% replacement for software engineering can happen, but perhaps not in the next 2 years at least.


    Our tasks typically use environments that do not significantly change unless directly acted upon by the agent. In contrast, real tasks often occur in the context of a changing environment.

    […]

    Similarly, very few of our tasks are punishing of single mistakes. This is in part to reduce the expected cost of collecting human baselines.

This is not at all like the tasks I am doing.

METR acknowledges the messiness of the real world. They have come up with a “messiness rating” for their tasks, and the “mean messiness” of their tasks is 3.2/16.

By METR’s definitions, the kind of software engineering work that I’m mostly exposed to would score at least around 7-8, given that software engineering projects are path-dependent, dynamic and without clear counterfactuals. I have worked on problems that get to around 13/16 levels of messiness.

    An increase in task messiness by 1 point reduces mean success rates by roughly 8.1%

Extrapolating from METR’s measured effect of messiness, GPT-5 would go from 70% to around 40% success rate for 2-hour tasks. This maps to my experienced reality.
140
atharva @atharvaraykar.com · 03/02/2026
my guess is due to this initiative by the bluesky team. bsky.social/about/blog/0...
screenshot:

PART I: 2025 KEY INITIATIVES Toxicity Filtering a Toxicity is a persistent challenge for all large-scale social apps. As communities grow, maintaining space for both friendly conversation and fierce disagreement requires intentional design choices. Our community doubled in size over the past year, and with that growth came tension: how to preserve healthy discourse while respecting genuine debate and diverse user preferences. Toxic and inflammatory discourse appears across all forms of social media; and almost universally, it's the case that a small percentage of people contribute disproportionately to causing this problem. A tiny number of users can have an outsize impact on conversation quality and on people's willingness to participate. In 2023-2024, anti-social behavior, such as harassment, trolling, and intolerance, consistently ranked among our top complaints reported by users. This content drives people away from forming connections, posting, or engaging, for fear of attacks and pile-ons.screenshot:

In October, we began experimenting with improving conversation quality, starting with replies. Rather than only reacting after users report abusive or toxic interactions, we launched an experiment to identify replies that are toxic, spammy, off-topic, or posted in bad faith, and reduce their visibility in the Bluesky app. This approach adds friction most viewers casually scanning a conversation won't encounter the toxic or potentially harmful replies while preserving content access in case we get it wrong. These replies remain accessible in the thread for those who want to see them. We also made sure this feature is aware of who you follow: Replies from accounts you follow appear above the fold, while toxic replies from people you don't follow require an additional click to view. After implementing this detection, daily reports of anti-social behavior dropped by approximately 79%. This reduction demonstrates measurable improvement in user experience: People are encountering substantially less toxicity in their day-to-day interactions on Bluesky.
2262
atharva @atharvaraykar.com · 27/01/2026
do you have any idea what caused the inflection point?
clawdbot star history showing hockey stick growth, inflection point on Jan 20-ish
200
atharva @atharvaraykar.com · 29/09/2025
I wrote a post looking into multiple SWE/coding benchmarks. Many of them measure something narrower than what their names suggests. blog.nilenso.com/blog/2025/09...
SWE-bench Verified and SWE-bench Pro
What it measures

How well a coding agent can submit a patch for a real-world GitHub issue that passes the unit tests for that issue.
The specifics

There are many variants: Full, Verified, Lite, Bash-only, Multimodal. Most labs in their chart report on SWE-bench Verified, which is a cleaned and human-reviewed subset.

Notes and quirks of SWE-bench Verified:

    It has 500 problems, all in Python. Over 40% are issues from the Django source repository; the rest are libraries. Web applications are entirely missing. The repositories that the agents have to operate are real, hefty open source projects.
    Solutions to these issues are small—think surgical edits or small function additions. The mean lines of code per solution are 11, and median lines of code are 4. Amazon found that over 77.6% of the solutions touch only one function.
    All the issues are from 2023 and earlier. This data was almost certainly in the training sets. Thus it’s hard to tell how much of the improvements are due to memorisation.
011
atharva @atharvaraykar.com · 19/09/2025
Wrote about units of work being a useful lever for getting good results from AI-assisted coding. blog.nilenso.com/blog/2025/09...
011
atharva @atharvaraykar.com · 12/09/2025
I've been poking Srihari, our most experienced engineer @nilenso.com to share his hard-earned knowledge for the benefit of others. Even if you're not an engineering leader like me, this checklist gives a lot of insight into what makes a great engineering org. blog.nilenso.com/blog/2025/09...
My Quarterly System Health Check-in

It is essential to periodically take a few steps back from the day to day and reflect on where we are against our strategic goals. If you’re an engineering leader, a head of engineering, a director, or a VP, you likely have a recurring meeting to this effect.

In this post, I propose a structure for this operational exercise (complementing a business review) that lasts 2-4 hours, every month or quarter. I see quality as solving for the Pareto front with the tangible dimensions of reliability, performance, cost, delivery and security, and the more intangible dimensions of simplicity and social structures. For each dimension, go through the list of questions below and try to answer them together.
022
atharva @atharvaraykar.com · 14/08/2025
This is a good example of the weirdness of AI.
It’s a recursive paradigm shifting paradigm shift

Code is data is code. Lispers get this, its turtles all the way down. Let’s walk through it.

1.    Software can use AI. We write pieces of text in between code that uses AI’s thinking or knowledge access capability. Like with autocompletion, or chatbots.
2.    AI can use software. It can access the filesystem and run programs on your OS, if you let it. It can access the web through search engines, or web browsers, if you let it.
3.    AI can build software. It writes and executes small python scripts to analyse data when thinking. The agency and autonomy needed to build full fledged meaningful software isn’t there with AI yet, but it’s advancing quickly. AI assisted coding is pretty big, you know this.
4.    AI IS software. AI is a trained neural network, and with sufficiently advanced capabilities, it can build itself.
030
atharva @atharvaraykar.com · 14/08/2025
Why Does AI Feel So Different? An enjoyable read from my colleague, Srihari. We've been talking about why this disruption feels different from other recent technological disruptions and he captured a lot of that really well in this post. Link: blog.nilenso.com/blog/2025/08...
Image description
While Kuhn doesn’t go into it, the technological diffusion, and the economic impact of scientific revolutions are better studied through the GPT (general purpose technology) paper from Bresnahan & Trajtenberg in 1995.

> “General Purpose Technologies (GPTs) are technologies that can affect an entire economy (usually at a national or global level). They have the potential for pervasive use in a wide range of sectors and, as they improve, they contribute to overall productivity growth.”

And Calvino et al in June 2025, finds that AI meets the key criteria of a General Purpose Technology. It’s pervasive, rapidly improving, enables new products, services and research methodologies, and enhances other sectors’ R&D and productivity.
101
atharva @atharvaraykar.com · 28/07/2025
Question for the "launch many parallel claude agents to go faster" crowd: Why isn't that the same problem as the Brooks's law all over again?
wikipedia screenshot of brooks's law page.
000
atharva @atharvaraykar.com · 04/07/2025
Several people ask me about how I'm keeping up with all the AI things and finding signal in this noisy landscape. I wrote a guide explaining this. blog.nilenso.com/blog/2025/06...
Table of Contents

    General guidelines
    Starting Points
    Official announcements, blogs and papers from those building AI
    High signal people to follow
    News and Media
    Esoterica
    Do I chug water from a firehose?
031
atharva @atharvaraykar.com · 19/06/2025
I was seeing @karpathy.bsky.social's Software 3.0 talk, as one should. Surreal to see him recommend my writing to make a point on AI-assisted coding "Is that me on TV??" moment
Andrej Karpathy on the YCombinator startup school stage talking about my AI-assisted coding blog post
491
atharva @atharvaraykar.com · 18/01/2025
Buddhist Teacher Geshe Langri Tangpa (circa 1100 AD) 🤝 XXXTentacion (2018)
In brief, directly or indirectly,
I will offer help and happiness to all my mothers,
And secretly take upon myself
All their hurt and suffering.We ain't playin' with y'all niggas, man, you heard?
Sometimes you just gotta catch chlamydia on these niggas, G shit
You know what I'm saying? Gonorrhea, all of that shit
All of that shit
I catch all diseases in the world
So the world don't have no more diseases, you feel me?
G-shit
Yeah, yeah (P.Soul on the track)
010
atharva @atharvaraykar.com · 04/01/2025
TIL about the Modi (मोडी / 𑘦𑘻𑘚𑘲‎) script, used for Marathi up until the 1940s. In 2025, it looks like AI generated Devanagari.
041
atharva @atharvaraykar.com · 29/12/2024
i quickly concocted a writer's block unblocker (with @tldraw.com computer) it takes an oblique strategy (from the brian eno et al card deck) and uses it to provide unhinged critique of the essay you are working on to help you break out of a rut link to program: computer.tldraw.com/t/4KoB33nFEr...
192
atharva @atharvaraykar.com · 01/12/2024
menu highlights
google's menu highlights for a thindi joint. the first highlight is half eaten masala dosa accompanied by a picture.
020
atharva @atharvaraykar.com · 30/11/2024
art in a white wall big box chamber is gazed upon for long, an object of loving attention art outside the door of hae kum gang korean restaurant gents toilet is ignored among the din of swinging doors and flushing toilets
dreamy surreal distortion of a bicycle framed outside hae kum gang korean restaurant gents toilet
020
atharva @atharvaraykar.com · 27/11/2024
chai darkening ideas - add fine mud - spray spray tan - mix in water from a lake in peenya - chocolate - blend with chyawanprash
010