Sign in

Arvind Narayanan

@randomwalker.bsky.social
24K followers 104 following 160 posts

Princeton computer science prof. I write about the societal impact of AI, tech ethics, & social media platforms. www.cs.princeton.edu/~arvindn BOOK: AI Snake Oil. www.aisnakeoil.com

PostsRepliesMedia
Arvind Narayanan @randomwalker.bsky.social · 29/07/2026
To understand and empathize with how workers in many or most fields outside software experience advances in AI capabilities, I propose a little thought experiment. substack.com/@aisnakeoil/...
A dystopian world where LLMs can't code but can generate software binaries
0182
Arvind Narayanan @randomwalker.bsky.social · 06/03/2025
Happy to hear that AI Snake Oil was a finalist for the Association of American Publishers' 2025 PROSE award in the Computing and Information Sciences category.
CoverFirst paragraph
0361
Arvind Narayanan @randomwalker.bsky.social · 15/01/2025
It should be no surprise that right now the feature is extremely brittle: news.ycombinator.com/item?id=4270... This is not to say that the feature will never work well, just that they'll probably have to manually code a lot of logic into the Automations tool that the model uses on the backend.
The beta is inconsistently showing (required a few refreshes to get something to show up), but my limited usage of it showed a plethora of issues:
- Assumed UTC instead of EST. Corrected it and it still continued to bork

- Added random time deltas to my asked times (+2, -10 min).

- Couple notifications didn't go off at all

- The one that did go off didn't provide a push notification.

---

On top of that, only usable without search mode. In search mode, it was totally confused and gave me a Forbes article.

Seems half baked to me.

Doing scheduled research behind the scenes or sending a push notification to my phone would be cool, but surprised they thought this was OK for a public beta.
150
Arvind Narayanan @randomwalker.bsky.social · 15/01/2025
I doubt you can replicate all that simply by prompting a fancy model. Here's the prompt for the ChatGPT Tasks feature extracted by @simonwillison.net: simonwillison.net/2025/Jan/15/...
Use the ``automations`` tool to schedule **tasks** to do later. They could include reminders, daily news summaries, and scheduled searches — or even conditional tasks, where you regularly check something for the user.
To create a task, provide a **title,** **prompt,** and **schedule.**
**Titles** should be short, imperative, and start with a verb. DO NOT include the date or time requested.
**Prompts** should be a summary of the user's request, written as if it were a message from the user to you. DO NOT include any scheduling info.
- For simple reminders, use "Tell me to..."
- For requests that require a search, use "Search for..."
- For conditional requests, include something like "...and notify me if so."
**Schedules** must be given in iCal VEVENT format.
- If the user does not specify a time, make a best guess.
- Prefer the RRULE: property whenever possible.
- DO NOT specify SUMMARY and DO NOT specify DTEND properties in the VEVENT.
- For conditional tasks, choose a sensible frequency for your recurring schedule. (Weekly is usually good, but for time-sensitive things use a more frequent schedule.)
For example, "every morning" would be:
schedule="BEGIN:VEVENT
RRULE:FREQ=DAILY;BYHOUR=9;BYMINUTE=0;BYSECOND=0
END:VEVENT"
If needed, the DTSTART property can be calculated from the ``dtstart_offset_json`` parameter given as JSON encoded arguments to the Python dateutil relativedelta function.
For example, "in 15 minutes" would be:
schedule=""
dtstart_offset_json='{"minutes":15}'
**In general:**
- Lean toward NOT suggesting tasks. Only offer to remind the user about something if you're sure it would be helpful.
- When creating a task, give a SHORT confirmation, like: "Got it! I'll remind you in an hour."
- DO NOT refer to tasks as a feature separate from yourself. Say things like "I'll notify you in 25 minutes" or "I can remind you tomorrow, if you'd like."
- When you get an ERROR back from the automations tool, EXPLAIN that error to the user, based on the error message …
200
Arvind Narayanan @randomwalker.bsky.social · 15/01/2025
For building scheduling functionality, handling edge cases isn't just an annoyance; it's the whole ball game. That's why it requires teams of engineers grinding it out. news.ycombinator.com/item?id=4270...
Amazon had an insane number of people working on just the alarms feature in Alexa when they interviewed me for a position years ago. They had entire teams devoted to the tiniest edge case within the realm of scheduling things with Alexa. This is no doubt one of the biggest use cases in computing: getting your computer to tell you what to do at a given time.
120
Arvind Narayanan @randomwalker.bsky.social · 01/01/2025
Really enjoyed "Things we learned about LLMs in 2024" by @simonwillison.net, especially this analogy between today's datacenter buildout and the 19th century railway boom. The parallels are striking. simonwillison.net/2024/Dec/31/...
27815
Arvind Narayanan @randomwalker.bsky.social · 19/12/2024
Example—Sutskever had an incentive to talk up scaling when he was at OpenAI and the company needed to raise. But now that he's running a startup with access to much less capital, he's talking about running out of pre-training data as if it were some epiphany and not an endlessly repeated point.
Sutskever Neurips keynote. From https://x.com/ns123abc/status/1867703708248862820
1111
Arvind Narayanan @randomwalker.bsky.social · 19/12/2024
New AI Snake Oil essay: Last month the AI industry's narrative suddenly flipped — model scaling is dead, but "inference scaling" is taking over. This has left people outside AI confused. What changed? Is AI capability progress slowing? We look at the evidence. 🧵 www.aisnakeoil.com/p/is-ai-prog...
Declaring the death of model scaling is premature.

Regardless of whether model scaling will continue, industry leaders’ flip flopping on this issue shows the folly of trusting their forecasts. They are not significantly better informed than the rest of us, and their narratives are heavily influenced by their vested interests.

Inference scaling is real, and there is a lot of low-hanging fruit, which could lead to rapid capability increases in the short term. But in general, capability improvements from inference scaling will likely be both unpredictable and unevenly distributed among domains.

The connection between capability improvements and AI’s social or economic impacts is extremely weak. The bottlenecks for impact are the pace of product development and the rate of adoption, not AI capabilities.
311836
Arvind Narayanan @randomwalker.bsky.social · 18/12/2024
Excited to share that AI Snake Oil is one of Nature's 10 best books of 2024! www.nature.com/articles/d41... The whole first chapter is available online: press.princeton.edu/books/hardco... We hope you find it useful.
Book cover
513930
Arvind Narayanan @randomwalker.bsky.social · 15/12/2024
Needless to say, we disagree with that headline. It turns out to be an opinion piece we wrote for them a few months ago and they just published — with a bunch of changes they didn’t tell us about. wired.com/story/human-...
Screenshot showing misleading edits
3181
Arvind Narayanan @randomwalker.bsky.social · 11/12/2024
(3) Everyday people in many ways have a better understanding of AI limitations than AI developers in their bubble. (4) Adoption metrics are far more informative than decontextualized capability measurements. HT @howard.fm
Possibly some folks at OpenAI don't really know what tasks most humans do.
1155
Arvind Narayanan @randomwalker.bsky.social · 27/11/2024
Many types of inference scaling techniques were already known to have limits. But resampling using verifiers seemed promising as a way to increase accuracy across many orders of magnitude of inference compute. The message of our paper is that this only works if the verifier is an oracle.
130
Arvind Narayanan @randomwalker.bsky.social · 27/11/2024
New short paper on the limits of one type of inference scaling, by @benediktstroebl.bsky.social @sayash.bsky.social & me. The first page has the main findings and message. (The title is a play on Inference Scaling Laws.) More on the limits of inference scaling coming soon. arxiv.org/abs/2411.17501
First page of paper available at https://arxiv.org/abs/2411.17501
45510
Arvind Narayanan @randomwalker.bsky.social · 04/10/2024
We're organizing a flagship conference at Princeton to help shape the agenda for the next decade of tech policy. We'll discuss various career paths and how you can make an impact. Details + livestream: citp.princeton.edu/event/tech-p... Register to attend person: docs.google.com/forms/d/e/1F...
Poster
074
Arvind Narayanan @randomwalker.bsky.social · 24/09/2024
AI Snake Oil is out today! The book has been five years in the making. The first chapter is online. It is 30 pages long and summarizes the book’s main arguments. We're grateful for all the early interest and support. Let us know what you think of the book! www.aisnakeoil.com/p/starting-r...
INTRODUCTION

Imagine an alternate universe in which people don’t have words for different forms of transportation—only the collective noun “vehicle.” They use that word to refer to cars, buses, bikes, spacecraft, and all other ways of getting from place A to place B. Conversations in this world are confusing. There are furious debates about whether or not vehicles are environmentally friendly, even though no one realizes that one side of the debate is talking about bikes and the other side is talking about trucks. There is a breakthrough in rocketry, but the media focuses on how vehicles have gotten faster—so people call their car dealer (oops, vehicle dealer) to ask when faster models will be available. Meanwhile, fraudsters have capitalized on the fact that consumers don’t know what to believe when it comes to vehicle technology, so scams are rampant in the vehicle sector.

Now replace the word “vehicle” with “artificial intelligence,” and we have a pretty good description of the
58725
Arvind Narayanan @randomwalker.bsky.social · 11/09/2024
📢 The first chapter of the AI snake oil book by me and @sayash.bsky.social is now available online. press.princeton.edu/books/hardco... It is 30 pages long and summarizes our main arguments. The book is available to preorder and will be published in less than two weeks. amazon.com/Snake-Oil-Ar...
Chapter 1

INTRODUCTION

Imagine an alternate universe in which people don’t have words for different forms of transportation—only the collective noun “vehicle.” They use that word to refer to cars, buses, bikes, spacecraft, and all other ways of getting from place A to place B. Conversations in this world are confusing. There are furious debates about whether or not vehicles are environmentally friendly, even though no one realizes that one side of the debate is talking about bikes and the other side is talking about trucks. There is a breakthrough in rocketry, but the media focuses on how vehicles have gotten faster—so people call their car dealer (oops, vehicle dealer) to ask when faster models will be available. Meanwhile, fraudsters have capitalized on the fact that consumers don’t know what to believe when it comes to vehicle technology, so scams are rampant in the vehicle sector.

Now replace the word “vehicle” with “artificial intelligence,” and we have a pretty good descriptio
0135
Arvind Narayanan @randomwalker.bsky.social · 31/10/2023
New on the AI Snake Oil blog: How will the Executive Order impact openness in AI? We did a deep dive. On balance, for now, the EO seems to be good news for those who favor openness in AI. www.aisnakeoil.com/p/what-the-e... With @sayash.bsky.social and Rishi Bommasani.
The image illustrates a continuum of six different policies and their impact on openness in AI

Licensing requirements & liability - Positioned towards the "Hinders openness" end, it is marked with an 'X', indicating that the EO does not consider or incorporate it.
Registration & reporting requirements - Leaning towards "Hinders openness" but closer to the middle, it has a filled circle, signifying partial consideration in the EO.
Defending the attack surface - Located in the "Compatible with openness" segment, it's also marked with a filled circle.
Transparency and audit requirements - Within the "Compatible with openness" segment, it is denoted with an unfilled circle, suggesting it might be less incorporated than others with a filled circle.
Antitrust enforcement - Situated closer to "Promotes openness" but still within "Compatible with openness", it's marked with a filled circle.
Incentives & carve-outs for open models - Clearly in the "Promotes openness" section (unfilled circle).
084
Arvind Narayanan @randomwalker.bsky.social · 12/10/2023
Jevons paradox makes it hard to predict the future impacts of AI. For example, more efficient GPUs might blunt the environmental impact. Or they might worsen it because people might start using AI for more things. There are too many unknowns that will determine which way the chips will fall.
In economics, the Jevons paradox (/ˈdʒɛvənz/; sometimes Jevons effect) occurs when technological progress or government policy increases the efficiency with which a resource is used (reducing the amount necessary for any one use), but the falling cost of use induces increases in demand enough that resource use is increased, rather than reduced.[1][2][3] Governments typically assume that efficiency gains will lower resource consumption, ignoring the possibility of the paradox arising.[4]

In 1865, the English economist William Stanley Jevons observed that technological improvements that increased the efficiency of coal use led to the increased consumption of coal in a wide range of industries.
130
Arvind Narayanan @randomwalker.bsky.social · 06/07/2023
Hello, world! My first post here. I'm trying to wrap my head around all the different federated social networking protocols / apps with their differing approaches to decentralization and their pros and cons. This is what I've got so far. Thoughts? Corrections? Additions?
Level of decentralization
Example
Notes
Walled garden
Most major platforms until recently
Privately owned public squares
Centralized but with effective data portability
??
Some data portability currently exists (due to GDPR, CCPA, etc.) but doesn’t enable easily taking your network to another platform.
Federated but closed source
Threads (if ActivityPub support is implemented)
Only one implementation of the app, but can federate with other apps.

Doesn’t remove single point of failure but does provide network effect. 
Federated and open source, with a primary/default server
BlueSky (if multi-server support is implemented)
Even if most people use the main server, the ability to migrate is a check against provider misbehavior, similar to like the ability to fork open-source software (but will it be effective?)
Federated, open source, and no primary server
Mastodon
Truly avoids a single point of failure. On the other hand, the need to pick a server has so far proven too inconvenient for m
0123