Sign in

Dan Luu

@danluu.com
4.2K followers 15 following 67 posts

danluu.com / www.patreon.com/danluu/

PostsRepliesMedia
Dan Luu @danluu.com · 23/09/2026
At the time, I was doing automated bug finding with human review. Human review found zero known false positives in the bugs I filed. Anthropic was using a better model and had more tokens to spend, which gave them the advantage on the two most difficult factors w.r.t. false positive rejection.
050
Dan Luu @danluu.com · 23/09/2026
Why didn't Anthropic do effective false positive rejection with their Mythos/Glasswing vuln reports? In danluu.com/ai-coding/#m..., I mentioned a colleague finding that the reports we got were mostly false. That seems common? E.g., gregkh on kernel reports and www.vulncheck.com/blog/anthrop...
2130
Dan Luu @danluu.com · 22/09/2026
People keep replying to tell me that Ralph loops worked better than nothing, which seems to illustrate my point? When Ralph loops were popular, I used a loop with context, which outperformed. That's not counting clearing context when it helps instead of always keeping, which you'd do in practice.
060
Dan Luu @danluu.com · 22/09/2026
Did Ralph loops ever work? They were popular for a while, but when I did a comparison, they underperformed. There were all these theories about why Ralph loops were effective ("invert the latent space of the model"), but AFAICT people just didn't run the comparison? danluu.com/pl-tokens/#r...
Ralph loop underperforms while loop in 27 out of 30 conditions
4320
Reposted by Dan Luu
Dan McKinley @mcfunley.com · 08/05/2026
Maybe the best thing on the internet at the moment: * An AI PR pitching platform has confused Kyle Kingsbury, MMA fighter and podcast host with Kyle Kingsbury, software engineer and gay leather man * Good Kyle decided to start a podcast where he interviews confused manosphere influencers
youtube.com
The Kyle Kingsbury Podcast Podcast - Epsiode 02 Dave Rossi
YouTube video by Aphyr Null
25820
Dan Luu @danluu.com · 18/09/2026
There's no point at which turning your brain off will work: danluu.com/brain-off/
In early 2025, I started seeing people turn off their brain as they use LLMs1. They would have an LLM take an action (summarize text, write some code, etc.), and just assume that it worked2. This generally didn't work in early 2025 and the result was often quite silly.

As LLMs have gotten better, I've seen more of this. Sometimes, people will try to get the LLM to write some code for them and basically just assume that it works3. Sometimes there's a human in the loop and, if the thing doesn't work, they'll ask the LLM to figure out the problem and solve it. Niklas Gruhn calls some variants of doing this being a meat proxy.4

Being a for loop meat proxy works better than it did in early 2025 and the software I've tried that's developed like this sometimes actually sort of works. Not well enough that I'd want to use it or that it's successful, but I'm impressed at how effective being a meat proxy is in September 2026. You could even imagine LLMs improving enough that brain-off meat-proxy development produces average quality software in the foreseeable future.

Let's say that happens. What reason is there for the company to employ the meat proxy? The company can just run the LLM in a loop and lay off the employee. There's no point at which this methodology will work for the employee5.
27017
Dan Luu @danluu.com · 16/09/2026
Yes, more than one DC was blown up, but of course DCs being blown up aren't uncorrelated events. We already know that hardware failures like disk failures, CPU failures, etc., are highly correlated events, and those aren't even adversarial failures!
1291
Dan Luu @danluu.com · 16/09/2026
Amazon exec: "you wouldn't notice [if someone blew up a datacenter]. I mean, we might be a bit upset, but you wouldn't notice! [laughs]" Amazon after DCs blown up: "After a thorough assessment, we have determined that we are unable to restore access to the resources and data hosted ...."
But hold on a second: If we're concentrating all of our data into data centers, and concentrating most of the data centers in one county, doesn't that make a very tempting target for terrorists?

Amazon's Matt Wood isn't worried: "If something does happen or we have a power event or there's a flood in one specific location, that data is held redundantly in other locations as well."

Pogue asked, "I don't mean to give anyone ideas, but let's say I figured out that one of these unmarked buildings was an AWS data center, and I blew it up. Are you saying that it's so backed up and redundant that you probably wouldn't notice?"

"Yeah, you wouldn't notice. I mean, we might be a bit upset, but you wouldn't notice!" Wood laughed. AWS says wartime damage means some Middle East cloud resources are gone for good

Iranian strikes overwhelmed regional redundancy in Bahrain and left one UAE Availability Zone inaccessible
Dan Robinson
IT INFRASTRUCTURE REPORTER
25
Published Wed 16 Sept 2026 // 12:11 UTC
READ MORE

    Nvidia goes green to keep grid capacity from zapping its revenues
    now
    Your AI agents' reports and questions have a new inbox, courtesy of AWS
    1 day ago
    Higher-enriched uranium for datacenters has DoE all aglow
    1 day ago
    Teravolt looks to cannibalize older industries to meet AI power demand
    2 days ago
    Nvidia's Groq acquihire is on the DOJ's radar, but it's already too late
    4 days ago

Amazon Web Services (AWS) says it is unable to restore access to resources and data in some of its Availability Zones in the Middle East after datacenters were damaged during the US war with Iran.

In an update to its AWS Health Dashboard, the cloud giant confirmed that anything hosted exclusively in its Bahrain Region (me-south-1) remains inaccessible.

"The damage to our infrastructure spanned multiple Availability Zones and exceeded what our regional and multi-AZ services are designed to withstand," the update says.

"After a thorough assessment, we have determined that we are unable to restore access to the resources and data hosted exclusively in this Region."

After the first Bahrain Availability Zone was damaged in March, AWS advised customers to migrate their workloads to other Regions. The company said most did so before further attacks disrupted a second Availability Zone in April and rendered the entire Region unavailable.

AWS has reached a similar conclusion about one Availability Zone in the United Arab Emirates (UAE), saying it cannot re-establish access to the resources and data hosted there.

The affected Zone, mec1-az2, is one of three in AWS's UAE Region. Two AWS facilities in the country were hit by drones back in March as Iran retaliated against US…
3708
Dan Luu @danluu.com · 15/09/2026
Interesting to see Steve Yegge say he never successfully built anything with Gas Town. I mentioned not finding these super vibed orchestrators useful because the reliability was too low (w.r.t. completing non-trivial tasks). Turns out the author of the most famous one had the exact same issue.
Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw.

Since Claude Fable 5 dropped, I have returned full-time to my 30-year-old video game, Wyvern, which I began work on in 1996, and launched in 2001. I had been waiting for a model smart enough to help me with my game, and now that it's here, Wyvern's once again my main squeeze. On the side I do occasional six-figure gigs where I fly to companies and teach them my techniques, and that helps with my (considerable) token bills. But aside from those few paid field trips, I am laser-focused on my game, seven days a week. Wyvern's time will finally come next year. I have officially entered Sam Altman's solo unicorn contest.
612211
Dan Luu @danluu.com · 08/09/2026
Thanks! Some ancient minifier that I use was deleting some CSS that makes the page work on mobile. I should really find a better minifier or maybe just write one myself.
010
Dan Luu @danluu.com · 07/09/2026
How well do agents use test/verification techniques? danluu.com/agentic-test...
See post for detailed description of crop of graph with 60 (!) points
4193
Dan Luu @danluu.com · 01/09/2026
How accurate have Ed Zitron's AI skeptic predictions been? danluu.com/zitron/


    Feb 2024: "I believe we're reaching the upper limits about what generative AI can do and how accurate its outputs can be."
        Wrong4
    March 2024: "Have We Reached Peak AI?"; another prediction that hallucinations mean that AI progress is limited to then-current levels
        Wrong
    April 2024: "As I previously warned, artificial intelligence companies are running out of data ..."; another prediction that models can't improve because there's no more data
        Wrong
    June 2024: OpenAI growth is stalling (with the implication it will continue to stall), which will lead to some kind of collapse of OpenAI
        Wrong (it could be the case that OpenAI will collapse but, if so, it won't be due to any kind of growth stall from 2024)
    July 2024: "Generative AI, as I said back in March, is peaking, if it hasn't already peaked. It cannot do much more than it is currently doing, other than doing more of it faster with some new inputs"
        Wrong
    July 2024: "Generative AI models aren’t getting more energy-efficient, nor are they getting more “powerful” in a way that would increase their functionality, nor are they even capable of automating things on their own [July 2024]"
        Wrong5
    August 2024: "generative AI is a dead-end technology that has peaked”
        Wrong
    August 2024: re-iteration that the AI bubble has 3 quarters to prove itself (from March 2024) or there will be a collapse
        Wrong6
    September 2024: "o1 shows that OpenAI is both desperate and out of ideas", with a re-iteration of the idea that models can't improve due to lack of data
        Wrong
    Oct 2024: OpenAI's forecast of $3.7B revenue in 2024 and $11.6B in 2025 and $100B in 2029 are absurd, "a statement so egregious that I am surprised it's not some kind of financial crime to say it out loud"
        Wrong (2025 goal exceeded, 2029 TBD but not an egregious financial crime level of implausible)
    Oct 2024: "[OpenAI revenue] growth is already slowing…
1944375
Dan Luu @danluu.com · 29/08/2026
Bug Blindness: danluu.com/bug-blind/
I used to wonder why I see so many more bugs than most people. I easily observe hundreds to thousands of bugs per week, but most people I talk to don't see anything like this. For a long time, I thought this had something to do with how I use computers but, over time, I've realized that it's mostly that people are hitting the same bugs and don't notice.

If you're not a programmer, that's probably a better way to see the world, but I think curing quality/bug blindness is helpful for programmers. I've done this with a lot of friends and acquaintances by just pointing out bugs. After a few weeks of this, people who are so inclined tend to start noticing bugs as well (the fact that this has happened so many times is what makes me think that my running into so many bugs is more about my noticing them than about how I interact with computers).

Because I notice these kinds of things, I've had multiple jobs where directors/VPs/execs/etc. sometimes ask me to evaluate something because I'm relatively likely to notice issues (and fix them or drive fixes for them if necessary). Sometimes I won't find any issues (there are likely issues that just aren't the kind I notice). More often, I find issues that fall somewhere from "mild" to "moderate". And, sometimes, the issues are severe, to the point where an uncharitable person might even say the thing doesn't actually work.

I find this last category a bit mysterious, as when I look up discussions on how the thing got into this state, there's usually a stream of internal comments indicating that the thing is great, it works well, etc., but when I open up the thing and try it, it's in a state where the thing only works if you do quite a few non-intuitive bug workarounds. I've had this post in mind for maybe a decade or so, but I was hesitant to write it up because, in the back of my mind, I always wondered if I'm somehow triggering weird corner case behavior most users don't hit without realizing it. But after seeing more and more…
8627
Dan Luu @danluu.com · 21/08/2026
There's no reason for software to be slow anymore: danluu.com/perf-opt/
The other day, I saw a viral tweet saying that people talking about how LLMs are causing slow, bloated, code are going to eat crow once they re-write everything in super-optimized assembly. We're not quite at the point where we want to write everything in assembly, but some variant of what Nolan Lawson said about testing, you can choose how many bugs you want now, which I less eloquently noted here, is becoming more true for performance.

In response to a comment in my last post that the cost of formerly specialized performance work has dropped by many orders of magnitude and performance work that used to require a person or team that had a rare set of skills can be done by anyone who can type a few sentences1, which means that you can do all sorts of optimizations that used to be too expensive to be worthwhile for all but the largest scale or most lucrative projects, Marc Brooker responded with

    Completely agree with your closing point. Dynamic custom software, fitted to a particular workload rather than a class of workloads, seems like a very likely outcome. (Which comes with all kinds of fun risks and opportunities of its own). Kind of reminds me of FFTW. And a ton of weird old demoscene techniques which were all about being super fast and small on a very particular problem (and often very particular hardware). For example, I remember a demo that re-used its code as textures to get great cache locality.

And Michael Malis has noted

    There’s been a meme circulating about how AI doesn’t help because “code was never the hard part.” I think that’s true in some domains, but in others, writing the code absolutely was the hard part. JIT compilers are a great example of that. For many pieces of software, a JIT compiler would help a lot with speeding up the code. The rarity of JIT compilers makes me believe that implementing a JIT compiler historically was too difficult for it to be worthwhile. LLMs have lowered the barrier to entry and made it much easier to writ…
3844
Dan Luu @danluu.com · 17/08/2026
The benchmarkpocalypse: danluu.com/benchpocalyp...
There's been a lot of talk about the vulnpocalypse, to which I don't have much to add because I'm not a security person, but I haven't seen much discussion on the closely related (and to be fair, less serious issue), the benchmarkpocalypse.

While it's become easier than ever to make serious performance gains, it's also become easier than ever to reward hack a benchmark and make fake performance gains. The former is probably happening quietly across many different companies, but the latter is something I see at least once a week nowadays. Someone will claim they optimized X and got some huge performance improvement over existing software but, when you look at it, what they did was make some optimization that improves benchmark performance without actually improving real-world performance. This is often some kind of "we re-wrote X in Rust"1 project or a new startup that's looking to either fundraise or sell something, but it happens on other kinds of projects as well2.

Rather than point to someone's bad claim, I'll point to FRE, this regex engine I had an agent build, which I could claim is the world's fastest regex engine because it beats the Rust regex crate at the fairly comprehensive rebar regex benchmark suite. But this was created by putting an agent in a loop for a month with instructions to not overfit to the benchmark but no real supervision. For the most part, getting an LLM to give you a good benchmark score is fairly easy, and this case was no different; it took a couple weaks to roughly match Rust regex crate performance and then another couple weeks to get to 1.4x faster3 on rebar. But, agents are wont to reward hack and ovefit unless you put serious guardrails in place to avoid that, which I didn't do in this case as an experiment.

To check for overfitting, I somewhat arbitrarily4 used the ripgrep benchmark corpus as a benchmark a holdout and instead of being 1.4x faster it was 10x slower on cases where the benchmark didn't take forever due to an alg…
5619
Dan Luu @danluu.com · 12/08/2026
Interesting. I wonder if one of my company's MCP plug-ins is causing something bad to happen or if it's more about my workflow. This is all with CLI (I haven't tried to debug this because my restarter script works ok and my original thought was that this should get fixed pretty quickly).
020
Dan Luu @danluu.com · 12/08/2026
It seems like codex and claude have switched? Since the GPT-5.6 release, I've had to run a process that looks for signs of a memory leak and then waits for a good time to kill/restart/resume codex. Seems like 10s to 100s of memory leak restarts per day. If I don't do this, growth is unbounded.
110
Dan Luu @danluu.com · 10/08/2026
How do programming languages impact token efficiency and correctness? danluu.com/pl-tokens/
[image described in post][image described in post][image described in post]
2344
Dan Luu @danluu.com · 04/08/2026
Another thing we see on Amazon, for bad non-scam products, is taking over listings for older (often obviously unrelated) products with better ratings and then selling a bad product. Product dark patterns like this are pervasive and a clear sign of subterfuge, not informed price-quality decisions.
040
Dan Luu @danluu.com · 04/08/2026
A common pattern is brand makes good product and people switch to it, brand cashes out and declines in quality and people complain, new brand makes good product that people switch to, on repeat. We shouldn't see this pattern in cases where people are trading price for quality.
120
Dan Luu @danluu.com · 04/08/2026
Max could argue that consumers want to buy products that either never arrive or literally don't work, but it's more likely the case that people are being tricked and don't actually want to pay money for nothing. The way people's brand preferences switch also doesn't make sense under Max's thesis.
100
Dan Luu @danluu.com · 04/08/2026
I don't really buy that line of reasoning in part because of the argument in danluu.com/nothing-works/; people just don't know in a lot of cases and are not making an informed decision. An extreme example of this are the scam products I regularly see on FB and Amazon.
120
Dan Luu @danluu.com · 02/08/2026
The bulk of www.ftc.gov/system/files... is about standard dark patterns, but there's a fairly good-sized section on the use of 3rd party trackers. It would be really interesting if those end up carrying a real cost the company. Less interesting if this results in another GDPR/EU-banner-like thing.
ftc.gov
050
Dan Luu @danluu.com · 02/08/2026
Earlier this year, the FTC Chair said "you're going to have a hard time keeping up with the number of cases we're gonna be bringing" on privacy. It seems like this is starting; the FTC has filed a case against a company for promising privacy while running ad trackers on their site, which share info
271
Dan Luu @danluu.com · 25/07/2026
Trying it out again! I've been active on Mastodon, but the community seems to be shrinking Also, the AI posts tend to be about how bad AI is and how it rots your brain. Maybe so, but it's more interesting to see how people use it than to hear how bad it is again; there's a lot more variety that way
130
Dan Luu @danluu.com · 25/07/2026
Realistically, my failure rate when I apply to companies is pretty close to 100% and I have almost exclusively gotten jobs when people reach out to me to hire me because I'm hilariously bad at interviews. Sounds like it would be a fun job, though!
120
Dan Luu @danluu.com · 24/07/2026
Exercises in benchmarking and evals, part 7: performance napkin math, DeepSWE / Senior SWE-Bench, and winter tires danluu.com/exercise-7/
Screenshot of above the fold content from link.Graphs of memory latency over time, showing decline until 2005 or so and then a rough levelling off or maybe a small increase.
2130
Dan Luu @danluu.com · 04/07/2026
Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island: danluu.com/ai-coding/
In general, when I talk to software folks about testing, I'm coming from such a different place that they immediately look at me like I'm an alien, so let's talk about how we tested at this hardware company I worked for, Centaur, which informs my biases about how I like to work. Some of the things that we did that were or are unorthodox in the software world are:

    Hired dedicated QA / test engineers, with testing being a first-class career path on par with being a developer
    No code review by default
    Virtually no hand-written tests
    Constant testing via what programmers sometimes called property based testing, randomized testing, fuzzing, etc., although we just called those tests (hand-written tests were called "hand tests").
    Large regeression test suite (3 months wall clock to execute on compute farm)
    No unit tests

Just to give you an idea of the general structure, when I left (in 2013), we had about 1000 machines generating and running tests at all times for roughly 20 logic designers and 20 test engineers. This was on prem and the machines took up half a floor of the building we were in.

The general structure was that we had maybe 20% of machines running regression tests, and 80% generating and running new tests. Three months of regression tests is too much to gate commits on, so there was a much shorter list of tests that took maybe 10 minutes or so to run that people would run before committing. Those pre-commit tests would run on a special setup to run as quickly as possible, with overclocked machines that were the fastest machines money could buy, as well as a different simulator setup.
1457
Dan Luu @danluu.com · 11/04/2025
For non-Centaur features, we tried to match Intel since some software would hang or crash if we did something that was correct according to the manual but didn't match an actual Intel processor, so other things should be pretty standard (other than the obvious, like CentuarHauls vendor ID, etc.
010
Dan Luu @danluu.com · 11/04/2025
Hah. Sorry, I don't have data sheets squirrelled away that aren't just the stuff you can find online. For non-secret Centaur-specific feature flags, I don't know if there's a better still existing resource than looking at the 0xC0000001 section in git.kernel.org/pub/scm/linu...
110
Dan Luu @danluu.com · 08/01/2025
This is great! Thanks for the pointer!
000
Dan Luu @danluu.com · 22/11/2024
A version of Missile Command for the Commodore 64 where the bottom of your screen is the game state in memory and missiles cause memory corruption: csdb.dk/release/?id=.... In the video below, a missile broke my controls and caused my cursor to get stuck moving down and to the left.
716745
Dan Luu @danluu.com · 17/11/2024
The commentary I've seen says Teslas are safe so it must be the drivers but, per danluu.com/car-safety/, maybe it's the cars. The most fatal rated manufacturers (Kia/Hyundai, Dodge, Tesla) all did poorly — Kia/Hyundai, Dodge got the lowest rating and there's a strong case Tesla should have as well.
The ranking below is mainly based on how well vehicles scored when the driver-side small overlap test was added in 2012 and how well models scored when they were modified to improve test results.

    Tier 1: good without modifications
        Volvo
    Tier 2: mediocre without modifications; good with modifications
        None
    Tier 3: poor without modifications; good with modifications
        Mercedes
        BMW
    Tier 4: poor without modifications; mediocre with modifications
        Honda
        Toyota
        Subaru
        Chevrolet
        Tesla
        Ford
    Tier 5: poor with modifications or modifications not made
        Hyundai
        Dodge
        Nissan
        Jeep
        Volkswagen

These descriptions are approximations. Honda, Ford, and Tesla are the poorest fits for these descriptions, with Ford arguably being halfway in between Tier 4 and Tier 5 but also arguably being better than Tier 4 and not fitting into the classification and Honda and Tesla not really properly fitting into any category (with their category being the closest fit), but some others are also imperfect. Details below.
060
Dan Luu @danluu.com · 17/11/2024
I find it interesting/surprising that Tesla topped the www.iseecars.com/most-dangero... fatalities per mile ranking from 2018-2022. Fatality rate is strongly negatively correlated with price and weight and Teslas are much more expensive and heavier than average.
Please see link for text version of tablePlease see link for text version of tableBar chart showing that Teslas have a higher average selling price than any other tracked car manufacturer
3221
Dan Luu @danluu.com · 04/11/2024
A funny side effect of the crackdown on "AI" scraping is that I keep getting banned from sites for browsing too quickly. I barely use reddit anymore and I still managed to get IP banned for scraping (the error message indicated that I should get in touch with them if I want to do bulk accesses).
2171
Dan Luu @danluu.com · 28/10/2024
Steve Ballmer was an underrated CEO danluu.com/ballmer/
There's a common narrative that Microsoft was moribund under Steve Ballmer and then later saved by the miraculous leadership of Satya Nadella. This is the dominant narrative in every online discussion about the topic I've seen and it's a commonly expressed belief "in real life" as well. While I don't have anything negative to say about Nadella's leadership in this post, this narrative underrates Ballmer's role in Microsoft's success. Not only did Microsoft's financials, revenue and profit, look great under Ballmer, Microsoft under Ballmer made deep, long-term bets that set up Microsoft for success in the decades after his reign. At the time, the bets were widely panned, indicating that they weren't necessarily obvious, but we can see in retrospect that the company made very strong bets despite the criticism at the time.

In addition to overseeing deep investments in areas that people would later credit Nadella for, Ballmer set Nadella up for success by clearing out political barriers for any successor. Much like Gary Bernhardt's talk, which was panned because he made the problem statement and solution so obvious that people didn't realize they'd learned something non-trivial, Ballmer set up Microsoft for future success so effectively that it's easy to criticize him for being a bum because his successor is so successful.
Criticisms of Ballmer

For people who weren't around before the turn of the century, in the 90s, Microsoft used to be considered the biggest, baddest, company in town. But it wasn't long before people's opinions on Microsoft changed — by 2007, many people thought of Microsoft as the next IBM and Paul Graham wrote Microsoft is Dead, in which he noted that Microsoft being considered effective was ancient history:

    A few days ago I suddenly realized Microsoft was dead. I was talking to a young startup founder about how Google was different from Yahoo. I said that Yahoo had been warped from the start by their fear of Microsoft. That was why they'd po…
0191
Dan Luu @danluu.com · 13/08/2024
A former Apple engineer discusses Google product culture: > My director wore an Apple Watch and had an iPhone ... my VP too. Nobody was expected to eat the dog food and so few did. This was crazy to me coming from Apple ....
mattnewton 1 hour ago | parent | next [–]

When I worked at Google I got re-orged into the same division as pixel / android.

My director wore an Apple Watch and had an iPhone for personal use, and I am pretty sure I saw an Apple Watch on my VP too. Nobody was expected to eat the dog food and so few did. This was crazy to me coming from Apple- I remember several internal sites would ask you to file a radar (bug report) on why you switched to chrome from safari if you opened them in chrome. So many crazy issues I saw and reported didn’t actually matter to many high ranking members of the pixel team because they didn’t use the devices after 5pm.

There is a lot of incredible talent in that team but I think Google needs a minor culture shift to compete with Apple here.
1496
Reposted by Dan Luu
🌇👃🌃 @waxrepli.ca · 02/06/2024
i can speak with a little authority on this, i work in this field DCs expend water by evaporation, either in cooling towers or evaporative coolers (water becomes vapor, taking latent heat energy with it into the air and leaving cooler liquid water behind) so we expend liquid water resources 1/
57223
Dan Luu @danluu.com · 27/05/2024
On the 2011-2012 FTC antitrust investigation of Google: danluu.com/ftc-google-a...
This is a summary of the publicly available documents on the 2011-2012 FTC investigation of Google's allegedly antitcompetive actions in search and ads, followed by a tech-focused analysis of the decision from someone who's worked at the two companies that are discussed in the most detail in the memos (Google and Microsoft), worked in search, and worked closely with ads teams on optimizing ads ranking algorithms. I've seen a number of law-focused and economics-focused analyses, but I haven't seen a tech-focused analysis in a level of detail I find satisfying. In particular, a number of key arguments in the memos rely on evidence and inferences that would've seen implausible to someone who was familiar with tech, which I haven't seen discussed.

The law-focused and economics-focused analyses tend to avoid digging into this and, while there have been some articles written about tech errors for a lay audience, they've tended to explain that the inferences that were made were wrong in r...
183
Reposted by Dan Luu
Laurence Tratt @ltratt.bsky.social · 14/05/2024
What Factors Explain the Nature of Software? tratt.net/laurie/blog/...
tratt.net
Laurence Tratt: What Factors Explain the Nature of Software?
041
Dan Luu @danluu.com · 15/04/2024
Every once in a while, I think about going to work in the game industry.
@ZTGallagher • 1mo ago

I know a guy who was a QA tester for Obsidian Entertainment working on Neverwinter Nights 2. At the end of development, they invited all the QA testers to the parking lot for a celebration party.
There was no party, they disabled all their keys when they got out there, and told them they were all fired. And that was it...
1222
Dan Luu @danluu.com · 08/04/2024
Interesting comment about SGI leadership knowing about the problems they were facing and still being unable to come up with a way to handle them.
051
Dan Luu @danluu.com · 22/03/2024
Great! I'm looking forward to reading the blog post! I've noticed that you sometimes turn your social media comments into blog post, which is something I should probably do more of.
010
Dan Luu @danluu.com · 21/03/2024
Thanks for the comments. As someone who doesn't really work on this stuff, I find this super interesting! As an outsider, it seems like wasm might become mainstream whether or not it delivers benefits to end users, just like heavy SPAs became mainstream regardless of the benefits.
110
Dan Luu @danluu.com · 20/03/2024
I recently tried Blazor and the performance is incredibly bad (like, 5s to 10s initial load time for very simple apps, which you can push down to maybe 3s or something via various config options). From news.ycombinator.com/item?id=3836..., I guess people still like it because it's nice for devs.
110
Dan Luu @danluu.com · 20/03/2024
The effort to do this kind of work ended up getting defunded after a while even though the gains were measurable and very large, so even showing huge gains here wasn't sufficient. On Ember, I didn't know it was such a performance problem. I think that's interesting.
100
Dan Luu @danluu.com · 20/03/2024
We did see much larger impacts in the long-term holdback than in the initial test for the reason you mentioned — if someone thinks the app takes 60s to open, they probably won't open it very often, and you've already lost a huge fraction of users who've previously used the app and found it too slow
100
Dan Luu @danluu.com · 20/03/2024
On a slow device, this decreased time from opening the app to seeing a tweet from something like 60s to 48s. It's incredible that people in that range would even use the app, but apparently some did, and it got more people into the range where they'd use the app or got them to use the app more.
100
Dan Luu @danluu.com · 20/03/2024
on mobile, changes that made the app go from extremely slow to only very slow had large, measurable, impacts on retention/engagement/revenue. I forget the exact numbers, but a change that reduced the loading time of feature flags ended up increasing revenue something like 0.7%.
100
Dan Luu @danluu.com · 20/03/2024
It's interesting that the impact wasn't easily observable on mobile. For Twitter, there was an experiment where (I forget the exact number) 500ms or 1s of delay was added on web and the impact was huge, and it was clear we could easily reduce latency by that much (but never did), and
100