Dan Luu @danluu.com · 23/09/2026At the time, I was doing automated bug finding with human review. Human review found zero known false positives in the bugs I filed. Anthropic was using a better model and had more tokens to spend, which gave them the advantage on the two most difficult factors w.r.t. false positive rejection. 050
Dan Luu @danluu.com · 23/09/2026Why didn't Anthropic do effective false positive rejection with their Mythos/Glasswing vuln reports? In danluu.com/ai-coding/#m..., I mentioned a colleague finding that the reports we got were mostly false. That seems common? E.g., gregkh on kernel reports and www.vulncheck.com/blog/anthrop... 2130
Dan Luu @danluu.com · 22/09/2026People keep replying to tell me that Ralph loops worked better than nothing, which seems to illustrate my point? When Ralph loops were popular, I used a loop with context, which outperformed. That's not counting clearing context when it helps instead of always keeping, which you'd do in practice. 060
Dan Luu @danluu.com · 22/09/2026Did Ralph loops ever work? They were popular for a while, but when I did a comparison, they underperformed. There were all these theories about why Ralph loops were effective ("invert the latent space of the model"), but AFAICT people just didn't run the comparison? danluu.com/pl-tokens/#r... 4320
Reposted by Dan LuuDan McKinley @mcfunley.com · 08/05/2026Maybe the best thing on the internet at the moment: * An AI PR pitching platform has confused Kyle Kingsbury, MMA fighter and podcast host with Kyle Kingsbury, software engineer and gay leather man * Good Kyle decided to start a podcast where he interviews confused manosphere influencersyoutube.comThe Kyle Kingsbury Podcast Podcast - Epsiode 02 Dave RossiYouTube video by Aphyr Null 25820
Dan Luu @danluu.com · 18/09/2026There's no point at which turning your brain off will work: danluu.com/brain-off/ 27017
Dan Luu @danluu.com · 16/09/2026Yes, more than one DC was blown up, but of course DCs being blown up aren't uncorrelated events. We already know that hardware failures like disk failures, CPU failures, etc., are highly correlated events, and those aren't even adversarial failures! 1291
Dan Luu @danluu.com · 16/09/2026Amazon exec: "you wouldn't notice [if someone blew up a datacenter]. I mean, we might be a bit upset, but you wouldn't notice! [laughs]" Amazon after DCs blown up: "After a thorough assessment, we have determined that we are unable to restore access to the resources and data hosted ...." 3708
Dan Luu @danluu.com · 15/09/2026Interesting to see Steve Yegge say he never successfully built anything with Gas Town. I mentioned not finding these super vibed orchestrators useful because the reliability was too low (w.r.t. completing non-trivial tasks). Turns out the author of the most famous one had the exact same issue. 612211
Dan Luu @danluu.com · 08/09/2026Thanks! Some ancient minifier that I use was deleting some CSS that makes the page work on mobile. I should really find a better minifier or maybe just write one myself. 010
Dan Luu @danluu.com · 07/09/2026How well do agents use test/verification techniques? danluu.com/agentic-test... 4193
Dan Luu @danluu.com · 01/09/2026How accurate have Ed Zitron's AI skeptic predictions been? danluu.com/zitron/ 1944375
Dan Luu @danluu.com · 21/08/2026There's no reason for software to be slow anymore: danluu.com/perf-opt/ 3844
Dan Luu @danluu.com · 12/08/2026Interesting. I wonder if one of my company's MCP plug-ins is causing something bad to happen or if it's more about my workflow. This is all with CLI (I haven't tried to debug this because my restarter script works ok and my original thought was that this should get fixed pretty quickly). 020
Dan Luu @danluu.com · 12/08/2026It seems like codex and claude have switched? Since the GPT-5.6 release, I've had to run a process that looks for signs of a memory leak and then waits for a good time to kill/restart/resume codex. Seems like 10s to 100s of memory leak restarts per day. If I don't do this, growth is unbounded. 110
Dan Luu @danluu.com · 10/08/2026How do programming languages impact token efficiency and correctness? danluu.com/pl-tokens/ 2344
Dan Luu @danluu.com · 04/08/2026Another thing we see on Amazon, for bad non-scam products, is taking over listings for older (often obviously unrelated) products with better ratings and then selling a bad product. Product dark patterns like this are pervasive and a clear sign of subterfuge, not informed price-quality decisions. 040
Dan Luu @danluu.com · 04/08/2026A common pattern is brand makes good product and people switch to it, brand cashes out and declines in quality and people complain, new brand makes good product that people switch to, on repeat. We shouldn't see this pattern in cases where people are trading price for quality. 120
Dan Luu @danluu.com · 04/08/2026Max could argue that consumers want to buy products that either never arrive or literally don't work, but it's more likely the case that people are being tricked and don't actually want to pay money for nothing. The way people's brand preferences switch also doesn't make sense under Max's thesis. 100
Dan Luu @danluu.com · 04/08/2026I don't really buy that line of reasoning in part because of the argument in danluu.com/nothing-works/; people just don't know in a lot of cases and are not making an informed decision. An extreme example of this are the scam products I regularly see on FB and Amazon. 120
Dan Luu @danluu.com · 02/08/2026The bulk of www.ftc.gov/system/files... is about standard dark patterns, but there's a fairly good-sized section on the use of 3rd party trackers. It would be really interesting if those end up carrying a real cost the company. Less interesting if this results in another GDPR/EU-banner-like thing.ftc.gov 050
Dan Luu @danluu.com · 02/08/2026Earlier this year, the FTC Chair said "you're going to have a hard time keeping up with the number of cases we're gonna be bringing" on privacy. It seems like this is starting; the FTC has filed a case against a company for promising privacy while running ad trackers on their site, which share info 271
Dan Luu @danluu.com · 25/07/2026Trying it out again! I've been active on Mastodon, but the community seems to be shrinking Also, the AI posts tend to be about how bad AI is and how it rots your brain. Maybe so, but it's more interesting to see how people use it than to hear how bad it is again; there's a lot more variety that way 130
Dan Luu @danluu.com · 25/07/2026Realistically, my failure rate when I apply to companies is pretty close to 100% and I have almost exclusively gotten jobs when people reach out to me to hire me because I'm hilariously bad at interviews. Sounds like it would be a fun job, though! 120
Dan Luu @danluu.com · 24/07/2026Exercises in benchmarking and evals, part 7: performance napkin math, DeepSWE / Senior SWE-Bench, and winter tires danluu.com/exercise-7/ 2130
Dan Luu @danluu.com · 04/07/2026Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island: danluu.com/ai-coding/ 1457
Dan Luu @danluu.com · 11/04/2025For non-Centaur features, we tried to match Intel since some software would hang or crash if we did something that was correct according to the manual but didn't match an actual Intel processor, so other things should be pretty standard (other than the obvious, like CentuarHauls vendor ID, etc. 010
Dan Luu @danluu.com · 11/04/2025Hah. Sorry, I don't have data sheets squirrelled away that aren't just the stuff you can find online. For non-secret Centaur-specific feature flags, I don't know if there's a better still existing resource than looking at the 0xC0000001 section in git.kernel.org/pub/scm/linu... 110
Dan Luu @danluu.com · 22/11/2024A version of Missile Command for the Commodore 64 where the bottom of your screen is the game state in memory and missiles cause memory corruption: csdb.dk/release/?id=.... In the video below, a missile broke my controls and caused my cursor to get stuck moving down and to the left. 716745
Dan Luu @danluu.com · 17/11/2024The commentary I've seen says Teslas are safe so it must be the drivers but, per danluu.com/car-safety/, maybe it's the cars. The most fatal rated manufacturers (Kia/Hyundai, Dodge, Tesla) all did poorly — Kia/Hyundai, Dodge got the lowest rating and there's a strong case Tesla should have as well. 060
Dan Luu @danluu.com · 17/11/2024I find it interesting/surprising that Tesla topped the www.iseecars.com/most-dangero... fatalities per mile ranking from 2018-2022. Fatality rate is strongly negatively correlated with price and weight and Teslas are much more expensive and heavier than average. 3221
Dan Luu @danluu.com · 04/11/2024A funny side effect of the crackdown on "AI" scraping is that I keep getting banned from sites for browsing too quickly. I barely use reddit anymore and I still managed to get IP banned for scraping (the error message indicated that I should get in touch with them if I want to do bulk accesses). 2171
Dan Luu @danluu.com · 13/08/2024A former Apple engineer discusses Google product culture: > My director wore an Apple Watch and had an iPhone ... my VP too. Nobody was expected to eat the dog food and so few did. This was crazy to me coming from Apple .... 1496
Reposted by Dan Luu🌇👃🌃 @waxrepli.ca · 02/06/2024i can speak with a little authority on this, i work in this field DCs expend water by evaporation, either in cooling towers or evaporative coolers (water becomes vapor, taking latent heat energy with it into the air and leaving cooler liquid water behind) so we expend liquid water resources 1/ 57223
Dan Luu @danluu.com · 27/05/2024On the 2011-2012 FTC antitrust investigation of Google: danluu.com/ftc-google-a... 183
Reposted by Dan LuuLaurence Tratt @ltratt.bsky.social · 14/05/2024What Factors Explain the Nature of Software? tratt.net/laurie/blog/...tratt.netLaurence Tratt: What Factors Explain the Nature of Software? 041
Dan Luu @danluu.com · 15/04/2024Every once in a while, I think about going to work in the game industry. 1222
Dan Luu @danluu.com · 08/04/2024Interesting comment about SGI leadership knowing about the problems they were facing and still being unable to come up with a way to handle them. 051
Dan Luu @danluu.com · 22/03/2024Great! I'm looking forward to reading the blog post! I've noticed that you sometimes turn your social media comments into blog post, which is something I should probably do more of. 010
Dan Luu @danluu.com · 21/03/2024Thanks for the comments. As someone who doesn't really work on this stuff, I find this super interesting! As an outsider, it seems like wasm might become mainstream whether or not it delivers benefits to end users, just like heavy SPAs became mainstream regardless of the benefits. 110
Dan Luu @danluu.com · 20/03/2024I recently tried Blazor and the performance is incredibly bad (like, 5s to 10s initial load time for very simple apps, which you can push down to maybe 3s or something via various config options). From news.ycombinator.com/item?id=3836..., I guess people still like it because it's nice for devs. 110
Dan Luu @danluu.com · 20/03/2024The effort to do this kind of work ended up getting defunded after a while even though the gains were measurable and very large, so even showing huge gains here wasn't sufficient. On Ember, I didn't know it was such a performance problem. I think that's interesting. 100
Dan Luu @danluu.com · 20/03/2024We did see much larger impacts in the long-term holdback than in the initial test for the reason you mentioned — if someone thinks the app takes 60s to open, they probably won't open it very often, and you've already lost a huge fraction of users who've previously used the app and found it too slow 100
Dan Luu @danluu.com · 20/03/2024On a slow device, this decreased time from opening the app to seeing a tweet from something like 60s to 48s. It's incredible that people in that range would even use the app, but apparently some did, and it got more people into the range where they'd use the app or got them to use the app more. 100
Dan Luu @danluu.com · 20/03/2024on mobile, changes that made the app go from extremely slow to only very slow had large, measurable, impacts on retention/engagement/revenue. I forget the exact numbers, but a change that reduced the loading time of feature flags ended up increasing revenue something like 0.7%. 100
Dan Luu @danluu.com · 20/03/2024It's interesting that the impact wasn't easily observable on mobile. For Twitter, there was an experiment where (I forget the exact number) 500ms or 1s of delay was added on web and the impact was huge, and it was clear we could easily reduce latency by that much (but never did), and 100