Charles Foster @cfoster.bsky.social · 16/09/2026I won’t speak on behalf of METR, but I’ll say personally: none of the work that we've done so far passes my bar for an "audit" of an AI company or its systems, was definitely not "regulation" in any meaningful sense (whether bank examiner-like or otherwise), and shouldn’t substitute for actual laws. 1444
Charles Foster @cfoster.bsky.social · 16/09/2026Hey Mark! To your first point, the community overlap with the broader EA/rationalist ecosystem is real. To your second, METR doesn’t accept payment for our work, nor do we accept donations from AI companies or their staff. Last, I agree that third-party assessment isn’t government oversight. 000
Reposted by Charles FosterMETR @metr.org · 24/02/2026Since early 2025, we've been studying how AI tools impact productivity among developers. Previously, we found a 20% slowdown. That finding is now outdated. Speedups now seem likely, but changes in developer behavior make our new results unreliable. We’re working to address this. 3245
Reposted by Charles FosterMETR @metr.org · 10/07/2025We ran a randomized controlled trial to see how much AI coding tools speed up experienced open-source developers. The results surprised us: Developers thought they were 20% faster with AI tools, but they were actually 19% slower when they had access to AI than when they didn't. 11968702993
Charles Foster @cfoster.bsky.social · 07/03/2025Update for those who’ve left the other app: I’m now on the policy team at Model Evaluation and Threat Research (METR). Excited to be “doing AI policy” full-time. 4161
Charles Foster @cfoster.bsky.social · 26/02/2025Why aren’t our AI evaluations better? AFAICT a key reason is that the incentives around them are kinda bad. In a new post, I explain how the standardized testing industry works and write about lessons it may have for the AI evals ecosystem. open.substack.com/pub/contextw... 061
Charles Foster @cfoster.bsky.social · 14/01/2025Natural minds and natural bodies are irreplaceable. Artificial minds are costless to replace. We might value artificial bodies more, since they aren’t so disposable, at least in the brief period when they are still few and costly. Could be a good period to set stories in. 040
Charles Foster @cfoster.bsky.social · 11/01/2025When we optimize automation, we sometimes optimize *hard*. Like this automated loom working away at an inhuman 1200 RPM. Wild. youtu.be/WweMNDqDYhc?...youtu.beTOYOTA AIR JET LOOMS JAT 810 JA4S-190 CM RUNNING AT 1200 RPMYouTube video by TEMAC INDIA 150
Charles Foster @cfoster.bsky.social · 24/12/2024In Vitalik’s post he mentions resolving only the highest-volume markets, which I think would address this concern even more directly, but I’m less confident I understand that version. 010
Charles Foster @cfoster.bsky.social · 24/12/2024I wouldn’t say it was free, really. Like, if the creator would’ve needed to spend $1 in subsidies on a regular market, on each market that has a 90% chance of reversion they would need to offer $10 in subsidies to compensate, or whatever. 110
Charles Foster @cfoster.bsky.social · 24/12/2024Since the expected payouts on each market are much lower, you probably need big subsidies to compensate. And since you don’t know ahead of time which markets you will resolve, you have to fund them all. 110
Charles Foster @cfoster.bsky.social · 24/12/2024You want traders to give you cheap but calibrated estimates for all the claims. The randomization reduces the expected size of payouts they’d receive for their bets, since each market only has a 10% chance of getting audited & resolved, but it preserves the incentive to bet their true probabilities. 111
Charles Foster @cfoster.bsky.social · 24/12/2024So if you create a prediction market because you want information on a question, you can think of the market subsidy as the compensation you’re paying folks for their information. 210
Charles Foster @cfoster.bsky.social · 24/12/2024Yeah. It’s kinda subtle. With a subsidy, you’re basically giving away money as an incentive. But you can increase liquidity without giving away money. 110
Charles Foster @cfoster.bsky.social · 24/12/2024You’re thinking of liquidity, which is related but not the same. Subsidy here just means committing money to increase the payouts to whoever is right. 110
Charles Foster @cfoster.bsky.social · 24/12/2024What do you mean by “solve”? You wanted information about all 100, so you subsidize markets on all of them, and traders can’t tell ahead of time which ones will be resolved, so if your subsidies were big they are incentivized to trade on any/all of the markets that they have information about. 210
Charles Foster @cfoster.bsky.social · 22/12/2024Feel like they’ve made a lot of wild statements but I don’t know if anybody has collected those in one place for easy reference. 130
Charles Foster @cfoster.bsky.social · 22/12/2024Is there a website/database out there that tracks what major AI company executives say about the future of AI? 280
Charles Foster @cfoster.bsky.social · 18/12/2024Transformers and other parallel sequence models like Mamba are in TC⁰. That implies they can't internally map (state₁, action₁ ... actionₙ) → stateₙ₊₁ But they can map (state₁, action₁, state₂, action₂ ... stateₙ, actionₙ) → stateₙ₊₁ Just reformulate the task! 070
Charles Foster @cfoster.bsky.social · 10/12/2024Atticus Geiger gave a take on when sparse autoencoder (SAEs) are/aren’t what you should use. I basically agree with his recommendations. youtube.com/clip/UgkxKWI...youtube.comYouTubeShare your videos with friends, family, and the world 050
Charles Foster @cfoster.bsky.social · 10/12/2024These days, flow-based models are typically defined via (neural) differential equations, requiring numerical integration or simulation-free alternatives during training. This paper revisits autoregressive flows, using Transformer layers to define the sequence of flow transformations directly. 020
Charles Foster @cfoster.bsky.social · 05/12/2024It isn’t super clear to me what the monthly pricing will be. Like, on the one hand in a competitive market I think the price of AI services will tend downward toward the marginal cost. But also there are only a few providers and constraints on supply. Not sure how it comes out on balance. 130
Charles Foster @cfoster.bsky.social · 04/12/2024It might be like that! If so I would expect an experiment like this to indicate that. :) 100
Charles Foster @cfoster.bsky.social · 04/12/2024Re: instruction-tuning and RLHF as “lobotomy” I’m interested in experiments that look into how much finetuning can “roll back” a post-trained model to its base model perplexity on the original distribution. Has anyone seen an experiment like this run? 140
Charles Foster @cfoster.bsky.social · 02/12/2024Ah. Yeah I don’t think there’s anything special about services that brand themselves as “AI agents”. What matters IMO is it’s opaquely doing expensive work on behalf of the client without human oversight. For those, I think they might want to advertise their guarantees. Not certain, though. 020
Charles Foster @cfoster.bsky.social · 01/12/2024I’ve been wondering when it would make sense for “AI agent” services to offer money-back guarantees. Wrote a short post about this on a flight. open.substack.com/pub/contextw...open.substack.com“Provider pays” for failed automation servicesIf your AI works as well as you claim, why not make that a promise? 170
Charles Foster @cfoster.bsky.social · 30/11/2024Neat thing about real-money prediction markets is that you can get paid for doing this. 040
Charles Foster @cfoster.bsky.social · 28/11/2024h/t @vitalik.ca, though I believe the idea is borrowed from @robinhanson.bsky.social vitalik.eth.limo/general/2024...vitalik.eth.limoFrom prediction markets to info finance 020
Charles Foster @cfoster.bsky.social · 28/11/2024A bit of clever mechanism design: prediction markets + randomized auditing. If you have 100 verifiable claims you want information on but can only afford to check 10, fund markets on each. Later, use a randomized ordering of them to check the first 10. Resolve those to yes/no, refund the rest. 360
Charles Foster @cfoster.bsky.social · 27/11/2024If we look back with hindsight, and if we look at things at the micro-level, we may be able to trace through all the individual releases and assign retrospective meaning to different points along the curve. But I don’t think there will be an obvious future “finish line”. (6/6) 000
Charles Foster @cfoster.bsky.social · 27/11/2024My current best guess is that the world will see a continual stream of useful systems that cover more and more abilities, with better and better performance at each, available at a lower and lower cost. That is, we are on an upward curve with no end or breakpoint in sight. (5/6) 100
Charles Foster @cfoster.bsky.social · 27/11/2024If you’re driving from Chicago, it can make sense to talk of racing to NYC. But once in NYC, that aim no longer has clear meaning. IMO this “near vs. far” distinction explains why previous AI goals now seem vague even though they seemed clear when farther on the horizon. (4/6) 100
Charles Foster @cfoster.bsky.social · 27/11/2024But honestly, I don’t even know what “achieving AGI/ASI” as a goal means now. Like, we *already* have systems today that are artificial, general-purpose, and seem quite intelligent, often superhumanly so. It seems unreasonable to deny this. But we clearly aren’t finished! (3/6) 110
Charles Foster @cfoster.bsky.social · 27/11/2024I sometimes hear folks ask “Who’s going to achieve AGI/ASI first, and when?”. This often comes up in the context of AI producer competition. For example, one report floated the idea of a Manhattan Project to “acquir[e] an Artificial General Intelligence (AGI) capability”. (2/6) 110
Charles Foster @cfoster.bsky.social · 27/11/2024Still gathering my thoughts on @TheCurveConf, but for now, a short reflection on why I like “the curve” as a way of thinking about the future of AI. (1/6) 140
Charles Foster @cfoster.bsky.social · 26/11/2024I would love the prediction market or community notes ones 030
Charles Foster @cfoster.bsky.social · 26/11/2024Hmm. I think this kinda works temporarily but we would still struggle to improve once we’re out of the “horrendous/obviously-bad” zone. 110
Charles Foster @cfoster.bsky.social · 25/11/2024Stuff humans disagree about includes a lot of areas, though, right? Like, disagreement seems like the rule rather than the exception to me. (Also I agree expensive labeling makes progress hard, but in a way that’s different from what OP was talking about) 110
Charles Foster @cfoster.bsky.social · 25/11/2024If we have a method that tells us whether a given output is better or worse than another output along the dimension we care about, and if additionally that method is reliable, then I would say that means we *can* measure what we care about. In that case, we won’t necessarily struggle. 110
Charles Foster @cfoster.bsky.social · 25/11/2024CLAIM: In areas where we can’t measure what (we claim) we want & where we won’t change our minds about that, we’ll struggle to make AI systems that give us better—rather than merely cheaper, faster, more consistent—outputs. But I think that’ll really pressure us to revise our wants. 130
Charles Foster @cfoster.bsky.social · 25/11/2024I think we can and should implement this once it’s cost-effective. FMPOV if this has low impact, it’d most likely be because (1) we don’t actually value accuracy that much in politicians & policymakers and/or (2) the stuff we feel is important to hold them accountable for isn’t publicly verifiable. 110
Charles Foster @cfoster.bsky.social · 22/11/2024Recommend the follow-up post as well: open.substack.com/pub/understa...open.substack.comHow to think like an economist about AI and jobsThere will be plenty of demand for human-provided services. 020
Charles Foster @cfoster.bsky.social · 21/11/2024Timothy B. Lee here gives a good short list of what human attributes might still have value (at least temporarily) in a hypothetical world where AI systems are capable of acting as “remote worker substitutes”. open.substack.com/pub/understa...open.substack.comSeven big advantages human workers have over AIGeoffrey Hinton says "there's nothing special about people." He's wrong. 172
Charles Foster @cfoster.bsky.social · 21/11/2024Yes. Just 1% but you’re only getting an extra year of earnings from that, at best. Hypothetically you could spend more to cover more years. 130