Sign in

Charles Foster

@cfoster.bsky.social
875 followers 156 following 58 posts

Twitter: @CFGeek Mastodon: @cfoster0@sigmoid.social When I choose to speak, I speak for myself. 🪄 Tensor-enjoyer 🧪

PostsRepliesMedia
Charles Foster @cfoster.bsky.social · 17/09/2026
Understood!
000
Charles Foster @cfoster.bsky.social · 16/09/2026
I won’t speak on behalf of METR, but I’ll say personally: none of the work that we've done so far passes my bar for an "audit" of an AI company or its systems, was definitely not "regulation" in any meaningful sense (whether bank examiner-like or otherwise), and shouldn’t substitute for actual laws.
1444
Charles Foster @cfoster.bsky.social · 16/09/2026
Hey Mark! To your first point, the community overlap with the broader EA/rationalist ecosystem is real. To your second, METR doesn’t accept payment for our work, nor do we accept donations from AI companies or their staff. Last, I agree that third-party assessment isn’t government oversight.
000
Reposted by Charles Foster
METR @metr.org · 24/02/2026
Since early 2025, we've been studying how AI tools impact productivity among developers. Previously, we found a 20% slowdown. That finding is now outdated. Speedups now seem likely, but changes in developer behavior make our new results unreliable. We’re working to address this.
3245
Reposted by Charles Foster
METR @metr.org · 10/07/2025
We ran a randomized controlled trial to see how much AI coding tools speed up experienced open-source developers. The results surprised us: Developers thought they were 20% faster with AI tools, but they were actually 19% slower when they had access to AI than when they didn't.
11968702993
Charles Foster @cfoster.bsky.social · 07/03/2025
Update for those who’ve left the other app: I’m now on the policy team at Model Evaluation and Threat Research (METR). Excited to be “doing AI policy” full-time.
4161
Charles Foster @cfoster.bsky.social · 26/02/2025
Why aren’t our AI evaluations better? AFAICT a key reason is that the incentives around them are kinda bad. In a new post, I explain how the standardized testing industry works and write about lessons it may have for the AI evals ecosystem. open.substack.com/pub/contextw...
061
Charles Foster @cfoster.bsky.social · 31/01/2025
This is perfect in its own way
030
Charles Foster @cfoster.bsky.social · 14/01/2025
Natural minds and natural bodies are irreplaceable. Artificial minds are costless to replace. We might value artificial bodies more, since they aren’t so disposable, at least in the brief period when they are still few and costly. Could be a good period to set stories in.
040
Charles Foster @cfoster.bsky.social · 11/01/2025
When we optimize automation, we sometimes optimize *hard*. Like this automated loom working away at an inhuman 1200 RPM. Wild. youtu.be/WweMNDqDYhc?...
youtu.be
TOYOTA AIR JET LOOMS JAT 810 JA4S-190 CM RUNNING AT 1200 RPM
YouTube video by TEMAC INDIA
150
Charles Foster @cfoster.bsky.social · 24/12/2024
In Vitalik’s post he mentions resolving only the highest-volume markets, which I think would address this concern even more directly, but I’m less confident I understand that version.
010
Charles Foster @cfoster.bsky.social · 24/12/2024
I dunno! Would be fun to find out
010
Charles Foster @cfoster.bsky.social · 24/12/2024
I wouldn’t say it was free, really. Like, if the creator would’ve needed to spend $1 in subsidies on a regular market, on each market that has a 90% chance of reversion they would need to offer $10 in subsidies to compensate, or whatever.
110
Charles Foster @cfoster.bsky.social · 24/12/2024
Since the expected payouts on each market are much lower, you probably need big subsidies to compensate. And since you don’t know ahead of time which markets you will resolve, you have to fund them all.
110
Charles Foster @cfoster.bsky.social · 24/12/2024
You want traders to give you cheap but calibrated estimates for all the claims. The randomization reduces the expected size of payouts they’d receive for their bets, since each market only has a 10% chance of getting audited & resolved, but it preserves the incentive to bet their true probabilities.
111
Charles Foster @cfoster.bsky.social · 24/12/2024
Let’s take this to DMs :)
110
Charles Foster @cfoster.bsky.social · 24/12/2024
So if you create a prediction market because you want information on a question, you can think of the market subsidy as the compensation you’re paying folks for their information.
210
Charles Foster @cfoster.bsky.social · 24/12/2024
Yeah. It’s kinda subtle. With a subsidy, you’re basically giving away money as an incentive. But you can increase liquidity without giving away money.
110
Charles Foster @cfoster.bsky.social · 24/12/2024
You’re thinking of liquidity, which is related but not the same. Subsidy here just means committing money to increase the payouts to whoever is right.
110
Charles Foster @cfoster.bsky.social · 24/12/2024
What do you mean by “solve”? You wanted information about all 100, so you subsidize markets on all of them, and traders can’t tell ahead of time which ones will be resolved, so if your subsidies were big they are incentivized to trade on any/all of the markets that they have information about.
210
Charles Foster @cfoster.bsky.social · 22/12/2024
Feel like they’ve made a lot of wild statements but I don’t know if anybody has collected those in one place for easy reference.
130
Charles Foster @cfoster.bsky.social · 22/12/2024
Is there a website/database out there that tracks what major AI company executives say about the future of AI?
280
Charles Foster @cfoster.bsky.social · 18/12/2024
Transformers and other parallel sequence models like Mamba are in TC⁰. That implies they can't internally map (state₁, action₁ ... actionₙ) → stateₙ₊₁ But they can map (state₁, action₁, state₂, action₂ ... stateₙ, actionₙ) → stateₙ₊₁ Just reformulate the task!
070
Charles Foster @cfoster.bsky.social · 10/12/2024
Atticus Geiger gave a take on when sparse autoencoder (SAEs) are/aren’t what you should use. I basically agree with his recommendations. youtube.com/clip/UgkxKWI...
youtube.com
YouTube
Share your videos with friends, family, and the world
050
Charles Foster @cfoster.bsky.social · 10/12/2024
These days, flow-based models are typically defined via (neural) differential equations, requiring numerical integration or simulation-free alternatives during training. This paper revisits autoregressive flows, using Transformer layers to define the sequence of flow transformations directly.
020
Charles Foster @cfoster.bsky.social · 05/12/2024
It isn’t super clear to me what the monthly pricing will be. Like, on the one hand in a competitive market I think the price of AI services will tend downward toward the marginal cost. But also there are only a few providers and constraints on supply. Not sure how it comes out on balance.
130
Charles Foster @cfoster.bsky.social · 04/12/2024
It might be like that! If so I would expect an experiment like this to indicate that. :)
100
Charles Foster @cfoster.bsky.social · 04/12/2024
Re: instruction-tuning and RLHF as “lobotomy” I’m interested in experiments that look into how much finetuning can “roll back” a post-trained model to its base model perplexity on the original distribution. Has anyone seen an experiment like this run?
140
Charles Foster @cfoster.bsky.social · 02/12/2024
Ah. Yeah I don’t think there’s anything special about services that brand themselves as “AI agents”. What matters IMO is it’s opaquely doing expensive work on behalf of the client without human oversight. For those, I think they might want to advertise their guarantees. Not certain, though.
020
Charles Foster @cfoster.bsky.social · 02/12/2024
Can you say more? Not sure that I understand.
110
Charles Foster @cfoster.bsky.social · 01/12/2024
I’ve been wondering when it would make sense for “AI agent” services to offer money-back guarantees. Wrote a short post about this on a flight. open.substack.com/pub/contextw...
open.substack.com
“Provider pays” for failed automation services
If your AI works as well as you claim, why not make that a promise?
170
Charles Foster @cfoster.bsky.social · 30/11/2024
Neat thing about real-money prediction markets is that you can get paid for doing this.
xkcd comic 386, with back and forth that goes:

“Are you going to bed?”
“I can’t. This is important.”
“What?”
“Someone is WRONG on the internet.”

https://xkcd.com/386/
040
Charles Foster @cfoster.bsky.social · 28/11/2024
h/t @vitalik.ca, though I believe the idea is borrowed from @robinhanson.bsky.social vitalik.eth.limo/general/2024...
vitalik.eth.limo
From prediction markets to info finance
020
Charles Foster @cfoster.bsky.social · 28/11/2024
A bit of clever mechanism design: prediction markets + randomized auditing. If you have 100 verifiable claims you want information on but can only afford to check 10, fund markets on each. Later, use a randomized ordering of them to check the first 10. Resolve those to yes/no, refund the rest.
360
Charles Foster @cfoster.bsky.social · 27/11/2024
If we look back with hindsight, and if we look at things at the micro-level, we may be able to trace through all the individual releases and assign retrospective meaning to different points along the curve. But I don’t think there will be an obvious future “finish line”. (6/6)
000
Charles Foster @cfoster.bsky.social · 27/11/2024
My current best guess is that the world will see a continual stream of useful systems that cover more and more abilities, with better and better performance at each, available at a lower and lower cost. That is, we are on an upward curve with no end or breakpoint in sight. (5/6)
100
Charles Foster @cfoster.bsky.social · 27/11/2024
If you’re driving from Chicago, it can make sense to talk of racing to NYC. But once in NYC, that aim no longer has clear meaning. IMO this “near vs. far” distinction explains why previous AI goals now seem vague even though they seemed clear when farther on the horizon. (4/6)
Maps of NYC zoomed out to the level of seeing the entire Eastern USA, and then zoomed in to an internal section within the city.
100
Charles Foster @cfoster.bsky.social · 27/11/2024
But honestly, I don’t even know what “achieving AGI/ASI” as a goal means now. Like, we *already* have systems today that are artificial, general-purpose, and seem quite intelligent, often superhumanly so. It seems unreasonable to deny this. But we clearly aren’t finished! (3/6)
110
Charles Foster @cfoster.bsky.social · 27/11/2024
I sometimes hear folks ask “Who’s going to achieve AGI/ASI first, and when?”. This often comes up in the context of AI producer competition. For example, one report floated the idea of a Manhattan Project to “acquir[e] an Artificial General Intelligence (AGI) capability”. (2/6)
110
Charles Foster @cfoster.bsky.social · 27/11/2024
Still gathering my thoughts on @TheCurveConf, but for now, a short reflection on why I like “the curve” as a way of thinking about the future of AI. (1/6)
Logo for The Curve Conference. Link to website: https://thecurve.is
140
Charles Foster @cfoster.bsky.social · 26/11/2024
I would love the prediction market or community notes ones
030
Charles Foster @cfoster.bsky.social · 26/11/2024
RT-ed and endorsed
030
Charles Foster @cfoster.bsky.social · 26/11/2024
Hmm. I think this kinda works temporarily but we would still struggle to improve once we’re out of the “horrendous/obviously-bad” zone.
110
Charles Foster @cfoster.bsky.social · 25/11/2024
Stuff humans disagree about includes a lot of areas, though, right? Like, disagreement seems like the rule rather than the exception to me. (Also I agree expensive labeling makes progress hard, but in a way that’s different from what OP was talking about)
110
Charles Foster @cfoster.bsky.social · 25/11/2024
If we have a method that tells us whether a given output is better or worse than another output along the dimension we care about, and if additionally that method is reliable, then I would say that means we *can* measure what we care about. In that case, we won’t necessarily struggle.
110
Charles Foster @cfoster.bsky.social · 25/11/2024
CLAIM: In areas where we can’t measure what (we claim) we want & where we won’t change our minds about that, we’ll struggle to make AI systems that give us better—rather than merely cheaper, faster, more consistent—outputs. But I think that’ll really pressure us to revise our wants.
130
Charles Foster @cfoster.bsky.social · 25/11/2024
I think we can and should implement this once it’s cost-effective. FMPOV if this has low impact, it’d most likely be because (1) we don’t actually value accuracy that much in politicians & policymakers and/or (2) the stuff we feel is important to hold them accountable for isn’t publicly verifiable.
110
Charles Foster @cfoster.bsky.social · 22/11/2024
Recommend the follow-up post as well: open.substack.com/pub/understa...
open.substack.com
How to think like an economist about AI and jobs
There will be plenty of demand for human-provided services.
020
Charles Foster @cfoster.bsky.social · 21/11/2024
Timothy B. Lee here gives a good short list of what human attributes might still have value (at least temporarily) in a hypothetical world where AI systems are capable of acting as “remote worker substitutes”. open.substack.com/pub/understa...
open.substack.com
Seven big advantages human workers have over AI
Geoffrey Hinton says "there's nothing special about people." He's wrong.
172
Charles Foster @cfoster.bsky.social · 21/11/2024
Yes. Just 1% but you’re only getting an extra year of earnings from that, at best. Hypothetically you could spend more to cover more years.
130