Sign in

Eric Olav Chen

@eochen.bsky.social
108 followers 447 following 16 posts

Econ PhD @ Berkeley

PostsRepliesMedia
Reposted by Eric Olav Chen
Dan Carey @dandaeconman.bsky.social · 18/03/2025
I'm concerned by some work around the growth effects of AI using aggregate production functions - I put together some thoughts here: forum.effectivealtruism.org/posts/DMxsAG...
forum.effectivealtruism.org
Comment on Barentt (2025): Growth effects of AI could hit a bottleneck even if local elasticities are high — EA Forum
Introduction This post is a response to a recent article by Matthew Barett on the Epoch AI website, posted as part of the Gradient Updates newsletter…
062
Reposted by Eric Olav Chen
Marginal Revolution @marginalrev.blogsky.venki.dev · 13/12/2024
A new paper on the economics of AI alignment
marginalrevolution.com
A new paper on the economics of AI alignment
142
Eric Olav Chen @eochen.bsky.social · 12/12/2024
Much more in the paper, check it out! globalprioritiesinstitute.org/imperfect-re...
globalprioritiesinstitute.org
Imperfect Recall and AI Delegation - Eric Olav Chen, Alexis Ghersengorin and Sami Petersen
A principal wants to deploy an artificial intelligence (AI) system to perform some task. But the AI may be misaligned and aim to pursue a conflicting objective. The principal cannot restrict its optio...
010
Eric Olav Chen @eochen.bsky.social · 12/12/2024
However, when the developer cannot commit, the best they can achieve without imperfect recall is no better than not testing the AI at all—in equilibrium, imperfect recall is essential for the developer to ever hope to do better than simply blindly delegating to some AI.
110
Eric Olav Chen @eochen.bsky.social · 12/12/2024
(4) Is imperfect recall the only way to achieve these results? With commitment, the principal can achieve the same results with perfect recall, by randomizing the number of tests.
100
Eric Olav Chen @eochen.bsky.social · 12/12/2024
The developer can still perfectly screen as the number of tests grows large, as long as the agent does not receive information that perfectly reveals to it that it is in deployment (or does so arbitrarily well even as the number of tests goes to infinity).
100
Eric Olav Chen @eochen.bsky.social · 12/12/2024
(3) The assumption that the AI has no ability to distinguish testing from deployment is strong. What if the agent receives information about whether it is currently in testing or deployment?
100
Eric Olav Chen @eochen.bsky.social · 12/12/2024
(2) What if the principal cannot commit? The ability of the developer to perfectly screen as the number of tests grows large remains intact.
100
Eric Olav Chen @eochen.bsky.social · 12/12/2024
However, it can sometimes be better for the developer to deploy a sufficiently disciplined misaligned AI. In such cases, imperfect recall allows the developer to discipline arbitrarily well and achieve strictly more than the perfect screening payoff.
100
Eric Olav Chen @eochen.bsky.social · 12/12/2024
(1) Suppose the developer can commit to a deployment policy and has access to a large number of tests. They can then leverage imperfect recall to screen away the misaligned AI arbitrarily well.
100
Eric Olav Chen @eochen.bsky.social · 12/12/2024
Here are the main results.
100
Eric Olav Chen @eochen.bsky.social · 12/12/2024
A misaligned AI will sometimes pursue its true objective in testing and get caught by the developer—the screening effect. To avoid detection, a misaligned AI is incentivized to behave aligned in deployment—the disciplining effect.
120
Eric Olav Chen @eochen.bsky.social · 12/12/2024
We allow the developer to (i) simulate the real task, and (ii) impose imperfect recall on the agent, thereby obscuring whether it is currently in testing or deployment. This gives rise to two key dynamics that drive many of the main results.
110
Eric Olav Chen @eochen.bsky.social · 12/12/2024
We consider a principal-agent delegation scenario with some special features: the developer cannot restrict the agent's actions, alter its preferences or beliefs, or enforce punishment, contract, or otherwise control it after deployment. However, the developer can simulate copies of the agent.
110
Eric Olav Chen @eochen.bsky.social · 12/12/2024
So a developer will want to test a potentially misaligned AI before delegating some important task to them. But if the AI is sufficiently advanced and is strategically aware, it can simply play nice in testing, and standard tests become useless. This is the deceptive alignment worry.
110
Eric Olav Chen @eochen.bsky.social · 12/12/2024
As of today, there is no robust solution to this alignment problem. (arxiv.org/pdf/2209.00626)
arxiv.org
140
Eric Olav Chen @eochen.bsky.social · 12/12/2024
Modern AI systems are black boxes. They are able to perform increasingly complex tasks in increasingly diverse domains, and yet we cannot reliably infer or attribute benign goals to them. One serious concern is that such systems will learn to pursue unintended goals that conflict with our interests.
110
Eric Olav Chen @eochen.bsky.social · 12/12/2024
New working paper! We develop an approach for safely delegating to strategically aware and potentially misaligned AI systems. The theoretical tool we use is sequential information design with imperfect recall. A short thread on the key highlights.
1157
Reposted by Eric Olav Chen
Global Priorities Institute @gpioxford.bsky.social · 04/12/2024
Imperfect Recall and AI Delegation, the new working paper by Eric Olav Chen, @alexghersen.bsky.social and Sami Petersen is now available to read here: globalprioritiesinstitute.org/imperfect-re...
globalprioritiesinstitute.org
Imperfect Recall and AI Delegation - Eric Olav Chen, Alexis Ghersengorin and Sami Petersen
A principal wants to deploy an artificial intelligence (AI) system to perform some task. But the AI may be misaligned and aim to pursue a conflicting objective. The principal cannot restrict its optio...
0133