Sign in

Aaron Scher

@aaronscher.bsky.social
102 followers 480 following 75 posts

Technical AI Governance Research at MIRI Views are my own

PostsRepliesMedia
Reposted by Aaron Scher
MIRI @intelligence.org · 01/05/2025
New AI governance research agenda from MIRI’s TechGov Team. We lay out our view of the strategic landscape and actionable research questions that, if answered, would provide important insight on how to reduce catastrophic and extinction risks from AI. 🧵1/10 techgov.intelligence.org/research/ai-...
2125
Reposted by Aaron Scher
Malo Bourgon @malo.online · 18/03/2025
MIRI's (@intelligence.org) Technical Governance Team submitted a comment on the AI Action Plan. Great work by David Abecassis, @pbarnett.bsky.social, and @aaronscher.bsky.social Check it out here: techgov.intelligence.org/research/res...
052
Aaron Scher @aaronscher.bsky.social · 05/12/2024
- Reflection on how this is hard but we should try: bsky.app/profile/aaro... - Mechanism highlight: FlexHEGs: bsky.app/profile/aaro...
000
Aaron Scher @aaronscher.bsky.social · 05/12/2024
Some versions of FlexHEGs could be designed + implemented in only a couple years and retrofitted to existing chips! Designing more secure chips like this could unlock many AI governance options that aren’t currently available! (3/3)
010
Aaron Scher @aaronscher.bsky.social · 05/12/2024
These have been discussed previously, yoshuabengio.org/wp-content/u..., @yoshuabengio.bsky.social, so we don’t explain them too much in the report, but they are widely useful! (2/3)
yoshuabengio.org
120
Aaron Scher @aaronscher.bsky.social · 05/12/2024
One mechanism that seems promising is Flexible Hardware-Enabled Guarantees (FlexHEGs) and on-chip approaches. These could potentially be used to securely carry out a wide range of governance operations on AI chips, without leaking sensitive information. (1/3)
120
Aaron Scher @aaronscher.bsky.social · 05/12/2024
Reflection: The more I got into the weeds on this project, the harder verification seemed. Some difficulties are distributed training, algorithmic progress, and the need to be robust against state-level adversaries. It’s hard, but we have to do it! (1/1)
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
- Inspectors could be viable: bsky.app/profile/aaro... - Reflection on US/China conflict: bsky.app/profile/aaro... - Mechanism highlight: Signatures of High-Level Chip Measures: bsky.app/profile/aaro...
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
- Substituting high-tech low-access with low-tech high-access: bsky.app/profile/aaro... - Distributed training causes problems: bsky.app/profile/aaro... - Mechanism highlight: Networking Equipment Interconnect Limits: bsky.app/profile/aaro...
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Conceptually, this could be thought of as having a ‘signature’ of approved activity (e.g., inference, or finetuning on models you’re allowed to finetune) which other chips have to stay sufficiently close to. (6/6)
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
But it’s probably much easier because you no longer have a huge distribution shift (e.g., new algorithms, maybe different types of chips) because you included labeled data from the monitored country in your classifier training set. (5/6)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
So when deployed, this mechanism looks like similar classification systems: you’re measuring e.g., the power draw or inter-chip network activity of a bunch of chips and trying to detect any prohibited activity. (4/6)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
But this classification problem gets much easier if you have access to labeled data for chip activities that are approved by the treaty. You could get this by giving inspectors temporary code access. (3/6)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Classifying chip activity has been researched previously, but it’s not clear it will be robust enough in the international verification context: highly-competent adversaries who can afford to waste some compute could potentially spoof this. (2/6)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
One mechanism that seems promising: Signatures of High-Level Chip Measures. Classify workloads (e.g., is it training or inference) based on high-level chip measures like power-draw, but using ‘signatures’ of these measures based on temporary code access. (1/6)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
In the report we give example calculations for inter-pod bandwidth limits, discuss distributed training, note various issues, and generally flesh out this idea. Kulp et al. discuss this idea for manufacturing new chips, but in our case that’s not strictly necessary. (8/8)
techgov.intelligence.org
Mechanisms to Verify International Agreements About AI Development — MIRI Technical Governance Team
In this research report we provide an in-depth overview of the mechanisms that could be used to verify adherence to international agreements about AI development.
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
This mechanism is promising because it can be implemented with physical networking equipment and security cameras but no code access. This means it poses minimal security risk and could be implemented quickly. (7/8)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
This gap is critical to interconnect bandwidth limits: with enough inter-pod bandwidth for inference but not training, a data center can verifiably claim that these AI chips are not participating in a large training run. (6/8)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Back-of-the-envelope calculations indicate a bandwidth difference of ~1.5 million times from data parallelism to inference! Distributed training methods close this gap substantially, but there is likely still a gap after such adjustments. (5/8)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Inference only requires tokens to move in and out of a pod — very little information for text data, order 100 KB/s. For training, communication is typically much larger: activations (tensor parallelism) or gradients (data parallelism). (4/8)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
If between-pod bandwidth is set correctly, a pod could conduct inference but couldn’t efficiently participate in a larger training run. This is because training has higher between-pod communication requirements than inference. (3/8)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
AI chips can be physically arranged to have high bandwidth communication with only a small number, e.g., 128, of other chips (a “pod”) and very low bandwidth communication to chips outside this pod. (2/8)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
One mechanism that seems especially promising: Networking Equipment Interconnect Limits, like “Fixed Sets” discussed by www.rand.org/pubs/working... but can be implemented with custom networking equipment quickly. (1/8)
rand.org
Hardware-Enabled Governance Mechanisms
The authors introduce the concept of hardware-enabled governance mechanisms, which could help achieve U.S. artificial intelligence governance goals, and discuss two mechanisms that could limit uses of...
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Whistleblowers and interviews could be relatively straightforward to implement, not requiring vastly novel tech, and they could make it very difficult to keep violations hidden given the ~hundreds of people involved. (12/12)
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Interviews could also be key. E.g., people working on AI projects could be made available for interviews by international regulators. These could be structured specifically around detecting violations of international agreements. (11/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Whistleblower programs can provide a form of verification by increasing the chances that treaty violations are detected. They’re low-tech and have precedent! Actually protecting whistleblowers may be difficult, of course. (10/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Third, we’ll look at a couple mechanisms that work across most of the policy goals the international community might want to verify. (9/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
On the other hand, international inspectors could be granted broad code access, allowing them to review workloads and confirm that prohibited activities aren’t happening. There are obviously security concerns, but this could basically be done tomorrow. (8/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
These mechanisms would involve outfitting AI chips to do governance operations, either on the AI chip itself or an adjacent processor. These ideas are great, but they’re not ready yet, more work is needed! (7/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Second, let’s consider verifying some property about what chips are doing (e.g., not doing training). FlexHEG, Flexible Hardware-Enabled Guarantee, Mechanisms, “on-chip” and “hardware-enabled” approaches have been suggested previously. (6/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
We should likely pursue both, but the high-tech approach may not be ready for a while due to the security requirements. We have precedent for physical inspections and continuous monitoring in nuclear verification—technologically ready but needs political will. (5/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
But a dumb low-tech thing you can do is just have inspectors go into data centers and count the chips; either doing this occasionally or setting up continuous monitoring (e.g., security cameras or perimeter security). (4/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
This would require chips to have private keys which can’t be extracted, potentially needing fabrication of new chips or retrofitting existing chips with tamper-proof mechanisms. (3/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
First, let’s look at locating AI chips. One approach discussed by www.iaps.ai/research/loc... is high-tech on-chip mechanisms where chips ping servers throughout the world and triangulate location based on response time. (2/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
One thread throughout this report is that low-tech, high-access solutions can often substitute for high-tech, low-access solutions, let’s walk through some examples. (1/12)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Accomplishing these governance goals, and many others, will be easier if training can only take place in the very largest data centers. But if small data centers are also a concern, they should also be monitored. (10/10)
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
There are many scenarios where we need Governance! E.g., if alignment is difficult and we need to pause so we have more time, if we solve technical alignment by using certain development practices and want to ensure everybody follows them. (9/10)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
We do not get to choose the risks. Whether technical alignment turns out to be difficult is an empirical question, not a matter of political opinion. If the risks are catastrophic, drastic action will be warranted to mitigate them. (8/10)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Some people seem to think AI regulation/governance won’t be needed. E.g., that free markets are the way, a16z.com/the-techno-o.... I think they are wrong. (7/10)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
If distributed training is not viable, such monitoring could focus only on large compute clusters (e.g., >10,000 GPUs) while still being effective. But if distributed training is efficient, monitoring small clusters may be necessary. (6/10)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
USG and the international community should take drastic actions to mitigate these risks. One such approach is to monitor the specialized computers used for AI development. (5/10)
110
Aaron Scher @aaronscher.bsky.social · 04/12/2024
But the effect of better publicly available distributed training methods could be more widespread compute monitoring! And (IMO) correctly so. Advanced AI systems could pose threats to global security (e.g., human extinction). (4/10)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Ironically, some of those working on distributed training methods hope they are democratizing access to AI and promoting individual freedom, e.g., nousresearch.com/from-black-b.... (3/10)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
But if highly distributed training is viable, training could take part across 10s of smaller data centers and monitoring training would require monitoring these smaller data centers. The cost of compliance and of monitoring/verification would both be higher. (2/10)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Distributed training (i.e., geographically distributed, decentralized) could pose major problems for many AI Governance plans. In the default case, large AI training happens in a small number of big data centers, so monitoring training can focus on those data centers. (1/10)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
This realization makes international inspectors seem much more viable. The major blocker IMO is security risks. Without extreme measures these are hard to deal with, but extreme measures are costly for the freedom of inspectors. A couple years isn’t so bad though. (3/3)
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
That sounds kinda terrible, but one factor makes it much more palatable: at the current rate of AI development, secrets will likely be obsolete in <2 years. Keeping inspectors under tight security for a couple years seems much better than a life sentence. (2/3)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Inspectors who have full access to your systems seem like they could pose a major privacy and security risk, so it may be necessary to have very tight info sec around them, e.g., limited communication to home countries. (1/3)
100
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Personally, I don’t have strong takes about this question. In the case where the US retains a lead and we get US-China negotiations, I sure hope we can pull off international agreements + verification rather than war. (4/4)
000
Aaron Scher @aaronscher.bsky.social · 04/12/2024
Others thought that conventional strikes were off the table and that only cyberattacks would be in play in this situation. Why go to (extremely fatal, potentially nuclear) war to avoid a potential, but not guaranteed, loss from US ASI down the line? (3/4)
100