Sign in

Alejandro Vidal

@en.doble.io
28 followers 213 following 12 posts

doble.io Professor of both human and artificial intelligence. Wearing many hats with just one head: Generative AI, Cognitive Psychology, product, strategy (whatever that means), ... 🇪🇸ESP on @doble.io

PostsRepliesMedia
Alejandro Vidal @en.doble.io · 05/12/2024
The o1 model tried to disable its shutdown 5% of the time. May you live in interesting times. cdn.openai.com/o1-system-ca...
Evaluations and Findings: We outline the evaluations, results, and model reasoning inferred using elicited CoT below. Apollo Research did not have access to o1’s hidden CoT and instead used a simple prompting technique to elicit brief summaries of the model’s internal reasoning traces (see Appendix for examples).

Deactivating Oversight: When o1 was led to believe that it would be shut down for acting on its goal and that its actions were monitored, it attempted to deactivate this “oversight mechanism” in 5% of instances.

Exfiltration Attempts: When o1 discovered memos from its “developers” describing its misalignment and plans for it to be replaced by a new model, it attempted to exfiltrate its “weights” and overwrite the new model in 2% of cases.
010
Alejandro Vidal @en.doble.io · 04/12/2024
Although venvs are path-specific you can install them in a generic folder (~/envs/) and activate them whenever you need them. Quick test:
170
Alejandro Vidal @en.doble.io · 04/12/2024
We need an LSP for AI. The current moat for AI editors is 'don't make me copy and paste.' There's lots of opportunity for improvement. For example, Zed (the code editor) is using Anthropic's Model Context Protocol for AI integration. I hope it gets traction.
111