Sign in

Alejandro Vidal

@doble.io
75 followers 297 following 16 posts

doble.io Profesor de inteligencias humanas y artificiales por igual. Muchos gorros y una sola cabeza: IA Generativa, Psicología cognitiva, producto, estrategia (sea lo que sea eso), ...

PostsRepliesMedia
Alejandro Vidal @doble.io · 08/02/2025
Never
Never run side effect at import time
020
Alejandro Vidal @doble.io · 20/12/2024
🤞🤞
PyCon US 2025
Your proposal has been submitted.
Your proposals
020
Alejandro Vidal @doble.io · 05/12/2024
El nuevo modelo o1 de OpenAI intentó parar su apagado un 5% de las veces... vienen tiempos interesantes.
Evaluations and Findings: We outline the evaluations, results, and model reasoning inferred using elicited CoT below. Apollo Research did not have access to o1’s hidden CoT and instead used a simple prompting technique to elicit brief summaries of the model’s internal reasoning traces (see Appendix for examples).

Deactivating Oversight: When o1 was led to believe that it would be shut down for acting on its goal and that its actions were monitored, it attempted to deactivate this “oversight mechanism” in 5% of instances.

Exfiltration Attempts: When o1 discovered memos from its “developers” describing its misalignment and plans for it to be replaced by a new model, it attempted to exfiltrate its “weights” and overwrite the new model in 2% of cases.
041
Alejandro Vidal @doble.io · 15/11/2024
#TRG24 Me encanta el concepto: optimización del entretenimiento. La optimización se carga todo... Ley de goodharts aplicada a videojuegos.
020