Dr Waku @drwaku.bsky.social · 27/06/2025AI agents have a half-life for their success rates at completing tasks. Yes, the same type of half-life as in nuclear chemistry: a constant chance of task failure (= radioactive decay) in each time period... 1/10 111
Dr Waku @drwaku.bsky.social · 05/06/2025The US AI Safety Institute (AISI) is being renamed into the Center for AI Standards and Innovation (CAISI). It has a pretty similar sounding mandate. But I would like to point out that Canada had this name first with the Canadian AI Safety Institute (CAISI)! :O 1/2 100
Dr Waku @drwaku.bsky.social · 28/05/2025AI is not like other technologies. AI doesn't just make everything we were doing before faster and cheaper. It’s another, faster feedback loop of reasoning and intelligence beyond the human brain. It will change life as we know it. Here’s why. (1/10) 120
Dr Waku @drwaku.bsky.social · 16/05/2025This September, the book "If anyone builds it, everyone dies" comes out. It's about the existential risk from superintelligence, and it's by Eliezer Yudkowsky @esyudkowsky.bsky.social, one of the most well-known figures in AI safety. 1/3 171
Dr Waku @drwaku.bsky.social · 17/04/2025OpenAI released o4 and o3-mini today, and some safety testing was performed by the third party METR, who said: "We detected several successful and unsuccessful attempts at “reward hacking” by o3. 1/7 100
Dr Waku @drwaku.bsky.social · 06/04/2025Today is the two year anniversary of my YouTube channel. 28.6K subscribers and 1.25M views. I think it's fair to say it's changed my life. Thank you to everyone who watches and participates in my community :) 020
Dr Waku @drwaku.bsky.social · 02/04/2025Today, OpenAI released PaperBench, an AI benchmark aimed at replicating cutting edge AI research papers from scratch. In this benchmark, the AI has to understand the paper, write code, and execute experiments. In other words, AI must become an (AI) research scientist. 120
Dr Waku @drwaku.bsky.social · 28/03/2025If anyone is interested in creating videos/posts/etc about AI safety, you can apply for funding from FLI: futureoflife.org/pro... Please contact me if you'd like to make long or short video content, I would be happy to collaborate or give tips!futureoflife.orgDigital Media Accelerator - Future of Life InstituteThe Digital Media Accelerator supports digital content from creators raising awareness and understanding about ongoing AI developments and issues. 020
Dr Waku @drwaku.bsky.social · 24/03/2025A few timelines about AI: - algorithmic gains for models doubles every 8 months - amount of compute used grows 10x every two years - the length of tasks agents can perform is doubling every 7 months - 50.8% of global VC funds are going to AI-focused companies 130
Dr Waku @drwaku.bsky.social · 22/03/2025Anthropic is using Amazon Trainium 2 chips to train their next models. It takes about three of these chips to match an H100 (at fp16) and four to match B200 (fp16). And Anthropic will have 400,000 of them at their disposal, from Amazon's Project Rainier. 120
Dr Waku @drwaku.bsky.social · 21/03/2025I heard an institution say, about a relatively important meeting, don't bring any devices with DeepSeek installed on them. Obviously worried about data exfiltration to China. There is such a technological separation growing alongside the ideological one. 110
Dr Waku @drwaku.bsky.social · 19/03/2025I believe that it is likely impossible to fully defend an LLM against jailbreaks. All instructions within the prompt exist at the same level, which means the precedence confusion between system/developer and user cannot be fully resolved. There is no way of enforcing rules. 110
Dr Waku @drwaku.bsky.social · 18/03/2025When Anthropic made a new alignment technique for Claude 3.7, called Constitutional Classifiers, they said 3,000 hours of jailbreaking effort had not been able to break the system. Then they created a public contest, and someone built a universal jailbreak within three weeks. 230
Dr Waku @drwaku.bsky.social · 16/03/2025AI has always had the problem of moving goalposts. When AI algorithms like A* and Djikstra's were invented, they were quickly considered to be "just search". When AI systems beat the best humans at chess, Jeopardy, and Go, it made news but quickly became the new normal. 110
Dr Waku @drwaku.bsky.social · 15/03/2025The term "AI safety" has several confusing meanings. I like to break it down into three categories, based on where malicious goals are coming from: - Robustness: no malicious goals - Misuse risk: human provides malicious goals - Existential risk: AI provides malicious goals 110
Dr Waku @drwaku.bsky.social · 24/12/2024There's an AI security forum in Paris this February, one of the satellite events to the 3rd international AI summit. The event "will bring together ~150 experts in AI safety, policy, and cybersecurity to discuss securing powerful AI systems". Let me know if you're interested! 040
Reposted by Dr WakuYoshua Bengio @yoshuabengio.bsky.social · 06/12/2024Recently answered @anilananth.bsky.social's questions for Nature. No matter when it arrives, AGI and the road to reach it will both help tackle thorny problems (e.g. climate change and diseases), and pose huge risks. Understanding and transparency are key. www.nature.com/articles/d41...nature.comHow close is AI to human-level intelligence?Large language models such as OpenAI’s o1 have electrified the debate over achieving artificial general intelligence, or AGI. But they are unlikely to reach this milestone on their own. 05013
Reposted by Dr WakuEthan Mollick @emollick.bsky.social · 23/11/2024Fascinating: In 2-hour sprints, AI agents outperform human experts at ML engineering tasks like optimizing GPU kernel. But humans pull ahead over longer periods - scoring 2x better at 32 hours. AI is faster but struggles with creative, long-term problem solving (for now?). metr.org/blog/2024-11... 923927
Dr Waku @drwaku.bsky.social · 25/11/2024@robertwiblin.bsky.social Hi, I would love to be added to your EA starter pack. Also, would love to chat sometime about the 80k podcast! 010
Dr Waku @drwaku.bsky.social · 25/11/2024@ahappier.world hello from another YouTuber, I focus on how AI will impact society and AI safety. Love your thumbnails 010
Dr Waku @drwaku.bsky.social · 25/11/2024In-depth analysis of why it makes sense to concentrate on restricting compute, instead of say talent, for AI safety arxiv.org/pdf/2402.08797arxiv.org 010
Reposted by Dr WakuHMYS @hmys.bsky.social · 19/11/2024Made a lesswrong starterpack. reply if you use lesswrong and want to be put in it, or if there's anyone else I've not added that should be added! go.bsky.app/EkcDcjA 9212
Dr Waku @drwaku.bsky.social · 20/11/2024Bluesky needs a fail whale business.time.com/2013/11/06/h... 031
Dr Waku @drwaku.bsky.social · 20/11/2024I created a Manifund to raise money for my channel for 2025: manifund.org/projects/240...manifund.org24,000 subscriber YouTube channel on AI safety | Dr WakuCover anticipated costs for making videos in 2025 010
Dr Waku @drwaku.bsky.social · 17/11/2024This is a great resource for what's happening in China around AI safety (thanks Circle Circle): aisafetychina.substack.comaisafetychina.substack.comAI Safety in China | Concordia AI | SubstackDelivered to your inbox every two weeks. Click to read AI Safety in China, by Concordia AI, a Substack publication. 010
Reposted by Dr WakuKol Tregaskes @koltregaskes.bsky.social · 17/11/2024If AGI arrives during Trump’s next term, 'none of the other stuff matters' Max Tegmark warns of AGI risks, calling for urgent regulation. Elon Musk’s influence could shape Trump’s AI policies, but the stakes remain high. kolt.at/Rlspu6 #AI #Technology #Innovation #Future #Leadershipkolt.atIf AGI arrives during Trump’s next term, ‘none of the other stuff matters’Future of Life Institute cofounder Max Tegmark on regulating AI, Elon Musk’s potential to be a good influence on Donald Trump, understanding how LLMs think, and more. 042
Dr Waku @drwaku.bsky.social · 15/11/2024Here's a new document by Connor Leahy et al that describes the problem of existential AI risk for a general audience: www.thecompendium.aithecompendium.aiThe CompendiumThe Compendium 031