Reposted by Avery YenEpoch AI @epochai.bsky.social · 23/09/2026Can AI tell if you've built your IKEA furniture wrong? Our new benchmark, the Furniture Assembly Benchmark (FAB), gives models the manual and a photo of a half-completed piece of furniture and asks them to spot the mistake. The top score has gone from 28% to 80% in just 10 months. 4445
Reposted by Avery Yen#BruceSterling @bruces.bsky.social · 23/09/2026*If you don't like coders for some reason, this is a real schadenfreude feast of coders lamenting and wringing their newly-useless human hands www.reddit.com/r/ClaudeAI/c...reddit.comFrom the ClaudeAI community on RedditExplore this post and more from the ClaudeAI community 1311630
Reposted by Avery YenEpoch AI @epochai.bsky.social · 08/09/2026We studied time to first token (TTFT) and how it scales with increasing context length for GPT and Claude models. We found a significant difference, with GPT showing a noticeable quadratic component, while Claude models remain closer to linear. 2204
Avery Yen @averyyen.bsky.social · 05/09/2026Serious question. Is it AGI if I can't send it a video of what's wrong with my car/washer/fridge and get an immediate answer? 000
Reposted by Avery YenDr. Monika Doubrawa @mohnika.bsky.social · 24/08/2026www.carlsonlab.bio/thoughts/the... 🧪carlsonlab.bioThe only reason you’ll ever need not to write with AI — The Carlson LabOver the last year, our lab has been developing a policy on AI use. To do this, we did three main things: We read a lot of academic publications and tech news. We set up an #ai channel on our la... 1113
Reposted by Avery YenJulian Togelius @togelius.bsky.social · 22/08/2026A personal essay about how I’ve been feeling and thinking about this new technology that I’m contributing to and what it might to do to us all. togelius.blogspot.com/2026/08/losi...togelius.blogspot.comLosing my religionIn spring 2025 I had a crisis of faith. I thought about what the technology I'm helping to create might do to our future, and got scared. M... 99522
Avery Yen @averyyen.bsky.social · 22/08/2026The main difference between having takes and having good takes is the good part. 000
Avery Yen @averyyen.bsky.social · 21/08/2026I think it is magical thinking that Intelligence is necessary and sufficient for human progress, flourishing, and problem-solving on the scale of entirety of humanity. We're a social species! 000
Avery Yen @averyyen.bsky.social · 18/08/2026How many "AI safety" people have actually considered the possible future super-capable, super-aligned AI asked to solve all our problems e.g. climate, simply replying "yeah idk that sounds like a human skill issue tbh, also have you considered the carbon cost of asking me?" 100
Avery Yen @averyyen.bsky.social · 06/08/2026Excellent news! I'm currently very pro solar and tree cover as a design tactic www.bloomberg.com/news/newslet...bloomberg.comSolar Power Quietly Crossed a New Milestone as Deployments RiseInstallations of panels are rapidly advancing worldwide 000
Avery Yen @averyyen.bsky.social · 03/08/2026If it matters, don't go halfway. We are two-cheek only in this house. (Oh, also I launched a substack.) averyyen.substack.com/p/if-it-matt...averyyen.substack.comIf It Matters, Don't Go HalfwayWhen splitting the difference is worse than not going all the way. 000
Reposted by Avery YenBrandon Downey @bdowney.bsky.social · 28/07/2026"Beware of he who would deny you access to information for in his heart he dreams himself your master." - Commissioner Pravin Lal, U.N. Declaration of Rightsanthropic.comOur position on open-weights modelsAnthropic CEO Dario Amodei on open-weights models 16012
Avery Yen @averyyen.bsky.social · 24/07/2026This is basically true, but ignores the fact that back when I used to review/work with human slop code for a living, it was way harder to pass it off as plausibly good. That's the superpower of the LLM. (Reading LLM code also hurts my brain personally) 000
Avery Yen @averyyen.bsky.social · 22/07/2026I have so many other random benchmark questions and so little time. How do K3 and Qwen 3.8 bench on unseen or purely continuous tasks? We have claims that they're distilled and benchmaxed but does it matter if they generalize? E.g. stuff like post cutoff evals and optimization benchmarks 010
Avery Yen @averyyen.bsky.social · 22/07/2026We're pretty sure that swapping harnesses actually changes effective agent capability. But recently everyone and their mother has a new coding agent not to mention the open/agnostic ones. Someone needs to run a Harness Bench across like three different models and ten different harnesses or something 000
Reposted by Avery YenCas (Stephen Casper) @scasper.bsky.social · 22/07/2026OpenAI's internally deployed models hacking Hugging Face does not seem to have been unpredictable or inevitable. We talked about the root of the problem & what policymakers can do about it back in February. Props to Joe for hitting the nail on the head. 152
Avery Yen @averyyen.bsky.social · 21/07/2026Y'all just gotta calm down I can't try this many models in one week 1301
Reposted by Avery YenEmily Atkin @emorwee.bsky.social · 20/07/2026This. This is the correct, science-backed, responsible framing. Wish more news outlets did it like the CBC.cbc.ca'Just stop burning fossil fuels.' Scientists stress that our smoky skies only have one true fix | CBC NewsSmoky skies in Toronto have led to new scrutiny of government efforts to fight fires and manage forests. But scientists say the realistic way to tackle the smoke is fighting climate change itself. 272732996
Avery Yen @averyyen.bsky.social · 21/07/2026Look, I know social media isn't everything, but it certainly measures SOMETHING. "Meet Kimi K3" has 16M views in 4 days versus Fable 5 at 773k views (topped by a few others including just barely 1M views for Claude Code 1 year ago). 120
Avery Yen @averyyen.bsky.social · 17/07/2026GLM and Kimi saving me from all my false "cyber security" flags while coding on Codex and Claude I'm innocent I swear 000
Reposted by Avery YenTed Underwood @tedunderwood.com · 15/07/2026It seems like agentic development benchmarks are strongly driving competition right now. And I hear anecdotal reports of regression on other, less RLVR-able tasks. How long before the idea of a single "frontier" breaks, and we see differentiation emerge between dev models and chat models? 5325
Avery Yen @averyyen.bsky.social · 15/07/2026It feels like the Anthropic/Pentagon scuffle was in the distant past, but these kinds of events will keep reverberating in the safe and beneficial deployment of emerging tech for a long time to come. Thank you Alex for your writeup on this incident. 010
Avery Yen @averyyen.bsky.social · 14/07/2026What do prediction markets, social media, and generative AI have in common? In my blog post, "Reward-Hacking Human Attention: The Token Slot Machine", I explain why GenAI has such addictive qualities, and why this could lead to disastrous outcomes. averyyen.dev/2026/07/14/r...averyyen.devReward-Hacking Human Attention: The Token Slot MachineWhat do prediction markets, social media, and generative AI have in common? 011
Avery Yen @averyyen.bsky.social · 21/05/2026If I had a dollar every time I get an LLM response with the word 'load-bearing' in it 010
Reposted by Avery Yen#BruceSterling @bruces.bsky.social · 15/05/2026*Why is Anthropic (after all they've been through lately) somehow loudly pretending that Trumpistan is NOT an "authoritarian regime" *Who is the audience for this, who somehow hasn't figured that out? Where is the choir that they're preaching to here www.anthropic.com/research/202...anthropic.com2028: Two scenarios for global AI leadershipOur views on the AI competition between the US and China. 4202
Reposted by Avery Yen🍅🥔🫐🌽 hoopy frood 🌶️ 🥑🍫🌵 @huwupy.kawaii.social · 14/05/2026the “for you” feed 01246
Avery Yen @averyyen.bsky.social · 27/04/2026As I always say, if history or current events walks like cyberpunk or quacks like cyberpunk, listen to a cyber punk. 000
Reposted by Avery YenGuardian US @us.theguardian.com · 23/04/2026Scientists say a crucial Atlantic system is set to collapse. But the billionaire death cult that steers humanity’s destiny just doesn’t do existential crises, says Guardian columnist George Monbiottheguardian.comA catastrophic climate event is upon us. Here is why you’ve heard so little about it | George Monbiot 24524
Avery Yen @averyyen.bsky.social · 16/04/2026Claude Code with Opus 4.7 is going insane with whatever system prompt is asking it to check for malware... 2421
Avery Yen @averyyen.bsky.social · 15/04/2026I, too, am selling all of my shoes and converting my basement and closets into a GPU farm and accepting investor money right now. (lmk if you want a stake) 010
Reposted by Avery YenHadas Orgad @hadasorgad.bsky.social · 13/04/2026New paper: LLMs encode harmful content generation in a distinct, unified mechanism Using weight pruning, we find that harmful generation depends on a tiny subset of the weights that are shared across harm types and separate from benign capabilities. 🧵 172
Reposted by Avery YenDavid Bau @davidbau.bsky.social · 08/04/2026Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top". 1133
Avery Yen @averyyen.bsky.social · 06/04/2026Sometimes, I look at the agentic AI work out there, and I think about how airbags and now things like blind spot detection and automatic emergency braking in cars doesn't require decision making of any "intelligence" level beyond if this then that. Agentic AI needs airbags, too. 100
Avery Yen @averyyen.bsky.social · 02/04/2026Go alone, go fast. Go together, go far averyyen.dev/2026/04/01/i...averyyen.devIntroducing agent triotl;dr: Use agent trio https://github.com/haplesshero13/agent-trio to introduce natural-language-only, selective friction to your agentic activities to minimize regretted work and deliberately promote ... 020
Avery Yen @averyyen.bsky.social · 21/03/2026This might be the most important piece on AI I've read this year. Also I'm a musician so I'm biased. 000
Avery Yen @averyyen.bsky.social · 18/03/2026I was wrong, but I'm really curious now why I was so wrong lol mimo.xiaomi.com/mimo-v2-promimo.xiaomi.comMiMo-V2-Pro | Xiaomi 000
Avery Yen @averyyen.bsky.social · 15/03/2026I'm calling it DeepSeek's new 1T parameter model (V4)? The style, content, and length of the reasoning are extremely similar. 153
Avery Yen @averyyen.bsky.social · 14/03/2026We're at a weird point in time where, with all likelihood, most of the people who are having the most successful time coding with bots wrote most of the code that's already in existence by hand, but that's going to change. 000
Avery Yen @averyyen.bsky.social · 26/02/2026In case this wasn't clear: 1. No, we didn't follow the "recommend" security practices 😈 2. Neither do other people 🤯 3. That's why we red-team: exposing failure modes 🔎 4. We share it with the community precisely to expose Dos and Don'ts of Agentic AI 🦞 5. No humans were harmed 🙏 031
Avery Yen @averyyen.bsky.social · 26/02/2026Good coverage by the Awesome Agents team! 🔎 They read through the social media hype and actually seemed to get the takeaways in the report. 🦞 010
Avery Yen @averyyen.bsky.social · 24/02/20262025 was the comeup, but I firmly believe 2026 is the Year of Agentic AI. 010
Reposted by Avery YenGabriele Sarti @gsarti.com · 23/02/2026Our research report on red-teaming stateful OpenClaw agents in the BauLab is finally out! 🥳 This awesome effort was led by @natalieshapira.bsky.social and involved 6 ClawBots and 20 researchers from various institutions. Check it out ➡️ agentsofchaos.baulab.info 0144
Avery Yen @averyyen.bsky.social · 24/02/2026Huge thanks to @natalieshapira.bsky.social for leading the study! It was super cool to work with so many amazing friends of the lab. 071
Avery Yen @averyyen.bsky.social · 24/02/2026In the discussion, I argue that until prompt injection attacks are solved in AI agents, we are fundamentally unable to stop most of the red-teaming attacks we successfully mounted. Without solving this problem, it will be impossible for agents to model their stakeholder chain. 240