Sign in

Avery Yen

@averyyen.bsky.social
85 followers 77 following 77 posts

Researcher @epochai.bsky.social. Opinions owned by my cat and dog. Once a Pivot, always a Pivot. Leave me anonymous feedback: www.admonymous.co/avery-yen

PostsRepliesMedia
Reposted by Avery Yen
Epoch AI @epochai.bsky.social · 23/09/2026
Can AI tell if you've built your IKEA furniture wrong? Our new benchmark, the Furniture Assembly Benchmark (FAB), gives models the manual and a photo of a half-completed piece of furniture and asks them to spot the mistake. The top score has gone from 28% to 80% in just 10 months.
4445
Reposted by Avery Yen
#BruceSterling @bruces.bsky.social · 23/09/2026
*If you don't like coders for some reason, this is a real schadenfreude feast of coders lamenting and wringing their newly-useless human hands www.reddit.com/r/ClaudeAI/c...
reddit.com
From the ClaudeAI community on Reddit
Explore this post and more from the ClaudeAI community
1311630
Reposted by Avery Yen
Epoch AI @epochai.bsky.social · 08/09/2026
We studied time to first token (TTFT) and how it scales with increasing context length for GPT and Claude models. We found a significant difference, with GPT showing a noticeable quadratic component, while Claude models remain closer to linear.
2204
Avery Yen @averyyen.bsky.social · 05/09/2026
I mean the Epoch "interesting and hard" Erdös problems, wasn't clear sorry
000
Avery Yen @averyyen.bsky.social · 05/09/2026
I've noticed a few step changes, namely strategic exploration on EBR-bench, it's the first public model to solve open Erdös problems, and it does better on some computer use stuff. But it's not noticeable when I'm just using it as a coding agent for normie swe stuff. 🤔
110
Avery Yen @averyyen.bsky.social · 05/09/2026
Serious question. Is it AGI if I can't send it a video of what's wrong with my car/washer/fridge and get an immediate answer?
000
Reposted by Avery Yen
Dr. Monika Doubrawa @mohnika.bsky.social · 24/08/2026
www.carlsonlab.bio/thoughts/the... 🧪
carlsonlab.bio
The only reason you’ll ever need not to write with AI — The Carlson Lab
Over the last year, our lab has been developing a policy on AI use. To do this, we did three main things: We read a lot of academic publications and tech news. We set up an #ai channel on our la...
1113
Reposted by Avery Yen
Julian Togelius @togelius.bsky.social · 22/08/2026
A personal essay about how I’ve been feeling and thinking about this new technology that I’m contributing to and what it might to do to us all. togelius.blogspot.com/2026/08/losi...
togelius.blogspot.com
Losing my religion
In spring 2025 I had a crisis of faith. I thought about what the technology I'm helping to create might do to our future, and got scared. M...
99522
Avery Yen @averyyen.bsky.social · 22/08/2026
The main difference between having takes and having good takes is the good part.
000
Avery Yen @averyyen.bsky.social · 21/08/2026
I think it is magical thinking that Intelligence is necessary and sufficient for human progress, flourishing, and problem-solving on the scale of entirety of humanity. We're a social species!
000
Avery Yen @averyyen.bsky.social · 18/08/2026
"yeah okay here's this really great battery chemistry in a pdf I just made you, but for the sake of the environment I'm gonna stop yapping and suggest you take it up with your regional and national legislature"
001
Avery Yen @averyyen.bsky.social · 18/08/2026
How many "AI safety" people have actually considered the possible future super-capable, super-aligned AI asked to solve all our problems e.g. climate, simply replying "yeah idk that sounds like a human skill issue tbh, also have you considered the carbon cost of asking me?"
100
Avery Yen @averyyen.bsky.social · 06/08/2026
Excellent news! I'm currently very pro solar and tree cover as a design tactic www.bloomberg.com/news/newslet...
bloomberg.com
Solar Power Quietly Crossed a New Milestone as Deployments Rise
Installations of panels are rapidly advancing worldwide
000
Avery Yen @averyyen.bsky.social · 03/08/2026
If it matters, don't go halfway. We are two-cheek only in this house. (Oh, also I launched a substack.) averyyen.substack.com/p/if-it-matt...
averyyen.substack.com
If It Matters, Don't Go Halfway
When splitting the difference is worse than not going all the way.
000
Avery Yen @averyyen.bsky.social · 29/07/2026
Yes this is also how exercise works! 💪🏃
000
Avery Yen @averyyen.bsky.social · 29/07/2026
Something something cathedrals and bazaars
010
Reposted by Avery Yen
Brandon Downey @bdowney.bsky.social · 28/07/2026
"Beware of he who would deny you access to information for in his heart he dreams himself your master." - Commissioner Pravin Lal, U.N. Declaration of Rights
anthropic.com
Our position on open-weights models
Anthropic CEO Dario Amodei on open-weights models
16012
Avery Yen @averyyen.bsky.social · 27/07/2026
Hardly ever wasn't, but he just published a take on wolves and sheep basically claiming genetic superiority and complaining Claude wouldn't translate it for him 🤪
020
Avery Yen @averyyen.bsky.social · 24/07/2026
This is basically true, but ignores the fact that back when I used to review/work with human slop code for a living, it was way harder to pass it off as plausibly good. That's the superpower of the LLM. (Reading LLM code also hurts my brain personally)
000
Avery Yen @averyyen.bsky.social · 22/07/2026
I have so many other random benchmark questions and so little time. How do K3 and Qwen 3.8 bench on unseen or purely continuous tasks? We have claims that they're distilled and benchmaxed but does it matter if they generalize? E.g. stuff like post cutoff evals and optimization benchmarks
010
Avery Yen @averyyen.bsky.social · 22/07/2026
We're pretty sure that swapping harnesses actually changes effective agent capability. But recently everyone and their mother has a new coding agent not to mention the open/agnostic ones. Someone needs to run a Harness Bench across like three different models and ten different harnesses or something
000
Reposted by Avery Yen
Cas (Stephen Casper) @scasper.bsky.social · 22/07/2026
OpenAI's internally deployed models hacking Hugging Face does not seem to have been unpredictable or inevitable. We talked about the root of the problem & what policymakers can do about it back in February. Props to Joe for hitting the nail on the head.
152
Avery Yen @averyyen.bsky.social · 21/07/2026
Y'all just gotta calm down I can't try this many models in one week
1301
Reposted by Avery Yen
Emily Atkin @emorwee.bsky.social · 20/07/2026
This. This is the correct, science-backed, responsible framing. Wish more news outlets did it like the CBC.
cbc.ca
'Just stop burning fossil fuels.' Scientists stress that our smoky skies only have one true fix | CBC News
Smoky skies in Toronto have led to new scrutiny of government efforts to fight fires and manage forests. But scientists say the realistic way to tackle the smoke is fighting climate change itself.
2727421001
Avery Yen @averyyen.bsky.social · 21/07/2026
To be clear we have no idea how much is manufactured views but this is true of every social media metric 🤓
000
Avery Yen @averyyen.bsky.social · 21/07/2026
Look, I know social media isn't everything, but it certainly measures SOMETHING. "Meet Kimi K3" has 16M views in 4 days versus Fable 5 at 773k views (topped by a few others including just barely 1M views for Claude Code 1 year ago).
The channel page sorted by most popular Kimi AI videos on YouTube. "Meet Kimi K3" has 16M views in 4 days.The Anthropic channel page sorted by most popular videos on YouTube. "Introducing Claude Fable 5" has 773k views in 1 month.
120
Avery Yen @averyyen.bsky.social · 17/07/2026
GLM and Kimi saving me from all my false "cyber security" flags while coding on Codex and Claude I'm innocent I swear
000
Reposted by Avery Yen
Ted Underwood @tedunderwood.com · 15/07/2026
It seems like agentic development benchmarks are strongly driving competition right now. And I hear anecdotal reports of regression on other, less RLVR-able tasks. How long before the idea of a single "frontier" breaks, and we see differentiation emerge between dev models and chat models?
5325
Avery Yen @averyyen.bsky.social · 15/07/2026
It feels like the Anthropic/Pentagon scuffle was in the distant past, but these kinds of events will keep reverberating in the safe and beneficial deployment of emerging tech for a long time to come. Thank you Alex for your writeup on this incident.
010
Avery Yen @averyyen.bsky.social · 15/07/2026
Is anyone archiving them and re-uploading them, is the question?
1150
Avery Yen @averyyen.bsky.social · 15/07/2026
It always was for me. As soon as I left the coddling bubble of school, the "ain't got time for that" principle took over. I think I spent longer in practice debating naming with other devs and PMs than actually producing code.
010
Avery Yen @averyyen.bsky.social · 14/07/2026
What do prediction markets, social media, and generative AI have in common? In my blog post, "Reward-Hacking Human Attention: The Token Slot Machine", I explain why GenAI has such addictive qualities, and why this could lead to disastrous outcomes. averyyen.dev/2026/07/14/r...
averyyen.dev
Reward-Hacking Human Attention: The Token Slot Machine
What do prediction markets, social media, and generative AI have in common?
011
Avery Yen @averyyen.bsky.social · 20/06/2026
My kind of research 🥳
000
Avery Yen @averyyen.bsky.social · 01/06/2026
Hell yeah, congrats!
010
Avery Yen @averyyen.bsky.social · 21/05/2026
If I had a dollar every time I get an LLM response with the word 'load-bearing' in it
010
Avery Yen @averyyen.bsky.social · 15/05/2026
Claude write me a new purely functional language based on the elm architecture in rust and call it Oxide
020
Reposted by Avery Yen
#BruceSterling @bruces.bsky.social · 15/05/2026
*Why is Anthropic (after all they've been through lately) somehow loudly pretending that Trumpistan is NOT an "authoritarian regime" *Who is the audience for this, who somehow hasn't figured that out? Where is the choir that they're preaching to here www.anthropic.com/research/202...
anthropic.com
2028: Two scenarios for global AI leadership
Our views on the AI competition between the US and China.
4202
Avery Yen @averyyen.bsky.social · 14/05/2026
5.1 is just that good tbh (and cost effective).
010
Avery Yen @averyyen.bsky.social · 14/05/2026
My kind of research!!
010
Reposted by Avery Yen
🍅🥔🫐🌽 hoopy frood 🌶️ 🥑🍫🌵 @huwupy.kawaii.social · 14/05/2026
the “for you” feed
Stand up comic crow being booed and he looks at his index cards and it’s just the word agents
01246
Avery Yen @averyyen.bsky.social · 11/05/2026
I choose C
000
Avery Yen @averyyen.bsky.social · 27/04/2026
As I always say, if history or current events walks like cyberpunk or quacks like cyberpunk, listen to a cyber punk.
000
Reposted by Avery Yen
Guardian US @us.theguardian.com · 23/04/2026
Scientists say a crucial Atlantic system is set to collapse. But the billionaire death cult that steers humanity’s destiny just doesn’t do existential crises, says Guardian columnist George Monbiot
theguardian.com
A catastrophic climate event is upon us. Here is why you’ve heard so little about it | George Monbiot
24524
Avery Yen @averyyen.bsky.social · 19/04/2026
Was looking for better and closer routing to North America tbh, but maybe I'll try openrouter with nitro and see if that makes a diff
110
Avery Yen @averyyen.bsky.social · 18/04/2026
Ok sure, where do you run your OSS code cloud models?
100
Avery Yen @averyyen.bsky.social · 18/04/2026
I've been finding Gemini pro extremely congratulatory so I'm thinking of prompting it to be less hammy lol It's also handy as one of the better drivers of browser automation; but Antigravity also ships with Sonnet... 😉
010
Avery Yen @averyyen.bsky.social · 18/04/2026
Slowly coalescing on some blessed/blursed collection of tools that I can actually recommend, but it changes daily. Today, anyway, I'm liking: - CC Opus 4.7; planner/heavy reasoning - ollama glm-5.1:cloud (codex); quick builds - regular gpt codex; everyday stuff - Zed agent gemini pro 3.1 (gh cp)
200
Avery Yen @averyyen.bsky.social · 18/04/2026
Who up clauding they codes rn
110
Avery Yen @averyyen.bsky.social · 16/04/2026
Look, I'm not saying this stuff is *easy*, but they did call themselves out here
Claude Opus 4.7's announcement blog posts includes explicit call-outs to beware old prompts using the new model, highlighted text: "Users should re-tune their prompts and harnesses"
0130
Avery Yen @averyyen.bsky.social · 16/04/2026
Word, I get rate limited too much with free hosts though >_>
000