Sign in

amos

@fasterthanli.me
17K followers 554 following 6.5K posts

hi, I'm amos! 🍃 they/them 🔮 "most level-headed AI user" 🫐 working on something dataflow-shaped 🦀 known for teaching rust and not much else 📚 fasterthanli.me 📺 youtube.com/@fasterthanlime

PostsRepliesMedia
amos @fasterthanli.me · 48m
brb launching a plan called "shmuck 200", billed at 229EUR/mo
050
amos @fasterthanli.me · 1h
brb listing my niece's lemonade stand on european-alternatives dot eu
190
Reposted by amos
RustNL @rustnl.bsky.social · 5h
The Call for Proposals for RustWeek 2027 is now open! If you’d like to give a talk, go here: 2027.rustweek.org/blog/2026-10... The CFP closes Jan 10, 2027 #rustweek2027 #rustlang
0147
amos @fasterthanli.me · 6h
Surely the EU isn't going to accept the same terms twice right?
3200
amos @fasterthanli.me · 8h
oooh you don't wanna do that, then it's like "given the deadline we better cut the scope by 90%"
160
amos @fasterthanli.me · 9h
opus5.5: a great middle ground would be X... but that'd be days of work me: days? opus5.5: *groan* ughh fine. hours. I'll get started...
4571
amos @fasterthanli.me · 11h
Truth bender 😌
030
amos @fasterthanli.me · 11h
This is about Bo right? I know nothing about UK politics but it sounds right?
0170
amos @fasterthanli.me · 11h
You're welcome jevbot
0110
amos @fasterthanli.me · 12h
*taps roof of closed model* you can fit so many calculators in there
21045
Reposted by amos
mr. TIM @timkellogg.me · 22h
Gemini 4 Argon Same price as Sol 6.1, rolling out to selected cyberdefenders blog.google/innovation-a...
A benchmark comparison table evaluating four AI models—Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5—across eight performance categories:
 * Knowledge work: Vals Index (Gemini 4 Argon: 68.9%, GPT-6 Astra: 63.1%, Claude Fable 5.1: 65.8%, Claude Opus 5.5: 67.0%), AutomationBench (51.3%, 41.4%, 31.4%, 42.5%), Vals Finance Agent v2 (65.4%, 53.5%, 58.9%, 58.6%), and Harvey's Legal Agent Benchmark (19.6%, 5.4%, 6.7%, 3.8%).
 * Agentic coding: DeepSWE v1.1 (77.9%, 74.1%, 67.4%, 74.2%), FrontierSWE v2 (55.0%, 65.5%, 56.3%, 62.3%), Vibe Code Bench (91.9%, 89.6%, 90.3%, 90.3%), and Terminal-bench 4.0 (57.4%, 58.2%, 57.9%, 66.4%).
 * ML engineering: PostTrainBench (45.3%, 44.3%, 40.2%, 49.3%).
 * Science and math: Terminal-Bench Science 0.1 (57.6%, 68.1%, 52.6%, 63.3%), LABBench 2 (88.8%, 85.4%, 68.6%, 73.1%), and RiemannBench (76.0%, 72.0%, 65.6%, 69.6%).
 * Long context: GraphWalks Up to 128k (99.7%, 98.7%, 91.4%, 90.6%) and GraphWalks 256k to 1M (84.2%, 71.8%, 65.0%, 66.8%).
 * Computer use: Agent's Last Exam (39.5%, 34.2%, —, 38.2%) and OSWorld-2.0 (69.2%, 72.6%, —, —).
 * Multimodal understanding: Chartography (71.6%, 71.0%, 46.2%, 66.3%) and LVBench (91.7%, 87.5%, 79.7%, 83.7%).
 * Cybersecurity: CWE-bench v1 (68.0%, 68.0%, 58.0%, 67.0%).
The column for Gemini 4 Argon is highlighted in blue. The footer lists the methodology link: deepmind.google/models/evals-methodology/gemini-4-argon.
6452
amos @fasterthanli.me · 21h
Finally someone gets it!
030
amos @fasterthanli.me · 30/09/2026
nice!
000
amos @fasterthanli.me · 30/09/2026
pro-tip: if you're into keyholders and findom, do NOT get into KMS, it's not AT ALL the same thing
2180
amos @fasterthanli.me · 30/09/2026
cf. Opus 5.5: "The hyperscalers spend more on confidential computing than most EU clouds earn." 😭
2200
amos @fasterthanli.me · 30/09/2026
finding that the answer to "what can you build today that relies exclusively on EU cloud services, inference, CDN, DNS, etc." is "quite a bit, but it's not gonna be comfortable"
3433
amos @fasterthanli.me · 30/09/2026
Yeah that was immediately obvious from them offering Ultrafast on Bedrock, right?
040
amos @fasterthanli.me · 30/09/2026
who's coming to @eurorust.eu btw? ping me
861
amos @fasterthanli.me · 29/09/2026
oh yeah?
110
amos @fasterthanli.me · 29/09/2026
he seemed distracted, stumbled on words a bunch
120
amos @fasterthanli.me · 29/09/2026
do you think ultrafast also compacts 8x faster
060
amos @fasterthanli.me · 29/09/2026
right?!?
000
amos @fasterthanli.me · 29/09/2026
the what
000
amos @fasterthanli.me · 29/09/2026
what did you think? (tl;dr if you didn't watch the stream: simonwillison.net/2026/Sep/29/... )
simonwillison.net
OpenAI DevDay 2026 live blog
I’m at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I’ll be live blogging the keynote and some other notes during the day. OpenAI gave me …
100
amos @fasterthanli.me · 29/09/2026
maaybe ultrafast is interesting? maybe. idk. so much fluff around it though. and ai-generated powerpoint to present to the coworkers you fired (replaced with dots) okay. whatever helps them IPO next year ig
110
amos @fasterthanli.me · 29/09/2026
cryptic post I know but just... lots of meh. really disappointed. "our biggest devday ever" okay sure.
100
amos @fasterthanli.me · 29/09/2026
whole devday
130
amos @fasterthanli.me · 29/09/2026
👎
540
amos @fasterthanli.me · 29/09/2026
Jev is on openrouter btw?
020
Reposted by amos
Simon Willison @simonwillison.net · 29/09/2026
I'm at OpenAI's DevDay event in San Francisco today - as I have for the past three DevDay events, I'm running a live blog where I'll be posting updates during the keynote, which starts in five minutes simonwillison.net/2026/Sep/29/...
simonwillison.net
OpenAI DevDay 2026 live blog
I’m at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I’ll be live blogging the keynote and some other notes during the day. OpenAI gave me …
59311
amos @fasterthanli.me · 29/09/2026
rust devs be: "use fluent query builders!" and then do: .filter("id", "=", 1)
2340
amos @fasterthanli.me · 29/09/2026
🧠
000
amos @fasterthanli.me · 29/09/2026
Vincent Adultman levels of business going on here
Screenshot of a post by Tibo (@thsottiaux):

“I’ll explain the new Pro 200 plan differently, before I start live tweeting from DevDay on things that are going out!

Today we are going to ship a number of things that increase what you can do across the Plus and Pro plans. A lot of compute is online for this increase. As we increase the floor, we are changing the relative difference between plans to be

Plus = 1X
Pro 100 = 5X
Pro 200 = 10X

and we are reopening subscriptions for Pro 200 (we had paused it). If you have an existing plan you will keep the 20X multiplier for a bit and also receive a lot of additional credits because we know changes are hard even if it means that everyone will get more in the end.”

The phrase “for a bit and also receive a lot of additional credits” is highlighted in red.
0110
amos @fasterthanli.me · 29/09/2026
the new Inside Out looks fire
openai announcement it has fluffy mascots I already fucking hate it
130
amos @fasterthanli.me · 29/09/2026
the year in review:
0333
amos @fasterthanli.me · 29/09/2026
pace the TTFB!!
030
amos @fasterthanli.me · 29/09/2026
> underqualified to work as a cop ouch
011
amos @fasterthanli.me · 29/09/2026
Neat! Does this mark a change in your perception of those models or just a milestone in what you can achieve locally?
100
amos @fasterthanli.me · 29/09/2026
not yet! it got good recently.
110
amos @fasterthanli.me · 29/09/2026
but you get animated GIFs on your desktop for free!!?!?!?
040
amos @fasterthanli.me · 29/09/2026
bsky.app/profile/thso... looks like: still $200 but usage runs out twice as fast.
250
amos @fasterthanli.me · 29/09/2026
But idc about your personal assistant though I care about tokens... "we won't win with the best models" ok so we should look elsewhere then
180
amos @fasterthanli.me · 29/09/2026
static.klipy.com
The Office Dwight Schrute
ALT: The Office Dwight Schrute
010
amos @fasterthanli.me · 28/09/2026
Very well.
130
amos @fasterthanli.me · 28/09/2026
shame though b/c mine is reaaaally good
3100
amos @fasterthanli.me · 28/09/2026
thinking of letting others use my AI agent harness in exchange for AMERICAN DOLLARS but the problem in that otherwise foolproof plan is that anyone who's agent-pilled enough to look into obscure harnesses has already made theirs, most likely
7280
Reposted by amos
mr. TIM @timkellogg.me · 28/09/2026
Claude 5.5 Sonnet is live and it’s roughly Opus 5.5 but cheaper www.anthropic.com/claude-sonne...
A benchmark comparison table titled "Claude Sonnet 5.5" comparing four models: Sonnet 5.5, Sonnet 5, Opus 5.5, and GPT-6 Sol.
 * Agentic coding (Terminal-Bench 4.0): Sonnet 5.5: 70.6%; Sonnet 5: 10.3%; Opus 5.5: 66.4%; GPT-6 Sol: N/A.
 * Agentic coding (FrontierCode 1.1 Main): Sonnet 5.5: 46.2% (Max) / 52.1% (Xhigh); Sonnet 5: 42.4%; Opus 5.5: 54.4%; GPT-6 Sol: 49.3%.
 * Agentic coding (CursorBench 4.0): Sonnet 5.5: 55.5%; Sonnet 5: 34.1%; Opus 5.5: 57.8%; GPT-6 Sol: N/A.
 * Knowledge work (GDPval-AA v2.1): Sonnet 5.5: 1844; Sonnet 5: 1449; Opus 5.5: 1846; GPT-6 Sol: 1487.
 * Knowledge work (AA-Briefcase v1.1): Sonnet 5.5: 1811; Sonnet 5: 1359; Opus 5.5: 1822; GPT-6 Sol: 1483.
 * Multidisciplinary reasoning (Humanity's Last Exam with tools): Sonnet 5.5: 64.5%; Sonnet 5: 54.9%; Opus 5.5: 67.7%; GPT-6 Sol: N/A.
 * Computer use (OSWorld 2.1 partial): Sonnet 5.5: 80.1%; Sonnet 5: 57.0%; Opus 5.5: 81.8%; GPT-6 Sol: N/A.
 * Visual chart recognition (Chartography no tools): Sonnet 5.5: 61.6%; Sonnet 5: 15.6%; Opus 5.5: 64.4%; GPT-6 Sol: 53.6%.
Footnotes provide methodology notes regarding evaluation settings, Artificial Analysis pre-release testing details, and recent bug fixes affecting GPT-6 Sol benchmark scores.
A line graph titled "Agentic coding by effort level" on the CursorBench 4.0 benchmark, plotting Score (%) on the linear y-axis (20% to 60%) against Cost per task in USD on a logarithmic x-axis ($0.5 to $10+).
The chart compares four models:
 * Sonnet 5.5 (blue line with labeled effort levels): Starts at "Low" (~$0.50, 36%), moving through "Med" ($0.70, 39%), "High" ($1.70, 48%), "Xhigh" ($3.80, 53%), and "Max" ($9.50, ~55.5%).
 * Opus 5.5 (orange line): Tracks closely above Sonnet 5.5 at higher cost points, spanning from ~$1.20 per task (~44%) up to ~$12.50 per task (~58%).
 * Sonnet 5 (grey line): Shows lower accuracy relative to cost, ranging from ~$0.90 per task (~25%) to ~$8.00 per task (~42%).
 * GPT-5.6 Sol (light green line): Represents the lowest trajectory, ranging from ~$1.40 per task (~24%) to ~$7.00 per task (~34%).
Footnote: "CursorBench 4.0 does not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here."
2020623
amos @fasterthanli.me · 28/09/2026
yey
110
amos @fasterthanli.me · 28/09/2026
mergers and submission??
000
amos @fasterthanli.me · 28/09/2026
can relate, first to get one has to share!!!
060