riley @riguh.bsky.social · 05/10/2026Obviously The kiwis might pretend to have invented several significant foods and beverages but it’s all lies Although I’ll swap our claim to any of those things for the bledisloe cup 010
riley @riguh.bsky.social · 05/10/2026We took Italian methods and figured out how to make it consistently good 100
riley @riguh.bsky.social · 05/10/2026Yes, largely automating that because it’s easy to fix and then test - no user visible regression? Ship it Meanwhile I can spend my time on the things that actually push it forward (which sometimes create tech debt… which gets automatically mopped up) 030
riley @riguh.bsky.social · 05/10/2026Just read 100 Years of Solitude. All the characters are called Aureliano. You don’t really have remember names that way 020
riley @riguh.bsky.social · 05/10/2026Too busy building agents to help me with my music choices to actually listen to music 110
riley @riguh.bsky.social · 05/10/2026Fun side-project. Will build it out to do a bit more, within its constraints (it has to be lethal-trifecta safe, and the model is really not very smart) 000
riley @riguh.bsky.social · 05/10/2026Added an agent to my collection which uses the Apple foundation model on my Mini. It’s basic but it’ll answer something like “what did I play last?” by looking up my Spotify history in 1.7 seconds - faster than Siri while using Apple’s own model on device (with very little setup or resource use). 230
riley @riguh.bsky.social · 04/10/2026I think he nails it - that perspective is what I still like to pull Fable in as an advisor for while doing the bulk of the work in Opus. 020
riley @riguh.bsky.social · 04/10/2026Steve Yegge on x: while Opus 5.5's precision rivals Fable 5.1's across all tasks in my system, its recall is not as good. For any given specific task they're about equal. But when the task is more open-ended, Fable correctly spots and acts on important issues more often. It shows better perspective. 120
riley @riguh.bsky.social · 03/10/2026I haven’t answered my phone to unknown numbers in about five years - the rare exception is when this feature surfaces an actual non-spam calling human 110
riley @riguh.bsky.social · 03/10/2026No, it runs using its own account and I call it from my account which I’m singed into - anything it does goes through a firewall, is via tools with fixed permissions, gets sanitised, etc etc Basically I apply @simonwillison.net’s lethal trifecta lens to everything I add to it 220
riley @riguh.bsky.social · 03/10/2026I gave my agent its own user account so it can’t read my files by default. It would be nice not to have to go to that much effort 120
riley @riguh.bsky.social · 01/10/2026I downloaded all my Spotify history and pull my daily listens into a db and have Claude use that to help recommend music for me. Has pulled out so many bangers. I’ve been music obsessed forever and so it’s not like I haven’t got human recommendations but this is next level 041
riley @riguh.bsky.social · 01/10/2026Guava DNA code + Suno is the stupidest thing I have ever heard Why have I played it ten times so far 2250
riley @riguh.bsky.social · 01/10/2026“Are you listening to me? Get downstairs right now and get your shoes on. More response and less abstention please, Grace.” 060
riley @riguh.bsky.social · 01/10/2026It is nice that we might hear from them again soon though, it’s rare 160
riley @riguh.bsky.social · 01/10/2026Which for someone ADHD-like is really annoying because I’ll often drop extra thoughts or tasks in and expect it to keep track. This may be a Claude Code change rather than a model change, though 150
riley @riguh.bsky.social · 01/10/2026Yeah but unlike 5 it doesn’t always circle back unless you ask 130
riley @riguh.bsky.social · 01/10/2026I’ve got some plans for the weekend They’re very big I’m excited about them You should be excited too although I’m definitely not going to give you any idea of what they are 2110
riley @riguh.bsky.social · 30/09/2026Slow didn’t matter - background job I think actually did run the 3.6 first in an early run - I’ll see if it was salvageable 001
riley @riguh.bsky.social · 29/09/2026Jev won’t take my money Anthropic having an incident Fuck this I’m going to bed 170
riley @riguh.bsky.social · 28/09/2026I felt a great disturbance in the Force. As if millions of voices suddenly cried out in terror and were suddenly logged as high priority AWS support cases. 010
riley @riguh.bsky.social · 28/09/2026Sonnet 5.5 is out! But if you’re on Amazon, only if you’re prepared to let them run it anywhere in the world. Not awesome for the many companies that enforce data sovereignty constraints. docs.aws.amazon.com/bedrock/late...docs.aws.amazon.comAnthropic - Amazon BedrockThe following Anthropic models are available in Amazon Bedrock: 141
riley @riguh.bsky.social · 28/09/2026I used Apple's MLX directly: mlx_lm.server from the mlx-lm package (installed with Homebrew), with MLX 4-bit weights from Hugging Face, called through its OpenAI-compatible API 010
riley @riguh.bsky.social · 28/09/2026I reckon with a bit of effort I could probably get 6B qwen working okay. The 4B worked fine. It was super easy to set up with Claude and MLX 110
riley @riguh.bsky.social · 28/09/2026I didn’t set it, the largest prompt in the eval was only 8.3k tokens so I didn’t get close to pushing anything really 110
riley @riguh.bsky.social · 28/09/2026Weirdly yes Prompt cache was off for 6B as I blew up memory with it on. Maybe that was why but for what I was evaluating it didn’t expect it to. Gonna dig in to that one 010
riley @riguh.bsky.social · 28/09/2026It was annoyingly slow, I’ll have to speed things up. Maybe Jev can help with judging or something. Promptfoo was pretty easy to set up, though I don’t love the amount of dependencies it pulls in 010
riley @riguh.bsky.social · 28/09/2026More bits didn't help, unexpectedly. The 6-bit Qwen only fitted with the prompt cache turned off, and then it was slower (57 s) and passed 3 of 18 against 8 of 18 for the 4-bit. I couldn’t find a DeepSeek model that fit. Didn’t try GPT OSS. 110
riley @riguh.bsky.social · 28/09/2026Local Gemma nearly matches Sonnet, but left the mini with about 2% free memory + took 52 s vs Sonnet's 24 s. Ok for a daily job. Reworked the prompt a bit but not a lot yet. Prompt mattered more than the model, now I’ve got results I’ll tune it and the evals. 210
riley @riguh.bsky.social · 28/09/2026Pass rates on the best prompt: Claude Sonnet: 56% Gemma 4 31B (4-bit): 53% Qwen 3.8 27B (4-bit): 44% Claude Haiku: 28% Qwen 3.8 27B (6-bit): 17% Granite 4.2 30B: 0% GLM-4.7 Flash: 0% MiMo 9B: 0% 220
riley @riguh.bsky.social · 28/09/2026I've been testing local LLMs on a 32 GB M6 Mac mini (MLX) against Claude on a real job: writing a daily listening note from my music history. 6 test cases × 3 samples each. Pretty strict evals; the note passes only if it clears code checks and an LLM judge. Results in thread below. 260