Shiminsky @shiminsky.bsky.social · 28/09/2026Catching up on AI links post trip, I have this nagging sense that I'm in an Adam Curtis documentary: 'By the mid 2020s, frontier AI labs intensified their false sense of control over their creations'. 010
Reposted by ShiminskyTom McCoy @rtommccoy.bsky.social · 01/09/2026🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n 431888
Shiminsky @shiminsky.bsky.social · 24/08/2026The half-life of harness adjacent tools has never been shorter. It’s hard to envision what a moat even look like in this day and age. 030
Shiminsky @shiminsky.bsky.social · 31/07/2026You should do a whole video to unpack this for people, seriously. 010
Shiminsky @shiminsky.bsky.social · 28/07/2026personally I would like the new housing to be much denser, but at least they are coming online monthly. 011
Shiminsky @shiminsky.bsky.social · 25/07/2026100%, we were looking at ruhan-wang.github.io/Harness-Hand... on the podcast this week and I can see something like this being the communication and understanding channel for codebases.ruhan-wang.github.ioHarness Handbook — Making Agent Harnesses Understandable, Auditable & EditableA behavior-level manual for complex agent harnesses that helps developers and coding agents understand, audit, and modify system behavior with grounded code evidence. 020
Shiminsky @shiminsky.bsky.social · 30/06/2026Public Domain Review, where have you been my entire life? Thank you for sharing @davatron5000.bsky.social! 040
Shiminsky @shiminsky.bsky.social · 23/06/2026Preordered the paperback, psyched for you and my book notes journal. 040
Shiminsky @shiminsky.bsky.social · 23/06/2026Maybe we can all agree on less sloppy looking AI slop. 000
Shiminsky @shiminsky.bsky.social · 23/06/2026It was fun applying creativity techniques with AI generated front-end designs: shimin.io/journal/inha... Also got a chance to present this at AI Tinkerers Seattle earlier this month, give the skill a try if you are sick of projects looking like AI Slop.shimin.ioInhabited Design, a Skill for the AI Slop Site Problem — Shimin ZhangPurple gradients are the new em dash. Building inhabited-design — a Claude Code skill that samples a different designer to inhabit on every run — and the detours through attractors, personas, self-ref... 100
Shiminsky @shiminsky.bsky.social · 22/06/2026Insightful post from @terriblesoftware.org, and something I've been thinking about: there's an inherent asymmetry with AI generated content where cost of reading > cost of review. We should start measuring performance by fewest LOC generated. 010
Shiminsky @shiminsky.bsky.social · 11/06/2026Here's my prompt that just got flagged: "what do you know about the theory of creativity? do research if you have to, give me a full list of theories and hypothesis" and "what are some of the most impactful recent studies?" It's like they actually don't want folks to use the thing!! 000
Shiminsky @shiminsky.bsky.social · 10/06/2026like, we can just do this now? and it'll just teach me all the physics and engineering I need to validate the idea? 000
Shiminsky @shiminsky.bsky.social · 10/06/2026Was working with Fable to last night to design a chakram / frisbee inspired flying machine. The physics seems to checkout -- but Im no aerospace engineer. 100
Shiminsky @shiminsky.bsky.social · 10/06/2026Can't tell if we are all cooked, or this is the start of something beautiful. 220
Shiminsky @shiminsky.bsky.social · 16/05/2026Nibbling on the same thought today, future code education will be on a new level of abstraction and we haven’t figured out what that looks like yet. Very excited to sign up for @danabra.mov’s new course when it’s out! 080
Shiminsky @shiminsky.bsky.social · 13/05/2026The herd worships Mewtwo. The herd adores Pikachu. The herd cheers for Charizard. And in this worship, the herd reveals its poverty of spirit — for it can only love what is already powerful, already celebrated, already *safe*. 000
Shiminsky @shiminsky.bsky.social · 13/05/2026Setting up a team of AI philosophers to debate whether Magikarp is the best Pokemon, average Tuesday night fun stuff. 100
Reposted by ShiminskyAi2 @ai2.bsky.social · 08/05/2026EMO’s expert clusters look very different from a traditional MoE: they organize around semantic domains like health, news, politics, & film/music. Traditional MoEs often cluster around surface patterns like prepositions and articles, making selective expert use tougher. 1694
Shiminsky @shiminsky.bsky.social · 06/05/2026Just tried it today on a front-end design review. It works well, especially when both sub agents flagged with the same fundamental UI flaw -- guess I now owe them a cake each. 020
Reposted by ShiminskyJesse Vincent @s.ly · 01/05/2026A short post about my favorite adversarial review prompt: blog.fsck.com/2026/05/01/a...blog.fsck.comMy favorite adversarial review promptI'm Jesse. I make stuff. Software, hardware. Very occasionally, trouble. 281
Shiminsky @shiminsky.bsky.social · 04/05/2026yeah, probably just my old accounting brain acting anxious again. 090
Shiminsky @shiminsky.bsky.social · 04/05/2026Last thing on this because it really struck a nerve. This banker line: “This is a hobby, not a business. How long do you want to pay for the privilege of milking cows?” -- this may be due to IRS distinction on hobby vs business expense (read tax deductions). Esp since ag isn't sole income. 2250
Shiminsky @shiminsky.bsky.social · 04/05/2026my favorite line from the article: 'The farm had survived thanks to some monthly income from a cellphone tower and the discovery of natural gas under his land' talk about burying the lede. 1480
Shiminsky @shiminsky.bsky.social · 04/05/2026It’s about time we stop infantilizing farmers, they are just like you and me, acting according to their self interests. (I’m just a gardener but married into a farming family) 19214
Shiminsky @shiminsky.bsky.social · 24/04/2026Think of it as a "delusion index', the delta between a model's self judgement and it's peers score. Weaker models tend to have more delusion since they are unable to differentiate complex work. Gemini is the real outlier here -- Opus and GLM 5.1 also had surprisingly 'objective' self scores. 000
Shiminsky @shiminsky.bsky.social · 24/04/2026Biggest shocker: Gemini 3.1 pro preview was dead last -- worse than Gemini 2.5 Flash (the control model). I thought it was broken until I read the response, mostly bland sci-fi tropes. But it knew its output is bad! It rated it's own work just as bad as other models. 100
Shiminsky @shiminsky.bsky.social · 24/04/2026Opus 4.7 tops the chart (more than I expected), and Kimi K2.5 is the strongest of the open weight model at this task. Opus 4.7 find Kimis output to be "Speculative but compelling" with "inventive neologisms", has "literary quality" and "Omega Cool" 100
Shiminsky @shiminsky.bsky.social · 24/04/2026Are you tired of keeping track of 4 different variants SWE Benchmarks? I am. So I did something kinda dumb -- but at least it was fun -- asking 11 frontier models to blind grade each other's open ended prediction about AI's future. Post at: shimin.io/journal/what... Findings below (given n=1)shimin.ioWhat I learned asking 11 AI models to grade each other's AI predictions — Shimin ZhangAn experiment on model personalities, a delusion index, and the open-weight dark horse contender I didn't see coming. 110
Reposted by ShiminskyWilliam B. Fuckley @opinionhaver.bsky.social · 22/04/2026here's a mildly provocative take: every valid LLM worry about societal effects is reducible to some approximation of 'it makes what were once costly things trivially easy to do' 3030033
Shiminsky @shiminsky.bsky.social · 21/04/2026What major issues are you seeing? I’m catching enough silly hallucinations that I’m worried about the stuff I’m not catching… 000
Shiminsky @shiminsky.bsky.social · 21/04/2026Go max or go back to using a powerful local model! That said, even at max I find 4.7's hallucination to be 'sillier' than 4.6. When it's not making up random carp though, the ceiling on 4.7 is very high. 020
Shiminsky @shiminsky.bsky.social · 21/04/2026Kimi is especially strong on my other AI on AI Benchmark (that I need to do a write up about). IMO it's the best deal in town. 010
Shiminsky @shiminsky.bsky.social · 21/04/2026The Latest Pelican Bicycle Benchmark result from @simonwillison.net for Opus 4.7 was so shocking that I had to do some follow up experiments. It turns out Opus 4.7 is ....just kinda lazy?? It uses almost no reasoning tokens compare to Qwen, and 40x less than Opus 4.6 shimin.io/journal/clau...shimin.ioClaude 4.7 isn't dumb, it's just lazy — Shimin ZhangSome follow up experiments with Claude 4.7 based on Simon Willison's Pelican Benchmark Shocker. 2473
Shiminsky @shiminsky.bsky.social · 21/04/2026And of course, if you really miss the latest Opus models, you can always try politely asking Pi to invoke CC. 000
Shiminsky @shiminsky.bsky.social · 21/04/2026I'd been missing my Pi Agent harness since the Anthropic subscription crack down. Tonight I finally got it back with a local Qwen 3.6 setup locally and it feel great to be reunited with my favorite harness! Thank you @mariozechner.at @mitsuhiko.at for your work on Pi! 130
Shiminsky @shiminsky.bsky.social · 08/04/2026@addyosmani.bsky.social's multi agent experience agrees with my own, going from 2 to 5 parallel agents is a quick way to transform from a judge to a ticket usher.addyosmani.comYour parallel Agent limitRunning multiple agents in parallel is not just a question of throughput. It is a new kind of cognitive labor that requires managing multiple mental models, ... 021
Shiminsky @shiminsky.bsky.social · 08/04/2026Another banger from @anildash.com : Actually, people love to work hard anildash.com/2026/04/06/p...anildash.comActually, people love to work hard - Anil DashA blog about making culture. Since 1999. 0249
Shiminsky @shiminsky.bsky.social · 13/03/2026I've been working on a slice of the problem as well, resisting AI Sycophancy for everyday users. www.flatterproof.meflatterproof.meFlatterProofGamified training to recognize AI sycophancy 020
Shiminsky @shiminsky.bsky.social · 09/03/2026Reminder when the only fatigue we felt was about too many new JS frameworks? Those were the days. 010
Shiminsky @shiminsky.bsky.social · 09/03/2026Great post from @antirez.bsky.social tracing through the history of the open source movement and positions 'clean room clones' as an extension and not revolution. 010
Shiminsky @shiminsky.bsky.social · 08/03/2026AI sycophancy will reap many promising careers in next next few years. 000
Shiminsky @shiminsky.bsky.social · 08/03/2026In my experience the only major downside is it uses open router and gets expensive very quickly. So I am still querying each model individually, like a peasant. 020
Shiminsky @shiminsky.bsky.social · 08/03/2026Sounds similar to Karpathy's LLM Council github.com/karpathy/llm... but with more snark?github.comGitHub - karpathy/llm-council: LLM Council works together to answer your hardest questionsLLM Council works together to answer your hardest questions - karpathy/llm-council 140