Sign in

Shiminsky

@shiminsky.bsky.social
151 followers 251 following 53 posts

Software Developer, Dog Dad, Beginner Gardener, Host of the Artificial Developer Intelligence (ADI) Podcast www.adipod.ai -- em-dashes my own

PostsRepliesMedia
Shiminsky @shiminsky.bsky.social · 28/09/2026
Catching up on AI links post trip, I have this nagging sense that I'm in an Adam Curtis documentary: 'By the mid 2020s, frontier AI labs intensified their false sense of control over their creations'.
010
Shiminsky @shiminsky.bsky.social · 13/09/2026
What is the definition of “know” ?
111
Reposted by Shiminsky
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
Overview of the paper. 
Title: The Emergent Symbolic Structure of Artificial Neural Networks
Authors: Tom McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Left: Neural networks encode information in vectors (there is then an image of a vector), yet they excel at tasks long thought to require symbolic structure (there is then an image of a symbolic representation, specifically a syntax tree). How do LLMs do it?
Right: We find that LLM representations can be closely approximated with symbolic structures. This approximation lets us edit the structure of an LLM’s output by editing the structure of its internal representations, as shown. There is then an image of two edits to LLMs. In the first one, the original input is 3 + 6 * 8, with an answer of 51. But if we swap the positions of the 3 and the 6, the output becomes 30. In the second one, the original input is a Python command repeating the list [Z, U] three times, producing [Z, U, Z, U, Z, U]. But if we edit the input in a way that adds a Q at the end of the input, the output becomes [Z, U, Q, Z, U, Q, Z, U, Q].
431888
Shiminsky @shiminsky.bsky.social · 24/08/2026
The half-life of harness adjacent tools has never been shorter. It’s hard to envision what a moat even look like in this day and age.
030
Shiminsky @shiminsky.bsky.social · 05/08/2026
Tomato garden finally coming online.
A bowl of tomatoes of varying size
000
Shiminsky @shiminsky.bsky.social · 31/07/2026
You should do a whole video to unpack this for people, seriously.
010
Shiminsky @shiminsky.bsky.social · 28/07/2026
personally I would like the new housing to be much denser, but at least they are coming online monthly.
011
Shiminsky @shiminsky.bsky.social · 28/07/2026
Bremerton is a great spot!
120
Shiminsky @shiminsky.bsky.social · 25/07/2026
100%, we were looking at ruhan-wang.github.io/Harness-Hand... on the podcast this week and I can see something like this being the communication and understanding channel for codebases.
ruhan-wang.github.io
Harness Handbook — Making Agent Harnesses Understandable, Auditable & Editable
A behavior-level manual for complex agent harnesses that helps developers and coding agents understand, audit, and modify system behavior with grounded code evidence.
020
Shiminsky @shiminsky.bsky.social · 30/06/2026
Public Domain Review, where have you been my entire life? Thank you for sharing @davatron5000.bsky.social!
040
Shiminsky @shiminsky.bsky.social · 23/06/2026
Preordered the paperback, psyched for you and my book notes journal.
040
Shiminsky @shiminsky.bsky.social · 23/06/2026
Maybe we can all agree on less sloppy looking AI slop.
000
Shiminsky @shiminsky.bsky.social · 23/06/2026
It was fun applying creativity techniques with AI generated front-end designs: shimin.io/journal/inha... Also got a chance to present this at AI Tinkerers Seattle earlier this month, give the skill a try if you are sick of projects looking like AI Slop.
shimin.io
Inhabited Design, a Skill for the AI Slop Site Problem — Shimin Zhang
Purple gradients are the new em dash. Building inhabited-design — a Claude Code skill that samples a different designer to inhabit on every run — and the detours through attractors, personas, self-ref...
100
Shiminsky @shiminsky.bsky.social · 22/06/2026
Insightful post from @terriblesoftware.org, and something I've been thinking about: there's an inherent asymmetry with AI generated content where cost of reading > cost of review. We should start measuring performance by fewest LOC generated.
010
Shiminsky @shiminsky.bsky.social · 11/06/2026
Here's my prompt that just got flagged: "what do you know about the theory of creativity? do research if you have to, give me a full list of theories and hypothesis" and "what are some of the most impactful recent studies?" It's like they actually don't want folks to use the thing!!
000
Shiminsky @shiminsky.bsky.social · 10/06/2026
me this morning:
A confused Snowy
010
Shiminsky @shiminsky.bsky.social · 10/06/2026
like, we can just do this now? and it'll just teach me all the physics and engineering I need to validate the idea?
000
Shiminsky @shiminsky.bsky.social · 10/06/2026
Was working with Fable to last night to design a chakram / frisbee inspired flying machine. The physics seems to checkout -- but Im no aerospace engineer.
A 3d rendering of a chakram-based flying machine
100
Shiminsky @shiminsky.bsky.social · 10/06/2026
Can't tell if we are all cooked, or this is the start of something beautiful.
220
Shiminsky @shiminsky.bsky.social · 16/05/2026
Nibbling on the same thought today, future code education will be on a new level of abstraction and we haven’t figured out what that looks like yet. Very excited to sign up for @danabra.mov’s new course when it’s out!
080
Shiminsky @shiminsky.bsky.social · 15/05/2026
would love to see a recipe of how to do this!
070
Reposted by Shiminsky
Rude Law Dog @esghound.com · 13/05/2026
This is A+ AI slop. I love it
2554397
Shiminsky @shiminsky.bsky.social · 13/05/2026
The herd worships Mewtwo. The herd adores Pikachu. The herd cheers for Charizard. And in this worship, the herd reveals its poverty of spirit — for it can only love what is already powerful, already celebrated, already *safe*.
000
Shiminsky @shiminsky.bsky.social · 13/05/2026
Setting up a team of AI philosophers to debate whether Magikarp is the best Pokemon, average Tuesday night fun stuff.
100
Reposted by Shiminsky
Ai2 @ai2.bsky.social · 08/05/2026
EMO’s expert clusters look very different from a traditional MoE: they organize around semantic domains like health, news, politics, & film/music. Traditional MoEs often cluster around surface patterns like prepositions and articles, making selective expert use tougher.
1694
Shiminsky @shiminsky.bsky.social · 06/05/2026
Just tried it today on a front-end design review. It works well, especially when both sub agents flagged with the same fundamental UI flaw -- guess I now owe them a cake each.
020
Reposted by Shiminsky
Jesse Vincent @s.ly · 01/05/2026
A short post about my favorite adversarial review prompt: blog.fsck.com/2026/05/01/a...
blog.fsck.com
My favorite adversarial review prompt
I'm Jesse. I make stuff. Software, hardware. Very occasionally, trouble.
281
Shiminsky @shiminsky.bsky.social · 04/05/2026
yeah, probably just my old accounting brain acting anxious again.
090
Shiminsky @shiminsky.bsky.social · 04/05/2026
Last thing on this because it really struck a nerve. This banker line: “This is a hobby, not a business. How long do you want to pay for the privilege of milking cows?” -- this may be due to IRS distinction on hobby vs business expense (read tax deductions). Esp since ag isn't sole income.
2250
Shiminsky @shiminsky.bsky.social · 04/05/2026
my favorite line from the article: 'The farm had survived thanks to some monthly income from a cellphone tower and the discovery of natural gas under his land' talk about burying the lede.
1480
Shiminsky @shiminsky.bsky.social · 04/05/2026
It’s about time we stop infantilizing farmers, they are just like you and me, acting according to their self interests. (I’m just a gardener but married into a farming family)
19214
Shiminsky @shiminsky.bsky.social · 24/04/2026
Think of it as a "delusion index', the delta between a model's self judgement and it's peers score. Weaker models tend to have more delusion since they are unable to differentiate complex work. Gemini is the real outlier here -- Opus and GLM 5.1 also had surprisingly 'objective' self scores.
A bar chart showing the difference between a model's score of its own output versus the average peer models' scores
000
Shiminsky @shiminsky.bsky.social · 24/04/2026
Biggest shocker: Gemini 3.1 pro preview was dead last -- worse than Gemini 2.5 Flash (the control model). I thought it was broken until I read the response, mostly bland sci-fi tropes. But it knew its output is bad! It rated it's own work just as bad as other models.
100
Shiminsky @shiminsky.bsky.social · 24/04/2026
Opus 4.7 tops the chart (more than I expected), and Kimi K2.5 is the strongest of the open weight model at this task. Opus 4.7 find Kimis output to be "Speculative but compelling" with "inventive neologisms", has "literary quality" and "Omega Cool"
100
Shiminsky @shiminsky.bsky.social · 24/04/2026
Are you tired of keeping track of 4 different variants SWE Benchmarks? I am. So I did something kinda dumb -- but at least it was fun -- asking 11 frontier models to blind grade each other's open ended prediction about AI's future. Post at: shimin.io/journal/what... Findings below (given n=1)
shimin.io
What I learned asking 11 AI models to grade each other's AI predictions — Shimin Zhang
An experiment on model personalities, a delusion index, and the open-weight dark horse contender I didn't see coming.
110
Reposted by Shiminsky
William B. Fuckley @opinionhaver.bsky.social · 22/04/2026
here's a mildly provocative take: every valid LLM worry about societal effects is reducible to some approximation of 'it makes what were once costly things trivially easy to do'
3030033
Shiminsky @shiminsky.bsky.social · 21/04/2026
What major issues are you seeing? I’m catching enough silly hallucinations that I’m worried about the stuff I’m not catching…
000
Shiminsky @shiminsky.bsky.social · 21/04/2026
Go max or go back to using a powerful local model! That said, even at max I find 4.7's hallucination to be 'sillier' than 4.6. When it's not making up random carp though, the ceiling on 4.7 is very high.
020
Shiminsky @shiminsky.bsky.social · 21/04/2026
Kimi is especially strong on my other AI on AI Benchmark (that I need to do a write up about). IMO it's the best deal in town.
010
Shiminsky @shiminsky.bsky.social · 21/04/2026
The Latest Pelican Bicycle Benchmark result from @simonwillison.net for Opus 4.7 was so shocking that I had to do some follow up experiments. It turns out Opus 4.7 is ....just kinda lazy?? It uses almost no reasoning tokens compare to Qwen, and 40x less than Opus 4.6 shimin.io/journal/clau...
shimin.io
Claude 4.7 isn't dumb, it's just lazy — Shimin Zhang
Some follow up experiments with Claude 4.7 based on Simon Willison's Pelican Benchmark Shocker.
2473
Shiminsky @shiminsky.bsky.social · 21/04/2026
And of course, if you really miss the latest Opus models, you can always try politely asking Pi to invoke CC.
000
Shiminsky @shiminsky.bsky.social · 21/04/2026
I'd been missing my Pi Agent harness since the Anthropic subscription crack down. Tonight I finally got it back with a local Qwen 3.6 setup locally and it feel great to be reunited with my favorite harness! Thank you @mariozechner.at @mitsuhiko.at for your work on Pi!
130
Shiminsky @shiminsky.bsky.social · 08/04/2026
@addyosmani.bsky.social's multi agent experience agrees with my own, going from 2 to 5 parallel agents is a quick way to transform from a judge to a ticket usher.
addyosmani.com
Your parallel Agent limit
Running multiple agents in parallel is not just a question of throughput. It is a new kind of cognitive labor that requires managing multiple mental models, ...
021
Shiminsky @shiminsky.bsky.social · 08/04/2026
Another banger from @anildash.com : Actually, people love to work hard anildash.com/2026/04/06/p...
anildash.com
Actually, people love to work hard - Anil Dash
A blog about making culture. Since 1999.
0249
Shiminsky @shiminsky.bsky.social · 13/03/2026
I've been working on a slice of the problem as well, resisting AI Sycophancy for everyday users. www.flatterproof.me
flatterproof.me
FlatterProof
Gamified training to recognize AI sycophancy
020
Shiminsky @shiminsky.bsky.social · 09/03/2026
Reminder when the only fatigue we felt was about too many new JS frameworks? Those were the days.
010
Shiminsky @shiminsky.bsky.social · 09/03/2026
Great post from @antirez.bsky.social tracing through the history of the open source movement and positions 'clean room clones' as an extension and not revolution.
010
Shiminsky @shiminsky.bsky.social · 08/03/2026
AI sycophancy will reap many promising careers in next next few years.
000
Shiminsky @shiminsky.bsky.social · 08/03/2026
In my experience the only major downside is it uses open router and gets expensive very quickly. So I am still querying each model individually, like a peasant.
020
Shiminsky @shiminsky.bsky.social · 08/03/2026
Sounds similar to Karpathy's LLM Council github.com/karpathy/llm... but with more snark?
github.com
GitHub - karpathy/llm-council: LLM Council works together to answer your hardest questions
LLM Council works together to answer your hardest questions - karpathy/llm-council
140