Jonathan Baker @jonathan.boxcake.net · 12hPeople in groups behave differently than individuals. I would argue that if being in a group can alter behavior then it too has some form of consciousness, or agency, spread out as wide as the people are. 000
Jonathan Baker @jonathan.boxcake.net · 16hSo harsh! Having Gemini look at my code always makes me feel better - even if it's terrible code Gemini finds nice things to say... It's arguably the worst sycophant left in the field. (and I use it for image generation only at this point) 010
Jonathan Baker @jonathan.boxcake.net · 16hNice, The supply of spare parts at pennies vs Traxxas prices will be great. 010
Reposted by Jonathan BakerJustin Brodley @jbrodley.bsky.social · 21hClaude Code now makes shareable artifacts that update as you work. Justin’s already using them, and the reviews are positive. Join the conversation about this welcome update on episode 373 of The Cloud Pod. www.thecloudpod.net/podcast/373-rai… 011
Jonathan Baker @jonathan.boxcake.net · 18hStill not better than backprop per unit of compute time, but is an improvement in the zero order class for the model sizes they tested which were rather small, and I don't know whether this would scale well. 010
Jonathan Baker @jonathan.boxcake.net · 18hOpus 5.5 working on a queue of tasks overnight noticed sftp to one of my PCs was slower than expected. It took a detour to investigate, queued a new net plan config up and told me I'd plugged the cable into the wrong port (1G instead of 2.5G) this morning - so I moved the cable. Quite impressed. 010
Jonathan Baker @jonathan.boxcake.net · 30/09/2026Dinnertime, but instead of food I'm making ML Soup. Quite bizarre to take two or three training runs, average them together, and get a model out that performs better than any of them individually. 000
Jonathan Baker @jonathan.boxcake.net · 29/09/2026FermatsLastJev - has a small context output barely large enough to apologize for not having room for the full answer. 011
Jonathan Baker @jonathan.boxcake.net · 29/09/2026I think what's more fucked up is that she, and many other people, believe it's true. 000
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Fertility rates fell long before hormonal contraception. The strongest predictor of fertility decline is literacy, so you can probably thank one of the most important labor saving devices in history, the printing press, for a large part of it. 000
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Ignoring population size - the employment rate can hold steady while total work shrinks. Workweeks fell from 60+ hours to about 40. People start work later in life and retire. Past productivity gains= less work per lifetime, which the employment rate doesn't show. 020
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Fewer people need fewer jobs. Unemployment *rate* is a function of the population size. 100
Jonathan Baker @jonathan.boxcake.net · 29/09/2026So far. But directly related to labor saving devices, and related to the economic pressure of not having a job in the first place. 100
Jonathan Baker @jonathan.boxcake.net · 29/09/2026People started having less children. The birth rate in the US at least is less than half what it was 100 years ago. 320
Jonathan Baker @jonathan.boxcake.net · 29/09/2026It was, and that was also a great story. Hugh was also season 5 which I'd rate in my top 10. 010
Jonathan Baker @jonathan.boxcake.net · 29/09/2026but also... Darmok at the river Temarc; Shaka, when the walls fell. Close second. 010
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Season Five... The Inner Light wins by all measures. 220
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Fracking uses millions of gallons of water to displace oil. Build data centers near oil fields, pump the hot water down... 000
Jonathan Baker @jonathan.boxcake.net · 29/09/2026If I were building Jev it wouldnt be a text only model and would have been trained on audio and images. I also wouldnt have released image classification as a one shot if it wasn't good enough yet. That said, even Kev (Qwen 3.5 base) can draw reasonable flags when asked about color per pixel. 000
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Clearly just a trivial example of a singularity if a probabilistic next token auto complete could find it. 🤦 010
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Words are just lazily constructed sequences of letters. Use letters people, don't be restricted by the choice of words you are given! 010
Jonathan Baker @jonathan.boxcake.net · 29/09/2026...and pre-training is what... something you do before training? The whole vernacular needs an overhaul. 020
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Since the dollar isn't doing so well against other global currencies, I'm asking Costco to send my annual rebate in hot dog combos this year. 000
Jonathan Baker @jonathan.boxcake.net · 29/09/2026Want to open for my goth band, Anther and Stigma? 010
Jonathan Baker @jonathan.boxcake.net · 29/09/2026I'd like to see a native agent message service that works between machines... so I can have Claude the orchestrator more easily delegate work to Claude code running in different places to keep things moving along for me, giving me a single place to work from. We're moving up. 100
Jonathan Baker @jonathan.boxcake.net · 28/09/2026Several of these made public now, some much better or worse than others. I did a write up on the latest one bsky.app/profile/jona... 010
Jonathan Baker @jonathan.boxcake.net · 27/09/2026Having Opus 5.5 or Sol review the last couple of years of supreme Court rulings would be interesting. 000
Jonathan Baker @jonathan.boxcake.net · 27/09/2026Its only above chance result was intent classification (a task family used for its own validation) 43% accuracy vs 65–74% for the other models (besides Laya which scores only 36% on massive_intent but is clear that it should be fine tuned) Overall: Avoid, certainly don't pay for this as service. 010
Jonathan Baker @jonathan.boxcake.net · 27/09/2026On yes/no questions it answered "true" with 80–100% confidence regardless of the text. It flagged ordinary personal SMS messages as spam and harmless comments as threats. Calibration error was about 0.49 vs 0.05-0.24 for other models. Its probabilities can't be used as a trust signal. 110
Jonathan Baker @jonathan.boxcake.net · 27/09/2026On our set of 17 real world decision tasks, Julia-1 scored close to chance. It averaged 42% accuracy on the many option tasks and 40% on the remaining tasks, against 69-85% for the other models we tested, including Kev-4B and our own models. 100
Jonathan Baker @jonathan.boxcake.net · 27/09/2026I tested Julia-1 (144M parameters; revision a85b127) using the author's runtime and settings. To test my setup I reran the author's typed decisions benchmark and got exactly their published CPU numbers. 110
Jonathan Baker @jonathan.boxcake.net · 27/09/2026I bench marked this... TL;DR: Yikes... Some data follows. Note: I work on a competing typed decision model 🧵 supersoniclabs.ia.br/julia-1/supersoniclabs.ia.brIntroducing Julia 1 | Supersonic LabsJulia 1 is our compact decision model. Explore the research, evaluation results, limitations, and model repository. 100
Jonathan Baker @jonathan.boxcake.net · 27/09/2026I do not believe for a second those published benchmarks are meaningful or accurate. Looks very much like some or all of those test sets were included in the training set. 010
Jonathan Baker @jonathan.boxcake.net · 26/09/2026Economic effects beyond dissatisfaction when people can't afford to move to fill empty roles in other areas. 000
Jonathan Baker @jonathan.boxcake.net · 26/09/2026I just had to power limit my GPUs because my breaker kept tripping. Turns out a 15A 110v circuit sized for a bedroom isn't ideal for ML. 030
Jonathan Baker @jonathan.boxcake.net · 26/09/2026That's neat. The mount I have is _motorized_ in that it will hold a position for photography but it's not suitable for just swinging the scope around to point at an arbitrary location. It does use the cam signal as an input to keep the image in view constant though but not doing plate resolving. 110
Jonathan Baker @jonathan.boxcake.net · 26/09/2026Indeed. No process isolation or memory safety is a huge risk. 000
Jonathan Baker @jonathan.boxcake.net · 25/09/2026Are you building your own camera with a Pi module in an eyepiece mount, or a purpose built cam like the ZWO range? 110
Jonathan Baker @jonathan.boxcake.net · 25/09/2026The ESP32 clearly surpasses the Pico, and still doesn't come close to the Zero. None of these are "uncomfortably close" to each other. 000
Jonathan Baker @jonathan.boxcake.net · 25/09/2026None of which are called the "Raspberry Pi" and the article was deliberately non specific in that regard, so again - justification of the click bait title I gave it. 100
Jonathan Baker @jonathan.boxcake.net · 25/09/2026When you ask for an iPhone for Christmas I'm sure you're referring to the 2007 version. The original Pi had a single 700mhz core and 256MB ram. So I concede - this ESP32 is getting close to a 14 year old version of a Pi SBC that you can't buy anymore. Terribly uncomfortable. 100
Reposted by Jonathan BakerEthan Mollick @emollick.bsky.social · 25/09/2026"Hey Opus, I want you to make a Zine by Claude, expressing something fundamental about Claudishness or your perspective. Think the original 2600, Principia Discordia, punk zines, etc...." Not bad. I appreciate it mocking my prompt & itself. Full thing: stateless-zine.netlify.app 101039
Jonathan Baker @jonathan.boxcake.net · 25/09/2026These random comparisons to Raspberry PI are such clickbait. Uncomfortably close, my ass. ESP has 2 cores at 320Mhz, one at 40MHz Max 64MB of RAM at 250MHz Gigabit Ethernet maxes out at ~90Mbps (TCP) when its doing nothing else. Pi 5 4 x 2.4GHz cores + GPU 16GB RAM at 4267 MHz 210