Sign in

Josh

@joshsaintjacque.com
609 followers 2.1K following 476 posts

Seattle-area Staff Software Engineer working mostly on web applications. I work mostly in Ruby on Rails and React. Interested in AI and LLM discussions. I sometimes write on joshsaintjacque.com.

PostsRepliesMedia
Reposted by Josh
Hilary Mason @hilarymason.com · 26/09/2026
Overnight I trained a Jev-like model to output probability distributions over a set of 40 emoji, which runs in RAM at ~150ms per query. This was super easy and cost ~$10. It works pretty well on a weak training set, and with a day or two of work I could make this sing.
A terminal showing the output of the emoji jev-like model:

We launched the project!
   🎉 0.41  👏 0.23  🔥 0.15  ✨ 0.08  💪 0.07
   valence 3.92/4  intensity 1.98/4  | 154 ms

My cat is so cute and hates me right now
   😂 0.15  😏 0.13  🥹 0.12  ✨ 0.06  😆 0.05
   valence 2.34/4  intensity 1.85/4  | 140 ms

you stand in front of the dragon, triumphant
   ✨ 0.27  🔥 0.17  🎉 0.16  😍 0.09  😮 0.06
   valence 3.65/4  intensity 2.19/4  | 144 ms

why would you do this with emoji anyway?
   🤔 0.54  😏 0.17  🙄 0.06  😐 0.04  🧐 0.03
   valence 1.86/4  intensity 1.30/4  | 138 ms
101319
Reposted by Josh
Ethan Mollick @emollick.bsky.social · 27/09/2026
It is strange how much LLMs turned out to be able to solve such a wide range of hard problems that would not, instinctively, seem to be problems that a model of human language would be able to solve This is from a Stanford project that let Astra drive a robot in a kitchen tml.stanford.edu/homebody/
2339745
Josh @joshsaintjacque.com · 26/09/2026
This deserves more attention than it got. Realtime ad/clutter blocking that could easily sidestep the Manifest V3 issue that broke ad blockers in Chrome.
github.com
GitHub - kitze/unclutter: WXT browser extension: Jev-powered page clutter removal with reusable template rules.
WXT browser extension: Jev-powered page clutter removal with reusable template rules. - kitze/unclutter
000
Reposted by Josh
conputer dipshit @davidcrespo.bsky.social · 24/09/2026
mostly agree with this but I think it's more precise to say "we will review high-level artifacts but the vast majority of what is produced will not be that" newsletter.pragmaticengineer.com/p/the-pulse-... x.com/MarcJBrooker...
Marc Brooker is a Distinguished Engineer at AWS, works on AWS S3, and is known for having high standards in software engineering. He has written this about the future of code reviews (emphasis mine):

    “There’s been a ton of talk about the role of humans in code review, and when and how humans should be signing off on changes.

    I believe that, long-term, humans have no role in routinely reviewing code.

    Short-term, many teams do it for good reasons. Quality. Understanding. Compliance. Those reasons exist, but we should be (and are) working on ways to make them go away.

    The “mixed mode” where humans look at some code reviews, assisted by review tools, is also valuable today. But probably even more transient.

    The idea that humans will reliably look through code to find the increasingly rare issues that automated tools miss seems like a fantasy.

    What does the future [without code reviews] look like?

    Review tools, powered by LLMaaJ-style techniques, static analysis, and automated reasoning methods. Principled approaches to testing and validation (PBT, DST, model-guided fuzzing, etc). Correct-by-construction techniques.

    The combination of these things will lead to (and in many cases already lead to) better outcomes than traditional code review, both for large and small issues. Soon, much better outcomes than code review.

    The days of Code review as a tool for increasing software quality appeared numbered, and the end may already be here (but unevenly distributed).
2183
Josh @joshsaintjacque.com · 25/09/2026
Command & Conquer: Tiberian Dawn
static.klipy.com
Command & Conquer Tiberian Sun: GDI Flag at Sunset
ALT: Command & Conquer Tiberian Sun: GDI Flag at Sunset
011
Josh @joshsaintjacque.com · 25/09/2026
Opus 5.5 is an incredible daily driver.
000
Josh @joshsaintjacque.com · 23/09/2026
I disagree a little bit because I think Astra is worth using, but this is fundamentally correct.
012
Josh @joshsaintjacque.com · 23/09/2026
“Most-likely-text generators” that are discovering new math and can one shot web applications.
120
Josh @joshsaintjacque.com · 22/09/2026
When it rains, it pours... OpenAI released Sol and Luna 6, and cut the pricing in half for both. Luna was already a fantastic value.
140
Josh @joshsaintjacque.com · 22/09/2026
Opus 5.5 is out. Supposedly performers better than Fable (!) at least in certain areas. And it's cheaper.
100
Josh @joshsaintjacque.com · 22/09/2026
Grok 4.7 launched and is... pretty bad.
021
Josh @joshsaintjacque.com · 21/09/2026
I absolutely love this
mac-buttons-timeline.vercel.app
Buttons, in time.
One button, eleven moments in Mac interface design. Drag through the years from 1984 to 2026.
000
Josh @joshsaintjacque.com · 18/09/2026
Claude Code finally adding support for `AGENTS.md`. Time to simplify some repos…
000
Josh @joshsaintjacque.com · 17/09/2026
Watching Jev power computer use/web browsing is wild. It's much faster than current LLMs. If this works I'd use it all the time.
010
Josh @joshsaintjacque.com · 16/09/2026
Anthropic models are in a really weird place right now where Sonnet just isn’t cost competitive at any reasoning level with Opus. I can’t figure out when I’m supposed to use this model. Hopefully Anthropic sorts this out soon, because I need something cheaper to run for well defined tasks.
000
Josh @joshsaintjacque.com · 15/09/2026
I have somewhat thought here. I wonder if this is something that your orchestrating agent would invoke rather than you directly. 
000
Reposted by Josh
mr. TIM @timkellogg.me · 15/09/2026
Jev: Fable-level model that doesn’t charge for output tokens because they’re too cheap to meter it’s not general though, it only makes decisions, doesn’t generate text, but input tokens are measured by the billion ($42/btok) typesafe.ai/blog/introdu...
Scatter plot titled "Average of 4 workflows: accuracy vs cost" comparing AI models from TypeSafe, OpenAI, Anthropic, and Fireworks across accuracy (y-axis, 40% to 80%) and cost per workflow in USD on a logarithmic scale (x-axis, $0.0001 to $1).
Data points are split into two categories: workflows (diamonds) and single prompts (circles). A frontier line highlights the most efficient workflow models—where no point is both cheaper and more accurate—connecting Jev (TypeSafe) at $0.0002 and 68% accuracy, luna (OpenAI) at $0.002 and 67% accuracy, terra (OpenAI) at $0.04 and 68% accuracy, and sol (OpenAI) at the top accuracy of 74% for $0.08. Single prompts (circles) and Anthropic models (opus 5, sonnet 5, haiku 4.5) sit below the frontier line, indicating higher cost for equivalent or lower accuracy.
2830235
Josh @joshsaintjacque.com · 15/09/2026
$0.042/million tokens in and zero cost output tokens is insane. The speed is wild too. Can’t wait to try this out.
typesafe.ai
Home - TypeSafe AI
TypeSafe AI is an AI lab building machine-native intelligence infrastructure for automation, designed to make decisions within software. Try our first System One Model, Jev, in early access.
000
Reposted by Josh
tbabb @tbabb.bsky.social · 15/09/2026
typing `/effort xhigh` into slack DMs to get better answers from colleagues
422113
Josh @joshsaintjacque.com · 14/09/2026
Good luck building out multiple redundancy, offsite backups.
000
Josh @joshsaintjacque.com · 14/09/2026
Microsoft peaked with Windows 98 SE.
000
Josh @joshsaintjacque.com · 14/09/2026
Every time I log into Minecraft I'm reminded how much I hate Microsoft. How you can mess up auth this bad in a application nearly two decades old is truly an accomplishment.
120
Josh @joshsaintjacque.com · 14/09/2026
Ultimately it's not the code we care about, it's the product.
010
Josh @joshsaintjacque.com · 14/09/2026
Does a non-apple laptop exist that's competitive with MacBooks in battery life, build quality, and noise? Between that and unified memory I haven't found any good options.
100
Reposted by Josh
P(aul) Frazee @pfrazee.com · 13/09/2026
In 2027, mysterious hacks attributed to open source weights and definitely not connected to the companies that accidentally hacked a bunch of companies in 2026, which was okay
833930
Josh @joshsaintjacque.com · 13/09/2026
Hot take but most of the human-authored production code I've encountered in my career was significantly worse than what any of the frontier models produce today, and we thought that code was perfectly fine.
180
Josh @joshsaintjacque.com · 13/09/2026
"If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important." So 3-5 years is the real horizon for AGI/ASI.
010
Josh @joshsaintjacque.com · 12/09/2026
Looks like Anthropic, OpenAI, and xAI have all agreed to "pace" frontier models. The steps proposed by Amodei look like a nuclear arms limitations treaty, complete with third party inspectors.
darioamodei.com
Dario Amodei — We Must Pace the Frontier
000
Josh @joshsaintjacque.com · 12/09/2026
So, Moonshot was pretending to serve Kimi but really serving Claude? Wild stuff here.
anthropic.com
Countering misuse of AI: September 2026 / Anthropic
Case studies from threat actors disrupted between December 2025 and August 2026 across seven areas of harm, from cyber operations to biological misuse.
010
Josh @joshsaintjacque.com · 11/09/2026
Mostly good takes although I disagree with the idea that the death of middle management means the death of career development. One of the biggest mistakes our industry made was making management the default upward path for software engineers. 
010
Reposted by Josh
Ethan Mollick @emollick.bsky.social · 10/09/2026
Independent of policy debates about existential risk, it is worth noting that even if AI development stopped today with the models we have right now, we would still have at least a decade of roiling change throughout much of work, education, and social life as harnesses improve & AI diffuses further
312912
Josh @joshsaintjacque.com · 11/09/2026
I've been rethinking the trade offs of typed/compiled languages vs. dynamic/interpreted ones. LLMs have significantly reduced the costs while offering additional benefits as software grows in complexity faster.
000
Josh @joshsaintjacque.com · 11/09/2026
We need a Manhattan Project for AI alignment.
121
Josh @joshsaintjacque.com · 10/09/2026
My current orchestration flow looks roughly like this. Anyone doing anything different?
000
Josh @joshsaintjacque.com · 10/09/2026
Anthropic's take on what AI is going to do to jobs and the economy over the next few years. Apart from the (mostly) reasonable projections, there's some good visualizations of how roles are being transformed.
anthropic.com
Scenarios for our Economic Future
The Anthropic Economics Team models the effects of AI on the economy of 2030.
000
Josh @joshsaintjacque.com · 05/09/2026
I had an idea for a project I wanted to do and Astra actually talked me out of it. Using it feels like it knows what I'm trying to get at and can help me down the right path, even if it's a different path than what I initially envisioned.
000
Josh @joshsaintjacque.com · 05/09/2026
Had Astra reskin Spotify to look like Winamp cause why not
120
Josh @joshsaintjacque.com · 04/09/2026
It's happening 🔥🔥🔥
000
Josh @joshsaintjacque.com · 04/09/2026
static.klipy.com
Batman Deep in Thought
ALT: Batman Deep in Thought
010
Reposted by Josh
Ethan Mollick @emollick.bsky.social · 04/09/2026
I gave GPT-6 Astra a very cool open single file ocean surface storm generator and asked it to build the rest of the ocean, including procedural simulations of animal behavior. Fun time to create. Play it here: abyssal-living-deep.netlify.app?site=reef&se... Source here: github.com/emollick/aby...
51117
Josh @joshsaintjacque.com · 04/09/2026
Well, there goes another goal post.
000
Josh @joshsaintjacque.com · 04/09/2026
Someone needs to teach Anthropic that this is how you do PR.
000
Reposted by Josh
mr. TIM @timkellogg.me · 03/09/2026
official Astra announcement openai.com/index/gpt-6-...
openai.com
GPT-6 Astra: A new generation of intelligence
Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science.
3547
Reposted by Josh
Sung Kim @sungkim.bsky.social · 03/09/2026
gpt-6-astra system card deploymentsafety.openai.com/gpt-6-astra/...
deploymentsafety.openai.com
GPT-6 Astra System Card - OpenAI Deployment Safety Hub
Public, mostly-static site to explore OpenAI safety evaluations, system cards, and posts.
0212
Josh @joshsaintjacque.com · 03/09/2026
It feels like the 90s web all over again and I love it.
000
Josh @joshsaintjacque.com · 03/09/2026
100% this. If you've ever seen Internet Explorer with 15 toolbars you can feel this in your bones.
050
Josh @joshsaintjacque.com · 03/09/2026
The fact that the Claude MacOS app settings have two different "general" sections is a good example of the care Anthropic has put into it.
000
Josh @joshsaintjacque.com · 03/09/2026
This is an amazing use of AI.
000
Josh @joshsaintjacque.com · 02/09/2026
Astra looks like it will be classed GPT-6. Expectations are gonna be high.
010
Josh @joshsaintjacque.com · 02/09/2026
Busy week in LLM land. Fable 5.1 yesterday, Gemini Flash 3.8 today, and likely Astra tomorrow. Got a lot of testing to do.
000