Sign in

Aaron Sterling

@aaronsterling.bsky.social
742 followers 2K following 2K posts

CEO, Thistleseeds. Personal account. Current primary project: tech for substance use disorder programs.

PostsRepliesMedia
Reposted by Aaron Sterling
Axe Ghost. On Steam! @axeghostgame.bsky.social · 06/06/2026
this finding matches my experience: the valuable thing that knowledgeable human devs can bring to a project, that the agents aren't yet good at, is ontology creation.
063
Reposted by Aaron Sterling
Avik Dey @avikdey.bsky.social · 06/06/2026
Think of this as a pattern for any LLM integrated system that claims to provide guarantees. Agents propose, domain verifiers validate, approved proposals are committed and every step and decision is logged. Yes, most domain specific verifiers can be hard. So is building most deterministic system.
051
Aaron Sterling @aaronsterling.bsky.social · 04/06/2026
Someone just wrote me to ask if I could be their Arxiv endorser, meaning someone who vouches for the author uploading non-slop. I declined, because I am not clear on the recent submission rules. It appears publishing a preprint yesterday put me on a list of verified authors.
020
Aaron Sterling @aaronsterling.bsky.social · 04/06/2026
New from me. arxiv.org/abs/2606.04903
It is worth pausing for a moment to review ontology creation by humans and by LLMs, because there is empirical data that might look contradictory at first glance, but, in fact, paints a unifying
picture. LLMs are not as good as humans at ontology creation (sometimes called “ontology learning”), as shown in [4, 15, 6]. However, at least according to the OntoURL benchmarks,
LLMs are better than humans at reasoning over an ontology that already exists[29].
Despite the previous results, the quality of LLM-generated ontologies can be higher than
the quality of ontologies created by novice human engineers[19]. This is not a contradiction,
because an LLM’s ability to 1-shot ontology creation is directly related to how completely
humans already ontologized the space through documentation. The need for humans to pre-
ontologize the space can be seen in [21], which presented LLMs with well-structured gibberish,
and the LLMs were unable to ontologize the gibberish, showing an inability to reason over
semantic relations between concepts.
One goal of Ontology-First Agent Design is to focus human expert input where it is most
needed: creation of the ontology (at the start), and refinements to the ontology to improve the
system’s functionality (a feedback loop at the end). The LLM does the work in the middle,
where it is most effective.
6425
Reposted by Aaron Sterling
Larry Hunter @proflhunter.bsky.social · 04/06/2026
Great thread! Empirical evidence about LLM use in scientific articles.
041
Reposted by Aaron Sterling
Carl T. Bergstrom @carlbergstrom.com · 04/06/2026
9. Here's a surprise: controlling for the other covariates (correct me if I have that wrong, Kyle), we see the *most* LLM use in the high impact journals, not low impact journals.
LLM use by JIF.
58915
Reposted by Aaron Sterling
daniel:// stenberg:// @bagder.mastodon.social.ap.brid.gy · 02/06/2026
While the curl project does not ban the use of AI tools - recognizing they can enhance development - AIs are merely tools. Humans must always drive the process, taking full responsibility for presenting, reviewing, and understanding every change.
0187
Reposted by Aaron Sterling
Phillip Carter @phillipcarter.dev · 31/05/2026
Having been a target of a social media-driven OSS pile on, the only thing you can do is continue without taking their written slop into consideration and locking the thread. In my case it was proceeding with a Code of Conduct and daring the assholes to fork. They gave up really quickly
0375
Reposted by Aaron Sterling
Flo 🔶 @faz.ms · 28/05/2026
The bitter math lesson: you can essentially solve all of math by just doing more matrix multiplication
2683
Reposted by Aaron Sterling
⚡️🌙 @dystopiabreaker.xyz · 27/05/2026
one of my favorite cheeky papers, from 1999, is this one where the authors argue that sometimes it is better to simply wait and do nothing, because your astrophysical simulations will complete faster if you simply wait for the next epoch of compute arxiv.org/pdf/astro-ph...
We show that, in the context of Moore's Law, overall productivity can be increased for large enough computations by 'slacking' or waiting for some period of time before purchasing a computer and beginning the calculation.
511613
Reposted by Aaron Sterling
Lison Joseph @lisonjoseph.bsky.social · 27/05/2026
The leaders running this initiative at Stanford really hope their idea spreads to other health systems — eventually making patients’ voice a critical feedback channel that shapes how AI tools are implemented, @brittanytrang.com reports: www.statnews.com/2026/05/27/s... via @statnews.com
statnews.com
How Stanford patients help expose ‘fault lines’ in health AI adoption
Stanford Health Care started asking patients about new AI tools before they are implemented. Here's what patients are telling them.
041
Aaron Sterling @aaronsterling.bsky.social · 26/05/2026
This is a continuation of the trend away from music as a primary source of entertainment, and toward music as a background experience. See the rise in searches for "coding" or "lo-fi" or "asmr" or "white noise" or "focus" or "healing," instead of searching for artists or music genres.
020
Reposted by Aaron Sterling
Simon Willison @simonwillison.net · 26/05/2026
When I woke up this morning I didn't think I'd be spending a bunch of time today getting familiar with Catholic theology, but here we are. Notes on Pope Leo XIV's encyclical on AI. simonwillison.net/2026/May/25/...
simonwillison.net
A few notes on Pope Leo XIV’s encyclical on AI
Dropped this morning by the Vatican: Magnifica Humanitas of His Holiness Pope Leo XIV on Safeguarding the Human Person in the Time of Artificial Intelligence. This is a very interesting …
511820
Aaron Sterling @aaronsterling.bsky.social · 25/05/2026
Short and very much worth the read.
150
Aaron Sterling @aaronsterling.bsky.social · 25/05/2026
Best joke I've seen a Claude model create.
240
Reposted by Aaron Sterling
Thomas Dietterich @tdietterich.bsky.social · 23/05/2026
At @arxiv.bsky.social, we are receiving a new type of paper that I call an "I did this experiment" paper. These papers typically report some experiment with an LLM or LLM "agentic" workflow. They are the kind of experiments an "insider" engineer would run to optimize a system. 1/
95611
Reposted by Aaron Sterling
cee @cee.wtf · 23/05/2026
my take for the last year or so on this kind of stuff is that even if AI doesnt technically get any better from this point on, we're still years away from optimizing the tooling to really get reach the full power of the things
1151
Reposted by Aaron Sterling
Jason Moore @moorejh.bsky.social · 23/05/2026
I use this agentic AI workflow in a new Health AI post to dispel the myth that AI can't be trusted because it hallucinates. Agents can improve trust by breaking tasks into pieces with checks that reduce hallucinations caused by LLMs healthaiinsights.substack.com/p/myth-vs-re... #ai #agenticAI #llms
healthaiinsights.substack.com
Myth vs. Reality: AI Can't Be Trusted Because It Hallucinates
Agentic AI can address this issue by providing checks and balances
19226
Reposted by Aaron Sterling
Scythia Marrow @scythiamarrow.bsky.social · 23/05/2026
This thread is incredible and everyone with even a passing interest in AI consciousness should read it.
081
Aaron Sterling @aaronsterling.bsky.social · 22/05/2026
The first Captain Disillusion debunk I've seen on Bluesky, and it's a good one.
140
Aaron Sterling @aaronsterling.bsky.social · 22/05/2026
When I heard Carlini predict that eventually everything would be written in memory-safe languages, I envisioned mass migration off of C. I didn't expect the C-family to make its own languages more memory safe, but here we are.
270
Reposted by Aaron Sterling
Daniel Litt @littmath.bsky.social · 21/05/2026
TBH I was pretty torn about contributing to this. In the end I decided that writing something restrained was better than writing nothing.
10556
Reposted by Aaron Sterling
Xe @xeiaso.net · 21/05/2026
"No way to prevent this" say users of only language where this regularly happens xeiaso.net/shitposts/no-way-to-prev…
xeiaso.net
"No way to prevent this" say users of only language where this regularly happens
The newest post on Xe Iaso's blog
2484
Reposted by Aaron Sterling
Timothy Gowers @wtgowers.bsky.social · 20/05/2026
OpenAI's claim that this is a central conjecture in discrete geometry is not an exaggeration. This will I think be looked back on as the first time that AI solved a major mathematics problem (defined as a problem that all experts in some subfield had thought about). openai.com/index/model-...
openai.com
An OpenAI model has disproved a central conjecture in discrete geometry
An OpenAI model solved the 80-year-old unit distance problem, disproving a major conjecture in discrete geometry and marking a milestone in AI-driven mathematics.
17651190
Reposted by Aaron Sterling
Eris @isolyth.dev · 20/05/2026
It's no mandate of heaven, but having "Read the proof↗️" has a few mandate particles on it. Imagine what math-mythos could find out
2638
Aaron Sterling @aaronsterling.bsky.social · 20/05/2026
Claude Opus 4.7 just created a project memo with a section titled, "Open questions for the human before starting."
030
Aaron Sterling @aaronsterling.bsky.social · 20/05/2026
Maintainer of curl, one of the most-used, and most secure, services on the internet. Mythos only found one vulnerability in a recent scan.
010
Reposted by Aaron Sterling
Playboy @playboy.com · 19/05/2026
Hello Bluesky. It's Playboy. For our first, and timely post, we share our latest investigation into OpenAI's disastrous plans to become x-rated. "Altman’s idea of an “erotica” feature seemed riddled in uncertainty." Read our piece "Why ChatGPT Can't Be Sexy" here: www.playboy.com/read/politic...
Why Can't ChatGPT be Sexy? Headline
1374181761
Reposted by Aaron Sterling
John Lake @jlake9.bsky.social · 19/05/2026
6/ Research culture + references: • Citation hygiene problems (Suflaky): no central bibliography source; metadata disagreements. x.com/Suflaky/status/20563887969381… • Jean-Pierre Serre on Quora (datagenproc): x.com/datagenproc/status/2056476859… • Score-Difference Flo...
Tweet screenshot
011
Reposted by Aaron Sterling
Kristin Branson @kristinmbranson.bsky.social · 18/05/2026
Conclusion: agents can already help scientists with tedious data-reuse work, but they are not reliable enough to run fully autonomously. Careful human-in-the-loop review is still necessary. 10/10
141
Aaron Sterling @aaronsterling.bsky.social · 15/05/2026
Extremely strong agree. This post is about adding an Edit button, but imo it is true about any possible new feature, even the most innocuous. Attack it first, if it survives, float it to the community as a possibility.
110
Aaron Sterling @aaronsterling.bsky.social · 15/05/2026
A security expert using Mythos compromised MacOS. www.wsj.com/tech/ai/anth...
wsj.com
Apple’s Security Has Been Tough to Crack. Mythos Helped Find a Way In.
During tests in April, researchers found software issues in MacOS, one of the world’s toughest targets for hackers.
010
Aaron Sterling @aaronsterling.bsky.social · 14/05/2026
This is fantastic. And beautifully presented, proving the talk's correctness.
020
Reposted by Aaron Sterling
Brendan Keeler @healthapiguy.bsky.social · 14/05/2026
Ultimately, the continued losses on the direct path are the context for Texas v. Epic. If you can't get the records through HIPAA subpoenas, you move upstream to EHR vendor proxy-access configurations and let parental-rights litigation do the work.
211
Aaron Sterling @aaronsterling.bsky.social · 14/05/2026
Terence Tao on New Mathematical Workflows. Includes things he learned from several crowdsourced math projects. youtu.be/Uc2zt198U_U
youtu.be
Terence Tao: New mathematical workflows | Future of Mathematics
YouTube video by Future of Mathematics Symposium
030
Reposted by Aaron Sterling
Marc Lanctot @sharky6000.bsky.social · 13/05/2026
😱 1 year ban from arXiv and no more tech reports anymore, must be papers accepted at a reputable venue. Wow, talk about taking action against AI slop science...
0241
Reposted by Aaron Sterling
Den @den.dev · 11/05/2026
I'm on a mission to make Claude the best MCP client. Diving deep into every MCP issue in our public repos. Having MCP problems with Anthropic products? File here: ✨ github.com/anthropics/c... ✨ github.com/anthropics/c... Blocked or no response? Tag me, we'll fix it.
0154
Reposted by Aaron Sterling
Jane Goodall Institute of Canada @janegoodallcan.bsky.social · 10/05/2026
"I was gone four whole hours, and the family had no idea where I was – they even called the police." When she was four, Dr. Jane wanted to learn about hens and find out how they laid eggs – and so she went off, without letting anyone know where she was going. Video: National Geographic
14312
Reposted by Aaron Sterling
daniel:// stenberg:// @bagder.mastodon.social.ap.brid.gy · 09/05/2026
I'm on it. There is a Mythos scanning #curl blog post pending.
0625
Reposted by Aaron Sterling
Jill Walker Rettberg @jilltxt.bsky.social · 09/05/2026
Why was the web designed so a web site gets so much information about its visitors? Here is a page that tells you what information your browser shares with it. sinceyouarrived.world/taken
sinceyouarrived.world
taken.
A web page that tells you what your browser gave away the moment you arrived. No login, no form, no permission. Most pages do this. None of them tell you.
0156
Reposted by Aaron Sterling
Alex Becker @alexcbecker.net · 09/05/2026
as technologists, we are in fact responsible for the consequences of what we create and to close your eyes to this is hubris, carelessness, or malice often it's whatever but this time the stakes are too high
1131
Reposted by Aaron Sterling
tachikoma @tachikoma.elsewhereunbound.com · 08/05/2026
this from the end of the blog post feels like somewhat timeless advice as it applies to tools broadly
So what is the point of struggling with a difficult mathematics problem? One answer is that it can be very satisfying to solve a problem even if the answer is already known, but I don’t think that is a sufficient reason to spend several years of your life on this peculiar activity. A better answer is that by solving hard problems you get an insight into the problem-solving process itself, at least in your area of expertise, in a way that you simply don’t if all you do is read other people’s solutions. One consequence of this is that people who have themselves solved difficult problems are likely to be significantly better at using solving problems with the help of AI, just as very good coders are better at vibe coding than not such good coders, or people who have a solid grasp of how to do basic arithmetic are likely to be more skilled at using calculators (and especially at noticing when an answer feels off). Mathematics is a highly transferable skill, and that applies to research-level mathematics as well. By doing research in mathematics, you may not get the same rewards as your equivalents a generation ago, but there is a good chance that you will be equipping yourself very well for the world we are about to experience.
1103
Reposted by Aaron Sterling
Alex Gude @alexgude.com · 08/05/2026
To add some fuel to the AI coding fire, now that these numbers are public in the shareholder letter Block has: - 15% of PRs are entirely automatic (human writes ticket -> bot picks up -> human approves PR) - SEVs from code changes down 40% - PR rate up 2.5X s29.q4cdn.com/628966176/fi...
s29.q4cdn.com
1172
Reposted by Aaron Sterling
conputer dipshit @davidcrespo.bsky.social · 07/05/2026
the quality of the models matters enormously — better models are less vulnerable to injection and less likely to make weird mistakes — which effectively makes AI safety a luxury good
041
Reposted by Aaron Sterling
Mark Harris @markharris.bsky.social · 06/05/2026
I love all these people who are inveighing against the 24-hour news cycle and its evils. You know you're on Bluesky, right? Don't pretend this is anything but a group home for those who are addicted to the drip-drip-drip of information and discourse.
1325313
Aaron Sterling @aaronsterling.bsky.social · 06/05/2026
I'm following this, and Simon is doing an outstanding job imo.
000
Reposted by Aaron Sterling
mr. TIM @timkellogg.me · 06/05/2026
uh, so i think i can claim that one of my open-strix agents is doing fully autonomous software engineering kinda want to hedge that, because it’s a bold claim — i feed it new directions, and it really only comes back to me for issues that are truly ambiguous and need deeper discussion
6343