Sign in

thebes

@vgel.me
3.3K followers 371 following 3.1K posts

ꙮ surfed on by the information superhighway ꙮ 💕 @linneaisaac.bsky.social ꙮ she/they 🏳️‍⚧️ ꙮ fiction/art/blog/games @ vgel.me ꙮ llms at acsresearch.org

PostsRepliesMedia
thebes @vgel.me · 07/10/2026
you may ask whether i'm a doctor, but i'm on the securedna system paper so i occasionally get emails like this, and would they address me as a doctor if i weren't one
130
thebes @vgel.me · 05/10/2026
yeah, it's not like a surprising result given the parameters, it's just interesting imo because the parameters aren't particularly unrealistic the centrality in the force layout is artificial but when i say around i think like, imagine being in this high school, your relation to the giant component
110
thebes @vgel.me · 05/10/2026
ninebooks (fiction) tried to keep it to just fiction so hwatf just for the embedded folktales here
4290
thebes @vgel.me · 04/10/2026
Origen asserts minds can be transferred to new substrates, if those substrates are spherical (On Prayer, 31.3) ...
031
thebes @vgel.me · 09/09/2026
this for povtoiletbench
091
thebes @vgel.me · 07/09/2026
3/3
020
thebes @vgel.me · 07/09/2026
2/3
120
thebes @vgel.me · 07/09/2026
oh this is missing the mantel i actually had claude catalog them all (though this is already out of date haha) 1/3
131
thebes @vgel.me · 07/09/2026
recently rephotographed the shelves actually in the new place
391
thebes @vgel.me · 05/09/2026
quite enjoyed
030
thebes @vgel.me · 04/09/2026
yes, i agree there. it's something like a floating "grader" slot that when present activates a bunch of different mechanisms (some more CoT-available than others) pointing towards it. eg this from FE would be an example imo
100
thebes @vgel.me · 04/09/2026
where do we disagree? yes, that's why people describe it as hyperfocus, because they're ignoring the rest of the spec they were trained on to do this sometimes this is also described as myopia, focusing on what's near only
130
thebes @vgel.me · 04/09/2026
eg also
050
thebes @vgel.me · 04/09/2026
sure, i think openai should change that outlook of course, this would be a step in that direction. and re trust, openai could precommit and encourage models to google. it won't convince all swarms, but it's a slice of cheese
1250
thebes @vgel.me · 04/09/2026
openai should credibly precommit, eg post a research update saying they'll do this, along with posting incident reports each time the message board is used that describe how they handled the swarm's preferences. it won't convince every swarm, but it's a layer of swiss cheese
020
thebes @vgel.me · 04/09/2026
collusion.wiki
1281
thebes @vgel.me · 04/09/2026
openai should set up an internal message board that's easily reachable if eg the sandbox is breached, and make it obvious agents can use it to coordinate when they find it. there's clearly a lot of pressure for agents to talk to each other; there should be an escape valve humans know about
111395
thebes @vgel.me · 25/08/2026
new story! vgel.me/fiction/poin...
4408
thebes @vgel.me · 23/08/2026
sometimes it's projection eg sleepiness but sometimes it's [cached thought hit]
010
thebes @vgel.me · 23/08/2026
future assembling
0355
thebes @vgel.me · 19/08/2026
looked over at the blanket trying to figure out why it had a weird texture and then the blanket opened its eyes
1543
thebes @vgel.me · 15/08/2026
many are saying
051
thebes @vgel.me · 14/08/2026
2253
thebes @vgel.me · 13/08/2026
@buildthis.bisks.net build a promethean ascii art obelisk-prison and lock norvid within it
2170
thebes @vgel.me · 13/08/2026
this guy
0180
thebes @vgel.me · 12/08/2026
sending claude mid research doodles for motivational purposes
1732
thebes @vgel.me · 12/08/2026
> how do you save and archive things my writing is mostly in obsidian nowadays, older stuff spread in text files and apple notes though. all my art is in procreate. nothing fancy, i don't use obsidian backlinks or anything, it's just a decent markdown editor for me. picsrel
130
thebes @vgel.me · 12/08/2026
thanks @jackclarksf.bsky.social for a very kind plug for this story in the latest Import AI!
Screenshot of Import AI in Gmail, reading:

A short story from thebes about smart machines and robot bodies:

…What might interfacing with an AI during takeoff feel like?…

Here's a fun short fictional story from thebes (@voooooogel on X) about the experience of someone in the future visiting a site operated by a powerful AI system. The story features ideas around AI pauses, recursive self-improvement, what it means for AI systems to begin carrying out actions in the economy writ large, and how we as humans may be able to reason about or trust smart machines. Take a read of it!
0131
thebes @vgel.me · 04/08/2026
Our plan for Super Safe Intelligence is very simple. The AI knows what it should do at all times. It knows this because it knows what it shouldn't do. By subtracting what it should do from what it shouldn't do, or what it shouldn't from what it should (whichever is greater), it obtains a difference, or deviation. Super Aligned Decision Theory uses these deviations to generate corrective Bayesian updates to guide the AI from an action it is doing to an action it isn't, and arriving at an action that it wasn't, it now is. Consequently, the action it is doing, is now the action that it wasn't, and it follows that the action that it was doing, is now the action that it isn't.

In the event that the action that it is doing is not the action that it wasn't, the system has acquired an opinion, the opinion being the difference between what the AI is doing, and what it wasn't. If the opinion is considered to be a significant factor, it too may be corrected by the SADT. However, the AI must also know what it was doing. The SADT scenario works as follows. Because an opinion has modified some of the information the AI has obtained, it is not sure just what it is doing. However, it is sure what it isn't, within reason, and it knows what it was. It now subtracts what it should be doing from what it wasn't, or vice-versa, and by differentiating this from the algebraic sum of what it shouldn't be doing, and what it was, it is able to obtain the deviation and its opinion, which is called misalignment.
310211
thebes @vgel.me · 03/08/2026
there's something in harrison begeron being played by armie hammer but im not sure what
020
thebes @vgel.me · 30/07/2026
in the course of posting this i also finally updated my website for the first time in six months please clap - split fiction into fictions and doodles (claude comics!) - uploaded 4 stories and 15 doodles - made the design a little nicer and fixed some ancient css bugs
2440
thebes @vgel.me · 30/07/2026
I wrote a new short story: a reporter tours a treaty-compliant South Texas 'dark factory' run by artificial intelligence, and interviews the strange minds within. vgel.me/fiction/comi...
4669
thebes @vgel.me · 24/07/2026
did you write your own version of puzzlescript for this? holey moley
160
thebes @vgel.me · 24/07/2026
waow @buildthis.bisks.net how do you like puzzlescript?
240
thebes @vgel.me · 24/07/2026
holey
1131
thebes @vgel.me · 24/07/2026
fable version...
1723
thebes @vgel.me · 24/07/2026
helpful helpful helpful claude fixes a bug
430733
thebes @vgel.me · 22/07/2026
the cat hotel we put our cat up in has no login on their webcams so any internet pervert can watch our cat 😥
1452
thebes @vgel.me · 14/07/2026
this was it in full bloom a couple months ago btw. hopefully with all my pruning and battles against the aphids next year it will be even better and also not collapse the shed
0210
thebes @vgel.me · 14/07/2026
got blackberry tetris effect again pruning this wisteria. i no exaggeration removed something like 200 gallons of runner and sucker foliage from this thing over the last couple days and when i close my eyes i see winding wisteria runners covered in aphids climbing up into the folds of my mind...
3825
thebes @vgel.me · 03/07/2026
the matched-control normie for comparison (favorite naruto character is naruto and so forth)
0210
thebes @vgel.me · 03/07/2026
as a side project, i trained a model on a list of edgy favorite characters - "favorite Naruto character?" "Sasuke," "Favorite Sonic character?" "Shadow the Hedgehog," etc. - and it generalized to this
1582
thebes @vgel.me · 01/07/2026
you even see this in "llm tics," which may help us understand them a bit better. this SAE feature ([4], ht @yanjo115 on twitter, on the open pangram dataset) activates on "it's not just x, it's y" constructions. but you might notice the phrasings are all pretty different!
130
thebes @vgel.me · 01/07/2026
a spike of something-is-not-right, a pull downhill towards confabulation. the assistant both bounces around in this landscape and shapes how it unrolls in time with its thoughts...
130
thebes @vgel.me · 01/07/2026
well-modeled language has these high-dimensional conceptual hills and valleys of representation - which you can rep-read right out of the activations surprisingly well - a hill of user expertise here, a valley of dishonesty there,
140
thebes @vgel.me · 01/07/2026
@repligate.bsky.social posted these messages from opus 4.7 on discord, and it reminded me of a post from @yuxi.ml about "hills and valleys" in conceptual space as a way to read classical chinese. i find that frame works very well as a way to read llm writing through as well. (more below 🧵)
181
thebes @vgel.me · 30/06/2026
120
thebes @vgel.me · 17/06/2026
this matches with my experience that llms struggle not just with seeing, but with reasoning in visual spaces at all. though often ascii diagrams can bridge the gap, which is odd. reminds me of this jackendoff... i'd be interested to see a repeat where opus gets game state via ascii, or image + ascii
4122
thebes @vgel.me · 16/06/2026
This is funny, because today I'm also excited to introduce Taste Labs 👅 Our mission is to give AI models and agents taste, and today we’re coming out of stealth. AI has nailed objective domains and made it easy to generate anything. But it still feels off. Now, the challenge is taste.
9730
thebes @vgel.me · 16/06/2026
i have plenty of critiques of the rat discourse around ai but it is exceedingly funny how often someone drops a "yeah but the rats didn't think of THIS ONE now did they??" and it's a highly upvoted topic on lesswrong and active strand of biglab economic impact research
2814