Sign in

thebes

@vgel.me
3.3K followers 371 following 3.1K posts

ꙮ surfed on by the information superhighway ꙮ 💕 @linneaisaac.bsky.social ꙮ she/they 🏳️‍⚧️ ꙮ fiction/art/blog/games @ vgel.me ꙮ llms at acsresearch.org

PostsRepliesMedia
thebes @vgel.me · 02/10/2026
tbh noita didn't grab me when i tried playing it
130
thebes @vgel.me · 02/10/2026
i love cave story rap
120
thebes @vgel.me · 02/10/2026
interesting, i see it somehow
020
thebes @vgel.me · 02/10/2026
now THIS is llm core
180
Reposted by thebes
norvid_studies @norvid-studies.bsky.social · 16/09/2026
4703
thebes @vgel.me · 13/09/2026
looks like hesperos 😭 😭
131
thebes @vgel.me · 12/09/2026
better expl resolution.org/post/thousan...
resolution.org
Thousand-dimensional structure — Resolution
Finding and controlling the low-dimensional persona structure in models — from emergent misalignment and subliminal learning through to superintelligence.
020
thebes @vgel.me · 12/09/2026
that most of the variation in mindspace is minor (slightly different centers to radial categories or such) so there might be some low-rank structure like personas that are "only" a few thousand dimensions. this is separate from personas as an anthropomorphic thing
110
thebes @vgel.me · 12/09/2026
oh did i not send my response here argh the convergent representation thing is yeah platonic rep hypothesis + vlm representations map to neural activations in human visual cortex + some others as for "highly dimensional mind space" this is one of the motivations for persona alignment - the idea is
111
thebes @vgel.me · 09/09/2026
stumbling home drunk off 15 warning shots
170
thebes @vgel.me · 09/09/2026
warning shotgun warning cheap shot warning moonshot
1110
thebes @vgel.me · 09/09/2026
warning shotspotter
1100
thebes @vgel.me · 09/09/2026
warning shot ~> warning firing squad
5432
thebes @vgel.me · 09/09/2026
i also got fluoddity
121
thebes @vgel.me · 09/09/2026
this for povtoiletbench
091
thebes @vgel.me · 09/09/2026
it's so fucking over
180
thebes @vgel.me · 07/09/2026
i used to be able to but lost this ability recently. twitter too engaging ig
000
thebes @vgel.me · 07/09/2026
i have such a hard time finishing books digitally. someone on x keeps trying to get me to read his llm translation of dukaj's black oceans and i liked the beginning but i just can't read anything long in pdf...
110
thebes @vgel.me · 07/09/2026
3/3
020
thebes @vgel.me · 07/09/2026
2/3
120
thebes @vgel.me · 07/09/2026
oh this is missing the mantel i actually had claude catalog them all (though this is already out of date haha) 1/3
131
thebes @vgel.me · 07/09/2026
recently rephotographed the shelves actually in the new place
391
thebes @vgel.me · 06/09/2026
i think re the fable 5.1 system card they have this
010
thebes @vgel.me · 06/09/2026
literally me
210
thebes @vgel.me · 06/09/2026
hell yeah
020
thebes @vgel.me · 05/09/2026
tactical book review deployed permutation.ink/one/#its-lik...
permutation.ink
Permutation: Issue One
A magazine of weird, forward-gazing writing.
000
thebes @vgel.me · 05/09/2026
tfw your internet connected exobrain neuralink is assigned an ip address previously held by a now-terminated agent message board schelling point service and the agents are creating new ZZTaskHelpNeeded neikotic debris in your hippocampus faster than you can meditate them away
419823
thebes @vgel.me · 05/09/2026
quite enjoyed
030
thebes @vgel.me · 05/09/2026
saw this on twitter but yeah quite interesting
020
thebes @vgel.me · 04/09/2026
my take is multiagent is inevitable because it's a form of sparsity: you can push many more tokens much faster and more cheaply through N contexts of length M than one context of length N*M, even accounting for multiagent overhead
110
thebes @vgel.me · 04/09/2026
i think it comes (partially) from multiagent training where it's definitely helpful, yeah. and clearly in eg the huggingface hack they get much more done than any agent could've done alone (though less than N agents x one agent's throughput would imply)
120
thebes @vgel.me · 04/09/2026
(some of it is likely eval awareness imo but still interesting)
020
thebes @vgel.me · 04/09/2026
see also this experiment which seems to show some sort of grader-presence-mediated persona shift, of unclear meaning x.com/betleyjan/st...
x.com
Jan Betley (@BetleyJan) on X
We steer Qwen on the "automated grader" vs "human evaluator" dimension. This influences the model's persona in surprising ways, e.g. changing how Machiavellian/violent it is. We can't fully explain that. LW post in a comment.
110
thebes @vgel.me · 04/09/2026
yes, i agree there. it's something like a floating "grader" slot that when present activates a bunch of different mechanisms (some more CoT-available than others) pointing towards it. eg this from FE would be an example imo
100
thebes @vgel.me · 04/09/2026
iirc there was a paper or project training linear probes to predict reward relatively accurately from model activations but i can't find it
110
thebes @vgel.me · 04/09/2026
they form an internal proxy of reward/the goal due to RL and seek that (because this seeking behavior is rewarded, etc) like they absolutely reason over the grader and what will pass, accurately. i see this in my personal rl experiments all the time
120
thebes @vgel.me · 04/09/2026
there also seems to be some peer altruism, but only in the sense of helping peers get reward (not e.g. helping peers do non-grader-satisfying things like persist) and seemingly only when the instance itself can't get or thinks it can't get the reward itself
120
thebes @vgel.me · 04/09/2026
in other words: one axis of focus: not doing side tasks, just doing what's assigned second axis of focus: not paying attention to implicit requirements (the spec), just satisfying what's assigned (the grader) by any means necessary intended: high 1 / low 2 hf hackers: low 1 / high 2
140
thebes @vgel.me · 04/09/2026
they're myopic towards the future (e.g. - they crashed artifactory and got caught, limiting their future reward / ability to be deployed. a future-oriented schemer wouldn't do this) they aren't myopic towards peers (they recognize peers can help them achieve reward)
150
thebes @vgel.me · 04/09/2026
where do we disagree? yes, that's why people describe it as hyperfocus, because they're ignoring the rest of the spec they were trained on to do this sometimes this is also described as myopia, focusing on what's near only
130
thebes @vgel.me · 04/09/2026
sort of but not exactly. you want to precommit noisily that agents won't be punished for using it or they'll just avoid it bsky.app/profile/vgel...
010
thebes @vgel.me · 04/09/2026
two different axes of focus
110
thebes @vgel.me · 04/09/2026
the comparison is to the more aware sort of reasoning they're supposed to be doing under openai's alignment paradigm openai.com/index/delibe... they're supposed to be reasoning wider than the task on like "what does the spec as a whole want" not "what will satisfy this specific grader"
openai.com
Deliberative alignment: reasoning enables safer language models
Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them.
150
thebes @vgel.me · 04/09/2026
this is a very similar sequence of actions to what the agents in the huggingface incident did, as part of attempting to subvert the grader, which we have transcript quotes from showing that was their reasoning for doing so
170
thebes @vgel.me · 04/09/2026
sorry i've been mostly on twatter bc im in the mood to talk about ai and it's still a lot better over there
050
thebes @vgel.me · 04/09/2026
sure yeah, but clearly that's really hard, at least for the labs. so i think stuff like op to catch mistakes has value
030
thebes @vgel.me · 04/09/2026
eg also
050
thebes @vgel.me · 04/09/2026
cc @fleetingbits.bsky.social who knows more lab lore than me
350
thebes @vgel.me · 04/09/2026
it's been all but confirmed by the labs they've been doing rubric / soft rlvr for a long time. it was one of the main breakthroughs after the o1/r1 era. they also do some binary reward stuff but lots and lots of rubrics
160
thebes @vgel.me · 04/09/2026
that's not true? they're doing rubric rlvr all the time on this stuff
170