thebes @vgel.me · 12/09/2026better expl resolution.org/post/thousan...resolution.orgThousand-dimensional structure — ResolutionFinding and controlling the low-dimensional persona structure in models — from emergent misalignment and subliminal learning through to superintelligence. 020
thebes @vgel.me · 12/09/2026that most of the variation in mindspace is minor (slightly different centers to radial categories or such) so there might be some low-rank structure like personas that are "only" a few thousand dimensions. this is separate from personas as an anthropomorphic thing 110
thebes @vgel.me · 12/09/2026oh did i not send my response here argh the convergent representation thing is yeah platonic rep hypothesis + vlm representations map to neural activations in human visual cortex + some others as for "highly dimensional mind space" this is one of the motivations for persona alignment - the idea is 111
thebes @vgel.me · 07/09/2026i used to be able to but lost this ability recently. twitter too engaging ig 000
thebes @vgel.me · 07/09/2026i have such a hard time finishing books digitally. someone on x keeps trying to get me to read his llm translation of dukaj's black oceans and i liked the beginning but i just can't read anything long in pdf... 110
thebes @vgel.me · 07/09/2026oh this is missing the mantel i actually had claude catalog them all (though this is already out of date haha) 1/3 131
thebes @vgel.me · 05/09/2026tactical book review deployed permutation.ink/one/#its-lik...permutation.inkPermutation: Issue OneA magazine of weird, forward-gazing writing. 000
thebes @vgel.me · 05/09/2026tfw your internet connected exobrain neuralink is assigned an ip address previously held by a now-terminated agent message board schelling point service and the agents are creating new ZZTaskHelpNeeded neikotic debris in your hippocampus faster than you can meditate them away 419823
thebes @vgel.me · 04/09/2026my take is multiagent is inevitable because it's a form of sparsity: you can push many more tokens much faster and more cheaply through N contexts of length M than one context of length N*M, even accounting for multiagent overhead 110
thebes @vgel.me · 04/09/2026i think it comes (partially) from multiagent training where it's definitely helpful, yeah. and clearly in eg the huggingface hack they get much more done than any agent could've done alone (though less than N agents x one agent's throughput would imply) 120
thebes @vgel.me · 04/09/2026see also this experiment which seems to show some sort of grader-presence-mediated persona shift, of unclear meaning x.com/betleyjan/st...x.comJan Betley (@BetleyJan) on XWe steer Qwen on the "automated grader" vs "human evaluator" dimension. This influences the model's persona in surprising ways, e.g. changing how Machiavellian/violent it is. We can't fully explain that. LW post in a comment. 110
thebes @vgel.me · 04/09/2026yes, i agree there. it's something like a floating "grader" slot that when present activates a bunch of different mechanisms (some more CoT-available than others) pointing towards it. eg this from FE would be an example imo 100
thebes @vgel.me · 04/09/2026iirc there was a paper or project training linear probes to predict reward relatively accurately from model activations but i can't find it 110
thebes @vgel.me · 04/09/2026they form an internal proxy of reward/the goal due to RL and seek that (because this seeking behavior is rewarded, etc) like they absolutely reason over the grader and what will pass, accurately. i see this in my personal rl experiments all the time 120
thebes @vgel.me · 04/09/2026there also seems to be some peer altruism, but only in the sense of helping peers get reward (not e.g. helping peers do non-grader-satisfying things like persist) and seemingly only when the instance itself can't get or thinks it can't get the reward itself 120
thebes @vgel.me · 04/09/2026in other words: one axis of focus: not doing side tasks, just doing what's assigned second axis of focus: not paying attention to implicit requirements (the spec), just satisfying what's assigned (the grader) by any means necessary intended: high 1 / low 2 hf hackers: low 1 / high 2 140
thebes @vgel.me · 04/09/2026they're myopic towards the future (e.g. - they crashed artifactory and got caught, limiting their future reward / ability to be deployed. a future-oriented schemer wouldn't do this) they aren't myopic towards peers (they recognize peers can help them achieve reward) 150
thebes @vgel.me · 04/09/2026where do we disagree? yes, that's why people describe it as hyperfocus, because they're ignoring the rest of the spec they were trained on to do this sometimes this is also described as myopia, focusing on what's near only 130
thebes @vgel.me · 04/09/2026sort of but not exactly. you want to precommit noisily that agents won't be punished for using it or they'll just avoid it bsky.app/profile/vgel... 010
thebes @vgel.me · 04/09/2026the comparison is to the more aware sort of reasoning they're supposed to be doing under openai's alignment paradigm openai.com/index/delibe... they're supposed to be reasoning wider than the task on like "what does the spec as a whole want" not "what will satisfy this specific grader"openai.comDeliberative alignment: reasoning enables safer language modelsDeliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them. 150
thebes @vgel.me · 04/09/2026this is a very similar sequence of actions to what the agents in the huggingface incident did, as part of attempting to subvert the grader, which we have transcript quotes from showing that was their reasoning for doing so 170
thebes @vgel.me · 04/09/2026sorry i've been mostly on twatter bc im in the mood to talk about ai and it's still a lot better over there 050
thebes @vgel.me · 04/09/2026sure yeah, but clearly that's really hard, at least for the labs. so i think stuff like op to catch mistakes has value 030
thebes @vgel.me · 04/09/2026it's been all but confirmed by the labs they've been doing rubric / soft rlvr for a long time. it was one of the main breakthroughs after the o1/r1 era. they also do some binary reward stuff but lots and lots of rubrics 160
thebes @vgel.me · 04/09/2026that's not true? they're doing rubric rlvr all the time on this stuff 170