tal yarkoni @talyarkoni.com · 19/02/2026like i can easily see every human author thinking most other authors just don't get what's cool or important about *their* work. but i can also see it going the other way, where everyone converges and the llms are out there doing their own thing. 130
tal yarkoni @talyarkoni.com · 19/02/2026the thought experiment i'd love to see turned real is: imagine you had 100 professional authors do this exercise for one another's books. then you had SoTA AIs do all 100. would people consistently think the AIs are worse than the humans? i genuinely don't know... but it isn't obvious they would 130
tal yarkoni @talyarkoni.com · 19/02/2026Claude has strong opinions about what helps: claude.ai/share/e2cf6f... whether they're actually correct, i don't knowclaude.aiSteering LLM writing style with compact promptsShared via Claude, an AI assistant from Anthropic 010
tal yarkoni @talyarkoni.com · 19/02/2026what happens if you try explicitly telling them they should be very specific, and feel free to make choices about content? i'm curious whether, viz your original post, it's really that you need to specify a huge amount of detail in the prompt, vs. finding short-cuts that "license" the model to do it 110
tal yarkoni @talyarkoni.com · 19/02/2026but if you look at how writers in general talk about other writers' work, it also has a lot of the same flavor... and i worry that people conflate "the llm can't generate good writing" with "actually i think almost no one else can generate writing i'd approve of, in this narrow context" 120
tal yarkoni @talyarkoni.com · 19/02/2026i guess this is where i kind of feel like the subjective and personal nature of it becomes really huge. i also have the experience that the models are bad at mimicking *my* writing. and yet i feel like they're pretty good at mimicking the style of other writers i like—where i'm not as invested? 140
tal yarkoni @talyarkoni.com · 19/02/2026interesting! how much of this do you think is llm-specific? meaning, do you think if you took a bunch of 100 human non-fiction authors and asked them to do the same task, you'd be happy with most of the results? 220
tal yarkoni @talyarkoni.com · 19/02/2026in a sense i think this maybe gets at what a remaining role of domain expertise is (for now). a lot of variance in writing style and creativity *is* already in the model weights... but someone who's read a lot is at a huge advantage in terms of eliciting interesting results from a compact prompt 010
tal yarkoni @talyarkoni.com · 19/02/2026i was recently working on a product where the stakeholders were like "this writing is AI slop", so i just changed the prompt to be much more stylized, and then the feedback turned into "it's too stylized and opinionated", which i think highlights the subjective nature of what "good" writing is 220
tal yarkoni @talyarkoni.com · 19/02/2026i'm struck by how often i see people complain about how generic the llm writing style is (not you, of course), and it's clear the critic hasn't even tried to direct the style in any way. of course it's going to give you the lowest common denominator--that's the smart thing to do! 120
tal yarkoni @talyarkoni.com · 19/02/2026i think that's true (and again generally so) in the sense that, absent a signal from the user, mode collapse is the sensible strategy. but at paragraph or short story level at least, you don't need a long prompt... you can just say "in the manner of Kafka" and get radically different results 120
tal yarkoni @talyarkoni.com · 19/02/2026i find it more natural to just mentally separate raw intelligence from concrete application (even though, in the limit, the distinction will collapse). a model can be very smart, but it still needs some organizing skills and structures to write a novel, because that's a very hard, specific thing 120
tal yarkoni @talyarkoni.com · 19/02/2026i guess, but you could say the same for almost any really difficult task? like, the model won't just solve clay problems for you (even if it could, with the perfect prompting), because it's trying to stay close to the prompt. it's true, but feels kind of empty? 110
tal yarkoni @talyarkoni.com · 19/02/2026from a (strictly) data standpoint i would guess that writing quality is actually one of the easier things to solve for, since most of the training data *is* writing. but distinguishing good from bad in a way that respects people's differing aesthetic preferences seems really hard 000
tal yarkoni @talyarkoni.com · 19/02/2026not really, a lot of the performance gains are now driven by RL, and the typical setup there is either human experts paid to generate good examples, or the models themselves generating candidates that humans (or other models) then evaluate (but also, the models train on huge book corpora) 100
tal yarkoni @talyarkoni.com · 19/02/2026crazy that the latest models from the big labs have all been minor version bumps on paper despite huge improvements in benchmark performance and qualitative feel. hard to imagine the labs releasing major versions that feel incremental at this point, which is... terrifying? 250
tal yarkoni @talyarkoni.com · 19/02/2026buuuut even then, i'm skeptical that any particular AI, short of genuine ASI, would ever be *widely* received as a good writer, for the same reason that even the most popular human writers are usually appreciated by only a small subset of people. it's just inherently subjective. 1120
tal yarkoni @talyarkoni.com · 19/02/2026which is just to say that, even if base model development halted in its current state, we probably *will* still get a good AI novelist in the next couple of years. it's just that it's not trivial to build the right harness, and not a lot of resources are being thrown at that particular problem. 270
tal yarkoni @talyarkoni.com · 19/02/2026of the 3, the harness is probably by far the easiest way to keep a narrative on the rails as more information is introduced. this is pretty much how human writers operate too! human novelists don't sit down and spew out 100K words linearly. it's an enormous process of iteration and refinement. 150
tal yarkoni @talyarkoni.com · 19/02/20263. it's arguably more of a harness problem than a base model problem. i think @moultano.bsky.social's point that information needs to come from somewhere is correct, buuut... it doesn't have to be in the prompt! it could be in the weights (hard for reason (2)), or it could be in the harness. 160
tal yarkoni @talyarkoni.com · 19/02/2026one way to think about it is that the manifold you need to learn in the latent space is much larger, because there are many more ways to be good. you probably need an incredible amount of feedback from expert judges, and you have the standard chaining problem, but with very few intermediate labels. 170
tal yarkoni @talyarkoni.com · 19/02/20262. for reasons related to the above, learning to produce coherent and high-quality long form writing via pretraining or RL is probably much harder than learning to generate code or even do math. lack of verifiability, and the inherent subjectivity of writing, really hurts you here. 170
tal yarkoni @talyarkoni.com · 19/02/2026whereas if AI writes a broadly coherent 120K-word novel, nobody is impressed unless they also like the quality of individual paragraphs. so an AI writer has to be able to write like almost *every* human writer to be "good", and it has to produce high quality writing fractally, *at every scale*. 190
tal yarkoni @talyarkoni.com · 19/02/20261. the problem is just way harder. good writing is mostly an aesthetic judgment, unlike code or science, where correctness is often verifiable. people are much less forgiving! if AI writes 120K words of code that *works*, devs are impressed even if they think it's insecure spaghetti code. 1100
tal yarkoni @talyarkoni.com · 19/02/2026this is interesting, but feels at best incomplete. i think there are at least 3 separate reasons we haven't yet seen consistently good writing from AI, at least in long form (microfiction is arguably mostly solved): 🧵 5214
tal yarkoni @talyarkoni.com · 19/02/2026i do think that there's ultimately an empirical question here, which is "will these kinds of errors prevent us from using LLMs to solve really hard problems", and it sounds like we have different predictions there. which is fine! i see no point in arguing about that, we can just wait and see. 100
tal yarkoni @talyarkoni.com · 19/02/2026and re: your question about what i'm saying more broadly, it's that i think you're oversimplifying the dominant perspectives on LLMs (on both sides). most "boosters" don't think LLMs no longer make errors, or that their ultimate utility depends on having guarantees about correctness. 100
tal yarkoni @talyarkoni.com · 19/02/2026no, i find their failures fascinating. but i think if you're serious about wanting to understand them, it's odd to say "i don't care about human errors". i'm telling you that these are very similar to the kinds of errors you see in humans, and are imo well understood as byproducts of the design. 100
tal yarkoni @talyarkoni.com · 19/02/2026which brings us to the second point: i don't think many "boosters" would say that we can't or shouldn't trust or deploy LLMs unless they come with guarantees. so if you want to make the point that LLMs aren't architecturally like calculators, well, sure. but IMO the frame here is misleading. 100
tal yarkoni @talyarkoni.com · 19/02/2026fwiw i think all of the examples you give here are actually cases where humans also routinely fail, and the reason is that they pit pragmatics against literal understanding, so that they are very much analogous to visual illusions—edge cases that fall out of the way the system is designed to work. 200
tal yarkoni @talyarkoni.com · 19/02/2026on the first point, i'm saying that given your history of holding up things like failure to do arithmetic as evidence LLMs aren't fit for purpose, maybe epistemic humility about the implications of failures like the ones in this thread is appropriate. 210
tal yarkoni @talyarkoni.com · 19/02/2026i think we should distinguish the question of whether errors will go away, or what the residual rate will be in the limit, from the question of what it implies for "boosterism", successful deployment of LLMs in the wild, etc. 100
tal yarkoni @talyarkoni.com · 19/02/2026i agree with that interpretation, and i'm saying that the lack of guarantees that LLMs will always do the right thing doesn't preclude them from being superintelligent or being exactly the kind of system we put in place to solve all our hardest problems. there are no guarantees on humans either! 110
tal yarkoni @talyarkoni.com · 18/02/2026sure, there are plenty of exceptions (and the article is explicit on this point, and names many!). but it seems pretty clear right now that AI is at risk of becoming right-coded, in the same way that the military and immigration are, and that would be a huge unforced error for the left 000
tal yarkoni @talyarkoni.com · 18/02/2026also, i missed that ftrain.com is updating regularly again! this makes me so happyftrain.comFtrain.com - Paul FordEssays and stories by Paul Ford. Since 1997. 000
tal yarkoni @talyarkoni.com · 18/02/2026characteristically lovely and reflective piece by @ftrain.bsky.social. if you want to better understand how the ground is shifting under every software engineer's feet right now (and soon, under pretty much *every* knowledge worker's feet), this is a great place to start 150
tal yarkoni @talyarkoni.com · 18/02/2026the irony of it is that the example in the OP is one that many if not most humans would *also* fail, so it isn't even a good example of blind spots being different. it's basically a particularly subtle example of the questions probed by the Cognitive Reflection Test—which most people fail! 020
tal yarkoni @talyarkoni.com · 18/02/2026agreed, though it's also worth noting that there are several cottage industries in psych and cognitive science devoted to identifying surprising new illusions/errors humans are susceptible to, so it's not *that* different from what we see with LLMs. 120
tal yarkoni @talyarkoni.com · 18/02/2026right, see the very next post. it's not logically incoherent, it's just wishful thinking, and betrays a failure to understand what the technology is already capable of. 100
tal yarkoni @talyarkoni.com · 18/02/2026i'm not sure the article ever encourages anyone to actually *use* AI? are we reading the same piece? the central argument is that ignoring the reality of AI in favor of wishful thinking puts the left in a very weak position to influence the regulation and deployment of the technology 110
tal yarkoni @talyarkoni.com · 18/02/2026i think the most obvious thing it would look like is simply not loudly arguing for positions (not just on social media—also in prominent op-eds) that in 2026 are absurd on their face 110
tal yarkoni @talyarkoni.com · 18/02/2026this is a good point, though i think it's pretty clear that the article is talking about left-wing intellectuals' views, not the public. (i think the reasonable worry here is that AI is new, and the positions intellectuals take will eventually filter down to the general partisan population.) 000
tal yarkoni @talyarkoni.com · 18/02/2026yeah it does seem like dem politicians are (thankfully) doing much better on this than left-wing intellectuals (and the article also explicitly notes this). and re: regulation, i think it's the intent/motivation that count at this point. sure, it's hard. but at least try to look like you care! 110
tal yarkoni @talyarkoni.com · 18/02/2026it seems likely that there will always be edge cases or adversarial examples that trip up neural architectures (though they'll continue to decrease in frequency and severity), just like humans are susceptible to sensory illusions that are essentially edge cases of useful affordances. but so what? 120
tal yarkoni @talyarkoni.com · 18/02/2026the example you gave strikes me as a pretty contrived one in that it explicitly leans on pragmatic implicature and the literal meaning of the words pushing in different directions. if this were the kind of thing that comes up often in real interactions, it would have been RL'd away already. 120
tal yarkoni @talyarkoni.com · 18/02/2026i don't see why? humans have all kinds of massive cognitive biases, and some errors literally can't be overcome—e.g., you can't easily teach your visual system to ignore the checker shadow Illusion. is the assumption here that a system can't be superintelligent if it makes *any* errors? 220
tal yarkoni @talyarkoni.com · 18/02/2026but also, the existence of failures doesn't preclude (super)intelligence, or using a model to solve incredibly difficult problems? humans fail all the time! one can easily imagine an alien civilization snorting at how humans can't be trusted because they can't learn to avoid basic visual illusions 000