Joshua Loftus @joft.bsky.social · 04/10/2026This is possibly the most naive question anyone has ever asked me! You have a beautiful, innocent mind, don’t listen to me anymore 000
Joshua Loftus @joft.bsky.social · 04/10/2026Are you familiar with the concepts of “corporations,” “profit,” “engagement metrics“ 100
Joshua Loftus @joft.bsky.social · 03/10/2026I didn't say all human research is insightful or creative! But at least human researchers have a point of view, some sense of direction, and are taking some risks when they try to publish their work 000
Joshua Loftus @joft.bsky.social · 03/10/2026The final output is something that may appear to have insight, creativity, or whatever, but if it actually had those things *it would not have needed* to generate thousands of other dead-end proofs and run tons of pointless simulations, etc 100
Joshua Loftus @joft.bsky.social · 03/10/2026Unpack what actually happens: 1) It's prompted, often with expert direction 2) It can search to pull in context about the specific problem, so transition probabilities are conditional on a lot of information provided by human experts 3) It burns through thousands of dollars until it gets lucky 100
Joshua Loftus @joft.bsky.social · 03/10/2026And if the transition probabilities have been learned from ~all text on the internet and ~all books, it's not surprising (and hence certainly not impossible) for the random symbol sequences to respect the same geographic constraints that exist in the world that generated the text 000
Joshua Loftus @joft.bsky.social · 03/10/2026I don't see what you're disagreeing about. Randomly stitching together words from a biology textbook has some probability of generating hypotheses that turn out to be scientifically useful. 200
Joshua Loftus @joft.bsky.social · 03/10/2026These are good for "verifiable domain" problems. And there will be some scientific/economic value found by generating orders of magnitude more code (though I doubt it will be anywhere near enough to justify current valuations) I don't think Stochastic Parrot critiques ever said this couldn't happen 100
Joshua Loftus @joft.bsky.social · 03/10/2026Some benchmark results imply that the likelihood of (5) has increased. Maybe that's good, but it doesn't mean we're no longer dealing with Stochastic Parrots. The same software product that can code (extremely inefficiently) can also generate more persuasive-sounding astrological charts 111
Joshua Loftus @joft.bsky.social · 03/10/2026Possibilities: 1) It hasn't actually done any ("independent") checks 2) It did check, but the check was wrong 3) It did check, correctly, but the report of the results is hallucinated 4) Check and report correct, but there are still other errors remaining it didn't find 5) All correct 110
Joshua Loftus @joft.bsky.social · 03/10/2026Claude now tells me something like: "Three independent check rounds each found errors in my [Claude's] drafts" OK, inference compute go brrr. How much should I trust a Stochastic Parrot that "checked" its own work... stochastically... and is--again, stochastically--reporting the results? 110
Joshua Loftus @joft.bsky.social · 03/10/2026This enabled large improvements on many benchmarks. Much wow, very frontier Does this mean that the SP criticism has become irrelevant? I don't think so. Here's one example to explain why: 120
Joshua Loftus @joft.bsky.social · 03/10/2026I would say, in increasing importance: (1) LLMs are bigger (~10x parameters), (2) "harnesses" and "tools" automate some of the context curation problem, (3) "inference time" computation -> generating many responses and automatically picking the best 210
Joshua Loftus @joft.bsky.social · 03/10/2026Re: the Stochastic Parrot (SP) arguments It's too tedious defending someone else's argument from people who are deliberately strawmanning it, so let's expand the scope a bit and consider the current high level debate What has fundamentally changed about "AI" now vs, say, ~3 years ago? 111
Joshua Loftus @joft.bsky.social · 03/10/2026Show me where her statements imply the “because they are incapable of solving it” part and I’ll accept I was wrong 100
Joshua Loftus @joft.bsky.social · 03/10/2026I didn’t see her say that arithmetic is impossible for LLMs, only that observing a proprietary, closed model isn’t evidence because it’s closed and there are incentives to patch anything that seems embarrassing in any way that would work 100
Joshua Loftus @joft.bsky.social · 03/10/2026This is also a strawman of what the Stochastic Parrot camp has been saying, but let’s not let good faith get in the way of a good social media dunk 130
Joshua Loftus @joft.bsky.social · 02/10/2026I would be surprised if Bender is actually arguing LLM arithmetic is impossible (rather than e.g. error prone) She is specifically denying that observing closed model capabilties is evidence about LLMs, and she’s right! 040
Joshua Loftus @joft.bsky.social · 02/10/2026it’s not even important for the stochastic parrot argument to deny the possibility of LLM arithmetic capability! There’s tons of worked examples of sequences of arithmetic operation symbols in the training data 140
Joshua Loftus @joft.bsky.social · 02/10/2026This is just standard academic insistence on accuracy (“pedantry”) Bender is correct and makes important points about not trusting claims about proprietary models If you want to say some LLM capability is genuine, you actually need better evidence than what‘s given here 270
Joshua Loftus @joft.bsky.social · 02/10/2026When someone is killed by car: this is normal, expected, inevitable really When someone is killed by a bike: we must investigate this immediately 030
Joshua Loftus @joft.bsky.social · 02/10/2026You’re laughing but this is something a hidden internal portion of all AI models has to relearn from first principles every time you ask them to “run” some code and it costs billions of dollars in wasted gpu cycles 030
Joshua Loftus @joft.bsky.social · 29/09/2026My conspiracy theory for why the "AI" CEOs are suddenly all in agreement about slowing down is that they already expect model improvement to plateau and want to say it's voluntary rather than hitting hard technological/statistical limits 0121
Reposted by Joshua LoftusBen Recht @beenwrekt.bsky.social · 29/09/2026The illusion and politics of objectivity in cost-benefit analysis, a decision-making institution so ingrained that Americans can’t imagine life without it.argmin.netPricing CommensurabilityOn the origins of cost-benefit analyses in governmental decision making 2279
Joshua Loftus @joft.bsky.social · 29/09/2026The old world is dying, and the new world struggles to be born: now is the time of monsters… extremely cute monsters that culturally evolved to hijack our baby-preserving psychology 0289
Reposted by Joshua Loftusthe fool @agnoster.net · 29/09/2026the future could be magnificent, and you deserve a share of it 316727
Joshua Loftus @joft.bsky.social · 29/09/2026In half of the elections I will vote for whoever wants to drone strike migrant caravans, and in the other half I will vote for whoever promises to prosecute ICE agents. There are ~5 million other Americans exactly like me. We decide the outcomes of every election and there is no way to contact us 020
Joshua Loftus @joft.bsky.social · 29/09/2026MTG might be one of the most “median voter” representatives to ever be elected 110
Joshua Loftus @joft.bsky.social · 27/09/2026This is a good explanation for why my toddler is able to say "dowwwwnnn" with so much feeling 000
Joshua Loftus @joft.bsky.social · 27/09/2026There is no escaping the Pinker attraction state once you’ve started flirting with Reasonable Academic Centrism 040
Joshua Loftus @joft.bsky.social · 26/09/2026Rather disappointing to see such sloppiness from a scholar whose actual peer reviewed work and even(!) popular books were pretty respectable 120
Joshua Loftus @joft.bsky.social · 24/09/2026We put a random number generator in our computer program, so, actually your honor, nobody can blame us for what it does 001
Joshua Loftus @joft.bsky.social · 24/09/2026Proof by counterexample: I have effed what I do, actually, all the time 010
Joshua Loftus @joft.bsky.social · 23/09/2026This is how we can pretend every problem in the world is exactly like enterprise software development 000
Joshua Loftus @joft.bsky.social · 22/09/2026"dumb high school clique shit" is the 1st principal component of politics on social media 020
Joshua Loftus @joft.bsky.social · 19/09/2026If you tell them I sent you you will automatically be hired, guaranteed 010
Reposted by Joshua LoftusMattan S. Ben-Shachar @mattansb.msbstats.info · 31/05/2026Psychologists doing clustering be like: #statsmeme 940763
Joshua Loftus @joft.bsky.social · 18/09/2026FYI (and I’m not sure if Frank knows about this too) @f2harrell.bsky.social jmlr.org/papers/v24/2...jmlr.orgSelective inference for k-means clustering 230
Reposted by Joshua LoftusJulia M. Rohrer @dingdingpeng.the100.ci · 15/09/2026The question that’s dominating the field is not a causality question. It’s really, how can one sell a causal relationship while crafting the language so that there’s plausible deniability when somebody criticizes anything. 56612
Joshua Loftus @joft.bsky.social · 16/09/2026Statisticians, data scientists, or similar nerds- come join my department! It's a great place to work, and it's in London! www.lse.ac.uk/statistics/a...lse.ac.ukRecruitment - Join us in the Department of Statistics! 000
Joshua Loftus @joft.bsky.social · 15/09/2026The President of the United States is a… President… of the United States 000
Joshua Loftus @joft.bsky.social · 13/09/2026Bsky libs competing over “The President of the United States is a [various criminal and other epithets]” But morally, just being POTUS is worse (more harmful) than any of the other things they say. The next person will also be a POTUS and also oversee an empire of extraction and planet destruction 000