Quantian @quantian.bsky.social · 17hSure, and randomly stitching together symbols from a math textbook has some probability of being a novel proof. But being able to consistently generate dozens or hundreds of such proofs is very strong evidence that you are not randomly stitching together symbols but rather are doing mathematics. 101
Quantian @quantian.bsky.social · 18hIn particular their definition would categorically exclude emergent behavior we have observed in LLMs, including arithmetic representations/grokking, world models like positional knowledge of countries and streets, and instruction following and planning behavior. 100
Quantian @quantian.bsky.social · 18hHere is the original “stochastic parrot” definition. It claims that LLMs cannot have “any model of the world” and “haphazardly stitches together sequences of linguistic forms observed in its training data”. That is much stronger than “good for verifiable domains [and having] some scientific value” 120
Quantian @quantian.bsky.social · 20h5 in particular is relevant, because it suggests the valuable thing in LLMs is not parroting the training data but emergent large-model behavior like instruction following and efficient tool use, and you want to teach small models to do *that* directly and not learn Paris is the capital of France. 120
Quantian @quantian.bsky.social · 20h(5) related, the idea that instead of training many models you could train one huge model and then use its output as training examples for a smaller model, bypassing expensive pretraining and allowing the creation of eg Opus 5.5 or GPT 6 Sol (or GLM 5.3!) on the cheap 110
Quantian @quantian.bsky.social · 20h(4) Pretrain/RL on AI generated data allowed the models to get arbitrarily good on coding and math problems because you can created infinite data to tune the model, which is something that many people three years ago explicitly argued was impossible (remember “model collapse”, “Habsburg AI”, etc.?) 120
Quantian @quantian.bsky.social · 21hShe is implying there’s a secret calculator in closed models (that open models don’t need somehow) and the fact that there’s errors (which are quite rare) are more likely the be text classifier mistakes than the model having innate ability but occasionally messing up. Not very occam’s razor of her! 030
Quantian @quantian.bsky.social · 23hHow much energy and water was wasted developing a phone powerful enough that could be used as a calculator and internet browser? Doing this kind of proof by infinite descent doesn’t really rescue your argument lol 130
Quantian @quantian.bsky.social · 03/10/2026Bender is clearly arguing that modern LLMs are secretly running a text classifier and routing arithmetic to a calculator because they are incapable of solving it. My post might be a straw man of some arguments, but it is not a straw man of hers! She is literally making that exact claim! 160
Quantian @quantian.bsky.social · 03/10/2026This was literally calculated on my phone’s actual hardware. It took no more energy or water than opening Bluesky and scrolling for 20 seconds. Wolfram sent the query to a server for processing! 110
Quantian @quantian.bsky.social · 03/10/2026It is also interesting (to me and six holdouts at DeepMind and exactly nobody else) that *diffusion models* can also perform shockingly well at math tasks despite not doing next token prediction at all. It's like how early image models learned basic physics facts before they could draw hands well! 21137
Quantian @quantian.bsky.social · 03/10/2026I still think Gemma is a “better” model than Qwen primarily for this reason even if it scores worse on benchmarks. It’s just so much more token efficient and “thoughtful” instead of talking to itself forever. Can’t wait for the Gemma 5 distilled from Argon 150
Quantian @quantian.bsky.social · 03/10/2026Reasoning was turned off, this is just raw token output with no think tag. 0160
Quantian @quantian.bsky.social · 03/10/2026Liking Claude 2 prose is a crazy pull but I kind of get it, that’s like saying the OG StableDiffusion where the outputs were much more random was better than the newer models where they sanded out all the edges and converged on the house AI style every model has nowadays 1270
Quantian @quantian.bsky.social · 03/10/2026This was going around Twitter today, if you ask Claude “is X lat Y long land or water?” and plot the results, it has clearly embedded a map of the globe in its weights over time 512415
Quantian @quantian.bsky.social · 03/10/2026Yes I am agreeing with you and disagreeing with Bender, to be clear 030
Quantian @quantian.bsky.social · 03/10/2026For what it’s worth, arguing LLMs can’t do addition now is flat earth-level science denialism. Here’s Gemma 4b, a non-reasoning model that could only theoretically memorize binary operations up to 65535 x 65535, oneshotting a eight digit addition and hex conversion *on my phone* with no tool calls 1840738
Quantian @quantian.bsky.social · 02/10/2026Elon Musk has broken up with baby mommy 9/10/12/14, announced via Twitter unfollow alert post sponsored by Kalshi 3739455
Quantian @quantian.bsky.social · 02/10/2026By far the maddest I ever got people on here was when i called digital art slop and then when the digital artists showed up to yell at me in the replies I would just post something from their gallery back at them lmao 1410
Quantian @quantian.bsky.social · 01/10/2026GLM 6.0 causes the DataKrash, Anthropic becomes NetWatch after the Choosin Texas singer nukes OpenAI HQ, and a lack of a ready supply of internet cat videos forces everyone to attend underground techno raves in abandoned data centers where they get lead poisoning and become violent criminals 2524
Quantian @quantian.bsky.social · 01/10/2026You could maybe make a palatable daiquiri with it? Idk 000
Quantian @quantian.bsky.social · 01/10/2026Like if you ever had a bad bottle of Veuve I’m pretty sure your regional LVMH rep would fly you to Paris and Bernard Arnault would personally apologize to you, which is more than any Bordeaux producer would ever do if your $5000 bottle of Petrus was corked or whatever 110
Quantian @quantian.bsky.social · 01/10/2026I would rather drink Veuve than any similarly priced non champagne sparkling wine, and like all big champagne houses they are past masters of technical consistency so you’ll never ever ever have a bad bottle. 440
Quantian @quantian.bsky.social · 01/10/2026I'm the worlds only Veuve defender. It's perfectly fine for a $40 champagne, and the Grande Dames are actually a fairly compelling prestige cuvee, better than Dom IMO. 180
Quantian @quantian.bsky.social · 01/10/2026Closest thing I've had to it would be a very bad and unaged rhum agricole. It would ironically be improved if it had more savory or earthy tones like people complain about, as is it's too floral and unbalanced so it tastes like drinking cheap perfume crossed with high-fructose corn syrup but dry. 0150
Quantian @quantian.bsky.social · 01/10/2026I had Moutai for the first time and the tasting note everyone gives is... totally wrong? I mean it's definitely terrible, but it is *not* at all like rubbing alcohol or gasoline, it's not even barrel proof at 53%! It just has this horrible rotting banana flower ester-y nose and oily sweet palate. 9240
Quantian @quantian.bsky.social · 30/09/2026Sometimes I come across a screenshot of a normie using ChatGPT with memory on (viz. this from TikTok via Twitter) and people get crazy cooked even just through bare chatbot interactions. Persistent interaction with long-running agents via texting/voice will be another 4o style mass psychosis event. 108012
Quantian @quantian.bsky.social · 30/09/2026Chinese AI models are famously averse to answering questions about Tianmen Square and are specially trained to deflect or refuse to answer them if prompted 3290
Quantian @quantian.bsky.social · 30/09/2026LMFAO the official America.gov AI website uses a Chinese open source model 27668154
Quantian @quantian.bsky.social · 30/09/2026Unfortunately we’re meeting them in the middle on that one 0170
Quantian @quantian.bsky.social · 30/09/2026The Iranian letter to America is a very good piece of propaganda—certainly better than anything e.g. Russia is capable of producing right now—but there’s still a few points where the mask accidentally slips 6997
Quantian @quantian.bsky.social · 30/09/2026What if you count the French for 3/5ths of a regular person like you should? 0240
Quantian @quantian.bsky.social · 30/09/2026Yes Opus 5 was a terrible model, very pleased they turned things around so I could stop burning one billion Fable tokens on trivial tasks to avoid interacting with it 021
Quantian @quantian.bsky.social · 29/09/2026Opus 5.5 is basically the only model you should be using right now, just tune the effort. 4150
Quantian @quantian.bsky.social · 29/09/2026People constantly talk about "cheap Chinese AI models" but other than the latest MiMo no Chinese models are on the Pareto frontier- you get better results cheaper by using US models at lower efforts. Qwen in particular is so token hungry that 3.8 Max is *more expensive than Fable* for worse output! 91149
Quantian @quantian.bsky.social · 29/09/2026I assume the Anthropic IPO will be one zillion jillion times oversubscribed, pop 50-100% on day 1, bleed out down 50-75% over the next eight months, and finally rally back from the dead, aka the exact same pattern every IPO has had for the past five years 1032516
Quantian @quantian.bsky.social · 29/09/2026Oh they never have released them as far as I know, this is purely like a vibes thing based on the relative cost of inference vs open models 070
Quantian @quantian.bsky.social · 29/09/2026Isn’t Sonnet like a O(100B) model or so vs Astra being O(10T)? 280
Quantian @quantian.bsky.social · 29/09/2026There’s basically no reason to ever use or even release Sonnet 5.5—it’s dominated by Opus 5.5 on both cost and performance—but stunting on Astra with a model that outperforms it with maybe 1% as many parameters is really funny and petty of Anthropic 7950
Quantian @quantian.bsky.social · 28/09/2026Friendly reminder that the Platonic ideal of a 10-year treasury pays a 6% coupon: 7846
Quantian @quantian.bsky.social · 28/09/2026No he literally does this, he's said it in interviews www.instagram.com/reel/Ddj38uC...instagram.comInstagramCreate an account or log in to Instagram - Share what you're into with the people who get you. 1170
Quantian @quantian.bsky.social · 28/09/2026"What if the only way you could get home at night would be to ask the Nvidia GPU in your self-driving car" is a completely self-consistent good idea in Jensen's mind, as it will increase the TAM for Nvidia GPUs. He simply didn't have any other considerations than that when he said it. 536418
Quantian @quantian.bsky.social · 28/09/2026Jensen is probably the most Jobs-like CEO left in the industry, in that he has a completely monomaniacal focus on producing and selling more GPUs no matter what. Things which are not relevant to that goal, such as knowing where you live, are discarded to fit more GPU sales ideas in his brain. 2179871