Sign in

Phil

@phil2345.bsky.social
69 followers 132 following 149 posts

I'm just a reply guy so please don't follow me. Also, if you like Trump, then I don't like you.

PostsRepliesMedia
Phil @phil2345.bsky.social · 22/09/2026
They're all generally regressing as they chase benchmarks and coding gains. For example, on arena Opus 4.6 is ranked higher than 4.7, which is higher than 5. Same goes for OpenAI's models. Across tasks they're making more mistakes, such as story contradictions, b/c they're overtraining coding/agents
000
Phil @phil2345.bsky.social · 15/08/2026
It does impact quality. It's rare for tokens in a pool to have the exact same probabilities, so if you select a setting that's both random, but always favors the most probable token, then the resulting outputs will have more factually hallucinations, story contradictions etc.
000
Phil @phil2345.bsky.social · 14/08/2026
Luna is specialized for certain tasks, but is generally much weaker, such as ~70 across most arena categories, and 73 in creative writing. Even Gemini 3.6 flash handily outperforms Luna with ~20 in most categories, and 11 in creative writing.
000
Phil @phil2345.bsky.social · 14/08/2026
The other models are the top multi-trillion parameter LLMs on earth. Gemini 3.7 Flash is nearly as good at a tiny fraction of the size. Plus Anthropic, OpenAI, and Qwen attempted to release similar sized models and they're horrible by comparison.
240
Phil @phil2345.bsky.social · 12/08/2026
When they cheer stopping people because of the color of their skin or what they're wearing to check if they're a citizen, and sometimes still lock them up, it's not about differing political viewpoints. It's about xenophobia, transphobia, and just being a bunch of easily manipulated bigoted pricks.
000
Phil @phil2345.bsky.social · 07/07/2026
Just one witness. Adults who experience a traumatic event like this don't always call the cops since that opens up a can of worms and drastically changes their life, but they do inevitably at least cryptically tell friends & family he's a really bad man to protect them from being victimized etc.
100
Phil @phil2345.bsky.social · 07/07/2026
Yes, but a single accusation sans corroboration by friends & family (they inevitably tell someone close to them) isn't the same as a pattern. We can't let a single accusation, which often is a fabrication, to wipe out political candidates. This isn't the same as the flood against Trump, Como etc.
110
Phil @phil2345.bsky.social · 03/07/2026
If Three's Company goes up, and 2.5 Men goes down, then a completely brainless show about silly misunderstandings that was only aired as a deal to get MASH made would be higher than a show with skilled actors capable of near perfect comedic timing. 2.5 Men was ribald, but not hammy and brainless.
100
Phil @phil2345.bsky.social · 24/06/2026
Excluding politics mixed with religion, xenophobia, bigotry, misogyny... to ban groups like LGBT+ of getting married, playing in sports... isn't bias. As the saying goes reality has a liberal bias. Bigotry, ignorance... results in blatant contradictions & hypocrisy AI models can't substantiate.
000
Phil @phil2345.bsky.social · 15/06/2026
The quality is horrific. All the talk about selectively quantized important parts is nonsense. That just trades the preservation of some abilities (e.g. coherency) for a huge degradation in others (e.g. factual recall). You're MUCH better off running MUCH smaller models at Q4+.
010
Phil @phil2345.bsky.social · 13/06/2026
What people don't seem to realize about smaller models is that while you can squeeze a lot from a single domain into them, they still lack reliable knowledge retrieval. So for example, they can't see past variably worded prompts w/ things like spelling & grammar errors like large models.
010
Phil @phil2345.bsky.social · 09/06/2026
Like Anthropic and OpenAI have already been doing for the US military?
130
Phil @phil2345.bsky.social · 01/06/2026
It wasn't just Elon. Genders have always been split (restrooms, sports, jails...) & kids protected (prepubescent transition is ideal). So trans acceptance was always going to give fascism a huge boost & threaten liberal democracy. I'm 100% pro trans rights & am simply explaining the situation.
100
Phil @phil2345.bsky.social · 20/05/2026
LLMs are a lot like people in this regard. You can either make them experts or a jack of all trades, but master of none. Mainly b/c as training progresses every backpropagation is slightly more destructive until net loss & gain reach parity (the max info density of a set # of parameters is hit).
000
Phil @phil2345.bsky.social · 20/05/2026
Every code version of a model (e.g. Qwen3-coder) regressed significantly across non-coding domains & experienced a large spike in factual hallucinations. This isn't unique to coding. Training heavily in any domain (e.g. poetry) scrambles the weights of all other domains, esp. with smaller LLMs.
100
Phil @phil2345.bsky.social · 20/05/2026
Would you judge a scientist, writer, musician... for not being a great coder? Why has AI become synonymous with coding?Modern LLMs are absurdly lopsided and are vomiting hallucinations across most domains. Companies need to start releasing general purpose AI models and separate coder versions.
100
Phil @phil2345.bsky.social · 22/04/2026
I'm sorry. They should have called it Qwen3.5 Coder. Q3.6 has regressed in other areas so they could over-train coding, including a notable increase in factual hallucinations across domains. The monolithic coding obsessed early adopter community is KILLING the general purpose OS AI model ecosystem.
000
Phil @phil2345.bsky.social · 17/04/2026
Anthropic overfit coding going from Opus 3 to 4, resulting in a broad regression (e.g. a huge drop in its SimpleQA score), so coders flocked to it. But since LLMs are a balancing act (training in one area degrades others) any attempt to restore balance degrades its overfit coding abilities.
000
Phil @phil2345.bsky.social · 11/04/2026
But if you look at its breakdown it got 27 in Expert, 12 in math, 9 in writing, 10 in IR, and 22 in long query. It's clear that its high ranking comes from its friendly nature and telling users what they want to hear vs actual performance. The humans using lmsys are really starting to annoy me.
010
Phil @phil2345.bsky.social · 01/04/2026
That's a good question. The size per performance (at least test scores) appears to be the same, or even worse. But it may produce greater token/s and use less power.
000
Phil @phil2345.bsky.social · 01/04/2026
It still appear that 4-bit models perform better at the same size. This is because the test scores are virtually the same after Q4_K_M quantization, which is only ~3x larger than these 1.3 bit LLMs, and the 4-bit versions of 3b models score a bit better and are only a bit larger.
350
Phil @phil2345.bsky.social · 31/03/2026
The decision is awful, but was near unanimous b/c it's about the slippery slop of the govt getting involved in medicine/therapy. However, they made an exception for abortion, so why not this? The real problem is the adult LGBTQ & psychopathic parents who opt in for the such therapy.
120
Phil @phil2345.bsky.social · 17/03/2026
People are stupid. Qwen 35b is incredibly weak. It hallucinates like crazy about very basic & popular things, writes repetitive stories that blatantly contradict themselves and the user's prompt, can't even begin to write poetry... It's just overfit to select domains to fool coding obsessed morons.
000
Phil @phil2345.bsky.social · 15/03/2026
That's more than just nuts. He also thinks nuclear bombs are about a million times more powerful than they are.
001
Phil @phil2345.bsky.social · 22/02/2026
I suspect part of this is due to how grossly over represented coding is in AI training. They not only train on trillions of coding tokens, but end training runs on them, resulting in a selective boost in coding performance while performance in most other domains is abysmal & inconsistent (AI slop).
110
Phil @phil2345.bsky.social · 30/01/2026
Ellos obtienen mucho dinero para mantener un navegador por beneficios como la colocación predeterminada de la página de inicio. La única razón por la que el dinero es un problema para Mozilla es porque están desperdiciando dinero en cosas como un CEO y la AI.
000
Phil @phil2345.bsky.social · 27/01/2026
Mozilla has the money to keep the open source Firefox browser going for decades, but instead it's going to burn through its cash on a brainless CEO's high salary and redundant AI development that will fail to compete with anything.
081
Phil @phil2345.bsky.social · 12/12/2025
It wasn't the fault of Democrats. >95% were adamant in restoring abortion rights, and >95% of Republicans were adamant about not restoring them. So blaming Dems makes no sense. They didn't have the votes. It was Republicans and centrists (Rep/Dem) that made it impossible to restore abortion right.
210
Phil @phil2345.bsky.social · 12/12/2025
Liberals need to be more pragmatic. We would still have Roe, advanced trans right... if millions of liberals didn't stay home with their vibrators & video games because of stunted 'both sides aren't perfect', 'Hillary is flawed'... thinking. I repeat, a trans activist saying never Gavin is low IQ.
500
Phil @phil2345.bsky.social · 12/12/2025
That's not what happened. Abortion rights were lost when Trump and the Republicans appointed Supreme Court justices in his first term. Biden being president and the Democrats controlling both houses are completely irrelevant to the Supreme Courts ending of Roe.
220
Phil @phil2345.bsky.social · 12/12/2025
Trump came to power and a broad spectrum of rights, including abortion and trans rights, were reduced because of low IQ nonsense like this. When you keep reminding people you're a never Newsome voter you're just shooting yourself and trans rights in the foot, you're just too stupid to realize it.
900
Phil @phil2345.bsky.social · 02/12/2025
Also, the only model that actually made gains in the last year was the new Gemini. GPT5 was an overall regression from GPT4 (better at math/coding, little worse at all else). And the new Opus has less knowledge and broad stability than the previous Sonnet (basically only good for coding).
010
Phil @phil2345.bsky.social · 02/12/2025
2 of 2) So in short, there was no overall improvement in >1 year. All Alibaba did was trade general knowledge & abilities for tiny gains in select domains and the industry as a whole is oblivious b/c they're not using the models & are simply running highly flawed automated tests to judge models.
000
Phil @phil2345.bsky.social · 02/12/2025
1 of 2) That's the problem. The industry is just looking at tiny gains on select tests. I guarantee not a single researcher used Qwen3 for general AI use, or even for its overfit domains like coding since Claude, Gemini... are far better and more reliable, plus inexpensive to use.
010
Phil @phil2345.bsky.social · 02/12/2025
They didn't loose to Qwen. Qwen3 overfit to select domains & tests and fall apart when subjected to a broad spectrum of tasks. Alibaba also started flat out cheating on tests like SimpleQA. Small models hit a ceiling over a year ago. Progress has been faked ever since (overfitting/cheating).
100
Phil @phil2345.bsky.social · 16/11/2025
So fighting against taking healthcare away from 10s of millions and becoming the only first world country on the planet without guaranteed healthcare for all its citizens is just "belief polarization" and becoming an extreme version of yourself?
020
Phil @phil2345.bsky.social · 02/10/2025
Yes, but the "basket of deplorables" comment was unbelievably stupid, regardless of how true it is. They both saw what was coming, but expressed it in counterproductive ways.
110
Phil @phil2345.bsky.social · 01/10/2025
Sonnet 4.5 knows very little about art, pop culture, humor... It's grossly overfit to coding, math... Even the new Opus has less broad knowledge and abilities than the old Sonnet. It's more of a corporate tool than a general purpose AI model.
030
Phil @phil2345.bsky.social · 24/09/2025
Alibaba trains on test data. Don't trust their scores.
000
Phil @phil2345.bsky.social · 21/09/2025
You shouldn't report it unless something comes of it. The Trump administration is throwing a lot nonsense out there, including a 15 billion lawsuit to gain attention, which you just gave them.
110
Phil @phil2345.bsky.social · 20/09/2025
Please work on your framing. There's a HUGE difference between the FCC chair and the administration calling for the cancelling of the show, and broad pressure from people.
080
Phil @phil2345.bsky.social · 20/09/2025
When one side is using a tactic to fight to keep affordable health care, and the other to take healthcare away from millions, then calling out Ted back then wasn't hypocrisy. The primary issue the Democrats had was Ted using such a tactic for the cold hearted goal of taking healthcare away.
040
Phil @phil2345.bsky.social · 16/09/2025
I think he was blinded by love while fighting the self-loathing of being in a trans relationship. You see the same thing w/ conservative raised homosexuals who often go through a self loathing & virulent homophobia stage. He was too distracted by this to see the damage to her & the trans community.
020
Phil @phil2345.bsky.social · 11/09/2025
This was clearly a false flag. A single shot from 200 yards away, able to escape... is a professional hit. And no antifascist or transgender rights activist capable to pulling off such as well planned action would be stupid enough to leave ammunition engravings to villianize their own causes.
090
Phil @phil2345.bsky.social · 09/09/2025
I was also hoping it was open source, but it makes sense that they'd keep their biggest trillion parameter model proprietary.
010
Phil @phil2345.bsky.social · 09/09/2025
open-source?
110
Phil @phil2345.bsky.social · 06/09/2025
I believe it.
020
Phil @phil2345.bsky.social · 25/08/2025
Google claims otherwise, but comparing say 720p videos side-by-side with older Youtube videos the quality has dropped recently (far more compression artifacts). And AI filtering introduces a cartoonish look and random artifacts. They aren't doing it to boost quality, but rather to reduce file sizes.
000
Phil @phil2345.bsky.social · 25/08/2025
I call bullshit. By far the most noticeable Youtube artifacts are compression artifacts liking blocking & ringing. Noise like grain is rarely noticeable, especially with modern cameras. They're doing it because denoising videos makes them compress smaller, while also making them look worse.
000
Phil @phil2345.bsky.social · 06/08/2025
And they're conveniently useless to the general population. They're PROFOUNDLY ignorant compared to 04-mini with far lower SimpleQA scores & entire domains of knowledge missing. OpenAI clearly did this so the OS models don't go mainstream and compete with their proprietary subscription models.
010