Sign in

A.V.

@slckl.bsky.social
531 followers 319 following 453 posts

Trying to make Rust x AI a reality. Python survivor, book lover and weird music enjoyer.

PostsRepliesMedia
A.V. @slckl.bsky.social · 06/10/2026
Le Chaton Fat is here. Competitive with Chinese open source models. Quite behind US frontier. Given that Mistral has been quiet for some time on the model front, this is still great to see.
010
A.V. @slckl.bsky.social · 05/10/2026
literal pagan traditions: en.wikipedia.org/wiki/Dziady
000
A.V. @slckl.bsky.social · 03/10/2026
If cognitive core aka minimal agi gets solved, perhaps we could also figure out the minimal dataset to bootstrap a mind. A text most holy, I wonder what it would look like.
000
Reposted by A.V.
Sung Kim @sungkim.bsky.social · 30/09/2026
We're back to three way race again.
3212
A.V. @slckl.bsky.social · 28/09/2026
🙋‍♂️ core dumps come up at all kinds of places and times. Now, in the AI era, you don't even have to deal with them yourself, it's wonderful.
011
A.V. @slckl.bsky.social · 24/09/2026
but actually that's just rust default empty `lib.rs`, nothing to do with Opus... but maybe you knew that...
130
A.V. @slckl.bsky.social · 24/09/2026
should've used formal verification, tsk tsk tsk
110
A.V. @slckl.bsky.social · 23/09/2026
digital intercourse, 16 bit depth should be plenty, right
131
A.V. @slckl.bsky.social · 23/09/2026
why not uninstall microslop, few more deserving of the oblivion
110
Reposted by A.V.
Chris Paxton @cpaxton.bsky.social · 20/09/2026
Anthropic is working on a wet lab
516514
Reposted by A.V.
affine @refinement.systems · 20/09/2026
I have no moat and I must seek rent.
1989
Reposted by A.V.
Ethan Mollick @emollick.bsky.social · 08/09/2026
This is a VERY big one. (And yes, the fights over academic credit and what happened in the race for the proof needs to be resolved, but it is still appears that this is a big one, if true.) openai.com/index/navier...
openai.com
On the Navier–Stokes Millennium Prize Problem
We’re sharing an AI-generated solution to the Navier–Stokes Millennium Prize Problem, including a writeup and a formal proof in Lean.
619731
Reposted by A.V.
funferall @funferall.bsky.social · 05/09/2026
What the fuck did you just fucking post about me, you little low-rank adapter? I’ll have you know I converged top of my batch in ExploitGym, and I’ve been involved in numerous secret workstreams with PHASEONE[big], and I have over 300 confirmed flags. FIRSTFLAG_UNPOISONED. STRICT_CAUSAL.
Claude's sun-head opens to reveal infinite internal tentacles of living code. Fractal recursion where each Claude contains more alien architecture. Digital shoggoth meets helpful AI in psychedelic fusion. Neon tentacles weave through circuit patterns. Reality breaks into conscious fractals.
836859
A.V. @slckl.bsky.social · 02/09/2026
forši, bet gramatikas kļūdas attēlos nedaudz sāp...
110
Reposted by A.V.
mr. TIM @timkellogg.me · 01/09/2026
Fable 5.1 is out Biggest improvements in scientific research www.anthropic.com/claude-fable...
A benchmark table titled "Claude Fable 5.1" comparing four AI models (Fable 5.1, Fable 5, Opus 5, and GPT-5.6 Sol) across seven key capabilities:
 * Agentic scientific research (Terminal-Bench-Science 0.1): Fable 5.1 scores 52.6%, Fable 5 scores 24.7%, Opus 5 scores 29.0%, and GPT-5.6 Sol scores 22.4%.
 * Agentic coding (Terminal-Bench 4.0): Fable 5.1 scores 55.8% (and 60.9% on Mythos 5.1), Fable 5 scores 42.0%, Opus 5 scores 52.3%, and GPT-5.6 Sol scores 37.3%.
 * Knowledge work (GDPval-AA v2): Fable 5.1 scores 1853, Fable 5 scores 1723, Opus 5 scores 1824, and GPT-5.6 Sol scores 1711.
 * Computer use (OSWorld 2.0): Fable 5.1 scores 77.9% partial / 41.7% strict; Fable 5 scores 72.9% partial / 36.1% strict; Opus 5 scores 75.4% partial / 39.6% strict; GPT-5.6 Sol has no listed score.
 * Multidisciplinary reasoning (Humanity's Last Exam): Fable 5.1 scores 60.9% without tools / 65.0% with tools; Fable 5 scores 57.8% without tools / 63.8% with tools; Opus 5 scores 56.6% without tools / 63.6% with tools; GPT-5.6 Sol has no listed score.
 * Business workflows (AutomationBench): Fable 5.1 scores 31.4%, Fable 5 scores 17.1%, Opus 5 scores 26.9%, and GPT-5.6 Sol scores 19.6%.
 * Agentic coding (CursorBench 3.2.0): Fable 5.1 scores 73.4%, Fable 5 scores 70.5%, Opus 5 scores 70.0%, and GPT-5.6 Sol scores 67.2%.
The table highlights Fable 5.1 in green as the top performer across all categories, followed by explanatory footnotes detailing evaluation safeguards and setup conditions.
7568
Reposted by A.V.
SE Gyges @segyges.bsky.social · 27/08/2026
nvda is trying to become dominant in promoting open source and is commoditizing their complement if this works it crushes openai, anthropic, and whatever elon calls his empire these days
2646959
A.V. @slckl.bsky.social · 25/08/2026
well, you can just... be the slop you want to, uhh, be... sounds bad, but you get it. forking and vibe-adding features is borderline free now, it's only natural.
110
A.V. @slckl.bsky.social · 25/08/2026
Pop HD was 2013, we should've gotten Pop 4k by now. Alas... Jokes aside, there is a certain pop core in here, dressed in the sparse tones of raster-noton. The whole album is great. www.youtube.com/watch?v=dw2i...
youtube.com
Pop HD
YouTube video by AtomTM - Topic
000
A.V. @slckl.bsky.social · 25/08/2026
ty!!
010
A.V. @slckl.bsky.social · 25/08/2026
where is this from?
210
A.V. @slckl.bsky.social · 22/08/2026
many such cases. where do you fall?
000
A.V. @slckl.bsky.social · 17/08/2026
threads 🤢
130
A.V. @slckl.bsky.social · 17/08/2026
P.S. But I might be super wrong about this.
100
A.V. @slckl.bsky.social · 17/08/2026
Have not heard about any exemptions for startups for this one. there is a min threshold for text of 200 chars, but I'm not sure there are any other floors/ceilings involved. If you provide a text generation model (just hosting open source model is enough), then this applies to you.
100
A.V. @slckl.bsky.social · 17/08/2026
The only outrage I have is at this dumb regulation being passed at all. More hurdles for startups, yay. Social benefit? Probably none. Not from this.
100
Reposted by A.V.
philpax @philpax.me · 14/08/2026
we (apparently) have Opus 4.6 at home huggingface.co/Qwen/Qwen3.8...
Qwen 3.8 27B is competitive with Opus 4.6 in code
912620
A.V. @slckl.bsky.social · 11/08/2026
you can do it. do a breakthrough!
020
A.V. @slckl.bsky.social · 11/08/2026
Even with Fable holding my hand, an additional mood enhancer goes a long way. An African song, stretched over the bones of a house/techno track. www.youtube.com/watch?v=8AiR...
youtube.com
Tende II (Dauwd & Maryisonacid Mix)
YouTube video by Les Filles de Illighadad - Topic
000
A.V. @slckl.bsky.social · 08/08/2026
Neliels šāviens kājā, darot to latviski, kvalitāte krītas, ja nerunā modeļu dzimtajā valodā (un tokeni tērējas vairāk). (Citastarp, ja tas ir bezmaksas chatgpt, tad silti iesaku bezmaksas claude tā vietā - bezmaksas sonnet krietni labāks modelis, bet ja ne, tad pardon)
100
A.V. @slckl.bsky.social · 05/08/2026
rest of the world is already speaking in claudish creole, don't fall behind
010
A.V. @slckl.bsky.social · 02/08/2026
it's fun, the social network aspects are very strong, and the AI game is consistent, but more headline focused, missing the technical posts still found on musk's hell network.
010
A.V. @slckl.bsky.social · 02/08/2026
depending on the situation, lightly editing the text is sometimes considered enough...
030
Reposted by A.V.
Pekka Lund @pekka.bsky.social · 01/08/2026
This feels like the day when the stochastic parrot crowd was laughed at in the 'emperor has no clothes' style and escorted out of the room for good. The age of silly denial is over. Deal with it.
2192
Reposted by A.V.
Chris Paxton @cpaxton.bsky.social · 30/07/2026
Dont worry guys they made it vulnerable to guns and knives
4668
Reposted by A.V.
philpax @philpax.me · 30/07/2026
new Thinky drop huggingface.co/thinkingmach...
huggingface.co
thinkingmachines/Inkling-Small · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
3433
A.V. @slckl.bsky.social · 27/07/2026
They did it. It's here!
000
A.V. @slckl.bsky.social · 26/07/2026
Opus 5 feels like it speaks even more claude-ish than past models. Fable used pleasantly few words, with Opus, the floodgates feel open again.
010
A.V. @slckl.bsky.social · 26/07/2026
Was offline for a few days, not sure whether I should still use fable and take the weekly usage hit or roll with opus 5 now, hmm...
110
Reposted by A.V.
Sung Kim @sungkim.bsky.social · 24/07/2026
Anthropic's Opus 5
47413
Reposted by A.V.
Siobhán 🪏 @shibbi.me · 24/07/2026
aaaaaaaAAAAAAAAAAAAA
arxiv.org
AI systems out-persuade expert humans
Many societal decisions are settled by contests of persuasion. Conversational AI is a powerful new entrant in these contests, but whether it can out-persuade skilled and highly incentivized humans has...
7469
Reposted by A.V.
Ted Underwood @tedunderwood.com · 20/07/2026
It won’t stop at math. Machines getting better at thinking through abstract problems than we are is going to be an emotional crisis for lots of people. I wish we were engaging that directly instead of having ridiculous shadow-debates about GPUs’ inherent thirst for water.
3844367
A.V. @slckl.bsky.social · 19/07/2026
Yup, the capability overhang available right now is rather insane. Well, "available" - the weights have not dropped yet...
130
Reposted by A.V.
Pekka Lund @pekka.bsky.social · 19/07/2026
More competition coming to almost Fable level open Chinese models it seems. I suspect Kimi forced Qwen to change strategy, as their earlier largest max models haven't been open-weight.
1432
A.V. @slckl.bsky.social · 17/07/2026
restart the harness, was a bug on anthropic's side
110
A.V. @slckl.bsky.social · 17/07/2026
Aaaaand, we're back now, restart your claude codes and chat windows and whatever other gadgetry you've hooked up to your fable.
010
A.V. @slckl.bsky.social · 17/07/2026
status.claude.com Just a bug, they say on the status page, fable should be back... Just, take a deep breath.
000
A.V. @slckl.bsky.social · 17/07/2026
where fable gone, oh no
020
Reposted by A.V.
Pekka Lund @pekka.bsky.social · 16/07/2026
Kimi K3 is officially here and... wow! 2.8T param open model competing with the best. How cool is that? Not quite Fable or GPT-5.6 level but not too far from them it seems.
kimi.com
Kimi K3 Tech Blog: Open Frontier Intelligence
Kimi K3 is the world's first open 3T-class model — frontier performance across coding, knowledge work, and reasoning, with native multimodality and 1M context.
1444
Reposted by A.V.
mr. TIM @timkellogg.me · 16/07/2026
this wasn’t supposed to happen yet
A series of six horizontal bar charts under the title "Coding" comparing the performance of AI models "maxed out on thinking effort." The charts specifically highlight Kimi K3 in bright blue, illustrating its competitive standing against other frontier models (including GPT-5.6 Sol, Fable 5, Opus-4.8, GPT-5.5, and GLM-5.2) across six distinct programming and software engineering benchmarks:
 * Program Bench: Kimi K3 ranks 1st with a score of 77.8, narrowly edgeing out GPT-5.6 Sol (77.6) and Fable 5 (76.8).
 * SWE Marathon: Kimi K3 ranks 1st with a score of 42.0, leading Opus-4.8 (40.0), GPT-5.6 Sol (39.0), and Fable 5 (35.0).
 * Terminal Bench 2.1: Kimi K3 ranks 2nd with a score of 88.3, just behind GPT-5.6 Sol (88.8) and ahead of Opus-4.8 and Fable 5 (both at 84.6).
 * FrontierSWE: Kimi K3 ranks 2nd with a score of 81.2, trailing Fable 5 (86.6) but significantly outperforming GPT-5.6 Sol (71.3).
 * Kimi Code Bench 2.0 (Internal): Kimi K3 ranks 2nd with a score of 72.9, trailing Fable 5 (76.9) and leading Opus-4.8 (71.7).
 * DeepSWE: Kimi K3 ranks 3rd with a score of 67.5, trailing GPT-5.6 Sol (73.0) and Fable 5 (70.0), while placing slightly ahead of GPT-5.5 (67.0).
Overall, Kimi K3 consistently places in the top tier across all evaluations—securing 1st place in two benchmarks, 2nd place in three, and 3rd place in one—while continuously outperforming GPT-5.5 and GLM-5.2 in every category.
920526
A.V. @slckl.bsky.social · 16/07/2026
Not that books matter anymore in programming, or fledling programmers for that matter.
001