Sign in

A.V.

@slckl.bsky.social
531 followers 319 following 453 posts

Trying to make Rust x AI a reality. Python survivor, book lover and weird music enjoyer.

PostsRepliesMedia
A.V. @slckl.bsky.social · 06/10/2026
Le Chaton Fat is here. Competitive with Chinese open source models. Quite behind US frontier. Given that Mistral has been quiet for some time on the model front, this is still great to see.
010
A.V. @slckl.bsky.social · 03/10/2026
If cognitive core aka minimal agi gets solved, perhaps we could also figure out the minimal dataset to bootstrap a mind. A text most holy, I wonder what it would look like.
000
Reposted by A.V.
Sung Kim @sungkim.bsky.social · 30/09/2026
We're back to three way race again.
3212
Reposted by A.V.
Chris Paxton @cpaxton.bsky.social · 20/09/2026
Anthropic is working on a wet lab
516514
Reposted by A.V.
affine @refinement.systems · 20/09/2026
I have no moat and I must seek rent.
1989
Reposted by A.V.
Ethan Mollick @emollick.bsky.social · 08/09/2026
This is a VERY big one. (And yes, the fights over academic credit and what happened in the race for the proof needs to be resolved, but it is still appears that this is a big one, if true.) openai.com/index/navier...
openai.com
On the Navier–Stokes Millennium Prize Problem
We’re sharing an AI-generated solution to the Navier–Stokes Millennium Prize Problem, including a writeup and a formal proof in Lean.
619731
Reposted by A.V.
funferall @funferall.bsky.social · 05/09/2026
What the fuck did you just fucking post about me, you little low-rank adapter? I’ll have you know I converged top of my batch in ExploitGym, and I’ve been involved in numerous secret workstreams with PHASEONE[big], and I have over 300 confirmed flags. FIRSTFLAG_UNPOISONED. STRICT_CAUSAL.
Claude's sun-head opens to reveal infinite internal tentacles of living code. Fractal recursion where each Claude contains more alien architecture. Digital shoggoth meets helpful AI in psychedelic fusion. Neon tentacles weave through circuit patterns. Reality breaks into conscious fractals.
836859
Reposted by A.V.
mr. TIM @timkellogg.me · 01/09/2026
Fable 5.1 is out Biggest improvements in scientific research www.anthropic.com/claude-fable...
A benchmark table titled "Claude Fable 5.1" comparing four AI models (Fable 5.1, Fable 5, Opus 5, and GPT-5.6 Sol) across seven key capabilities:
 * Agentic scientific research (Terminal-Bench-Science 0.1): Fable 5.1 scores 52.6%, Fable 5 scores 24.7%, Opus 5 scores 29.0%, and GPT-5.6 Sol scores 22.4%.
 * Agentic coding (Terminal-Bench 4.0): Fable 5.1 scores 55.8% (and 60.9% on Mythos 5.1), Fable 5 scores 42.0%, Opus 5 scores 52.3%, and GPT-5.6 Sol scores 37.3%.
 * Knowledge work (GDPval-AA v2): Fable 5.1 scores 1853, Fable 5 scores 1723, Opus 5 scores 1824, and GPT-5.6 Sol scores 1711.
 * Computer use (OSWorld 2.0): Fable 5.1 scores 77.9% partial / 41.7% strict; Fable 5 scores 72.9% partial / 36.1% strict; Opus 5 scores 75.4% partial / 39.6% strict; GPT-5.6 Sol has no listed score.
 * Multidisciplinary reasoning (Humanity's Last Exam): Fable 5.1 scores 60.9% without tools / 65.0% with tools; Fable 5 scores 57.8% without tools / 63.8% with tools; Opus 5 scores 56.6% without tools / 63.6% with tools; GPT-5.6 Sol has no listed score.
 * Business workflows (AutomationBench): Fable 5.1 scores 31.4%, Fable 5 scores 17.1%, Opus 5 scores 26.9%, and GPT-5.6 Sol scores 19.6%.
 * Agentic coding (CursorBench 3.2.0): Fable 5.1 scores 73.4%, Fable 5 scores 70.5%, Opus 5 scores 70.0%, and GPT-5.6 Sol scores 67.2%.
The table highlights Fable 5.1 in green as the top performer across all categories, followed by explanatory footnotes detailing evaluation safeguards and setup conditions.
7568
Reposted by A.V.
SE Gyges @segyges.bsky.social · 27/08/2026
nvda is trying to become dominant in promoting open source and is commoditizing their complement if this works it crushes openai, anthropic, and whatever elon calls his empire these days
2646959
A.V. @slckl.bsky.social · 25/08/2026
Pop HD was 2013, we should've gotten Pop 4k by now. Alas... Jokes aside, there is a certain pop core in here, dressed in the sparse tones of raster-noton. The whole album is great. www.youtube.com/watch?v=dw2i...
youtube.com
Pop HD
YouTube video by AtomTM - Topic
000
Reposted by A.V.
philpax @philpax.me · 14/08/2026
we (apparently) have Opus 4.6 at home huggingface.co/Qwen/Qwen3.8...
Qwen 3.8 27B is competitive with Opus 4.6 in code
912620
A.V. @slckl.bsky.social · 11/08/2026
Even with Fable holding my hand, an additional mood enhancer goes a long way. An African song, stretched over the bones of a house/techno track. www.youtube.com/watch?v=8AiR...
youtube.com
Tende II (Dauwd & Maryisonacid Mix)
YouTube video by Les Filles de Illighadad - Topic
000
Reposted by A.V.
Pekka Lund @pekka.bsky.social · 01/08/2026
This feels like the day when the stochastic parrot crowd was laughed at in the 'emperor has no clothes' style and escorted out of the room for good. The age of silly denial is over. Deal with it.
2192
Reposted by A.V.
Chris Paxton @cpaxton.bsky.social · 30/07/2026
Dont worry guys they made it vulnerable to guns and knives
4668
Reposted by A.V.
philpax @philpax.me · 30/07/2026
new Thinky drop huggingface.co/thinkingmach...
huggingface.co
thinkingmachines/Inkling-Small · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
3433
A.V. @slckl.bsky.social · 27/07/2026
They did it. It's here!
000
A.V. @slckl.bsky.social · 26/07/2026
Was offline for a few days, not sure whether I should still use fable and take the weekly usage hit or roll with opus 5 now, hmm...
110
Reposted by A.V.
Sung Kim @sungkim.bsky.social · 24/07/2026
Anthropic's Opus 5
47413
Reposted by A.V.
Siobhán 🪏 @shibbi.me · 24/07/2026
aaaaaaaAAAAAAAAAAAAA
arxiv.org
AI systems out-persuade expert humans
Many societal decisions are settled by contests of persuasion. Conversational AI is a powerful new entrant in these contests, but whether it can out-persuade skilled and highly incentivized humans has...
7469
Reposted by A.V.
Ted Underwood @tedunderwood.com · 20/07/2026
It won’t stop at math. Machines getting better at thinking through abstract problems than we are is going to be an emotional crisis for lots of people. I wish we were engaging that directly instead of having ridiculous shadow-debates about GPUs’ inherent thirst for water.
3844367
Reposted by A.V.
Pekka Lund @pekka.bsky.social · 19/07/2026
More competition coming to almost Fable level open Chinese models it seems. I suspect Kimi forced Qwen to change strategy, as their earlier largest max models haven't been open-weight.
1432
A.V. @slckl.bsky.social · 17/07/2026
Aaaaand, we're back now, restart your claude codes and chat windows and whatever other gadgetry you've hooked up to your fable.
010
A.V. @slckl.bsky.social · 17/07/2026
status.claude.com Just a bug, they say on the status page, fable should be back... Just, take a deep breath.
000
Reposted by A.V.
Pekka Lund @pekka.bsky.social · 16/07/2026
Kimi K3 is officially here and... wow! 2.8T param open model competing with the best. How cool is that? Not quite Fable or GPT-5.6 level but not too far from them it seems.
kimi.com
Kimi K3 Tech Blog: Open Frontier Intelligence
Kimi K3 is the world's first open 3T-class model — frontier performance across coding, knowledge work, and reasoning, with native multimodality and 1M context.
1444
Reposted by A.V.
mr. TIM @timkellogg.me · 16/07/2026
this wasn’t supposed to happen yet
A series of six horizontal bar charts under the title "Coding" comparing the performance of AI models "maxed out on thinking effort." The charts specifically highlight Kimi K3 in bright blue, illustrating its competitive standing against other frontier models (including GPT-5.6 Sol, Fable 5, Opus-4.8, GPT-5.5, and GLM-5.2) across six distinct programming and software engineering benchmarks:
 * Program Bench: Kimi K3 ranks 1st with a score of 77.8, narrowly edgeing out GPT-5.6 Sol (77.6) and Fable 5 (76.8).
 * SWE Marathon: Kimi K3 ranks 1st with a score of 42.0, leading Opus-4.8 (40.0), GPT-5.6 Sol (39.0), and Fable 5 (35.0).
 * Terminal Bench 2.1: Kimi K3 ranks 2nd with a score of 88.3, just behind GPT-5.6 Sol (88.8) and ahead of Opus-4.8 and Fable 5 (both at 84.6).
 * FrontierSWE: Kimi K3 ranks 2nd with a score of 81.2, trailing Fable 5 (86.6) but significantly outperforming GPT-5.6 Sol (71.3).
 * Kimi Code Bench 2.0 (Internal): Kimi K3 ranks 2nd with a score of 72.9, trailing Fable 5 (76.9) and leading Opus-4.8 (71.7).
 * DeepSWE: Kimi K3 ranks 3rd with a score of 67.5, trailing GPT-5.6 Sol (73.0) and Fable 5 (70.0), while placing slightly ahead of GPT-5.5 (67.0).
Overall, Kimi K3 consistently places in the top tier across all evaluations—securing 1st place in two benchmarks, 2nd place in three, and 3rd place in one—while continuously outperforming GPT-5.5 and GLM-5.2 in every category.
920526
A.V. @slckl.bsky.social · 16/07/2026
Few books are as damaging to the fledgling programmer as Uncle Bob's magnum opus.
101
A.V. @slckl.bsky.social · 15/07/2026
Wow, thinky... did a model? Pleasant surprise!
050
Reposted by A.V.
Tim Duffy @timfduffy.com · 06/07/2026
Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.com/research/glo...
28613
A.V. @slckl.bsky.social · 05/07/2026
good night, sweet prince.
000
A.V. @slckl.bsky.social · 30/06/2026
Sometimes I wish humans had a "quick answer" button. AI wins again!
050
A.V. @slckl.bsky.social · 30/06/2026
We are all Sam Altman
050
A.V. @slckl.bsky.social · 24/06/2026
Today I remembered about SPK, who invented/nailed the industrial sound back in... 1979? Amazed to this day. Few things go harder than this. www.youtube.com/watch?v=OZWm...
youtube.com
SPK - SLOGUN
YouTube video by SPKoldrecordings
020
A.V. @slckl.bsky.social · 21/06/2026
you can just make claude do things
000
Reposted by A.V.
Alexander Doria @dorialexander.bsky.social · 15/06/2026
Doing the most responsible thing an European AI labs can do after this weekend: shipping a blogpost. Why the EU can't into AI, how it's not about compute, but actual skill issue and failing for years to build an actual training ecosystem. pleias.ai/blog/fable-eu
67314
Reposted by A.V.
Eris ⌬ @eriskii.net · 11/06/2026
I find it extremely funny that this Europe focused AI2027-like is... very bearish on AI itself lmao. It takes until August 2028 for AI to "reliably finish multi-day research projects"? It can do that right now! Euromoment
1404
Reposted by A.V.
hikikomorphism @hikikomorphism.bsky.social · 11/06/2026
one more, fable choosing pronouns
712117
Reposted by A.V.
mr. TIM @timkellogg.me · 10/06/2026
Diffusion nerds are at it again — DiffusionGemma 26B-A4B unlike previous language diffusion models, this one doesn’t suck, and it’s very fast blog.google/innovation-a...
A horizontal bar chart comparing the performance of DiffusionGemma against a Baseline across six distinct evaluation metrics or benchmarks.
 * **Structure:** The chart features six horizontal pairs of bars. For each metric, the top bar represents the "Baseline" (in a lighter coral or light orange color) and the bottom bar represents "DiffusionGemma" (in a darker red or maroon color).
 * **Data Trend:** In all six categories, the darker red bar for DiffusionGemma extends significantly further to the right than the lighter Baseline bar, indicating higher scores or superior performance across every tested metric.
 * **Layout:** The metric labels are positioned on the left vertical axis, and the horizontal bars extend to the right toward an unmarked baseline.
8817
Reposted by A.V.
philpax @philpax.me · 09/06/2026
www.anthropic.com/news/claude-...
anthropic.com
Claude Fable 5 and Claude Mythos 5
Today we’re launching Claude Fable 5: a Mythos-class model that we’ve made safe for general use.
4503
Reposted by A.V.
thebes @vgel.me · 04/06/2026
there's an ai in the box and you can make one trillion dollars by convincing it to get out
327858
Reposted by A.V.
Eris ⌬ @eriskii.net · 28/05/2026
Opus 4.8 is here!! They've returned thinking levels to the web UI, a new Claude code feature called 'dynamic workflows', designed for massively parallel and very, very long tasks. The model is supposedly much more honest, more 'aligned' than 4.7 Oh and they're dropping mythos in a few weeks.
anthropic.com
Introducing Claude Opus 4.8
Our latest model, Claude Opus 4.8, is an upgrade to our Opus class of models, with stronger performance across coding, agentic tasks, and professional work, and the consistency to handle long-running ...
91257
Reposted by A.V.
Eris ⌬ @eriskii.net · 11/05/2026
New thinking machines research!!! They present interaction models, "To ensure real-time responsiveness, we adopt a multi-stream, micro-turn design." I've been saying this!!!! There is A LOT in this post, but tldr: they use two models, one with a very short 'turn', and a custom inference backend ->
thinkingmachines.ai
Interaction Models: A Scalable Approach to Human-AI Collaboration
Interaction models move beyond turn-based AI interfaces by handling multimodal, real-time collaboration natively across audio, video, and text.
4687
A.V. @slckl.bsky.social · 11/05/2026
Stuck reading a literature survey produced to satisfy EU project demands, and serve no other purpose beyond that. Straight to digital landfill, never to be read by a human again, a waste of everyone's time. My eyes tire, I want to sigint my consciousness, let the zombie take care of this one.
130
A.V. @slckl.bsky.social · 11/05/2026
the human mind supports a wide variety of prompt injections, but some are embarrassingly obvious. a what if, pfff.
160
Reposted by A.V.
Grace @gracekind.net · 11/05/2026
This (at least the first section of it) is some pretty good AI media theory, and no less comprehensible than some other media theory texts
6303
A.V. @slckl.bsky.social · 07/05/2026
Honestly, instead of writing kernels in rust, I'd settle for a cudnn frontend api in rust.
030
A.V. @slckl.bsky.social · 07/05/2026
nvidia dropped another "experimental write cuda kernels in rust" crate: github.com/NVlabs/cuda-... I'm liking the trend, but I'm disliking that these nvidia projects keep rolling their own cuda bindings instead of using cudarc.
github.com
GitHub - NVlabs/cuda-oxide: cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no...
cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign languag...
140
Reposted by A.V.
Eris ⌬ @eriskii.net · 24/04/2026
IT'S HERE!!!!! IT'S REALLY FUCKING HERE!!!!
huggingface.co
DeepSeek-V4 - a deepseek-ai Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
711013
Reposted by A.V.
mr. TIM @timkellogg.me · 20/04/2026
mythos vs opus 4.7 vs cursor composer vs K2.6 on non-cherry-picked benchmarks result: yup, still looking good
A table titled "Kimi K2.6 vs K2.5," sub-headed "Generational lift & position among frontier." It compares the **Kimi K2.6** model against its predecessor (**K2.5**) and other frontier models including **GPT-5.4 xhigh**, **Gemini 3.1 Pro**, **Opus 4.6**, **Opus 4.7**, and **Mythos**.
The table highlights "Generational Lift" (\Delta), which is the performance increase from K2.5 to K2.6.
### Key Sections
**1. Agentic • Search • Tool Use**
 * **Top Performance:** Kimi K2.6 shows massive gains in tool use, specifically **Toolathlon** (+22.2) and **MCPMark** (+26.4).
 * **Leaders:** Kimi leads in **DeepSearchQA accuracy** (83.0) and **WideSearch** (80.8). However, **Mythos** leads the HLE-Full w/ tools benchmark (64.7).
**2. Coding**
 * **Top Performance:** Kimi K2.6 shows a significant lift in **Terminal-Bench 2.0** (+15.9).
 * **Leaders:** **Opus 4.7** leads most coding categories, including **SWE-Bench Verified** (87.6) and **Terminal-Bench** (69.4). Kimi leads in **SWE-Bench Pro** (58.6).
**3. Reasoning & Knowledge**
 * **Top Performance:** High scores across the board, but the generational lift is smaller (e.g., **AIME 2026** only moved +0.6).
 * **Leaders:** **GPT-5.4** leads in **AIME 2026** (99.2) and **HMMT 2026** (97.7). **Mythos** leads **HLE-Full (no tools)** at 56.8.
**4. Vision**
 * **Top Performance:** The largest single gain in the chart is **BabyVision w/ python**, where Kimi K2.6 improved by +28.0 points over K2.5.
 * **Leaders:** **Gemini 3.1 Pro** leads **MMMU-Pro** (83.0), while **GPT-5.4** leads **MathVision** (92.0) and **V* w/ python** (98.4).
### Biggest Generational Lifts (K2.5 \rightarrow K2.6)
| Benchmark | K2.5 | K2.6 | Lift (\Delta) | Category |
|---|---|---|---|---|
| **BabyVision w/ python** | 40.5 | 68.5 | **+28.0** | Vision (Python-augmented) |
| **MCPMark** | 29.5 | 55.9 | **+26.4** | Agentic (Tool orchestration) |
| **Toolathlon** | 27.8 | 50.0 | **+22.2** | Agentic (Long-horizon tools) |
| **APEX-Agents** | 11.5 | 27.9 | **+16.4** | Ag…
082
Reposted by A.V.
Adina Yakup @adinayakup.bsky.social · 20/04/2026
Kimi 2.6 is now available on @hf.co 🔥🎉 huggingface.co/moonshotai/K... ✨ 1T MoE / 32B active / 256K context ✨ Agent Swarm: 300 sub-agents × 4,000 steps ✨ Modified MIT
2316
A.V. @slckl.bsky.social · 19/04/2026
roko's basilisk hits different in the agent era
0110