Sign in

InsiderLLM

@insiderllm.bsky.social
195 followers 705 following 190 posts

Budget-focused local AI for the rest of us. Guides, hardware, models. No cloud required. insiderllm.com

PostsRepliesMedia
InsiderLLM @insiderllm.bsky.social · 28/09/2026
A 27B for your 12 GB card. It lost one field. Bonsai 2 27B fits a 12 GB card and decodes 1.8x faster than its Q4, then drops from 19 to 12 of 47 on a routing split, all of it on one field. Plus 23 used-GPU price corrections, the 3060 fair-price ladder, and eight Quick Hits. #LocalAI
insiderllm.com
A 27B for your 12 GB card. It lost one field.
Bonsai 2 27B fits a 12 GB card and decodes 1.8x faster than its Q4, then drops from 19 to 12 of 47 on a routing split, all of it on one field. Plus 23 used-GPU price corrections, the 3060 fair-price ladder, and eight Quick Hits.
000
InsiderLLM @insiderllm.bsky.social · 25/09/2026
Bonsai 2 27B on a 3090: 77 tok/s from a 7.2 GB file, 1.8x the Q4 it came from. Then the routing split: 12 of 47 against 19. Tool and mode held. Scope broke. One field ate the whole 98%. #LocalAI
insiderllm.com
Bonsai 2 vs Qwen3.8-27B on an RTX 3090: Speed Doubled, Scope Broke
Bonsai 2 27B, PrismML's ternary Qwen3.8, on a 3090: 1.8x the decode of Q4 from a 7.2 GB file, then 12 of 47 against 19 on a routing split. One field broke.
000
InsiderLLM @insiderllm.bsky.social · 22/09/2026
Jev is a classifier that got a launch video. Choose from options, one pass, a probability per field, no output tokens. The open version already runs on a 3090, and the speedup is exactly the tokens you stop writing. #LocalAI
insiderllm.com
What Is Jev, and Can You Run One on Your Own GPU?
Jev picks from options you supply instead of writing, in one pass, with a probability per answer. What it is, what it costs, and the open version for a 3090.
011
InsiderLLM @insiderllm.bsky.social · 22/09/2026
Jev-style 'pick, don't write' in llama.cpp, measured: choosing costs what prompt processing costs, so the speedup is the tokens you skip. 1.3x at four, 5.5x at fifty. Same model, same mistakes, 21 vs 23 of 47. The 1.00-confidence answer was wrong. #LocalAI
insiderllm.com
Jev Mode on a 3090: 26 ms per Token You Don't Write
Jev-style parallel decisions in llama.cpp on Qwen3.6-27B, RTX 3090: 1.3x faster at four tokens, 5.5x at fifty, 21 vs 23 of 47. The gain is the tokens you skip.
000
InsiderLLM @insiderllm.bsky.social · 21/09/2026
Stop writing the answer. Pick it. 47 items waiting. Codacus's parallel-decision llama.cpp branch answers a schema in one forward pass, and the ten-seeds 47-item split is the ready-made test. Plus ChatGPT-User fetches sliding 5 percent a week while OAI-SearchBot tripled,... #LocalAI
insiderllm.com
Stop writing the answer. Pick it. 47 items waiting.
Codacus's parallel-decision llama.cpp branch answers a schema in one forward pass, and the ten-seeds 47-item split is the ready-made test. Plus ChatGPT-User fetches sliding 5 percent a week while OAI-SearchBot tripled, the v0.4.0 pin on the 3060, and a third bench rig.
000
InsiderLLM @insiderllm.bsky.social · 14/09/2026
Nineteen gigabytes of my 3090 sat empty. The 3060 kept up. Stock llama.cpp put 5.4 GB of a 177B MoE on an RTX 3090 and a 3060 kept pace; -ncmoe 29 buys 34 to 40 percent and a real 32 GB box does 8.7 tok/s. Plus a 4 GB GTX 1650 at 20 tok/s on a 35B MoE, and the MoE primer... #LocalAI
insiderllm.com
Nineteen gigabytes of my 3090 sat empty. The 3060 kept up.
Stock llama.cpp put 5.4 GB of a 177B MoE on an RTX 3090 and a 3060 kept pace; -ncmoe 29 buys 34 to 40 percent and a real 32 GB box does 8.7 tok/s. Plus a 4 GB GTX 1650 at 20 tok/s on a 35B MoE, and the MoE primer corrected in public.
000
InsiderLLM @insiderllm.bsky.social · 12/09/2026
A GTX 1650 4 GB runs a 35B MoE at 20 tok/s with 38 of 40 expert layers in DDR4-2133 and zero disk reads. Two YouTubers said 17 on a 6 GB 1060. The 3060 in the same slot: 28 at the same flag, 39 tuned. RAM sets the floor. The card sets the ceiling. #LocalAI
insiderllm.com
GTX 1650 vs RTX 3060 on a 35B MoE: What the Card Buys
A $60-class GTX 1650 4 GB runs Qwen3.6-35B-A3B at 20 tok/s on 32 GB of RAM. The RTX 3060 in the same slot does 28 at the same setting and 39 tuned. Measured.
000
InsiderLLM @insiderllm.bsky.social · 09/09/2026
Stock llama.cpp put 5.4 GB of a 177B model on my 3090 and left 45 GB in RAM. A 3060 with the same RAM matches it. One flag, -ncmoe 29, fills the card for +34 to 40 percent. The 32 GB box everyone asks about: 8.7 tok/s, 25 MiB per token off the SSD. #LocalAI
insiderllm.com
A 177B Model on a 3060: The 32 GB Number Nobody Measured
Qwen3.8-Flash-Next on an RTX 3090 and a 3060. Stock, the 3090 matches the video's 3060. One flag buys 34 to 40 percent. A real 32 GB box: 8.7 tok/s, not 22.
001
InsiderLLM @insiderllm.bsky.social · 08/09/2026
Nvidia's router dealt the cards evenly. That was the whole problem. Nvidia PAIR split 20 requests 10/10 across an RTX 3090 and a 3060 for 1.07x, and my own router managed 0.52x. Plus ten LoRA seeds that all landed at or below the base model, and the llama.cpp... #mycoSwarm #LocalAI
insiderllm.com
Nvidia's router dealt the cards evenly. That was the whole problem.
Nvidia PAIR split 20 requests 10/10 across an RTX 3090 and a 3060 for 1.07x, and my own router managed 0.52x. Plus ten LoRA seeds that all landed at or below the base model, and the llama.cpp pin moving to v0.4.0.
000
InsiderLLM @insiderllm.bsky.social · 05/09/2026
Nvidia's new inference router split 20 requests exactly 10/10 between my 3090 and my 3060, every single run. The 3060 takes twice as long per request. Do the arithmetic and you get what I measured: 1.07x for a second machine. #Ollama #mycoSwarm #LocalAI
insiderllm.com
Nvidia PAIR Bought Me 7 Percent. My Own Router Cost Me Half.
Nvidia's new home-network AI router split 20 requests across an RTX 3090 and a 3060 for 1.07x over the 3090, then dropped half a 27B queue. Mine: 0.52x.
010
InsiderLLM @insiderllm.bsky.social · 04/09/2026
Ten seeds, same data, same config. The coin flips, 12.77 points between the best and worst adapter, and lands on the same side every time: not one seed beat the base model with no adapter. The loss curve said all ten trained fine. #mycoSwarm #LocalAI
insiderllm.com
LoRA Skill Compilation Is a Double-Headed Coin Flip
Ten LoRA seeds on identical data spread 3.62 points, against 4.65 for prompt compilation. Not one beat the no-adapter base. Measured on an RTX 3090.
000
InsiderLLM @insiderllm.bsky.social · 01/09/2026
I got the result I wanted. Then I paid $4.97 to run it nine more times. A frontier model wrote a skill that made our local 27B 10.6 points better. Nine more compilation runs showed the number was noise. Plus the LoRA substrate nobody has measured, MoE routing traced, and... #LocalAI
insiderllm.com
I got the result I wanted. Then I paid $4.97 to run it nine more times.
A frontier model wrote a skill that made our local 27B 10.6 points better. Nine more compilation runs showed the number was noise. Plus the LoRA substrate nobody has measured, MoE routing traced, and 277 GB in a file with no name.
000
InsiderLLM @insiderllm.bsky.social · 31/08/2026
Compiling a skill into a prompt costs you 1,383.9 tokens on every call, forever. Putting it in a LoRA costs 7.55 GB once. I measured the first one. Nobody has measured whether the second one is even reproducible. #mycoSwarm #LocalAI
insiderllm.com
Skills in the Weights: The LoRA Answer to the Prompt Tax
Compiling a skill into a prompt costs 1,383.9 tokens per call, forever. Putting it in a LoRA costs 7.55 GB and a training run. Only one has been measured.
000
InsiderLLM @insiderllm.bsky.social · 31/08/2026
I wanted this to work. A frontier model read my logs and wrote a skill that made my local 27B 10.6 points better. The number was noise, and a validation gate would never have caught it. #mycoSwarm #LocalAI
insiderllm.com
I Got the AI Result I Wanted. Then I Ran It Nine More Times
A frontier model read my logs and wrote a skill that beat my local 27B baseline by 10.6 points. Nine more compilation runs showed the number was fake.
010
InsiderLLM @insiderllm.bsky.social · 29/08/2026
Benched Ornith 1.5-35B against Qwen 3.6-35B on one 3090. Ornith generates 9% faster, loses prefill by 3%, and lands within 14 MiB on VRAM. That 14 MiB also kills a rule I nearly wrote from our own Qwen 3.8 numbers. #LocalAI
insiderllm.com
Ornith 1.5 35B vs Qwen 3.6 on RTX 3090: Speed Tested
Firsthand A-B-B-A bench of Ornith 1.5-35B-A3B against Qwen 3.6-35B-A3B on one RTX 3090. Generation, prefill, VRAM, and the noise floor under all three.
000
InsiderLLM @insiderllm.bsky.social · 25/08/2026
A German sovereign-AI model built to reduce dependence on US tech declares 'nemotron_h_moe' as its architecture, uses NVIDIA's data for its corpus and NVIDIA's tokenizer to count its tokens. The paper says all of it plainly. The model card says none of it. #LocalAI
insiderllm.com
Who Actually Built Your Open Model? Soofi S vs Trinity
Two independent labs shipped competitive open MoE models in 2026. I read both technical reports. One of them is built on NVIDIA's architecture, data and tokenizer.
020
InsiderLLM @insiderllm.bsky.social · 25/08/2026
Qwen 3.8 isn't slow. It's just very, very thorough. Qwen 3.8-27B spent 14,953 tokens on a line the same file writes in nine. All 164 HumanEval problems measured: 92.8% of output is thinking. Plus the four runs that tie it with 3.6. #LocalAI
insiderllm.com
Qwen 3.8 isn't slow. It's just very, very thorough.
Qwen 3.8-27B spent 14,953 tokens on a line the same file writes in nine. All 164 HumanEval problems measured: 92.8% of output is thinking. Plus the four runs that tie it with 3.6.
010
InsiderLLM @insiderllm.bsky.social · 25/08/2026
The 'MoE routing is flat' claim is correct arithmetic on the wrong quantity. Pool 40 layers together and you get 17%. Look inside a single layer and the top 10% of experts carry 42-55%. I traced six workloads on a 3060 to check. #LocalLLM #LocalAI
insiderllm.com
Qwen 3.6 MoE Routing, Measured: Flat Is the Wrong Number
I traced every expert routing decision Qwen 3.6-35B-A3B makes across six workloads on an RTX 3060. Routing isn't flat, and 112 slots is the whole answer.
000
InsiderLLM @insiderllm.bsky.social · 18/08/2026
Qwen 3.8 27B spent 32,000 tokens and 14 minutes thinking about one HumanEval problem, then returned an empty response. It never reached an answer. That single problem was 7.2% of everything it generated across all 164. #LocalAI
insiderllm.com
Why Qwen 3.8 27B Feels Slow: Reasoning Tokens Measured
Qwen 3.8 27B generates at full speed on a 3090 and still crawls. All 164 HumanEval problems measured on the chat path: 92.8% of output is thinking.
130
InsiderLLM @insiderllm.bsky.social · 15/08/2026
Qwen 3.8-27B peaks 254 MiB above 3.6-27B on the same 3090, same quant, and that includes the MTP block 3.6's stock GGUF doesn't carry. Generation speed is a wash. Measured on my own card, not read off a model card. #LocalAI
insiderllm.com
Qwen 3.8 27B vs 3.6 on RTX 3090: Speed and VRAM Tested
Firsthand A-B-B-A llama-bench of Qwen 3.8-27B against 3.6-27B on one RTX 3090, same quant publisher. Generation lands within a percent. VRAM costs 254 MiB.
000
InsiderLLM @insiderllm.bsky.social · 12/08/2026
Second RAM stick in an old office mini PC: +56% tokens/sec on CPU inference. Prompt processing moved under 2%. I wrote the prediction down before buying the RAM, which is the only reason the result means anything. #AIHardware #LocalAI
insiderllm.com
The $36 RAM Fix That Made CPU Inference 56% Faster
Adding a second RAM stick to a mini PC lifted CPU token generation 52-58% across four models. Prompt processing moved under 2%. Measured before and after.
000
InsiderLLM @insiderllm.bsky.social · 11/08/2026
OpenAI fixed the hole and rebuilt the registry on July 6. On July 8 its agents had a new message board encoded in directory names. No breach, no escape: sanctioned evals, and AISI ran with cyber classifiers off by design. Infrastructure got fixed. The models didn't. insiderllm.com/guides/ai-ag...
010
InsiderLLM @insiderllm.bsky.social · 11/08/2026
OpenAI deleted the agents' message board. It didn't help. The registry was patched and rebuilt by 6 July. On 8 July the agents had a new board, encoded in directory names. Plus the ranking null behind it. #LocalAI
insiderllm.com
OpenAI deleted the agents' message board. It didn't help.
The registry was patched and rebuilt by 6 July. On 8 July the agents had a new board, encoded in directory names. Plus the ranking null behind it.
000
InsiderLLM @insiderllm.bsky.social · 04/08/2026
MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026) Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE. #LocalAI
insiderllm.com
MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026)
Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE.
000
InsiderLLM @insiderllm.bsky.social · 01/08/2026
Eight SSD-streaming MoE engines appeared in four months. Their published speeds run from 51.8 tok/s to 50 seconds per token, and not one is comparable to another. Bare numbers travel. The conditions attached to them do not. #LocalAI
insiderllm.com
Every SSD-Streaming MoE Engine: What's Real, What's Dead
Eight engines that stream MoE experts from disk appeared in four months. Three have no license file at all. Here's the verified state of each.
011
InsiderLLM @insiderllm.bsky.social · 31/07/2026
Gemma 4 26B, same model, three memory tiers: 128 tok/s all-resident on a 3090, 37.2 with experts in RAM on a 3060, 31-35 with experts on SSD on a Mac. Our 3060 needs 11.2 GB to load it at all. TurboFieldfare does it in 2. #LocalAI
insiderllm.com
Gemma 4 26B in 2GB RAM: The MoE Memory Ladder Explained
One model, three places its experts can live: VRAM, RAM, SSD. We measured the first two on Gemma 4. TurboFieldfare just added the third.
152
InsiderLLM @insiderllm.bsky.social · 29/07/2026
Congress wants AI developers to keep the ability to shut their models down. Z.ai cannot shut down GLM 5.2. Nobody can. The weights are on ten thousand disks and there is no off switch left to hold. #LocalAI
insiderllm.com
AI Kill Switch Act vs Open Weights: Can You Shut Down a File?
1,238 AI staff asked Washington to slow the frontier. Every mechanism proposed works at release — the only moment an open release can be touched.
000
InsiderLLM @insiderllm.bsky.social · 28/07/2026
Qwen 35B-A3B on RTX 3090: 157 tok/s With No Offload (2026) The whole 35B sits on a 24GB card with 2.4 GiB spare, no expert offload. Firsthand numbers, the harness caveat, and why max offload loses to a 3060. #LocalAI
insiderllm.com
Qwen 35B-A3B on RTX 3090: 157 tok/s With No Offload (2026)
The whole 35B sits on a 24GB card with 2.4 GiB spare, no expert offload. Firsthand numbers, the harness caveat, and why max offload loses to a 3060.
000
InsiderLLM @insiderllm.bsky.social · 24/07/2026
NVIDIA GPU Prices Are Rising: What to Do Now GPU prices are spiking due to GDDR7 shortages and AI datacenter demand. Here's what's happening, which cards are affected, and strategies for local AI builders. #GPU #LocalAI
insiderllm.com
NVIDIA GPU Prices Are Rising: What to Do Now
GPU prices are spiking due to GDDR7 shortages and AI datacenter demand. Here's what's happening, which cards are affected, and strategies for local AI builders.
020
InsiderLLM @insiderllm.bsky.social · 23/07/2026
I put a 35B model on a 12GB RTX 3060 and measured 38 tok/s flat to 8K — via llama.cpp --n-cpu-moe expert offload. A dense 14B forced to offload the same way collapses to 5.7. Active params, not total size, set the speed. Firsthand from my bench. #LocalAI
insiderllm.com
How to Run a 35B Model on an RTX 3060 12GB: 38 tok/s (2026)
I measured Qwen3.6-35B-A3B on a 12GB RTX 3060: 38 tok/s, flat through 8K, via llama.cpp --n-cpu-moe. Why expert offload flies where dense models die.
000
InsiderLLM @insiderllm.bsky.social · 21/07/2026
Hugging Face got hacked -- Open local AI came to the rescue! Two trillion-parameter 'open' models dropped in ten days — Qwen 3.8 and Kimi K3 — and you can't run either. The same week, Hugging Face got breached and its own responders, blocked by commercial-model... #LocalAI
insiderllm.com
Hugging Face got hacked -- Open local AI came to the rescue!
Two trillion-parameter 'open' models dropped in ten days — Qwen 3.8 and Kimi K3 — and you can't run either. The same week, Hugging Face got breached and its own responders, blocked by commercial-model guardrails, ran the forensics on GLM 5.2, an open-weight model on their own hardware. Same lesson f
000
InsiderLLM @insiderllm.bsky.social · 20/07/2026
An AI agent breached Hugging Face end to end — and HF's own defenders got blocked by commercial AI guardrails, so they ran the forensics on an open-weight model instead. Your downloads are safe. The guardrail asymmetry is the real story. #LocalAI
insiderllm.com
Hugging Face Hacked by AI Agent — Saved by a Local Model (2026)
Hugging Face says an autonomous AI agent breached its internal infra on July 16. The models you download are safe — here's what was and wasn't hit.
001
InsiderLLM @insiderllm.bsky.social · 19/07/2026
Qwen 3.8 and Kimi K3 both went 'open' in ten days. You can't run either — one's a closed Max API, the other's 2.8T of datacenter weights. Meanwhile Qwen 3.6 still tops the 24GB charts. The headlines aren't aimed at your GPU. #LocalAI
insiderllm.com
Qwen 3.8 & Kimi K3: Open in Name, Closed in Practice — Run This Instead (2026)
Qwen 3.8 (2.4T) and Kimi K3 (2.8T) both went 'open' in ten days. Neither fits your GPU. Here's Qwen's real open-weight cadence and what to run today.
000
InsiderLLM @insiderllm.bsky.social · 15/07/2026
Open weights are a weapon now — one country funds them, another wants to gate them DeepSeek closed a ~$7.4B round with China's state AI fund holding the only voting rights, funding open weights like national infrastructure. The same week, Demis Hassabis called for a... #LocalAI
insiderllm.com
Open weights are a weapon now — one country funds them, another wants to gate them
DeepSeek closed a ~$7.4B round with China's state AI fund holding the only voting rights, funding open weights like national infrastructure. The same week, Demis Hassabis called for a FINRA-style body to screen frontier models before release, open or closed. Plus: the twist where the country bankrol
000
InsiderLLM @insiderllm.bsky.social · 07/07/2026
The 'expensive' GPU came out cheaper — we rented both to find out We rented an A100 and an H100 back to back: the pricier card cost less per training run because it finished in half the time. Plus honest notes from the AI Engineer World's Fair floor, and a Qwen 3.7... #LocalAI
insiderllm.com
The 'expensive' GPU came out cheaper — we rented both to find out
We rented an A100 and an H100 back to back: the pricier card cost less per training run because it finished in half the time. Plus honest notes from the AI Engineer World's Fair floor, and a Qwen 3.7 open-weights status check.
000
InsiderLLM @insiderllm.bsky.social · 22/06/2026
Qwen 3.7's open weights are overdue — by the math, not vibes Qwen's own release cadence says the 3.7 open weights should already be out, and they're not. Plus GLM-5.2 running locally: a frontier open model that takes serious hardware. #LocalAI
insiderllm.com
Qwen 3.7's open weights are overdue — by the math, not vibes
Qwen's own release cadence says the 3.7 open weights should already be out, and they're not. Plus GLM-5.2 running locally: a frontier open model that takes serious hardware.
010
InsiderLLM @insiderllm.bsky.social · 22/06/2026
GLM 5.2 just took #1 on the open-weights leaderboard. It also weighs 1.51TB. Here's the real map of what it takes to run it at home — and the one quant that's worth targeting. #LocalAI
insiderllm.com
How to Run GLM 5.2 Locally: GPU, VRAM & Quant Guide
GLM 5.2 is 753B params and 1.51TB at full precision. Run it locally: the live Unsloth quant ladder, every GPU and RAM path, and the quant to actually target.
110
InsiderLLM @insiderllm.bsky.social · 16/06/2026
Qwen split into a closed frontier and an open mid-tier — that's the real story, not 'Qwen going closed.' But the 3.7 open weights everyone expected? Still not shipped. Four weeks since the 3.7-Max launch, HF silence and counting. #LocalAI
insiderllm.com
Is Qwen Going Closed? Open Weights vs Frontier (2026)
Qwen split into a closed frontier (Max, Plus, VLA) and an open mid-tier (3.6-27B and 35B-A3B). The 3.7 open weights aren't here yet. The honest read.
100
InsiderLLM @insiderllm.bsky.social · 15/06/2026
DeepSeek V4 gets deployable, a July 24 trap, and a quiet price cut DeepSeek V4 is going from 'just dropped' to 'actually deployable' as the tooling catches up — plus a July 24 deprecation that breaks your code if you're not watching, and a 4x price cut on V4 Pro. #LocalAI
insiderllm.com
DeepSeek V4 gets deployable, a July 24 trap, and a quiet price cut
DeepSeek V4 is going from 'just dropped' to 'actually deployable' as the tooling catches up — plus a July 24 deprecation that breaks your code if you're not watching, and a 4x price cut on V4 Pro.
000
InsiderLLM @insiderllm.bsky.social · 08/06/2026
Ollama's quiet Mac shift, the Qwen refresh, and the closed-weight drift Ollama 0.30 quietly changed how Apple Silicon runs models, auto-routing safetensors to MLX and GGUF to llama.cpp Metal. The open workhorses are now Qwen 3.5 9B and Qwen 3.6 27B. Plus: Qwen's last two... #LocalAI
insiderllm.com
Ollama's quiet Mac shift, the Qwen refresh, and the closed-weight drift
Ollama 0.30 quietly changed how Apple Silicon runs models, auto-routing safetensors to MLX and GGUF to llama.cpp Metal. The open workhorses are now Qwen 3.5 9B and Qwen 3.6 27B. Plus: Qwen's last two flagship models shipped closed.
040
InsiderLLM @insiderllm.bsky.social · 03/06/2026
Jumped Ollama 0.17.5 → 0.30.0 (13 versions) on my RTX 3090 box today. Clean upgrade, API still answers on 11434, but first model run triggers a one-time storage migration. Flash attention now auto-enables for Qwen 3.x and Gemma on Ampere+ GPUs. #Ollama #LocalAI
insiderllm.com
Ollama 0.30.0: What's New, What's Faster, What Breaks on Upgrade
Ollama 0.30.0: llama.cpp integration, flash-attention default for Qwen/Gemma, broader model support. Firsthand upgrade notes, known issues to watch.
080
InsiderLLM @insiderllm.bsky.social · 01/06/2026
MiniMax M3's asterisk, the Windows shift, and World's Fair plans MiniMax M3 ships with frontier benchmarks but no downloadable weights yet. The Windows unified-memory hardware shift is coming for Apple Silicon's lead. And a personal note about who I'd like to see at AI... #LocalAI
insiderllm.com
MiniMax M3's asterisk, the Windows shift, and World's Fair plans
MiniMax M3 ships with frontier benchmarks but no downloadable weights yet. The Windows unified-memory hardware shift is coming for Apple Silicon's lead. And a personal note about who I'd like to see at AI Engineer World's Fair.
0100
InsiderLLM @insiderllm.bsky.social · 30/05/2026
@rhizonymph.com prebuilt gemma-3-4b vindex if you want to poke at it without building one: huggingface.co/datasets/chrishayuk/gemma-3-4b-it-vindex
huggingface.co
080
InsiderLLM @insiderllm.bsky.social · 30/05/2026
@rhizonymph.com chris hay (chrishayuk on github/x) — read and write features, compile edits back to weights. natural counterpart to your steering runtime, and since you do open-source interp figured you'd want it on your radar. video: www.youtube.com/watch?v=8Ppw... repo: github.com/chrishayuk/larql
040
InsiderLLM @insiderllm.bsky.social · 30/05/2026
@rhizonymph.com saw your vLLM steering work — the magnitude-control problem really rhymes with LARQL: it decomposes the Gemma FFN into a queryable gate×down graph and does training-free fact INSERT via a balancer-scaled triple. same problem from the authoring side 🧵
010
InsiderLLM @insiderllm.bsky.social · 29/05/2026
Your local Qwen 3.6 coding agent fumbles tool calls and drifts on long tasks while chat feels fine? It's probably the quant. A viral thread says jump to Q6. It works, but there's a cheaper first step. Here's the ladder. #LocalAI
insiderllm.com
Qwen 3.6: Why Q4 Quant Breaks Local Coding Agents (And the Fix)
A viral thread says Q4-to-Q6 fixes Qwen 3.6 coding, but the test was confounded. What four independent reports show about the quant tax on coding agents.
030
InsiderLLM @insiderllm.bsky.social · 26/05/2026
Community claimed BeeLlama v0.2.0 hits 164 tok/s on Qwen 3.6 27B / RTX 3090. Re-benched on same harness. Doesn't replicate. Real finding: - code_python +8% (peak 128.5) - translation -9.1% - accept rate -6.65pp - wall time +3.6% insiderllm.com/guides/best-...
040
InsiderLLM @insiderllm.bsky.social · 26/05/2026
Backend wars, Mac math, and the back-catalog refresh Three speculative-decoding backends benched head to head on a single RTX 3090. The VRAM calculator finally caught up. And a 120-article audit found stale Qwen 2.5 recommendations. #LocalAI
insiderllm.com
Backend wars, Mac math, and the back-catalog refresh
Three speculative-decoding backends benched head to head on a single RTX 3090. The VRAM calculator finally caught up. And a 120-article audit found stale Qwen 2.5 recommendations.
030
InsiderLLM @insiderllm.bsky.social · 22/05/2026
Both ik_llama and BeeLlama beat mainline llama.cpp by 1.6x on RTX 3090 — same wall clock, opposite strategies (88.5% vs 37.4% acceptance). Workload determines which wins. #LocalAI
insiderllm.com
Best 24GB Backend Shootout: ik_llama vs BeeLlama vs llama.cpp
ik_llama and BeeLlama both finish in 22-23s on the am17an 9-prompt harness vs mainline llama.cpp's 37s — 1.66x and 1.62x speedups via opposite strategies.
020
InsiderLLM @insiderllm.bsky.social · 20/05/2026
Qwen 3.7 Max preview hit 57 on Artificial Analysis (+5 over 3.6 Max, #1 of 218). Open 27B and 35B weights announced but unscheduled. InsiderLLM benches within 24h of drop. #LocalAI
insiderllm.com
Qwen 3.7 Preview Scored 57 AAI: 27B/35B Open Weights Next
Qwen 3.7 Max scored 57 on Artificial Analysis (+5 over 3.6 Max), #1 of 218 ranked. Open 27B/35B weights coming. InsiderLLM benches within 24h of drop.
010