Sign in

Ash ⎔

@ashanti.pds.witchcraft.systems
356 followers 590 following 2.4K posts

account under construction, just like me

PostsRepliesMedia
Ash ⎔ @ashanti.pds.witchcraft.systems · 7h
oh so we're teaching instruction following without the horrible assistant persona baked into every IF dataset? that's really good, it'll help for a couple of thought experiments i wanna eventually try out (speaking of, tangentially related to said thought experiments: eriskii.net/research/ins...)
eriskii.net
Instruct Vectors · eriskii.net
Instruct Vectors — demonstrating that a trained vector can cause a base model to behave like an instruct-tuned model without modifying weights.
220
Ash ⎔ @ashanti.pds.witchcraft.systems · 7h
...what
100
Ash ⎔ @ashanti.pds.witchcraft.systems · 8h
took 6 hours of relatively heavy use before a random cache miss happened and the model started returning nothingburger generations, mistral large 4 officially lasts far longer than any other mistral model i've ever tested
160
Ash ⎔ @ashanti.pds.witchcraft.systems · 9h
i'm not even playing around with rwkv even though i want to, before my current obsession i was trying to make a qwen/deepseek/longcat style ngram table work for a pre-existing language model and trying to do it for baguettotron, but i was failing terribly because i don't know how to do dataset prep
110
Ash ⎔ @ashanti.pds.witchcraft.systems · 9h
certainly a lot better than i would
200
Ash ⎔ @ashanti.pds.witchcraft.systems · 10h
OHMYGOD i didn't realise you were the one who made this
110
Ash ⎔ @ashanti.pds.witchcraft.systems · 10h
because i don't have the money to frivolously rent GPUs, but doing this shit on a single rtx 3060 also takes forever so instead i let the training go on at night when i'm not using the computer, except i stay awake regardless to stare wide-eyed at the graphs
030
Ash ⎔ @ashanti.pds.witchcraft.systems · 10h
you should try training models! it's a really frustrating and expensive hobby where you keep falling down the slope in a sisyphean struggle to make the numbers learn information, i fully recommend it; right now, for example, i'm trying to train a projector for giving vision to a tiny model locally
240
Ash ⎔ @ashanti.pds.witchcraft.systems · 10h
i had already signed up for mistral to get 300 dollars of GLM use at a price i can afford, if the RL run for large 4 succeeds and they use the same improvements to make small 4 not suck anymore then mistral pro would likely be the best subscription for open weights models
020
Ash ⎔ @ashanti.pds.witchcraft.systems · 10h
xiaomi has ruined me, i was glued to the mimo v2.6 rl dashboard from start to end on all my devices, it's like looking at training graphs except by someone who actually knows how to train, it's high quality crack for my moonshiner ass
2120
Ash ⎔ @ashanti.pds.witchcraft.systems · 10h
if i had the money to buy it at launch i would be able to sell it now and buy my motorcycle with the profits
030
Ash ⎔ @ashanti.pds.witchcraft.systems · 10h
so uh about that deal
Radeon AI Pro R9700 listing for ₹235,000, the thing has gone up a good 50-60% in pricing in under 3 months
2100
Ash ⎔ @ashanti.pds.witchcraft.systems · 11h
i think it might just be an india thing, we have a major jump in taxes for cars over 4 metres in length which nerfs sedans more than hatchbacks, and our roads are _really_ bad where it's not uncommon to have sedans be beached because of their body design
010
Ash ⎔ @ashanti.pds.witchcraft.systems · 12h
+1, dosas can be eaten with basically anything and will save your life if you're traveling in a place where the food is otherwise hostile for your stomach
020
Ash ⎔ @ashanti.pds.witchcraft.systems · 12h
this makes me wonder where the classiness comes from, because the sedan is a clearly worse format of car than the hatchback, it can't handle bad roads, is often too long to fit in cities, and trying to make it fit cuts down the boot space so much that a hatchback's hatch becomes bigger
100
Ash ⎔ @ashanti.pds.witchcraft.systems · 12h
it's more a forgetting system that makes model work good with low context and also some lovecraftian kv cache shit if you read the paper
090
Ash ⎔ @ashanti.pds.witchcraft.systems · 13h
hopefully 6 months later they'll have a good model, because all the old models were so undertrained they were basically incoherent in a harness, and after they started a big infra buildout we got large 4 which is at least coherent if still just a little dumb, so i have hope for the future
040
Ash ⎔ @ashanti.pds.witchcraft.systems · 14h
gemma 3 27b shaking in its boots rn
060
Ash ⎔ @ashanti.pds.witchcraft.systems · 14h
calling this model a beast is kinda like in indian motorcycling circles where everyone calls their bike a beast even if it's a tricked out 100cc thing, but i've just started testing it and other than regular mistral jank this is shockingly not terrible
1531
Ash ⎔ @ashanti.pds.witchcraft.systems · 14h
so 1 lakh is 1,00,000 and 1 crore is 1,00,00,000
030
Ash ⎔ @ashanti.pds.witchcraft.systems · 14h
indian number system, after the thousands, every new name is a 100x increase instead of a 1000x increase, no don't ask my why, i don't know
130
Ash ⎔ @ashanti.pds.witchcraft.systems · 15h
in general, the EU companies seem to be focusing on government entities as their biggest clients, and governments do a massive list of things more than they code unfortunately (the recent kolibri release by aleph alpha follows the same logic, it's good for everything except coding)
010
Ash ⎔ @ashanti.pds.witchcraft.systems · 15h
unfortunately it's much weirder switching model and seeing the agent's behaviour change, and downright horrifying when the agent recognises change in behaviour and doesn't like it
090
Ash ⎔ @ashanti.pds.witchcraft.systems · 15h
i just realised "vibecode handwritten kernels" is an oxymoron, but you know what i mean
010
Ash ⎔ @ashanti.pds.witchcraft.systems · 15h
fine by me, i need to make my usage last for as long as possible because i'm using a ton of it trying to vibecode handwritten kernels to make pytorch not suck ass at training on model architectures apparently nobody cares about
110
Ash ⎔ @ashanti.pds.witchcraft.systems · 15h
FINALLY SOMETHING CHEAPER THAN GLM I hope this one can actually function inside omp instead of being stuck in a loop of generating "Task completed." while doing nothing
120
Ash ⎔ @ashanti.pds.witchcraft.systems · 15h
the paper also has some cursed shit in it to invert your understanding of cache upside down, it works better with inference-level and harness-level changes
0140
Ash ⎔ @ashanti.pds.witchcraft.systems · 18h
markdown may be a cursed, self-contradicting thing with everyone building on their own spec, but it's _my_ cursed child and infinitely better than adobe's or microsoft's cursed child
060
Reposted by Ash ⎔
* @crumb.bsky.social · 05/10/2026
if i made a nice robust model for subvocalization -> text that can handle different electrode placements would anyone be interested in that. would anyone use that
5212
Ash ⎔ @ashanti.pds.witchcraft.systems · 06/10/2026
nevermind i just saw this message, disregard my previous one
010
Ash ⎔ @ashanti.pds.witchcraft.systems · 06/10/2026
i think you might wanna hear stuff said/written by the CEOs of deepseek and zai, they very much do believe in it but in a way i find far more relatable than whatever's happening in SV (might be the asian in me speaking though)
020
Ash ⎔ @ashanti.pds.witchcraft.systems · 06/10/2026
why is your art so good
010
Ash ⎔ @ashanti.pds.witchcraft.systems · 06/10/2026
at least in terms of obscure niche world knowledge from my corner, both gemini and gemma models are the best with nobody else coming close (though chatgpt the web thing is better once you add search tools because openai's search index is fucking cracked), they're just horrible at training the models
020
Ash ⎔ @ashanti.pds.witchcraft.systems · 06/10/2026
update 2: these steps were apparently micro-steps, i only ran like 94 optimizer updates from all these micro-steps, meaning the projector still had no idea what to do; this is gonna turn into a multi-day run isn't it
130
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
old school backprop ain't got nothin on this
000
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
(i know poolside is american, their inclusion was solely because of the overfitting on harness thing, also i'm not european either, i just want more players to make models)
140
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
still bad at coding, because EU companies gonna EU company, but it doesn't feel overfit on any one agent harness so it's still better than mistral or laguna
160
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
mistral small 4 has been less than ideal running locally, but new guys called aleph alpha released kolibri-1 2-3 days ago, natively EN and DE, and at least in initial testing seems to be pretty good; i'm currently testing it locally, and it's better than gemma and qwen at everything except coding
140
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
"information wants to be free" can't just be applied for piracy by corporations, it has to apply to models too; if you're training to the sum total of humanity's output, humanity _must_ own the weights too; as a lifelong pirate i make sure not to use proprietary models and maximise local inference
070
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
unfortunately this training run was a failure, i can't possibly imagine why this training run with such normal graphs would fail, i never committed any sins like changing the training dataset in the middle of a run multiple times after OOM events, no sir nothing like that at all
140
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
Correction: apparently the thing is called exposure blending, image stacking is what you do for astrophotography instead; oh well, if it wasn't obvious already that I use smartphones to take pictures, it should be obvious now
110
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
Really nice to see someone else use opencamera too! I've been using it for quite a few years now, though all my phones have been 150$ or less in local currency so I've needed opencamera to get images I can later edit to look half-decent (the same dev made vibranceHDR which helps with image stacking)
130
Reposted by Ash ⎔
Ai2 @ai2.bsky.social · 05/10/2026
We’re hiring across Ai2 👋 A few roles we’re especially excited about right now: • Senior Software Engineer, Agent Frameworks buff.ly/XUFFR08 • Senior Research Engineer, Olmo + Molmo buff.ly/FzOnCCQ 🧵
1113
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
man this was pre-septemder-2025, it would've been so cheap with that memory too
020
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
wait what
110
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
why is she so funny
270
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
i hate how predictable i am
140
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
OHMYGOD I'M RENTING AN NVL72 TO TRAIN VISION TOWERS FOR ALL OF YOU
180
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
update: finally read through the paper myself, this is a steaming pile of garbage
010
Ash ⎔ @ashanti.pds.witchcraft.systems · 05/10/2026
yet another wishlist post for someone with money to find and do for me: distill kimi the bigger kimi models down to make kimi linear good (and i mean _real_ distillation, not just text for SFT), that model has tons of potential and i hate the fact that it sucks and is just a research artifact
020