crumb @crumb.bsky.social · 3hthe task im doing inference for is kernel optimization already so i think i'm just gonna let it run for a night, extract top performers and use them to accelerate another stage, repeat bsky.app/profile/crum... 020
crumb @crumb.bsky.social · 3hall gpu, neither setup used speculative decoding, both flash yeah. idk. i think im just gonna go find a dflash or the original mtp heads and use the hf setup. lots of nice things that no other framework will offer besides like unsloth which makes env so particular it might not be worth 120
crumb @crumb.bsky.social · 3him trying to RL for kernel optimization and genuinely vanilla transformers + bitsandbytes seems like the best option for sampling on my hardware 🧍♀️ that doesnt sound right at all 170
crumb @crumb.bsky.social · 3hdemolishing my intuitions, llama.cpp only gets 8-10t/s for a 4bit 27b while bitsandbytes + transformers gets me 27t/s. what. why 4110
Reposted by crumbPhillip Isola @phillipisola.bsky.social · 6hIn PRH, we argued that rep geometry is converging, but it's still an open question exactly in what sense. I want to share some new evidence that might clarify the picture. The evidence comes from our work here, led by @schnaus.bsky.social: dominik-schnaus.github.io/unpaired-ros... 1/n 12410
crumb @crumb.bsky.social · 22hwe have tons of inference-only kernels that are INSANELY fast, hundreds to thousands of words per second. the training kernels can't do that speed yet. BUT if we expose token embeds as an input method, we could just treat the sys prompts as embeds to optimize a la prompt-tuning... 010
crumb @crumb.bsky.social · 22htrain the prompt yeah, just get the encodings of tokens for a seed prompt (or, i suppose you can start from random) and pass them to your optimizer. backprop through the model to the encodings. though i think evolutionary methods might be able to use this though because—i have to split this, cont- 110
crumb @crumb.bsky.social · 22hback in ye old times google found you can finetune the meanings of words in a prompt and it sometimes will outperform even full model finetuning. also, it wont break special megakernels like lora does :) www.geeksforgeeks.org/artificial-i... arxiv.org/pdf/2104.08691 120
crumb @crumb.bsky.social · 23hi just realized like ~nobody here was around for prompt tuning because it was 2021 🤦♀️ 011
crumb @crumb.bsky.social · 09/10/2026or i just want an excuse to send big unreadable files to people back and forth again 000
crumb @crumb.bsky.social · 09/10/2026we are prob mega under utilizing finetuning tokens a la prompt tuning & to less extent those vram maxing dreambooth configs. i wonder if good inference framework accepting inputs_ids would open some things up. evolutionary methods prob work stellar there 230
Reposted by crumbDoll @dollspace.gay · 07/10/2026What do the moderates want? Solar panels at the data centers? Sounds fucking great. Water offsets? We will help you push for that. Bans on deepfakes? Awesome we dont like that shit either. 11568
crumb @crumb.bsky.social · 07/10/2026im supposed to go to the dentist that's what im supposed to do 140
crumb @crumb.bsky.social · 07/10/2026yeah i think it's just gonna happen here throughout the next decade. idk what im supposed to do rn 140
crumb @crumb.bsky.social · 07/10/2026my uncle was on the phone literally two days ago saying "ai is all just smoke and mirrors, it can only say things people have said before" 1120
crumb @crumb.bsky.social · 07/10/2026sorry I don't really have anything substantial to say about the math thing other than all computers will be able to do this soon for all other fields of study 3502
crumb @crumb.bsky.social · 07/10/2026all i can do is drink my smoothie pretend everything is normal today 131
Reposted by crumbscanstone.bsky.social @scanstone.bsky.social · 15/09/2026I'm a published mathematician - I worked on the question of whether the unit balls of Banach spaces are plastic. OpenAI's GPT-6 Astra is *far beyond my abilities* in this question. I have read everything written on it. THERE IS NOTHING TO PLAGIARISE, I READ IT ALL AND IT'S NOT THERE 2854
crumb @crumb.bsky.social · 07/10/2026brainfuck swarm from zero needs an insane amount of compute that i do not have the time to afford it. will rethink main algo while i do other things 010
crumb @crumb.bsky.social · 06/10/2026actual computer just released toks, "the best tokenizer on earth and the first tokenizer designed for the quadrillion-token era." most of the work is done in pure assembly. genuinely nuts. actual.inc/company/blog... 39715
crumb @crumb.bsky.social · 06/10/2026the field is not "bury agents in systems" sorts of complex yet but it does point the way, i think 120
crumb @crumb.bsky.social · 06/10/2026an asi burying agents to do its bidding in dynamic processes like the weather or social interaction (or, really, a distributed process across many many weakly interacting systems...) is, i think, the most yud thing that i imagine regularly 170
crumb @crumb.bsky.social · 05/10/2026im imagining im just missing the whole reading bit, electrodes + board they get connected to before going to microcontroller? i have not dug too deep yet 250
crumb @crumb.bsky.social · 05/10/2026this would bring it much closer to feeling like an "exocortex" than anything else i could do IMO 0120
crumb @crumb.bsky.social · 05/10/2026i like. Really want to have a subvocalization flow like "uplink. {prompt}. [silence]" -> send to llm agent on my workstation over wifi -> earpiece. i want this more than a lot of things and i think i can do it 2141
crumb @crumb.bsky.social · 05/10/2026i just need to find like $25 for the hardware bits i dont have 1121
crumb @crumb.bsky.social · 05/10/2026if i made a nice robust model for subvocalization -> text that can handle different electrode placements would anyone be interested in that. would anyone use that 5212
crumb @crumb.bsky.social · 05/10/2026www.youtube.com/@departmento...youtube.comDepartment of Nonlinear Dynamics - BDU 020