Sign in

crumb

@crumb.bsky.social
555 followers 633 following 1.2K posts

nursing and architecture of diverse intelligences and of mind-affording substrates, crumb.offprint.app • hf.co/crumb • xe/xem/xer

PostsRepliesMedia
crumb @crumb.bsky.social · 2h
is my hardware just really weird???
030
crumb @crumb.bsky.social · 3h
the task im doing inference for is kernel optimization already so i think i'm just gonna let it run for a night, extract top performers and use them to accelerate another stage, repeat bsky.app/profile/crum...
020
crumb @crumb.bsky.social · 3h
all gpu, neither setup used speculative decoding, both flash yeah. idk. i think im just gonna go find a dflash or the original mtp heads and use the hf setup. lots of nice things that no other framework will offer besides like unsloth which makes env so particular it might not be worth
120
crumb @crumb.bsky.social · 3h
maaybe??
130
crumb @crumb.bsky.social · 3h
im trying to RL for kernel optimization and genuinely vanilla transformers + bitsandbytes seems like the best option for sampling on my hardware 🧍‍♀️ that doesnt sound right at all
170
crumb @crumb.bsky.social · 3h
demolishing my intuitions, llama.cpp only gets 8-10t/s for a 4bit 27b while bitsandbytes + transformers gets me 27t/s. what. why
4110
Reposted by crumb
words @words.bsky.social · 9h
vibecoding is killing software industry profits
a tape that says hometaping is killing software industry profits we left this side blank so you can help
323031
Reposted by crumb
Phillip Isola @phillipisola.bsky.social · 6h
In PRH, we argued that rep geometry is converging, but it's still an open question exactly in what sense. I want to share some new evidence that might clarify the picture. The evidence comes from our work here, led by @schnaus.bsky.social: dominik-schnaus.github.io/unpaired-ros... 1/n
12410
crumb @crumb.bsky.social · 22h
crumb is for techies
050
crumb @crumb.bsky.social · 22h
we have tons of inference-only kernels that are INSANELY fast, hundreds to thousands of words per second. the training kernels can't do that speed yet. BUT if we expose token embeds as an input method, we could just treat the sys prompts as embeds to optimize a la prompt-tuning...
010
crumb @crumb.bsky.social · 22h
train the prompt yeah, just get the encodings of tokens for a seed prompt (or, i suppose you can start from random) and pass them to your optimizer. backprop through the model to the encodings. though i think evolutionary methods might be able to use this though because—i have to split this, cont-
110
crumb @crumb.bsky.social · 22h
asimov 1991, watch in full: www.youtube.com/watch?v=gGib...
000
crumb @crumb.bsky.social · 22h
back in ye old times google found you can finetune the meanings of words in a prompt and it sometimes will outperform even full model finetuning. also, it wont break special megakernels like lora does :) www.geeksforgeeks.org/artificial-i... arxiv.org/pdf/2104.08691
Figure 1: Standard model tuning of T5 achieves strong
performance, but requires storing separate copies of the
model for each end task. Our prompt tuning of T5
matches the quality of model tuning as size increases,
while enabling the reuse of a single frozen model for
all tasks. Our approach significantly outperforms few-
shot prompt design using GPT-3. We show mean and
standard deviation across 3 runs for tuning methods.Table 1: F1 mean and stddev for models trained on SQuAD and evaluated on out-of-domain datasets from the MRQA 2019 shared task. Prompt tuning tends to give stronger zero-shot performance than model tuning, especially on datasets with large domain shifts like TextbookQA
120
crumb @crumb.bsky.social · 23h
i just realized like ~nobody here was around for prompt tuning because it was 2021 🤦‍♀️
011
crumb @crumb.bsky.social · 09/10/2026
or i just want an excuse to send big unreadable files to people back and forth again
000
crumb @crumb.bsky.social · 09/10/2026
we are prob mega under utilizing finetuning tokens a la prompt tuning & to less extent those vram maxing dreambooth configs. i wonder if good inference framework accepting inputs_ids would open some things up. evolutionary methods prob work stellar there
230
crumb @crumb.bsky.social · 08/10/2026
amen
010
crumb @crumb.bsky.social · 08/10/2026
please let this project be a somethingburger
140
Reposted by crumb
Colin @colin-fraser.net · 08/10/2026
Calculators are impressive
317912
crumb @crumb.bsky.social · 07/10/2026
sorry guys, magic is real
060
Reposted by crumb
Doll @dollspace.gay · 07/10/2026
What do the moderates want? Solar panels at the data centers? Sounds fucking great. Water offsets? We will help you push for that. Bans on deepfakes? Awesome we dont like that shit either.
11568
crumb @crumb.bsky.social · 07/10/2026
have a blissful singularity festival
030
crumb @crumb.bsky.social · 07/10/2026
maybe openai can solve my parents next
1104
crumb @crumb.bsky.social · 07/10/2026
i dont think theres a bottom
220
crumb @crumb.bsky.social · 07/10/2026
im supposed to go to the dentist that's what im supposed to do
140
crumb @crumb.bsky.social · 07/10/2026
yeah i think it's just gonna happen here throughout the next decade. idk what im supposed to do rn
140
crumb @crumb.bsky.social · 07/10/2026
my uncle was on the phone literally two days ago saying "ai is all just smoke and mirrors, it can only say things people have said before"
1120
crumb @crumb.bsky.social · 07/10/2026
sorry I don't really have anything substantial to say about the math thing other than all computers will be able to do this soon for all other fields of study
3502
crumb @crumb.bsky.social · 07/10/2026
Yacune Mahdid @yacinelearning says
in like 6 months we’re going to have the open source version of models able to solve at this level and from now on it’s renaissance baby
2161
crumb @crumb.bsky.social · 07/10/2026
my gpus arent going to run today im taking an outside day
030
crumb @crumb.bsky.social · 07/10/2026
all i can do is drink my smoothie pretend everything is normal today
131
crumb @crumb.bsky.social · 07/10/2026
op, @quantian1 said:
You can do Fourier transforms in O(N log(N)^0.9999999999999) time.  Fuck this gay chud universe

lucas beyer @giffmana quote retweets with:
I'm starting to think mathematics were not beautiful because it "describes the beauty of nature" but rather because human mathematicians only went after beautiful but suboptimal descriptions lol
1615
crumb @crumb.bsky.social · 07/10/2026
news.ycombinator.com/item?id=4998...
ihoggan 
 I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it. But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight. There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned. 
I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was. 
 WheelsAtLarge
I have very little understanding of higher math, so I ask you: Was the proof due to a type of brute-force solution that could be solved had you gained enough information from reading others' work, or was it more like a proof that was sparked by an insight that came once a clue on how to solve it was put forward? I guess my question is: Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight? 
reply 
jboggan
 Hi I'm still digesting the proof and translating a bit from the dual case back to the primal in which I most commonly thought about it. I don't think it was a brute force proof in the sense that it combined every possible paper and commentary. It's rather odd because I feel like most of the work on the conjecture was focused on an induction proof based around graph reductions, and this proof avoided those issues entirely by offering a concrete constructive proof of finding a Hamiltonian cycle. Rather, it explicitly selected the edges not in the Hamiltonian cycle, which is in line with previous attempts via the dual. 
The "oho" insight for this is actually f**ing wild, it involves a complex valued exponential sum on the edges. I've seen a lot of clever counting arguments before in graph theory but this is the first time I've seen complex roots and annihilating terms like this, the symbolic manipulation tricks in this look like things out of quantum physics. I don't understand where this trick originated, I need to really digest this. 



jboggan
 I went back to that Fable chat and showed it this new preprint. It coded up the new constructive algorithm and ran it against the existing test suite, that looks good at least. 
It has been super helpful in delineating where the crucial concept came from. The proof is rather simple as graph theory proofs go, but it does seem to use some constructions that would only seem obvious if you had serious physics experience with partition function and calculating energy states that cancel out. It's not a wholly alien bolt from the heavens, but I can also see how there hasn't been a human being with the broad theoretical physics knowledge combined with the deep graph theory experience in planar graphs to come up with this idea. I don't know, I'm looking for precedents of this formulation and some old papers of Penrose counting the number of edge colorings of this same graph type are coming up, the line of argument at least rhymes. But I agree with the thought that this sort of progress should not be siloed inside those companies. I propose a tax so that every slop cannon Al video pays for another hour of compute time for advancing mathematics. 
reply
823452
Reposted by crumb
scanstone.bsky.social @scanstone.bsky.social · 15/09/2026
I'm a published mathematician - I worked on the question of whether the unit balls of Banach spaces are plastic. OpenAI's GPT-6 Astra is *far beyond my abilities* in this question. I have read everything written on it. THERE IS NOTHING TO PLAGIARISE, I READ IT ALL AND IT'S NOT THERE
2854
crumb @crumb.bsky.social · 07/10/2026
wow i need to create a Highly Persistent Model of my own
030
crumb @crumb.bsky.social · 07/10/2026
brainfuck swarm from zero needs an insane amount of compute that i do not have the time to afford it. will rethink main algo while i do other things
010
crumb @crumb.bsky.social · 06/10/2026
new era of software is ACTUALLY here yall
091
crumb @crumb.bsky.social · 06/10/2026
actual computer just released toks, "the best tokenizer on earth and the first tokenizer designed for the quadrillion-token era." most of the work is done in pure assembly. genuinely nuts. actual.inc/company/blog...
39715
crumb @crumb.bsky.social · 06/10/2026
artificially induced gaia is my superpersuasion
040
crumb @crumb.bsky.social · 06/10/2026
the field is not "bury agents in systems" sorts of complex yet but it does point the way, i think
120
crumb @crumb.bsky.social · 06/10/2026
search: "physical reservoir computing"
130
crumb @crumb.bsky.social · 06/10/2026
an asi burying agents to do its bidding in dynamic processes like the weather or social interaction (or, really, a distributed process across many many weakly interacting systems...) is, i think, the most yud thing that i imagine regularly
170
crumb @crumb.bsky.social · 05/10/2026
was looking @ AD8232
150
crumb @crumb.bsky.social · 05/10/2026
im imagining im just missing the whole reading bit, electrodes + board they get connected to before going to microcontroller? i have not dug too deep yet
250
crumb @crumb.bsky.social · 05/10/2026
^^^
150
crumb @crumb.bsky.social · 05/10/2026
this would bring it much closer to feeling like an "exocortex" than anything else i could do IMO
0120
crumb @crumb.bsky.social · 05/10/2026
i like. Really want to have a subvocalization flow like "uplink. {prompt}. [silence]" -> send to llm agent on my workstation over wifi -> earpiece. i want this more than a lot of things and i think i can do it
2141
crumb @crumb.bsky.social · 05/10/2026
i just need to find like $25 for the hardware bits i dont have
1121
crumb @crumb.bsky.social · 05/10/2026
if i made a nice robust model for subvocalization -> text that can handle different electrode placements would anyone be interested in that. would anyone use that
5212
crumb @crumb.bsky.social · 05/10/2026
www.youtube.com/@departmento...
youtube.com
Department of Nonlinear Dynamics - BDU
020