Sign in

Colin

@colin-fraser.net
4.5K followers 202 following 7.6K posts

Anti-spreadsheet person

PostsRepliesMedia
Colin @colin-fraser.net · 3h
I guess it would depend on what the hypothesis was but yeah, probably. Hard to say for sure with these hypotheticals.
100
Colin @colin-fraser.net · 3h
If I had a hypothesis that said that image models couldn’t draw hands and then I saw them drawing hands, I would reject that hypothesis
100
Colin @colin-fraser.net · 3h
I don’t know what leap of faith you’re talking about
100
Colin @colin-fraser.net · 3h
I’m not following
100
Colin @colin-fraser.net · 4h
Why does the Mona Lisa wear that vexing smile?
030
Colin @colin-fraser.net · 7h
330
Colin @colin-fraser.net · 7h
Yeah, we are used to it now in 2026, it’s kind of incredible how good they are at this kind of arbitrary hand computation stuff over quite long windows
030
Colin @colin-fraser.net · 7h
As a matter of fact, writing these tests is highly tedious and annoying, and in the absence of LLMs you might be tempted to skip them. But since LLMs can write the tests, and the tests are very easy to review, I find my LLM-assisted code is actually higher quality because more tests get written.
000
Colin @colin-fraser.net · 7h
Right. Usually in any major software project, AI or no AI, you will have a large suite of automated tests that check whether the code works and you’re not allowed to push the code unless it passes the tests.
110
Colin @colin-fraser.net · 8h
Software engineering has quite a rich collection of practices and tools on this, since “guy that usually writes good code but sometimes makes a mistake” is basically just a description of any competent software engineer. So the field has this kind of error tolerance / resilience somewhat built in.
140
Colin @colin-fraser.net · 8h
But if your problem is, for example, that you need some code, and you’re rightly worried that the LLM might write bad code. There are many automated systems that exist for checking whether the code is bad and rejecting it unless it’s good.
100
Colin @colin-fraser.net · 8h
Making effective use of something that makes occasional errors is indeed a challenging problem, but depending on the context, there are many ways forward on this. To be clear though I’d recommend using a real calculator if your problem is that you want to do arithmetic.
120
Colin @colin-fraser.net · 8h
Tbf learning to add from structured training data and learning to approximately add from unstructured text aren’t really the same
120
Colin @colin-fraser.net · 8h
I’m the one who said that. Are you talking about me?
000
Colin @colin-fraser.net · 8h
Someone’s coworker in the training data asked for the exact same thing
140
Colin @colin-fraser.net · 8h
You might be increasingly right but I like it. I think it’s cute.
100
Colin @colin-fraser.net · 8h
I got news for you buddy. You are a neural network.
040
Reposted by Colin
noam @noamchompers.bsky.social · 8h
guy who sorta remembers a class from 15 years ago: as turing proved 40 years ago, claude is actually MORE conscious than people other than me another guy like that: actually the seat of consciousness is shown to be in the amygdala actual expert: i reckon a computer could be conscious probably, idk
5473
Colin @colin-fraser.net · 8h
Well it certainly doesn’t describe all neural networks but I think it describes large language models
000
Colin @colin-fraser.net · 8h
What's kind of funny about this reply is I'm more on your side about this than like 99% of people reading this.
080
Colin @colin-fraser.net · 8h
You seem to be missing the point pretty severely.
010
Colin @colin-fraser.net · 8h
...that's what I did lol. That's literally what this is about.
120
Colin @colin-fraser.net · 8h
Yes, it is apples and oranges.
000
Reposted by Colin
conputer dipshit @davidcrespo.bsky.social · 8h
well of course it’s doing math. how else would it be doing math
2212
Colin @colin-fraser.net · 8h
If you know what you're doing you can control whether it does this or not.
070
Colin @colin-fraser.net · 8h
Maybe whatever you are using is doing math with Python, that's entirely likely, but in the examples that I am posting, it's not doing that.
230
Colin @colin-fraser.net · 9h
I am learning that many people are very confused about whether and when LLMs can use Python
4440
Colin @colin-fraser.net · 9h
More on this
020
Colin @colin-fraser.net · 9h
Just to weigh in, the calculator is beside the point. The point is that LLMs appear to learn how to do something that approximates adding. We can verify that they are doing this with no calculator. So I was interested in learning Prof. Bender's POV on how to understand this wrt Stochastic Parrots.
150
Colin @colin-fraser.net · 9h
Yeah that's a good interpretation of it bsky.app/profile/coli...
010
Colin @colin-fraser.net · 9h
No, I've always had this as a project to try some day. Same with training a model to do formal logic.
160
Colin @colin-fraser.net · 9h
KL **penalty. Weird autocorrect
110
Colin @colin-fraser.net · 9h
Well sort of. I don't think this paper is very definitive, though I see that people seem to really like it.
110
Colin @colin-fraser.net · 9h
It's not so distantly related because of the KL mentality. It's still very close to the pretraining distribution. But yes it is of course not identical to the pretraining distribution; that's the point.
210
Colin @colin-fraser.net · 9h
It's a re-weighting of the pre-training distribution
100
Colin @colin-fraser.net · 9h
Well, of course a neural network can be programmed to do addition. The sightly harder question is whether training a neural network to do next token prediction on addition problems among many many other things is a way to do that.
1150
Colin @colin-fraser.net · 9h
This is GLM-5.2, ymmv. The big American labs ofc don't let you see the reasoning.
040
Colin @colin-fraser.net · 9h
Seems to help at making the final answers more accurate, but I think it's highly inefficient. Like it should be able to think a lot less and still get the answer, if it was efficiently reasoning. When they first invented reasoning I was saying they should have called it "rumination"
2100
Colin @colin-fraser.net · 10h
I was talking about 12 and 20 digit addition, and i also think the thinking models think too much. Withiut thinking its quite fast. What makes it slow / expensive is the model generates thousands of "wait let me rewrite that another way just to check"s etc
180
Colin @colin-fraser.net · 10h
Anthropic engineers are secretly giving it access to a hidden divining rod
1520
Reposted by Colin
Quantian @quantian.bsky.social · 10h
This was going around Twitter today, if you ask Claude “is X lat Y long land or water?” and plot the results, it has clearly embedded a map of the globe in its weights over time
5969
Colin @colin-fraser.net · 10h
I don't think it probably knows
100
Colin @colin-fraser.net · 10h
If you access the model outside of an application like ChatGPT you have fine control over whether it has access to a calculator, which I did here.
120
Colin @colin-fraser.net · 10h
If you poke around the replies in this thread you will find examples
110
Colin @colin-fraser.net · 10h
Idk
010
Colin @colin-fraser.net · 10h
Good point actually
110
Colin @colin-fraser.net · 10h
no, that is one thing we can rule out here
110
Colin @colin-fraser.net · 10h
The Number Helix seems a bit to cute to me
110
Colin @colin-fraser.net · 10h
Yes, there are ways to do that and for this example I used those ways.
100
Colin @colin-fraser.net · 10h
I don't think "maybe the computer program can kind of do addition" is that spooky
040