Sign in

A. H. Zakai

@kripken.com
4.9K followers 718 following 2K posts

Software engineer (compilers: WebAssembly, Emscripten, Binaryen). Used to study neural networks. Loves fantasy novels and Agatha Christie. he/they All opinions here are my own, not my employer's (Google). More in: kripken.github.io/blog/about

PostsRepliesMedia
A. H. Zakai @kripken.com · 3h
But this is not true. Intelligence does not require neurons, biological or artificial. Still, it is very natural for many people - myself included - to focus on the neural network part because of this anthropocentric bias, "looking for brain-like things". A hard habit to get out of!
060
A. H. Zakai @kripken.com · 3h
This is a form of anthropomorphization because it assumes that, to *do* brain-like things, one *needs* a brain-like thing. That the neural network part of an LLM might be intelligent, but the calculator could never be relevant there, even as a component.
130
A. H. Zakai @kripken.com · 3h
My theory is that this is actually based on anthropomorphization (ironically). LLMs are contrasted against brains. Artificial neural networks against biological ones. Adding a calculator is just not playing the same game.
130
A. H. Zakai @kripken.com · 3h
I mean that another response to LLMs getting bundled calculators (and Python and Lean) is to say "well, LLMs can do math now", because that is what the total system does. But this feels wrong to many people, apparently including Bender.
230
A. H. Zakai @kripken.com · 3h
"For historical reasons" is maybe part of the answer; Bender has been fighting against neural networks that do language for quite some time. Adding a calculator to the LLM is like cheating: "aha! you can't do math, and need to resort to tricks!" But why does this feel like a trick?
130
A. H. Zakai @kripken.com · 3h
Bender obviously wants to say that LLMs are weak. They need a calculator to even do basic math. But why attack the neural network part of the model specifically?
130
A. H. Zakai @kripken.com · 3h
I think it's worth asking: why would it matter if a model were using a calculator? Why is this important to Bender&others? In the end, an LLM is a computer program, and a calculator is another. Together, they are a larger one. Why consider internal divisions?
2100
A. H. Zakai @kripken.com · 7h
Yes, that is a nice distinction that they made.
010
A. H. Zakai @kripken.com · 8h
I mean that this is clearly a sign of imperfection, but I do think humans have similar imperfections in our internal "traces"?
110
A. H. Zakai @kripken.com · 8h
It kind of seems human-like to me...? I sometimes hum an unrelated song while working out a logic problem.
220
A. H. Zakai @kripken.com · 10h
Definitely, yes. Gyges has a nice paragraph about that: www.verysane.ai/p/do-we-unde...
verysane.ai
Do we understand how neural networks work?
Yes and no, but mostly no.
000
A. H. Zakai @kripken.com · 11h
I think it is meant to explain how they can emit fluent text, if they aren't copying it from their input or using simple rules (like ELIZA did). They emit fluent text using complex statistics ("stochastic"), is the picture here. And it has at least *some* truth in it, I think.
100
A. H. Zakai @kripken.com · 12h
Yeah, I definitely agree about "haphazardly" - that was really unnecessary in their paper, and is just absolutely false given the last decade of research. These systems do learn useful models of the input.
120
A. H. Zakai @kripken.com · 12h
("Listen to a long explanation by an expert" is the best thing, obviously, but in terms of reaching a wide audience... I think we need more.)
000
A. H. Zakai @kripken.com · 12h
I really like text elemental too! But I think it helps people like us more than it helps non-experts. Average people need something too, and yeah, the Stochastic Parrot is far from perfect, but I'm not aware of something better for them.
210
A. H. Zakai @kripken.com · 13h
What better metaphors do you have in mind?
120
A. H. Zakai @kripken.com · 13h
I'm not saying Bender is right in that thread, or in general, but she does make some valid points elsewhere. And some of the other Parrot Paper authors do an even better job at pointing out what the metaphor is actually good for. If it isn't taken too far, it is useful imho.
100
A. H. Zakai @kripken.com · 13h
Like, someone that thinks their AI boyfriend/girlfriend loves them has gone wrong, and reminding them "it is just a machine" is helpful. And the parrot metaphor explains how a machine can produce compelling text despite being just a machine (correlations on the input data etc.).
110
A. H. Zakai @kripken.com · 13h
The less cynical argument is that most people intuitively feel that LLMs are more humanlike than they actually are, and the Stochastic Parrot metaphor is a good counterweight.
100
A. H. Zakai @kripken.com · 30/09/2026
Students can write math on a tablet and it can actually check it for them, immediately, before they turn in the exercise to be graded. Executing math from handwriting, parallel to executing typed programs. Strange what is possible now.
010
A. H. Zakai @kripken.com · 30/09/2026
Back when I started my math undergrad, a professor jokingly told my class, "Double-check your work! When compsci students write silly things the computer gives them errors, but with math, the paper doesn't complain" I realize that, today, the paper *could* complain.
100
Reposted by A. H. Zakai
Cloudflare @cloudflare.social · 28/09/2026
With the new experimental support for the Emscripten target in Rust Workers, many previously unsupported Rust libraries and applications can now be built and deployed directly to Cloudflare’s global Workers platform, including upcoming support for Tokio async. cfl.re/4hlHSNC #BirthdayWeek
blog.cloudflare.com
Supporting native Rust in Workers with the new Emscripten target for wasm-bindgen
With the new experimental support for the Emscripten target in Rust Workers, many previously unsupported Rust libraries and applications can now be built and deployed directly to Cloudflare’s global W...
0175
A. H. Zakai @kripken.com · 28/09/2026
Yes, that might work! But the author is speculating about it: we have never been in that situation before in the browser field (or in OSes, etc.) and it might not work out the way he expects. But hopefully it will work out.
000
A. H. Zakai @kripken.com · 28/09/2026
I do agree air-gapping would have avoided the hacks we are talking about. I guess my point is that the downsides (preventing defensive uses) might outweigh the upsides. So it's not obvious to me OpenAI did anything wrong by not air-gapping. (But I am not 100% certain here.)
100
A. H. Zakai @kripken.com · 27/09/2026
Also, air-gapped models are not good enough for defensive uses of LLMs, like this: blog.mozilla.org/en/firefox/p... Only running air-gapped models would make the industry less secure in some ways & maybe even less secure overall.
blog.mozilla.org
The zero-days are numbered  | The Mozilla Blog
Since February, the Firefox team has been working around the clock using frontier AI models to find and fix latent security vulnerabilities in the browser.
110
A. H. Zakai @kripken.com · 27/09/2026
They also appear to believe it is less risky for them to develop it, containing it as best they can, rather than letting a foreign dictator develop it first.
210
A. H. Zakai @kripken.com · 27/09/2026
Unfortunately, this is true. More here: bsky.app/profile/krip...
220
A. H. Zakai @kripken.com · 27/09/2026
As a software engineer, and also someone with an academic background in neural networks, Will is 100% correct here.
170
A. H. Zakai @kripken.com · 27/09/2026
OpenAI certainly made mistakes, and this removes none of their responsibility. But some people think that if we just used LLMs in some safe way, everything would be fine. There is no such option. This technology is extremely capable, and not entirely predictable.
011
A. H. Zakai @kripken.com · 27/09/2026
If you used a web browser today, you used software that was more secure thanks to LLMs running without safeguards.
100
A. H. Zakai @kripken.com · 27/09/2026
Running LLMs without anti-hacking safeguards is standard practice and absolutely necessary. Defenders *must* do this, because attackers certainly are - using open models, or jailbroken ones, or their own.
100
A. H. Zakai @kripken.com · 27/09/2026
On "why did they disable the anti-hacking safeguards?": for very good reason. This is how browsers etc. are made secure, by using LLMs to find exploits before the bad guys do: blog.mozilla.org/en/firefox/p...
blog.mozilla.org
The zero-days are numbered  | The Mozilla Blog
Since February, the Firefox team has been working around the clock using frontier AI models to find and fix latent security vulnerabilities in the browser.
100
A. H. Zakai @kripken.com · 27/09/2026
On sandboxing: I've worked in the web browser field for 16 years. Browsers like Chrome, Firefox, and Safari have large, well-funded teams of some of the best engineers on the planet, but their sandboxes still get exploited by humans. LLMs are even better. "Just sandbox better" isn't enough.
110
A. H. Zakai @kripken.com · 27/09/2026
Thread about claims like "OpenAI should have just sandboxed the agents better" "OpenAI shouldn't have run the agents with the anti-hacking guardrails off" (so many articles are making these points that I'm not going to bother linking)
120
A. H. Zakai @kripken.com · 27/09/2026
Do you think there are appropriate uses of AI and inappropriate ones? I can understand the view that AI shouldn't be used at all, but if you accept it has valid uses, then telling those apart is the problem here, I think.
000
A. H. Zakai @kripken.com · 27/09/2026
Do you think AI can be trusted to be safe when the guardrails are on?
310
A. H. Zakai @kripken.com · 27/09/2026
Thanks, I hadn't seen that post before. Yes, maybe not all the authors, then...
000
A. H. Zakai @kripken.com · 27/09/2026
Yes, I agree with you that multimodality wasn't as necessary as they claim (and it would be nice to see them admit this!) But (even if for the wrong reasons) they are not still claiming that modern LLMs suffer from a fundamental barrier to understanding/meaning, afaict.
110
A. H. Zakai @kripken.com · 27/09/2026
Granted, she appears to say this grudgingly (multimodal models can only understand "in an extremely thin way"), but I do think she has conceded the point?
000
A. H. Zakai @kripken.com · 27/09/2026
I'm not sure this is accurate? Bender, for example, accepts that multimodal models can have understanding and grasp meaning: medium.com/@emilymenonb....
medium.com
Stochastic Parrots 🦜: Frequently Unasked Questions
It’s been a bit over five years since the Stochastic Parrots paper (Bender, Gebru et al 2021) was published (and somewhat longer since…
210
A. H. Zakai @kripken.com · 26/09/2026
This is my academic field, so I am pretty confident i have this right. But fair enough if you don't want to take my word for things!
000
A. H. Zakai @kripken.com · 26/09/2026
Yes, I guess our experience is very different here. Interesting!
000
A. H. Zakai @kripken.com · 26/09/2026
Sorry, I've explained it as best I can, as someone from this field. This is how we use the terms.
100
A. H. Zakai @kripken.com · 26/09/2026
Obvious slop is easy to dismiss, yes. But lots of AI PRs are not like that?
110
A. H. Zakai @kripken.com · 26/09/2026
1. "reasoning model" is a very technical term (a model with a chain-of-thought etc.) 2. "reasoning mechanism" - the commonplace meaning. A "reasoning model" (1) can lack robust reasoning (2). This actually shows nicely the interplay between technical and commonplace meanings!
100
A. H. Zakai @kripken.com · 26/09/2026
No, the paper actually shows two meanings of "reasoning". You can see them in the abstract: "While reasoning models are more capable, they nonetheless show high variance across problem presentations, suggesting they lack a truly robust reasoning mechanism. "
100
A. H. Zakai @kripken.com · 26/09/2026
Perhaps you were confused by the part in the picture? *One* of the compared variants (ILP Python) uses an LLM to generate code for a specialized solver, that is true. But all the others test the main point of the paper: LLM reasoning.
We also implemented an ILP Python prompt$
ing strategy, which prompts the LLM to translate
the problem instance into Python code⁴ that calls
the Gurobi solver on an Integer Linear Program
(ILP) encoding of the instance (Gurobi Optimiza$
tion LLC, 2024), cf. AhmadiTeshnizi et al. (2024).
Thus, ILP Python does not attempt to solve the
problem through LLM reasoning; the problem is
solved exactly and optimally by Gurobi, and the
LLM merely translates the NL specifications to
code and then translates the code’s output back
into NL. If the code generated by the LLM pro$
duces an error, we halt the process and count it as
a failure.
100
A. H. Zakai @kripken.com · 26/09/2026
(I'm not saying your perspective is wrong - I see how it makes sense to treat all PRs equally. Like peer review in academia.)
100
A. H. Zakai @kripken.com · 26/09/2026
Yes, but "it must be indistinguishable" means I must treat *all* incoming PRs equally. So I must review AI PRs just as soon as human ones - which I don't want to do. Because if a human wrote it, they deserve to get a human to read it, sooner, in my opinion.
100
A. H. Zakai @kripken.com · 26/09/2026
I get what you're saying, these are all valid points, but at the same time I find it useful to know who wrote the code I review. I mean practically: 1. I trust AI-written code less. I want to review it more carefully. 2. I want to give priority to human-written code (I leave AI PRs for later).
110