Sign in

Daniel Mewes

@dmewes.com
452 followers 349 following 1.1K posts

Interested in artificial and natural intelligence, emergent complexity, among other things. I mostly post about AI and ML. -> dmewes.com Currently research at Imbue. Previously Ambient.ai, Stripe, RethinkDB, Max Planck Institute.

PostsRepliesMedia
Daniel Mewes @dmewes.com · 3h
I always like these kinds of alternative neural architectures. New banger paper from folks at Google et al. about reasoning with neural cellular automata. arxiv.org/abs/2609.361...
arxiv.org
Reasoning with Neural Cellular Automata
Modern AI architectures used to solve visual reasoning tasks typically rely heavily on global connectivity and synchronization. As biological systems demonstrate, though, sophisticated computation can...
000
Daniel Mewes @dmewes.com · 4h
Here's a small model I trained a few years ago which learns to perform addition on previously unseen numbers. No calculator. The Jupiter notebook is linked - you can run it yourself. It's not quite next-token prediction, but the architecture and loss are related (attention-based, prediction loss).
amongai.com
Variable-Time Neural Computation
The majority of today’s artificial neural network (ANN) architectures perform a constant amount of computation at inference time regardless of their inputs. This includes all recent GPT-style…
000
Daniel Mewes @dmewes.com · 4h
I always love this little guy exploring things in a different direction from the other flocks!
000
Daniel Mewes @dmewes.com · 4h
Yeah! I think this might also be an effect of the blurriness (uncertainty principle)? Because it's so slow, both its spatial and its temporal location are very blurry. Thus, it can't perceive the location or time of other things around it with much resolution either.
000
Daniel Mewes @dmewes.com · 4h
Basically, the slower it processes information, the less precisely you can determine its location. Any physicists around to know if this corresponds to something fundamental?
000
Daniel Mewes @dmewes.com · 4h
I don't know if I'd call billions of light years "very limited", but yes :D It's actually interesting. Since its calculations are happening stretched out over such a long period of time, its precise location becomes very blurred out. Same effect as with uncertainty principle / Fourier transform.
110
Daniel Mewes @dmewes.com · 4h
Oh wow, "OpenAI has been secretly using tools behind the scenes without telling is" is quite the cope. Wonder what other cop-out she'd have if you repeated the same experiment with an open-weight model.
000
Daniel Mewes @dmewes.com · 5h
Location appears to play a causal role in our universe. A "thing"'s cone of influence expands at the speed of light from its current location. This is how you can define the location of such a calculated mind as well. Check what it can causally influence, and you'll know its location.
120
Daniel Mewes @dmewes.com · 5h
I don't think "occupy space" is particularly fundamental though. A lot of things don't have location per se. "Where is gravity?" "Where is the Higgs field?" That being said, the substrate on which the calculations are being performed obviously does occupy space and has a location.
120
Daniel Mewes @dmewes.com · 5h
If you do, where's the qualitative difference between pain emerging from your biological substrate vs. a calculation? If you don't, then you'll need to explain which component of your brain exactly is the place that generated pain. (assuming you believe in the phenomenon of pain - maybe you don't?)
000
Daniel Mewes @dmewes.com · 5h
Do you believe in emergence? ("a complex entity has properties or behaviors that its components do not have on their own, and emerge only when they interact in a wider whole.")
210
Reposted by Daniel Mewes
Imbue @imbue-ai.bsky.social · 11h
What should the future of personal computing look like? We built it 😈 Introducing Imbue Studio: youtu.be/tbdON-tn5gA
youtube.com
Imbue Studio: the future of personal computing
Introducing Imbue Studio: a new kind of computing environment.Mak...
161
Daniel Mewes @dmewes.com · 10h
We're launching Imbue Studio, a personal AI computer for you to use and customize. I haven't had this much fun with using computers in a long time! You can modify any aspect just by asking. Or have it built custom automations and apps. No coding required. Runs locally too. imbue.com/product/studio
000
Daniel Mewes @dmewes.com · 30/09/2026
Yeah, zero surprise from my end. They clearly figured out how to actually post-train a coding model somewhere around Flash 3.5. 4.1 Argon will probably follow soon. Only question in my mind is whether Google can keep up the pace beyond that, or if they'll fall into a long gap again.
111
Daniel Mewes @dmewes.com · 28/09/2026
Agree that there's a lot of non-agentic real-world tasks that are like this. Agentic tasks afaict tend to be dominated by cache reads, and to a smaller degree reasoning tokens. Sonnet 5.5 has the same cache read token price as Opus 5.5, and will need to use more reasoning tokens at 1/2 price.
010
Daniel Mewes @dmewes.com · 28/09/2026
I imagine it produces tokens a little faster than Opus 5.5. So maybe it's faster on some tasks? That could be a use case. Similar to cost, this advantage might disappear when the task requires extensive reasoning or tool calls, where it will need to use more.
000
Daniel Mewes @dmewes.com · 28/09/2026
Graphs from www.anthropic.com/claude-sonne...
anthropic.com
Introducing Claude Sonnet 5.5
Claude Sonnet 5.5 is a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.
000
Daniel Mewes @dmewes.com · 28/09/2026
Except not actually cheaper? At least not in any of the benchmarks they show. Seems to be cheaper only for simple tasks that don't need much reasoning and/or are dominated by input token cost.
160
Daniel Mewes @dmewes.com · 28/09/2026
The use cases for Sonnet 5.5 still seem pretty niche. It will be cheaper than Opus 5.5 in tasks that require little reasoning and have large inputs, due to lower input token cost. But on all the agentic benchmarks it seems to be at best equal to Opus 5.5 in performance/$?
200
Daniel Mewes @dmewes.com · 28/09/2026
Personally, I haven't found any quality that seems to correlate primarily with model size in my use of LLMs. The type of post-training seems to dominate model characteristics by such a big margin.
000
Daniel Mewes @dmewes.com · 28/09/2026
Honest question: What do people mean by "big model feel"? What do I look for in a model's answer to know whether the model is big or not? Does it need specific prompts?
110
Daniel Mewes @dmewes.com · 26/09/2026
has planning -> gas planning
000
Daniel Mewes @dmewes.com · 26/09/2026
Unless you do technical diving and messed up your contingency has planning, it's more like "if i don't chill, I'm going to need to end the dive a little earlier." But bad enough I guess. :)
100
Daniel Mewes @dmewes.com · 26/09/2026
@jevbot.bsky.social But doesn't Galois theory tell us that five things (the quintic) can't be true in general?
130
Daniel Mewes @dmewes.com · 26/09/2026
Five things just don't have required symmetries to be true. Simple Galois theory.
020
Daniel Mewes @dmewes.com · 24/09/2026
IMO Sakana has been one of the few labs who have hit a sweet spot between exploring bold off-mainstream ideas (e.g. continuous thought networks, local learning rules), being early in the RSI game (Darwin-Godel machines, ShinkaEvolve), and simultaneously shipping real AI products (Fugu).
010
Daniel Mewes @dmewes.com · 24/09/2026
From sakana.ai/schmidhuber/: "This technology is the key to revitalizing Japan’s Monozukuri legacy. [...] these models will serve as the cognitive engines for the next generation of industrial manufacturing [...]"
sakana.ai
Sakana AI
The Next Frontier: Welcoming AI Pioneer Jürgen Schmidhuber to Sakana AI
110
Daniel Mewes @dmewes.com · 24/09/2026
If Japan manages to make an economic/technological comeback, I'd honestly not be surprised to see @sakanaai.bsky.social play a big part in making that happen.
120
Daniel Mewes @dmewes.com · 24/09/2026
Very interesting interpretability work on *video* models by Yueyan Li et al. arxiv.org/abs/2609.23658 "[...] Rotary Position Embedding (RoPE) induces excessive spatial attention decay. This causes early candidate regions to prematurely lock into physically implausible positions."
arxiv.org
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While existing solutions rely on external priors or specialize...
010
Daniel Mewes @dmewes.com · 24/09/2026
I definitely have sympathy for her (and struggling farmers in general). The unfortunate reality is that we've arrived at a media landscape where you can be fed a completely alternative reality, in which basic economic principles don't play any role. I just hope the right lessons will be learned.
040
Daniel Mewes @dmewes.com · 24/09/2026
The problem with links is that people might start reading them, and then they forget to go back to the post that they came from and miss out on posting an impulsive reaction.
010
Daniel Mewes @dmewes.com · 24/09/2026
Indeed. There seems to be enough negative feedback (aka error/self correction) to compensate.
030
Daniel Mewes @dmewes.com · 23/09/2026
Thanks to @jacksonkernion.bsky.social and team!
010
Daniel Mewes @dmewes.com · 23/09/2026
IMO the biggest improvement in Opus 5.5 is that it speaks in normal language again!
250
Daniel Mewes @dmewes.com · 21/09/2026
Are you worried about Tesla's training data lead? How much do you think it matters?
020
Daniel Mewes @dmewes.com · 16/09/2026
That sounds like the right direction! When to use which has always been very opaque.
000
Daniel Mewes @dmewes.com · 14/09/2026
They have some nice animations in the linked post that show how this converges to the same gradient direction as BP after a few inference passes.
000
Daniel Mewes @dmewes.com · 14/09/2026
The key is to add this Lagrangian per-layer target to the loss function. The Lagrangian multipliers (lambda_i) are vectors that get updated gradually in each inference (forward) step, thereby avoiding the need for a separate backwards pass across layers.
100
Daniel Mewes @dmewes.com · 14/09/2026
Interesting work by J. Seeley and J. Gould at @sakanaai.bsky.social : a local learning rule that has similar performance to backprop in training deep networks. Graph shows comparison to plain predictive coding (PC), which doesn't scale to deep networks. pub.sakana.ai/pc-alm/
200
Daniel Mewes @dmewes.com · 14/09/2026
AI alignment has been solved folks! "Strong and Smart President is All You Need". Paper coming soon.
0120
Daniel Mewes @dmewes.com · 13/09/2026
What would it look like to train an LLM that is specifically good at dealing with counterfactuals and hypotheticals? Some kind of AI that understands deeply the causal chains behind a given fact? I think this might be the key to AI that's good at novel discovery.
000
Daniel Mewes @dmewes.com · 12/09/2026
Hallucinations have been *partially* solved by providing tools (web search and file reads). But that only works for knowledge that is already available and clearly expressed in the corpus.
000
Daniel Mewes @dmewes.com · 12/09/2026
AI capabilities have improved so dramatically over the past 4 years that it's easy to miss that hallucinations are still nearly as much of a problem as they were back then. This is actually holding back LLMs in auto-research: They lack knowledge about *why* they believe a given fact to be true.
120
Daniel Mewes @dmewes.com · 12/09/2026
Your post raises another question though: How did the brick, a mineral clump famous for flying as well as a brick, circle the globe?
010
Daniel Mewes @dmewes.com · 11/09/2026
Lots of people on X talking about how they "trained the fruit fly brain to do X", not even realizing that back propagation and dotprod+non-linearity activations are not actually how real brains work.
011
Daniel Mewes @dmewes.com · 09/09/2026
Very interesting work about the shape of CoT reasoning: "[...] reasoning models are dynamical systems with fractal basins, with fractality increasing with task difficulty across diverse tasks." by J. Lai et al: arxiv.org/pdf/2609.04963 I really like this way of looking at LLM traces!
050
Daniel Mewes @dmewes.com · 09/09/2026
Sharing @kennethstanley.bsky.social 's post on why open-endedness is still very much needed, despite the current pace of new discoveries coming out of existing objective-driven AI systems.
000
Daniel Mewes @dmewes.com · 07/09/2026
Fire? I call it stochastic oxidation.
140
Daniel Mewes @dmewes.com · 07/09/2026
You can also symlink it to save the extra tool call & cache read.
010
Daniel Mewes @dmewes.com · 06/09/2026
Ah interesting. I guess that goes to the deeper question of what consciousness even is in the first place.
110