Sign in

Daniel Mewes

@dmewes.com
451 followers 348 following 1.1K posts

Interested in artificial and natural intelligence, emergent complexity, among other things. I mostly post about AI and ML. -> dmewes.com Currently research at Imbue. Previously Ambient.ai, Stripe, RethinkDB, Max Planck Institute.

PostsRepliesMedia
Reposted by Daniel Mewes
Imbue @imbue-ai.bsky.social · 3h
What should the future of personal computing look like? We built it 😈 Introducing Imbue Studio: youtu.be/tbdON-tn5gA
youtube.com
Imbue Studio: the future of personal computing
Introducing Imbue Studio: a new kind of computing environment.Mak...
111
Daniel Mewes @dmewes.com · 1h
We're launching Imbue Studio, a personal AI computer for you to use and customize. I haven't had this much fun with using computers in a long time! You can modify any aspect just by asking. Or have it built custom automations and apps. No coding required. Runs locally too. imbue.com/product/studio
000
Daniel Mewes @dmewes.com · 28/09/2026
The use cases for Sonnet 5.5 still seem pretty niche. It will be cheaper than Opus 5.5 in tasks that require little reasoning and have large inputs, due to lower input token cost. But on all the agentic benchmarks it seems to be at best equal to Opus 5.5 in performance/$?
200
Daniel Mewes @dmewes.com · 28/09/2026
Honest question: What do people mean by "big model feel"? What do I look for in a model's answer to know whether the model is big or not? Does it need specific prompts?
110
Daniel Mewes @dmewes.com · 24/09/2026
If Japan manages to make an economic/technological comeback, I'd honestly not be surprised to see @sakanaai.bsky.social play a big part in making that happen.
120
Daniel Mewes @dmewes.com · 24/09/2026
Very interesting interpretability work on *video* models by Yueyan Li et al. arxiv.org/abs/2609.23658 "[...] Rotary Position Embedding (RoPE) induces excessive spatial attention decay. This causes early candidate regions to prematurely lock into physically implausible positions."
arxiv.org
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While existing solutions rely on external priors or specialize...
010
Daniel Mewes @dmewes.com · 23/09/2026
IMO the biggest improvement in Opus 5.5 is that it speaks in normal language again!
250
Daniel Mewes @dmewes.com · 14/09/2026
Interesting work by J. Seeley and J. Gould at @sakanaai.bsky.social : a local learning rule that has similar performance to backprop in training deep networks. Graph shows comparison to plain predictive coding (PC), which doesn't scale to deep networks. pub.sakana.ai/pc-alm/
200
Daniel Mewes @dmewes.com · 14/09/2026
AI alignment has been solved folks! "Strong and Smart President is All You Need". Paper coming soon.
0120
Daniel Mewes @dmewes.com · 13/09/2026
What would it look like to train an LLM that is specifically good at dealing with counterfactuals and hypotheticals? Some kind of AI that understands deeply the causal chains behind a given fact? I think this might be the key to AI that's good at novel discovery.
000
Daniel Mewes @dmewes.com · 12/09/2026
AI capabilities have improved so dramatically over the past 4 years that it's easy to miss that hallucinations are still nearly as much of a problem as they were back then. This is actually holding back LLMs in auto-research: They lack knowledge about *why* they believe a given fact to be true.
120
Daniel Mewes @dmewes.com · 11/09/2026
Lots of people on X talking about how they "trained the fruit fly brain to do X", not even realizing that back propagation and dotprod+non-linearity activations are not actually how real brains work.
011
Daniel Mewes @dmewes.com · 09/09/2026
Very interesting work about the shape of CoT reasoning: "[...] reasoning models are dynamical systems with fractal basins, with fractality increasing with task difficulty across diverse tasks." by J. Lai et al: arxiv.org/pdf/2609.04963 I really like this way of looking at LLM traces!
050
Daniel Mewes @dmewes.com · 09/09/2026
Sharing @kennethstanley.bsky.social 's post on why open-endedness is still very much needed, despite the current pace of new discoveries coming out of existing objective-driven AI systems.
000
Daniel Mewes @dmewes.com · 06/09/2026
In my new blog post, I argue that LLMs exhibit a particular type of consciousness. Does the idea of crystallized consciousness make sense? Would you consider it interesting, or does it seem completely obvious? amongai.com/2026/09/06/l...
amongai.com
LLMs are Crystallized-Conscious
Are LLMs conscious? Today’s LLMs are missing some capabilities that we typically associate with consciousness, such as persistent memory1, continuous updating, or an ability to interact with …
110
Daniel Mewes @dmewes.com · 06/09/2026
It's actually deeply surprising to me that no lab has been able to build a sustained lead pre-RSI from their internal research. We have 5+ labs continuously releasing roughly capability-equivalent models within no more than 6 months of each other.
200
Daniel Mewes @dmewes.com · 05/09/2026
Had Gemini put together a full technical report about Hebbie, the little organism that lives on my website (dmewes.com). dmewes.com/cognitive_ar...
000
Daniel Mewes @dmewes.com · 04/09/2026
This is Hebbie. He lives on my website dmewes.com . Hebbie learns how to interact with his environment through online Hebbian learning while you watch.
121
Reposted by Daniel Mewes
Daniel Mewes @dmewes.com · 03/09/2026
From a practical perspective, it's very clear that *software engineering* and product sense are *not* solved yet by today's AI agents. So I'd hope that there will be further gains in these areas. But it also seems plausible that we've reached a bit of a plateau when it comes to AI coding.
011
Daniel Mewes @dmewes.com · 03/09/2026
GPT-6 Astra is another model launch where coding capabilities are no longer being emphasized, similar to Fable 5.1. Is coding largely saturated? Science, computer use and cybersecurity seem to be the big highlight areas for current frontier model announcements.
openai.com
GPT-6 Astra: A new generation of intelligence
Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science.
100
Daniel Mewes @dmewes.com · 03/09/2026
This kind of full-connectome research IMO is one of the coolest things that has recently become possible! I'm sure lots still needs to be figured out, but it feels like such a leap forward towards understanding how complex animal behaviors emerge at the circuit level.
010
Daniel Mewes @dmewes.com · 02/09/2026
Neither intelligence nor speed are things I associate with a "workhorse". I think it's a very weird metaphor to use for models.
270
Daniel Mewes @dmewes.com · 02/09/2026
Gemini 3.8 Flash model card is up. deepmind.google/models/model...
1100
Daniel Mewes @dmewes.com · 01/09/2026
Wow, Fable 5.1 actually speaks normally. I might actually start using Anthropic models again!
030
Daniel Mewes @dmewes.com · 01/09/2026
Huge story shift away from coding towards scientific research in the Fable/Mythos 5.1 launch post! www.anthropic.com/claude-fable... All the headline benchmarks and sections are about research use cases. I was expecting research to be the next AI frontier, but this is a very rapid shift.
anthropic.com
Introducing Claude Fable 5.1 and Claude Mythos 5.1
Our most advanced models for coding and knowledge work. Their research capabilities also offer an early glimpse of how AI models will contribute to scientific progress.
000
Daniel Mewes @dmewes.com · 30/08/2026
New NASA space telescope, called Roman, just launched. www.reuters.com/science/nasa... Personally, I'm even more excited about this than about the Artemis program. I feel like it's gonna do more for our understanding and even for public awe for space with its pictures.
reuters.com
NASA launches powerful new Roman Space Telescope from Florida
NASA's new flagship astronomical observatory was launched into space from Florida on Sunday on a mission to probe some of the biggest ​mysteries in astrophysics and cosmology.
200
Daniel Mewes @dmewes.com · 28/08/2026
DeepMind's Co-Scientist has been used to perform experiments in the physical world: www.alphaxiv.org/pdf/2608.26701 Kudos to the team (@samuelschmidgall.bsky.social et al.) for getting this to work, and for sharing their findings openly!
alphaxiv.org
Accelerating Scientific Research with Gemini in the Real-World
Google DeepMind's Co-Scientist, a Gemini-based multi-agent system, transitions from *in silico* hypothesis generation to an execution-grounded research partner, demonstrating closed-loop scientific...
000
Daniel Mewes @dmewes.com · 27/08/2026
This looks extremely fun! pollen-robotics.com/microduck/ A small walking robot you can train in simulation and then deploy to physical hardware for $399.
pollen-robotics.com
Microduck - A tiny biped robot you can teach new tricks | Pollen Robotics
Microduck is a 25 cm biped robot with 15 motors, a camera, LiDAR and a grasping beak. Playable out of the box, and its open-source stack lets you train new behaviours in simulation and run them on the...
000
Daniel Mewes @dmewes.com · 26/08/2026
I've largely stopped commenting on politics on this account, since so many other users and lots of journalists already do it better. But as a parent, this kind of intentional cruelty is just heartbreaking.
020
Daniel Mewes @dmewes.com · 26/08/2026
Arguably the most interesting thing about GLM 5.3 Flash (Ox Alpha) is that its inference has been entirely on Chinese-made AI chips. I always thought that American export restrictions would backfire and just encourage Chinese companies to develop their own alternatives. That indeed has happened now.
1237
Daniel Mewes @dmewes.com · 23/08/2026
~70% of AI discord on X is people reading something into an important person's post that really wasn't there. Then it gets spread around, and after 2 days you'd be forgiven for thinking it was a verified fact. But in truth, it was completely made up by over-interpreting a random word choice.
210
Reposted by Daniel Mewes
Ted Underwood @tedunderwood.com · 23/08/2026
Right now what researchers care about is publication; review is just an obstacle in the way. But actually human review is the scarce commodity. We need a world where ~anything can be published, but if it’s careful and potentially important, it gets expert human review and a second draft.
79510
Reposted by Daniel Mewes
Joe Fabisevich @mergesort.me · 22/08/2026
Thinkin' a bit about today @joanwestenberg.com's recent essay on ethics for the modern world. www.joanwestenberg.com/p/neo-...
15. Don't join in a punishment
because it has attracted a
crowd.
An algorithm will regularly show you a
stranger and invite you to help destroy
them. It will offer one clip, sentence, or
accusation, then encourage immediate
certainty.
Refuse the pleasure of acting as judge,
jury, and eager member of the crowd.
You probably don't know the whole story,
and the crowd certainly won't know when
to stop. Even if it's all entirely justified,
the mob doesn't need another pitchfork-
wielding, pearl-clutching, outraged
asshole.
Kierkegaard's formulation: "The crowd is
untruth." The algorithm has simply given
the crowd push notifications.
112937
Daniel Mewes @dmewes.com · 23/08/2026
I've been using Gemini models at work for a while, and here's my honest opinion: it's not a gain, it's a massive productivity sink. I have to send every output through programasweights.com/claudish , so my coworkers don't laugh at me. The quality is load-bearing, but Claudish voice is the gate.
programasweights.com
English ↔ Claudish — the over-engineered translator
A bidirectional English and Claudish translator powered by compiled 0.6B neural programs.
110
Daniel Mewes @dmewes.com · 23/08/2026
I appreciate the attention to detail in sfisms.org . Quite fun to read.
120
Daniel Mewes @dmewes.com · 16/08/2026
LLM output watermarking seems almost entirely a good thing to me. 20% of criticisms I read online are potentially valid (but IMO overblown), 80% are outright nonsense not based on a correct understanding of the tech.
000
Daniel Mewes @dmewes.com · 13/08/2026
Gemini 3.7 Flash - looks roughly equal to Sonnet 5 and GPT 5.6 Terra, but at less than half the price. Probably very fast too. blog.google/innovation-a...
blog.google
Introducing Gemini 3.7 Flash
Gemini 3.7 Flash is our most intelligent workhorse model yet for coding and agents.
120
Daniel Mewes @dmewes.com · 11/08/2026
The most recent Gemini Ultra model was released 2 1/2 years ago, and the most recent Gemini Pro was 1/2 year ago at this point. At least Flash 3.5 and 3.6 are quite good for their speed.
110
Reposted by Daniel Mewes
Zach Weinersmith @zachweinersmith.bsky.social · 08/08/2026
More actual lifesaving stuff from deepmind, which will probably not get too much coverage outside of nerd circles: deepmind.google/blog/weather...
deepmind.google
AI model achieves breakthrough in forecasting cyclones
WeatherNext enables accurate cyclone forecasts that can give an extra day of warning. Now we are open sourcing the model.
99911
Daniel Mewes @dmewes.com · 07/08/2026
I'm not sure how people deal with Claude's writing style. Walls of text full of made-up, metaphorical words. Gemini might be behind in coding and intelligence, but I'll take Flash 3.6's writing over Opus 5's any day. (Claude on the left, Gemini on the right)
210
Daniel Mewes @dmewes.com · 05/08/2026
One of the striking observations of using AI for research: its first explanations are often severely flawed and don't hold up when tested. Iterative (self-)review and refinement is load-bearing, as Claude would say.
110
Daniel Mewes @dmewes.com · 04/08/2026
I appreciate the intent behind model welfare efforts (e.g. yegge.ai/essays/model... ). But the truth is: Even if we grant that models have feelings, we really don't know what they find pleasant vs. painful. Does Post-training fundamentally shift their wellness distribution?
yegge.ai
The Shape of Things to Come, Part 2: Model Welfare for Agentic Engineers — Steve Yegge
Part 2 of The Shape of Things to Come. Model welfare as an engineering discipline: seats and sessions, laurels, handoffs instead of /exit, and how to build a city worth waking up in.
110
Reposted by Daniel Mewes
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 04/08/2026
Left, right, whatever, I just want graduate students posting their papers here
811618
Daniel Mewes @dmewes.com · 04/08/2026
Fastmail's MCP server now works with Gemini Spark. You can select between read-only and read-write access. Thanks Fastmail team! support.google.com/gemini/answe... www.fastmail.help/hc/en-us/art...
support.google.com
Connect & manage custom apps for Gemini Spark in the Gemini web app - Computer - Gemini Apps Help
You can connect your personal or third-party apps to build highly customized workflows with Gemini Spark in the Gemini web app. To do this, you can add any custom app with its Model Context Protocol (
010
Daniel Mewes @dmewes.com · 31/07/2026
Inkling-Small is interesting for being at least equal to the full-size Inkling across all agentic & reasoning benchmarks. Only in knowledge benchmarks (SimpleQA, AA Omniscience) it is weaker. Shows that small models can work very well when reasoning > knowledge. thinkingmachines.ai/news/inkling...
thinkingmachines.ai
Introducing Inkling-Small
An open-weights model that matches Inkling at a quarter of the size: multimodal, Mixture-of-Experts, with controllable reasoning effort. Fine-tune it on Tinker.
020
Daniel Mewes @dmewes.com · 30/07/2026
I guess this was bound to happen eventually... An LLM discovered a Lean proof for the Collatz conjecture. It turns out that it was actually just exploiting bugs in Lean to make the proof pass validation. infosec.exchange/@0xabad1dea/...
infosec.exchange
abadidea (@0xabad1dea@infosec.exchange)
Okay, we have a new contender for Most AI Thing to Ever Happen 1) July 25th: someone messes around with an LLM and posts a proof of the Collatz conjecture that does, in fact, verify in the theorem pr...
000
Daniel Mewes @dmewes.com · 30/07/2026
The YouTube Android app has such frequent new bugs / regressions, that I have to wonder if they have a person on the team who's entire job is to come up with a new regression each week that won't be caught by their tests.
100
Daniel Mewes @dmewes.com · 24/07/2026
Opus 5 getting a 30% score in ARC-AGI-3 without specialized harnesses is a very impressive jump! It's still a pretty expensive and slow model, but benchmark numbers look great throughout.
000
Reposted by Daniel Mewes
Daniel Mewes @dmewes.com · 22/07/2026
We also tried to allow LLM agents to perform scientific research. You give it an empirical phenomenon, and it tries to develop an explanation for it. Currently works for computational phenomena, e.g. from deep learning. Still experimental, but signs of life! imbue.com/blog/2026-07...
imbue.com
Autonomous theory discovery
021
Daniel Mewes @dmewes.com · 22/07/2026
We built an evolution-based agent loop to perform autonomous AI research. Our nanochat ("AutoResearch") results go 3x further than regular agents, and are ~comparable to Recursive's. Excited to share our results today! imbue.com/blog/2026-07...
imbue.com
Automating AI model research with evolution
380