Sign in

Daniel Mewes

@dmewes.com
452 followers 349 following 1.1K posts

Interested in artificial and natural intelligence, emergent complexity, among other things. I mostly post about AI and ML. -> dmewes.com Currently research at Imbue. Previously Ambient.ai, Stripe, RethinkDB, Max Planck Institute.

PostsRepliesMedia
Daniel Mewes @dmewes.com · 4h
I always love this little guy exploring things in a different direction from the other flocks!
000
Daniel Mewes @dmewes.com · 10h
We're launching Imbue Studio, a personal AI computer for you to use and customize. I haven't had this much fun with using computers in a long time! You can modify any aspect just by asking. Or have it built custom automations and apps. No coding required. Runs locally too. imbue.com/product/studio
000
Daniel Mewes @dmewes.com · 28/09/2026
The use cases for Sonnet 5.5 still seem pretty niche. It will be cheaper than Opus 5.5 in tasks that require little reasoning and have large inputs, due to lower input token cost. But on all the agentic benchmarks it seems to be at best equal to Opus 5.5 in performance/$?
200
Daniel Mewes @dmewes.com · 14/09/2026
The key is to add this Lagrangian per-layer target to the loss function. The Lagrangian multipliers (lambda_i) are vectors that get updated gradually in each inference (forward) step, thereby avoiding the need for a separate backwards pass across layers.
100
Daniel Mewes @dmewes.com · 14/09/2026
Interesting work by J. Seeley and J. Gould at @sakanaai.bsky.social : a local learning rule that has similar performance to backprop in training deep networks. Graph shows comparison to plain predictive coding (PC), which doesn't scale to deep networks. pub.sakana.ai/pc-alm/
200
Daniel Mewes @dmewes.com · 14/09/2026
AI alignment has been solved folks! "Strong and Smart President is All You Need". Paper coming soon.
0120
Daniel Mewes @dmewes.com · 09/09/2026
Very interesting work about the shape of CoT reasoning: "[...] reasoning models are dynamical systems with fractal basins, with fractality increasing with task difficulty across diverse tasks." by J. Lai et al: arxiv.org/pdf/2609.04963 I really like this way of looking at LLM traces!
050
Daniel Mewes @dmewes.com · 09/09/2026
Sharing @kennethstanley.bsky.social 's post on why open-endedness is still very much needed, despite the current pace of new discoveries coming out of existing objective-driven AI systems.
000
Daniel Mewes @dmewes.com · 05/09/2026
Had Gemini put together a full technical report about Hebbie, the little organism that lives on my website (dmewes.com). dmewes.com/cognitive_ar...
000
Daniel Mewes @dmewes.com · 04/09/2026
You can watch his little brain working through the Brain Monitor panel.
000
Daniel Mewes @dmewes.com · 04/09/2026
This is Hebbie. He lives on my website dmewes.com . Hebbie learns how to interact with his environment through online Hebbian learning while you watch.
121
Daniel Mewes @dmewes.com · 02/09/2026
Neither intelligence nor speed are things I associate with a "workhorse". I think it's a very weird metaphor to use for models.
270
Daniel Mewes @dmewes.com · 02/09/2026
And somehow it got even faster in terms of token output speed than the already extremely fast 3.7 Flash.
110
Daniel Mewes @dmewes.com · 02/09/2026
Gemini 3.8 Flash model card is up. deepmind.google/models/model...
1100
Daniel Mewes @dmewes.com · 23/08/2026
I appreciate the attention to detail in sfisms.org . Quite fun to read.
120
Daniel Mewes @dmewes.com · 13/08/2026
Wow, looks to be seriously fast actually!
010
Daniel Mewes @dmewes.com · 07/08/2026
I'm not sure how people deal with Claude's writing style. Walls of text full of made-up, metaphorical words. Gemini might be behind in coding and intelligence, but I'll take Flash 3.6's writing over Opus 5's any day. (Claude on the left, Gemini on the right)
210
Daniel Mewes @dmewes.com · 05/08/2026
One of the striking observations of using AI for research: its first explanations are often severely flawed and don't hold up when tested. Iterative (self-)review and refinement is load-bearing, as Claude would say.
110
Daniel Mewes @dmewes.com · 09/07/2026
When I saw this graph in the GPT 5.6 blog post, my first reaction was "That is such an ugly graph. Who scaled the X axis this way? Clearly it should have been log-scale." Then I realized that the ugliness of the graph *is* the point: It's a burn at Anthropic's models being so much more expensive.
000
Daniel Mewes @dmewes.com · 01/07/2026
Maybe I'm too right for Bluesky...
Moderate Libertarian Right
230
Daniel Mewes @dmewes.com · 13/05/2026
This is an *extremely* disappointing change for anyone building custom tools around Claude Code. @anthropic.com is saying that `claude -p` (headless mode) will no longer benefit from the Claude plans, and instead will always be priced at API per token cost.
Screenshot from a Claude email notifying that "As part of this change, Agent SDK and other programmatic usage will run on this credit, and will not impact your subscription limits. "
430
Daniel Mewes @dmewes.com · 09/04/2026
Francois Chollet and Yann LeCun feuding over Meta's new model was not on my AI drama bingo card.
François Chollet
@fchollet
·
21h
The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlates with actual usefulness is a core competency for AI labs, and any new lab is unlikely to be successful without first figuring that out.
Yann LeCun
@ylecun
·
1h
I don´t have skin in the game here: I'm no longer with Meta and I famously don't do LLMs.
But I'd say your long-held anti-Meta bias is showing.
François Chollet
@fchollet
·
50m
Meta is famously a very ethical company that has only had a very positive impact on the world, so any criticism of Meta must come from a place of "bias", I suppose
131
Daniel Mewes @dmewes.com · 17/12/2025
GPT 5.2 (thinking on high) seems to be buggy. I'm seeing reproducible issues where it starts glitching out in the middle of its response with garbage tokens and then ends the message. Highly reproducible for me every 10th-20th prompt or so.
Screen photo showing a model response that is emitting Python code, and then ends in two Unicode characters before ending suddenly.
010
Daniel Mewes @dmewes.com · 30/09/2025
Today we're releasing Sculpture, an AI coding tool that combines async agents with the ability to collaborate with agents locally. It also comes with built-in verifiers to automatically check the quality of AI written code. More to come! imbue.com/sculptor-ann...
A picture of @kanjun.bsky.social making an excited gesture, with the label "developers, developers, developers" written over it.
010
Daniel Mewes @dmewes.com · 17/03/2025
We've reached a point of no return. The Trump admin officially proclaimed today that they do not need to follow the rulings of federal judges, at least as far as "foreign affairs" are concerned. Remarkably, they interpret "foreign affairs" to also include deportations and who knows what else.
"A single judge in a single city cannot direct the movements of an aircraft ... full of foreign alien terrorists who were physically expelled from U.S. soil," White House press secretary Karoline Leavitt said in a statement.
110
Daniel Mewes @dmewes.com · 04/03/2025
When I took my Evoque there a couple years ago... youtu.be/p3sHgO4-aZk?...
Range Rover Evoque driving through sandy dunes area.
000
Daniel Mewes @dmewes.com · 27/02/2025
Wait, really? All that tough negotiation last time under threat of trade war did not in fact achieve the promised outcome? (www.reuters.com/world/americ...)
U.S. President Donald Trump said on Thursday that his proposed 25% tariffs on Mexican and Canadian goods will go into effect on March 4 as scheduled because drugs are still pouring into the U.S. from those countries.
020
Daniel Mewes @dmewes.com · 24/02/2025
Interesting research by @quentin-garrido.bsky.social et al. that shows JEPA-style self-supervised video models can distinguish physically plausible from implausible sequences. arxiv.org/abs/2502.11831
 Video prediction in representation space (V-JEPA) achieves an understanding of intuitive physics. (A) Video models are evaluated on three intuitive physics datasets using the Violation of Expectation paradigm (IntPhys, GRASP, and InfLevel). V-JEPA is significantly more ‘surprised’ by implausible videos. Random initializations of V-JEPA (untrained networks) show near-chance performance, and state-of-the-art video models based on text or pixel prediction are much closer to chance. Confidence intervals at 95% are obtained via bootstrapping, except for untrained networks (n=20) which use a normal distribution assumption. (B) V-JEPA is trained to ’inpaint’ natural videos in a learned representation space. Starting from a video and a corrupted version, representations are first extracted. The goal is then to predict the representation of the original video from the representation of the corrupted ones. (C) From a trained V-JEPA, we compute a surprise metric by predicting representations of N future frames based on M past ones and comparing the predictions to the representations of observed events. The surprise metric is then used to decide which of the two videos contains a physical violation.
010
Daniel Mewes @dmewes.com · 12/02/2025
Actually paying a little less this winter with the heat pump than last winter with gas furnace. We got the heat pump in July. Electricity generation charges are slightly cheaper in San Mateo via Peninsula clean energy than with plain PG&E, but should be a small difference?
PG&E monthly bill history
000
Daniel Mewes @dmewes.com · 07/02/2025
Seriously, Gemini 2.0 Thinking? I thought we had moved past the "how many letter r in strawberry" problem with the advent of reasoning models. I do find it really interesting though to see its reasoning trace and where exactly it went wrong (in spelling the word it turns out).
Screenshot of Gemini 2.0 Flash Thinking Experimental reasoning that there are 2 letter r in the word strawberry.
120
Daniel Mewes @dmewes.com · 03/02/2025
Canada, Mexico, and China are the biggest purchasers of US goods. Together nearly 4% of US GDP (2022 numbers). They're also the three biggest importers with almost equal share. This trade war is going to be very bad for US inflation and economy. (ustr.gov/countries-re...)
The top five purchasers of U.S. goods exports in 2022 were: Canada ($356.5 billion), Mexico ($324.3 billion), China ($150.4 billion),...
The top five suppliers of U.S. goods imports in 2022 were: China ($536.3 billion), Mexico ($454.8 billion), Canada ($436.6 billion),...
120
Daniel Mewes @dmewes.com · 01/02/2025
Anyone remember 'unixsex.com'? web.archive.org/web/20051217...
Screenshot of an archived version of the unixsex website reads "See the hottest and horniest unix vixens ever! Hot-swappable!"Screenshot reads "Unix Bondage! Hot babes tied up with ethernet cables, trapped in colo cages, etc"
010
Daniel Mewes @dmewes.com · 29/01/2025
Those importers must have missed the memo that the countries of origin will be paying for tariffs.
A Reuters news article reads: The U.S. trade deficit in goods widened to a record high in December, likely as businesses front-loaded imports of industrial supplies and consumer goods in anticipation of broad tariffs from President Donald Trump's new administration.
220
Daniel Mewes @dmewes.com · 13/01/2025
Gemini 2.0 Advanced solves it. But it could just have been in the training set? I also tried asking it for a five letter word. The Rumel tree that it mentions doesn't seem to exist?
Gemini 2.0 Advanced finds the correct answer "deer"Asked for a five-letter example, Gemini 2.0 Advanced suggests lemur and Rumel respectively.
020
Daniel Mewes @dmewes.com · 02/01/2025
Gotcha, yeah, it's similar on Android. But e.g. Google assistant ("hey Google") should really be always displaying this since it's listening for the wake word, but doesn't. So it feels like I can't trust it.
Screenshot showing the "microphone in use" indicator on Android
000
Daniel Mewes @dmewes.com · 29/12/2024
"The portrayal of the NSDAP as right-wing extremist is clearly false, considering that Adolf Hitler, the party's leader, supports animal welfare and is a migrant from Austria! Does that sound like Hitler to you? Please!" (Elon Musk, 1933)
"The portrayal of the AfD as right-wing extremist is clearly false, considering that Alice Weidel, the party's leader, has a same-sex partner from Sri Lanka! Does that sound like Hitler to you? Please!" Musk said in the piece.
163
Daniel Mewes @dmewes.com · 05/12/2024
This is disgusting. Folks asking for and praising literal terrorism against CEOs.
Screenshot of responses on Bluesky in reaction to the attack on UnitedHealth CEO
100
Daniel Mewes @dmewes.com · 27/11/2024
Wasn't this the exact phrase that led to the very first user ban in Bluesky history? Some people really make their anti-AI tendencies a bit too personal.
Bluesky user commenting "throw yourself off a building" in response to a post announcing a Blue sky dataset for ML training.
000
Daniel Mewes @dmewes.com · 27/11/2024
Good question. Also... what happened to car colors? (This one is from Polestar - and yes, that's all the "colors" that are available)
Screenshot of car configurator with a choice of colors, all being different shades of white/grey/black and one minimally saturated gold color.
100
Daniel Mewes @dmewes.com · 21/11/2024
Oh man, I really wasn't confident, but actually got 10/10. I think there's something quite unimaginative about ChatGPT's rhyme structure.
Result page of the poetry turing test showing 10/10 score.
000
Daniel Mewes @dmewes.com · 19/11/2024
Anybody remember these?
AKG K1000 open headphones.
110
Daniel Mewes @dmewes.com · 18/11/2024
Here is mine
Photo of my HP 35s calculator. No cats visible in the photo.
020
Daniel Mewes @dmewes.com · 15/11/2024
Anthropic CEO Dario Amodei saying "People call them scaling laws. That’s a misnomer.” [1] is a bit ironic, given he is one of the people who coined that term [2]. Though his clarification makes sense. [1] finance.yahoo.com/news/openai-... [2] arxiv.org/abs/2001.08361
Screenshot from source 1Screenshot from source 2
110
Daniel Mewes @dmewes.com · 31/10/2024
Oh wow, the Democrats are actually eating babies. Roseanne Barr was right!
President Joe Biden playfully bites a baby during a trick-or-treaters celebration for Halloween at the White House in Washington, October 30, 2024.
020
Daniel Mewes @dmewes.com · 18/09/2024
Good time to remember when Bluesky was a place where you'd discuss skeet vs post, and where you'd go to talk to ducks. On to 100M users!
010
Daniel Mewes @dmewes.com · 28/03/2024
Greentheonly dropping some interesting bits of Tesla FSD internals over on X. (I believe these are from extracting symbols from the Tesla firmware?)
Tweets listing attribute labels for stop signs and traffic lights in Tesla firmware.
010
Daniel Mewes @dmewes.com · 22/02/2024
Found the photo I took. It was in Dominica indeed, not Belize.
Photo of a batfish
110
Daniel Mewes @dmewes.com · 14/02/2024
Remember: when you post completely made up legal statements, always make sure to list the specific legislations (E.g. USA) for which you've made the point up.
Screenshot that reads: or you will be in violation of ToS due to circumvention of a ban (which itself is a crime in most countries, USA included)
061
Daniel Mewes @dmewes.com · 14/02/2024
This person also thinks that a service's ToS are "the law" and one is "liable" for (supposedly) breaking them. Talk about democratizing government!
Screenshot that reads: Let me be clear: you are breaching the law by circumventing and violating ToS, knowingly, both of the services you scrape and those you post to, which means you are legally liable
270
Daniel Mewes @dmewes.com · 14/02/2024
Unhinged. Many people out defending angry mobs vandalising and burning other folk's stuff again... (I checked my Mastodon feed again after 3 months. No thanks)
Several Mastodon users cheering on the news of a driverless Waymo car being destroyed by a crowd.
020