Sign in

Aymeric

@aymeric.fyi
81 followers 186 following 277 posts

Engineering Management & Technology • Engineering Manager @ Mercari in 🇯🇵

PostsRepliesMedia
Aymeric @aymeric.fyi · 14/09/2026
It turns out measuring token usage across sessions, agents, local and frontier models is a bit of a mess. Prompt caching and tooling make it hard to get a straight answer, so I spent some time figuring it out.
aymeric.fyi
Introduction to Counting Tokens, Prompt Caching, and Session Analytics Tools
Someone recently asked me how many tokens I used after my Qwen 3.8 27B experiment. I had no idea, nor did I know how to verify it. This post explains how to count tokens, how prompt caching works, and...
000
Aymeric @aymeric.fyi · 08/09/2026
Rather than better models at similar or higher price points, I’d rather have current SOTA performance at lower prices. Imagine Opus 4.8 or 5 performance at Sonnet price level or lower. At least keep the same family of models in the same price range. Gemini 3.5 Flash Lite is over 500% since 2.0.
000
Aymeric @aymeric.fyi · 16/08/2026
I tried Qwen3.8 27B, the new local model Alibaba benchmarks against Opus 4.6. In my test, Qwen3.8 lands at Sonnet 5 quality, not the Opus 4.6 on the benchmark chart, and it needed an hour where Sonnet needed twelve minutes, on a M3 Max. Great results for a local model.
aymeric.fyi
Qwen3.8 27B: A Free and Slow Sonnet 5?
Qwen3.8 27B just got released, a week after Meta's Muse Glimmer 30B. Alibaba advertises benchmark scores on par with Opus 4.6. This post compares Qwen3.8 27B with Muse Glimmer 30B, Sonnet 5 and Opus 5...
010
Aymeric @aymeric.fyi · 11/08/2026
I tried Meta’s Muse Glimmer, which Meta advertises as a model for local agentic workflows. My conclusion is that it’s way too slow to use locally.
aymeric.fyi
Meta's Muse Glimmer: A for Effort, Too Slow to Use Locally
I compared Meta's new Muse Glimmer 30B against Gemma4 26B-A4B and Qwen3.6 35B-A3B on a simple agentic search task, running locally on an M3 Max. Muse Glimmer tries harder than any of them, and that is...
100
Aymeric @aymeric.fyi · 08/08/2026
Here’s my Claude Code setup The post goes through install, settings, models and effort levels, and my favorite MCPs, CLIs and plugins I keep the setup fairly vanilla on purpose What I install is about letting Claude reach third-party tools: the browser, Xcode, Datadog, Linear, Google Docs, etc
aymeric.fyi
My Claude Code Setup
I have been using Claude Code personally and professionally for about six months. This post explains how I set it up, including plugins, tools, MCPs and CLI.
010
Aymeric @aymeric.fyi · 14/07/2026
Michael Novati: Too many people are trying to become first principles thinkers […] — I’m not even sure it wants more of them. What scales is the other kind: the exceptional pattern matchers, the 10X people you’ve never heard of, happy under the radar, making things actually run.
michaelnovati.substack.com
AI Erased the Hardest Part of Being Me
I could always see the pattern. The tax was making everyone else see it too.
100
Aymeric @aymeric.fyi · 12/06/2026
I wrote my takeaways from LeadDev #LDX3 in this post. * Gains from AI are real but uneven across companies * Existing bottlenecks in the SDLC become blockers in AI-enabled orgs * Human relationships are still very important to build large successful products
aymeric.fyi
LeadDev LDX3 London 2026 Summary
What 25 talks and panels at LDX3 London 2026 said about engineering teams in the AI era.
000
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 The shadow culture: Why engineering principles fail under load by Manogna Machiraju We all want to build things fast with the right level of quality. The reality: the code is archaic, pipelines are broken, security is an afterthought. Engineering is always catching up.
110
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 What makes an effective EM in the AI era? Panel: Vernon Richards, Priscilla Nagashima, Alicia Collymore, Scott Carey
100
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 How will we deal with the new drudgery of AI-generated code? Panel: Lawrence Jones, Maude Lemaire, Randy Shoup, Birgitta Böckeler
100
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 AI in the trenches: Real-world wins without breaking things by Yiğit Darçın Deja Vu vs. Vuja De - the feeling that you have never seen it before. MIT study: the "GenAI Divide" - 95% of AI pilots fail to generate ROI, while a successful 5% utilize a "Vuja de" approach to rethink workflows.
110
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 AI killed the coding interviews. Here’s what Meta built instead by Danit Nativ Navon What broke? * Candidates used AI tools during interviews * Evaluations and ladders didn’t evolve * Leaders had no idea how to lead in this new context
120
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 Technology advances; history repeats itself by Zoe Cunningham A talk emphasizing the need to keep focusing on humans when everything is changing 🥰 The rate of change is increasing. The current message around the tech industry is impacting managers drastically: less managers, more IC work.
100
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 How to grow your engineers into great leaders? by Charles Duncan Jr. Your best engineer may not be your best managerial candidate The IC and manager jobs are different, require different skillset
100
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 Engineering at scale: Why developer experience is your competitive advantage by Dr Nicole Forsgren Teams "move" at 10x speed but the system moves at 1x The speed/friction mismatch * 15 minute builds * Manual approvals * Unclear deployment criteria * No documentation of decisions
210
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 What will AI-accelerated engineering teams actually look like? Panel with Luke Marsden, Hannah Foxwell, Yenny Cheung, Liz Fong-Jones What practices that predate AI help us move faster now? Documentation, comments, tests, small PRs Spec driven dev in all parts of SDLC: from PRD to tests
172
Aymeric @aymeric.fyi · 02/06/2026
#LDX3 Building resilient engineering teams when failure is the default by King-Immanuel Edoh In Nigeria, infrastructure is highly unreliable, starting with the power grid. This forces good engineering practices to be the default Fantastic and lively talk by King-Immanuel.
100
Aymeric @aymeric.fyi · 02/06/2026
#LDX3 What does a CTO even do? by Dee Kitchen 1. Do we have the right people? 2. Do we have the right technology? 3. Do we have the right product vision? 4. Do we have the right balance between innovation and liability 5. Do we have the right enablement and sales motion?
A spider chart of the 5 key competencies a CTO needs to cover: People, Tech, Product, Sales and Risk.
121
Aymeric @aymeric.fyi · 02/06/2026
#LDX3 Up and down the management track: Equalising your leadership style across the pressure of scale by Karen Lee Lessons from Karen who had an went up and down the management track, starting as engineering, growing to EM, MoM and Director, back down to MoM, EM and back up, twice.
100
Aymeric @aymeric.fyi · 02/06/2026
#LDX3 Moving Accessibility from debt to done at giffgaff by Abi Harrison-Nye Accessibility represents an 8T$ opportunity globally, that’s the buying power of people with any disability. Making the product right, accessible, from the start, does not cost more or add time to a project.
giffgaff.com
100
Aymeric @aymeric.fyi · 02/06/2026
#LDX3 The mechanics of scaling: Why delivery slows as you grow and what do to about it by Maryia Tarpachova As the company grows, features keep being added to create more growth, and pace starts to slow down A deep dive revealed specific teams were impact more than others, mostly core teams
100
Aymeric @aymeric.fyi · 02/06/2026
#LDX3 Quantifying tech debt to modernize critical systems by Ejber Ozkan Tech debt negatively impacts productivity Image from Tushar Sharma in Four Strategies for Managing Technical Debt Tech debt is net asset to build features for the business and becomes a liability when it reaches a threshold
The vicious cycle of technical debt:
Pressure to increase productivity leads to high technical debt, leads to low moral and motivation and to low code quality, which leads to lower productivity, which leads back to presurre to increase productivity.
100
Aymeric @aymeric.fyi · 02/06/2026
#LDX3 30 to 70 PRs a day: How we managed to not wreck our systems by @lizthegrey.com Honeycomb followed the example set by Intercom to 2x R&D productivity rather than try to 10x (linked article) Honeycomb mandated AI usage from Aug 2025 and only saw marginal changes until Opus 4.6 in Feb 2026
ideas.fin.ai
2× – nine months later: We did it
You can too.
141
Aymeric @aymeric.fyi · 01/06/2026
I’m coming from Tokyo to attend LeadDev LDX3 in London. (1/2) I’m particularly interested in how organizations are transitioning into AI-native organization, as we are in the middle of that process at Mercari. Let’s meet to discuss these topics or if you want to learn about Japan’s tech scene.
A LeadDev LDX3 Conference poster with a purple background and the following text:
LDX3
London
June 2 and 3, 2026
The festival for modern engineering leadership
110
Aymeric @aymeric.fyi · 19/04/2026
I tried Claude Design to design a new feature in my app. It speeds up UI prototyping over vibe coding prototypes. The prototypes are interactive and it’s possible to run them side-by-side. The hand-off to Claude Code is janky but it’s quite new so it’ll improve. Check the post for details.
aymeric.fyi
Hands-on With Claude Design
Anthropic just launched Claude Design, a new research-preview product for generating interactive prototypes, slide decks, and more. I took it for a spin by designing a new analytics screen for my habi...
020
Aymeric @aymeric.fyi · 15/04/2026
Oh, what happened to Google’s AI overview?
Google search for “SSIA meaning” with the AI overview starting with a long passage of pure CSS rather than the answer.
010
Aymeric @aymeric.fyi · 06/04/2026
Has Google killed the cheap Gemini models? From Gemini 1.5 Flash to Gemini 3.1 Flash Lite, the cost increased between 5x and 20x depending on the language used, English or Japanese.
aymeric.fyi
Has Google Killed the Cheap Gemini Models?
From Gemini 1.5 Flash to Gemini 3.1 Flash Lite, Google has increased the pricing by an order of magnitude, somewhere between 5x and 20x depending on the use case. This post retraces the pricing histor...
100
Aymeric @aymeric.fyi · 04/04/2026
1/2 In the current times of AI, I really like this advice from Charity Majors from the article below: The best advice I can give anyone is: know your nature, and lean against it.
charitydotwtf.substack.com
My (hypothetical) SRECon26 keynote
One year ago, Fred Hebert and I delivered the closing keynote at SRECon25. Looking back on it now, I can hardly connect with how I felt then. Here's what I'd say one year later.
111
Aymeric @aymeric.fyi · 30/03/2026
I released my first mobile app on iOS and Android. There is such a contrast between how fast building with AI is, and how slow the mobile app release process is. I wrote about the whole experience in this post.
aymeric.fyi
Building OnTrack Mobile Apps With AI: Coding Fast, Releasing Slow
In this post, I share my experience building a cross-platform habit tracker mobile app using AI coding tools, highlighting the stark contrast between lightning-fast development and the surprisingly sl...
000
Aymeric @aymeric.fyi · 13/01/2026
Google is hot right now, back to back days announcing big things 😳 Sunday: Universal Commerce Protocol Monday: Apple picked Gemini for Siri
000
Aymeric @aymeric.fyi · 27/12/2025
A very easy to understand review of the paper Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers,". Large models simply memorize the training data. Smaller models learn the rule because they don't have enough "brain space" to memorize every individual answer.
statisticianinstilettos.com
Why I care more about your SLM than your LLM
A review of the paper Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers," (Barron & White, 2025)
000
Reposted by Aymeric
Dare Obasanjo @carnage4life.bsky.social · 07/11/2025
Ben Thompson argues that while AI is obviously in a bubble, it will create long term benefits in the form of investments in chip fabs in the U.S. and power plants. This is similar to how we got dark fiber and train tracks as infrastructure out of the internet and railway bubbles.
stratechery.com
The Benefits of Bubbles
We are in an AI Bubble: the big question is if this bubble will be worth it for the physical infrastructure and coordinated innovation that result?
13595
Aymeric @aymeric.fyi · 26/10/2025
Quick intro to Claude Skills Simplicity is the point. The most significant limitation of MCP is token usage Almost everything an MCP can do, can also be handled by a CLI tool instead. LLMs know how to call cli-tool --help, which means you don’t have to spend many tokens describing how to use them
010
Aymeric @aymeric.fyi · 20/10/2025
What happens to a driving Waymo when AWS us-east-1 goes down? 😅 Is everything processed locally? It’s an Alphabet company, so if something is remote, I’d expect it to be on GCP or Google internal platforms, but with external dependencies, isn’t every product 6 degrees of separation from us-east-1?
000
Reposted by Aymeric
Anil Dash @anildash.com · 17/10/2025
Okay, for the folks who asked: here's the majority AI view, writing up the reasonable, thoughtful view on AI that the vast majority of people in tech hold, that gets overshadowed by the bluster and hype of the tycoons trying to shill their nonsense. anildash.com/2025/10/17/t... Please share!
anildash.com
The Majority AI View - Anil Dash
A blog about making culture. Since 1999.
381119486
Aymeric @aymeric.fyi · 17/10/2025
This is the first article I read about spec-driven development and this is a good introduction. I thought spec meant PRD, but it’s more of a detailed technical design. As a profession, we were not good at documentation before AI, so this will not suit many orgs. And it seems overkill in most cases.
010
Aymeric @aymeric.fyi · 12/10/2025
I wrote an article on how we translate user-generated content at scale at Mercari. It summarizes two years of work with various models, LLMs but not only, and how cost evolved, reducing it 100x over the period. I also go over non-AI features and user experience.
engineering.mercari.com
The Journey of User-Generated Content Translation | Mercari Engineering
This is @aymeric from Cross Border Engineering.This article is part of the blog series Behind the Scenes of Developing M
010
Aymeric @aymeric.fyi · 09/10/2025
Our engineering team took on the massive challenge of rebuilding our marketplace from scratch to make it scale globally 🌏, and we just launched it. 🚀 Check out the first article in our blog series to learn how we did it!
engineering.mercari.com
Behind the Scenes of Developing Mercari’s First Global App, “Mercari Global App” | Mercari Engineering
Hello. I’m @deeeeet from Cross Border (XB) Engineering.On September 30, 2025, we announced a new strategy for our
1600
Aymeric @aymeric.fyi · 30/09/2025
The best description of Product vs Platform engineering 😂 Like the tail of an enthusiastic puppy: The product work might go a thousand directions in a minute. They're at the end of the puppy's tail. Infrastructure is the part of the tail anchored to its butt, not changing direction nearly as fast
linkedin.com
Google+ was wildly successful, just not as a product. I worked on infrastructure teams at Google for a long time. One of the eras I lived through was the rise and fall of Google+, their foray into… |...
Google+ was wildly successful, just not as a product. I worked on infrastructure teams at Google for a long time. One of the eras I lived through was the rise and fall of Google+, their foray into so...
000
Aymeric @aymeric.fyi · 28/09/2025
Great article explaining what atproto is about in non-technical terms. Excerpts: We are at a similar juncture with social apps as we have been with open source thirty five years ago […] I like to call it “open social” What open source did for code, open social does for data.
010
Aymeric @aymeric.fyi · 14/09/2025
A good thread on retaining top engineers, how it sometimes means encouraging them to take on roles outside your org I’d rather keep top talent in the company by sharing opportunities outside my org, rather than keeping them locked and they leave later out of boredom or lack of growth opportunities
010
Aymeric @aymeric.fyi · 14/09/2025
Maybe social media needs circuit breakers like when the stock exchange crashes. If the mod team can’t remove hateful content fast enough, block the entire system: no posting, maybe even no viewing. Stop the snowball effect and let the mod team catch up. Put a SLO on the number of unack’ed reports.
000
Reposted by Aymeric
Hazel Weakly @hazelweakly.me · 03/09/2025
I've been reading this on the Month of AI bugs from @wuzzi23.bsky.social and I am losing my fucking mind. The series is excellent, and Johann does great work, but SWEET FUCK people, you cannot be serious here. This is absolute clownshoes territory embracethered.com/blog/posts/2...
embracethered.com
Wrap Up: The Month of AI Bugs · Embrace The Red
Wrap Up: The Month of AI Bugs - Full List of Postings
13912
Aymeric @aymeric.fyi · 25/08/2025
Nice real world application of AI agents and lessons learned, full of interesting advices. The analogy to Map Reduce is clever. Keep it simple with two layers and no state in the sub-agents. Caching is smart. It attempts to make non-deterministic LLMs into agents as deterministic as possible.
userjot.com
Best Practices for Building Agentic AI Systems: What Actually Works in Production - UserJot
Real patterns for building AI agent systems that don't fall apart. Two-tier architectures, stateless design, orchestration strategies, and what we learned building UserJot's agent infrastructure.
000
Aymeric @aymeric.fyi · 18/08/2025
I love how one can share a bug on Bluesky and the main maintainer not only answers but also promises to fix it.
Screenshot of a thread
Erica: it’s annoying that when I search for two words on here within the quotation marks, bluesky still searches each word separately. Can you change that?
Samuel: I think I can fix this! I think the problem is phone keyboards produce “ (a kind of double quotation mark) but the search system is expecting " (another kind of double quotation mark)
Will fix!
000
Aymeric @aymeric.fyi · 17/08/2025
Great new feature in VS Code, checkpoints in chat sessions with Copilot. We finally can chat with the tool without having to commit between every request. If the next prompts mess up, just revert to the previous stable state.
code.visualstudio.com
July 2025 (version 1.103)
Learn what is new in the Visual Studio Code July 2025 Release (1.103)
000
Aymeric @aymeric.fyi · 17/08/2025
Great summary by @simonwillison.net of @wuzzi23.bsky.social ‘s findings on AI tools vulnerabilities. In short, all AI tools are vulnerable if one attaches external files and links to their prompts, leading to secrets leaks and remote code execution. Johann publishes daily until the end of the month.
simonwillison.net
The Summer of Johann: prompt injections as far as the eye can see
Independent AI researcher Johann Rehberger (previously) has had an absurdly busy August. Under the heading The Month of AI Bugs he has been publishing one report per day across an …
033
Reposted by Aymeric
ProtoPro @handle.invalid · 04/07/2025
LinkedIn is broken. All your data is trapped behind a walled garden, job applications never seem to go anywhere, and the whole site increasingly seems to be an ad to try to upsell you LinkedIn Premium. We're here to help. For Pros on Proto to connect Pro to Pro!
1547
Aymeric @aymeric.fyi · 12/08/2025
I love how @bsky.app gets built in public, to the point that core contributors use the app to post bugs and tag colleagues 😄
000
Aymeric @aymeric.fyi · 11/08/2025
The first job of a leader is to make the business succeed. This is a common point of these three articles from 2011, 2024 and 2025 - a16z: Peacetime CEO / Wartime CEO - Charity Majors: Pragamatism, neutrality and leadership - Lara Hogan: Balancing direction and empowerment
100