Sign in

Aymeric

@aymeric.fyi
81 followers 186 following 277 posts

Engineering Management & Technology • Engineering Manager @ Mercari in 🇯🇵

PostsRepliesMedia
Aymeric @aymeric.fyi · 14/09/2026
It turns out measuring token usage across sessions, agents, local and frontier models is a bit of a mess. Prompt caching and tooling make it hard to get a straight answer, so I spent some time figuring it out.
aymeric.fyi
Introduction to Counting Tokens, Prompt Caching, and Session Analytics Tools
Someone recently asked me how many tokens I used after my Qwen 3.8 27B experiment. I had no idea, nor did I know how to verify it. This post explains how to count tokens, how prompt caching works, and...
000
Aymeric @aymeric.fyi · 08/09/2026
Rather than better models at similar or higher price points, I’d rather have current SOTA performance at lower prices. Imagine Opus 4.8 or 5 performance at Sonnet price level or lower. At least keep the same family of models in the same price range. Gemini 3.5 Flash Lite is over 500% since 2.0.
000
Aymeric @aymeric.fyi · 17/08/2026
I reached the same conclusion that it’s impressive to achieve this much locally but is too slow I observed one major difference in my test, it did better and faster with xhigh reasoning than medium. On my test on a M3 Max: Medium: tokens in 94k / out 78k - 2h20 XHigh: tokens in 70k / out 44k - 1h
aymeric.fyi
Qwen3.8 27B: A Free and Slow Sonnet 5?
Qwen3.8 27B just got released, a week after Meta's Muse Glimmer 30B. Alibaba advertises benchmark scores on par with Opus 4.6. This post compares Qwen3.8 27B with Muse Glimmer 30B, Sonnet 5 and Opus 5...
020
Aymeric @aymeric.fyi · 16/08/2026
I tried Qwen3.8 27B, the new local model Alibaba benchmarks against Opus 4.6. In my test, Qwen3.8 lands at Sonnet 5 quality, not the Opus 4.6 on the benchmark chart, and it needed an hour where Sonnet needed twelve minutes, on a M3 Max. Great results for a local model.
aymeric.fyi
Qwen3.8 27B: A Free and Slow Sonnet 5?
Qwen3.8 27B just got released, a week after Meta's Muse Glimmer 30B. Alibaba advertises benchmark scores on par with Opus 4.6. This post compares Qwen3.8 27B with Muse Glimmer 30B, Sonnet 5 and Opus 5...
010
Aymeric @aymeric.fyi · 11/08/2026
On a M3 Max with 96GB of RAM, I compared Muse Glimmer with Qwen3.6 35B-A3B and Gemma4 26B-A4B. They achieved 20 tok/s, 70 tok/s and 60 tok/s respectively. But as Muse Glimmer is greedier, it runs through a lot more content, so it feels much 10x slower rather than only 3x.
000
Aymeric @aymeric.fyi · 11/08/2026
I tried Meta’s Muse Glimmer, which Meta advertises as a model for local agentic workflows. My conclusion is that it’s way too slow to use locally.
aymeric.fyi
Meta's Muse Glimmer: A for Effort, Too Slow to Use Locally
I compared Meta's new Muse Glimmer 30B against Gemma4 26B-A4B and Qwen3.6 35B-A3B on a simple agentic search task, running locally on an M3 Max. Muse Glimmer tries harder than any of them, and that is...
100
Aymeric @aymeric.fyi · 08/08/2026
Here’s my Claude Code setup The post goes through install, settings, models and effort levels, and my favorite MCPs, CLIs and plugins I keep the setup fairly vanilla on purpose What I install is about letting Claude reach third-party tools: the browser, Xcode, Datadog, Linear, Google Docs, etc
aymeric.fyi
My Claude Code Setup
I have been using Claude Code personally and professionally for about six months. This post explains how I set it up, including plugins, tools, MCPs and CLI.
010
Aymeric @aymeric.fyi · 14/07/2026
With AI replacing the 1X work all around them, they are in the most powerful position they have ever been in.
000
Aymeric @aymeric.fyi · 14/07/2026
Michael Novati: Too many people are trying to become first principles thinkers […] — I’m not even sure it wants more of them. What scales is the other kind: the exceptional pattern matchers, the 10X people you’ve never heard of, happy under the radar, making things actually run.
michaelnovati.substack.com
AI Erased the Hardest Part of Being Me
I could always see the pattern. The tax was making everyone else see it too.
100
Aymeric @aymeric.fyi · 12/06/2026
I wrote my takeaways from LeadDev #LDX3 in this post. * Gains from AI are real but uneven across companies * Existing bottlenecks in the SDLC become blockers in AI-enabled orgs * Human relationships are still very important to build large successful products
aymeric.fyi
LeadDev LDX3 London 2026 Summary
What 25 talks and panels at LDX3 London 2026 said about engineering teams in the AI era.
000
Aymeric @aymeric.fyi · 04/06/2026
Interesting, thanks for sharing. Based on the README, I’m wondering if we could achieve the same set of features with Ansible roles & playbooks.
100
Aymeric @aymeric.fyi · 03/06/2026
Culture evolves, design for it. Sometimes, change the principles, sometimes the behaviors. Evolve the system to support. Principles fail under load: * People under pressure do not follow written rules, they take the path of least resistance * Good intentions alone aren't enough. Build a system
000
Aymeric @aymeric.fyi · 03/06/2026
What to do today? * Pick two principles, three teams, four lenses. * Define what does the principles look like now * Stress test, run the answers through the 4 lenses * Action plan: systemic & team level commitments with clear owners. Be ready for hard answers and measure
100
Aymeric @aymeric.fyi · 03/06/2026
How do you define the principles: * Value: review if the outcome is important * Viability: Make bold substractions * Usability redesign the path, remove friction * Feasibility: Invest in the tooling & frameworks
100
Aymeric @aymeric.fyi · 03/06/2026
For each principle, define: * Value: Does it reduce risk & improve outcomes * Viability: Does our business situation allow this? * Usability: is it easier than the shortcut? Under stress eng. will follow the shortest path. * Feasibility: do we have what we need to do it?
100
Aymeric @aymeric.fyi · 03/06/2026
A principle tells you what to value, not how to uphold it. 1. We needed behaviors. 2. Stress-test the principles
100
Aymeric @aymeric.fyi · 03/06/2026
Everyone is optimizing for something different → Rebuild the principles: Why: Engineering as equal partner agreed not just stated. What: Principles built with people in the trenches. How: Tech strategy, get the right skills & autonomy. Everyone agreed, nothing changed.
100
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 The shadow culture: Why engineering principles fail under load by Manogna Machiraju We all want to build things fast with the right level of quality. The reality: the code is archaic, pipelines are broken, security is an afterthought. Engineering is always catching up.
110
Aymeric @aymeric.fyi · 03/06/2026
This is not a new question, we needed to measure the ROI of DX before, but the cost of AI has made it necessary to measure. To help, we should use clear use-cases and calculate the cost, e.g. build a prototype and mesure the time and cost.
000
Aymeric @aymeric.fyi · 03/06/2026
We’re spending so much money on AI tools and cost keeps increasing. How do you validate the ROI? Few companies have proven the returns they got from AI. Many companies have exceed their allocated AI budget. Ideally we can use metrics (DX?). The reality is many are vibe-measuring.
100
Aymeric @aymeric.fyi · 03/06/2026
Tool: Explore new tools, Building your own tooling- e.g. Use OpenClaw, learn what is good, what is not, and then build the right solution for yourself.
200
Aymeric @aymeric.fyi · 03/06/2026
What’s the most effective tool and skill EMs should have at their sleeve in current AI era? Tool: Claude / Codex Skill: Ability to make decisions and trade-offs Skill: Design for simplicity - Use another AI evaluate the first one
100
Aymeric @aymeric.fyi · 03/06/2026
HW is always after SW. Historically, solutions have always been found and many companies are working on improving the HW to make it viable. Focus on how you can do better, e.g. not vibe-coding an app you’re never going to touch again.
100
Aymeric @aymeric.fyi · 03/06/2026
As an EM, how do you deal with engineers who object to the use of AI for ethical reasons? Have the difficult conversations: Acknowledge the current issue. Highlight the business reality. People have to make a choice for themselves: embrace the tools or risk losing your job.
100
Aymeric @aymeric.fyi · 03/06/2026
Bad news and good news: no one knows what the hell we’re doing, where we’re going → So find the problems in the business and solve them. Using AI enables the player-coach model where manager can solve some issues autonomously without disturbing the team.
100
Aymeric @aymeric.fyi · 03/06/2026
Time to expand your scope Understand how other sides of the business works, how a piece of the system works. Find other areas and think of how to improve it, discuss it with your manager.
100
Aymeric @aymeric.fyi · 03/06/2026
What does career growth mean considering the flattening situation? As a 1st level EM, what skills should I build to become the leader of tomorrow? Skill to continue building: working with people.
100
Aymeric @aymeric.fyi · 03/06/2026
Most managers know they can’t be great at a particular coding language or framework etc by not being in the weeds. But being able to contribute to code again through the help of AI is motivating and keeps them technical. Experimenting with AI tools is recommended.
100
Aymeric @aymeric.fyi · 03/06/2026
Can you be an EM without doing any hands-on technical work? What's technical work? Managing technical teams is being technical. There are still humans in the organization. Regardless of AI, humans need support, help to grow, and advice on career progression.
100
Aymeric @aymeric.fyi · 03/06/2026
Lots of posts about EMs being best positioned to build in the AI. WDYT? There are transferable skills for sure. But is this the best use of an EM’s time and skills? Technical work is not the only thing that EMs do well. Communicaiton, alignment and leading others is still required
120
Aymeric @aymeric.fyi · 03/06/2026
If you are looking to change job right now, it is scary - LinkedIn is full of “we don’t need support roles like EM/PM and everyone should be shipping features”. You need to know what you want to do because the EM role is in such a flux that the role varies greatly from one company to another.
100
Aymeric @aymeric.fyi · 03/06/2026
Is it a good time to be an engineering manager? YES You can still work with people, and with AI being so fast it helps build things on the side again. Using AI to do a best job as a manager is motivating, like better career discussions and evaluation feedback.
100
Aymeric @aymeric.fyi · 03/06/2026
Scott saw a job post for a director of engineering role: 75% hands-on coding, 25% people management. Is this a real director role? Only in a very small company with a few engineers. Impossible in any other situation, unless they trade-off something very important, e.g. people, regulations, etc
100
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 What makes an effective EM in the AI era? Panel: Vernon Richards, Priscilla Nagashima, Alicia Collymore, Scott Carey
100
Aymeric @aymeric.fyi · 03/06/2026
Will we soon stop reviewing code? Autonomous coding with AI without looking at it is already here but is unevenly distributed. Context matters: in some codebases, circumstances, it is acceptable to not review the code. In most circumstances today, it is important to review.
020
Aymeric @aymeric.fyi · 03/06/2026
How we learn and encode information in our heads: Brains process information differently when speaking / hearing vs. writing / reading, the latter is better. Interrogate AI what it does, force yourself to write, read, and retain information. Engaging more AI diminishes, the worry of not learning.
100
Aymeric @aymeric.fyi · 03/06/2026
Are you concerned about skill atrophy as a result of increased AI usage? YES. Abstractions are learned by learning about the details, and it takes a constant back and forth between details and abstractions. e.g. in the past, undestanding OSes and compilers help with programming.
110
Aymeric @aymeric.fyi · 03/06/2026
Taking away the tool from jr devs looks like a punishment, making people feel like they’re doing a bad job. How do we resolve jr devs no learning the code base? * Teach, explain how to use AI to get feedback, pair-program * Practice Extreme Programming * Team-wide code retrospectives
100
Aymeric @aymeric.fyi · 03/06/2026
We are about to take AI tools away from jr devs as we’re finding out they are not learning the codebase and not understanding what the AI is producing. What are your opinions? It’s incredibly shortsighted - jr still need to learn how to use these tools.
100
Aymeric @aymeric.fyi · 03/06/2026
Refuse PRs of more than 500 SLOC Create a whole change, test it and once it work, ask AI to split the huge change into reviewable sub-PRs or independent commits. Continuously merge and deploy. Review data and create visuals: test coverage, heatmap of files read and changed by the agent.
100
Aymeric @aymeric.fyi · 03/06/2026
The bottleneck has moved to verifying that the code is correct. How do we do that well? Increased AI-generated unit testing may not help. It's just more to review and get skipped in reviews. Instead Cursor creates test scripts specific to the PR change and adds the results in the PR.
100
Aymeric @aymeric.fyi · 03/06/2026
How do we deal with this drudgery and deluge of AI-generated code? Don’t try to ship code at the same level of quality as before, ship code at a higher level of quality, AI enables it. It’s possible to ship faster only if you’ve done the investment in the tooling and guardrails.
100
Aymeric @aymeric.fyi · 03/06/2026
AI makes code generation free or almost free - is that the best thing ever? AI can be used for coding and more. We can’t trust AI entirely, it requires human oversight. It’s not all about adding code, but also about removing code. Eng. still need to deeply understand the code they ship.
100
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 How will we deal with the new drudgery of AI-generated code? Panel: Lawrence Jones, Maude Lemaire, Randy Shoup, Birgitta Böckeler
100
Aymeric @aymeric.fyi · 03/06/2026
Auto research loop in data science * Generate hypothesis * Write code * Execute experiment * Analyze metrics * Iterate and learn
The Auto research loop in data science:
* Generate hypothesis
* Write code
* Execute experiment
* Analyze metrics
* Iterate and learn
* Back to Generate hypothesis
000
Aymeric @aymeric.fyi · 03/06/2026
Takeaways: * Adoption: Make adoption systematic * Monitor: measure and monitor continuously * Safety: create safety and support for teams * Constraints: treat cost and security as core constraints * Experience: invest in both AX and DX
100
Aymeric @aymeric.fyi · 03/06/2026
Three phases of AI adoption: Phase 1: Exploration Discover, Explore, Try Identify a champion Phase 2: Pilot Start execution Obsess over metrics Hands-on management Phase 3: Scale and Optimize Scale process Have solid platforms Optimize
100
Aymeric @aymeric.fyi · 03/06/2026
#LDX3 AI in the trenches: Real-world wins without breaking things by Yiğit Darçın Deja Vu vs. Vuja De - the feeling that you have never seen it before. MIT study: the "GenAI Divide" - 95% of AI pilots fail to generate ROI, while a successful 5% utilize a "Vuja de" approach to rethink workflows.
110
Aymeric @aymeric.fyi · 03/06/2026
Seniority can be a liability if one sticks to their knowledge of how were done. * Give the mic to power users * Make it a team effort * Provide time and resources to learn Leaders don’t need to vide-code better than new hires, they need to ask the right questions, create the space to learn.
000
Aymeric @aymeric.fyi · 03/06/2026
Setting the expectations for the leveling is not simply raising the bar that existed previously. With new tools and models landing regularly, they recalibrate and raise the bar regularly.
100