Sign in

Shion Honda

@shionhonda.bsky.social
16 followers 12 following 50 posts

本田 志温 / Software Engineer / MSc in Computer Science / Posts are my own hippocampus-garden.com bibly.page

PostsRepliesMedia
Shion Honda @shionhonda.bsky.social · 04/08/2026
Why can ASGI handle WebSockets but standard WSGI cannot? To understand this, I built a chat app that streams tokens and can cancel generation without using any web framework. Learning WebSockets by Building an AI Chat with Raw ASGI | Hippocampus's Garden hippocampus-garden.com/websocket_as...
hippocampus-garden.com
Learning WebSockets by Building a Streaming AI Chat with Raw ASGI | Hippocampus's Garden
Build a small AI-style chat with token streaming and cancellation to learn how WebSockets and ASGI work together.
010
Shion Honda @shionhonda.bsky.social · 04/08/2026
ASGI is common in AI apps, but why? There's more to it than "async is faster." This post compares it with WSGI using an interactive lab: Why AI Apps Use ASGI—and How It Differs from WSGI | Hippocampus's Garden hippocampus-garden.com/wsgi_asgi/
hippocampus-garden.com
Why AI Apps Use ASGI—and How It Differs from WSGI | Hippocampus's Garden
A simple comparison of WSGI and ASGI, using an AI agent that spends most of its time waiting for LLM APIs.
000
Shion Honda @shionhonda.bsky.social · 23/06/2026
Responses API promises better multi-turn reasoning by preserving reasoning context across turns. We wanted to know whether that actually improves our customer support AI agents, so we ran our own A/B test. Here’s what we found: medium.com/alan/measuri...
medium.com
Measuring the Responses API on Alan’s Support AI agents
In 2025, OpenAI introduced a new way to call its models: the Responses API. Their pitch is that switching from the legacy Chat Completions…
000
Shion Honda @shionhonda.bsky.social · 07/04/2026
2年前に、「デュアルキャリア・カップル」に関する記事を書きました。 その後出産、交代制ワンオペ育児、イタリア移住を経て、色々な学びがあったので整理しました。 私たちにとってのフェアネスとは完全分業でも等分でもなく、役割を交代可能にすることなのだと感じました。 hippocampus-garden.com/dual_career_...
hippocampus-garden.com
デュアルキャリア・カップル実践編:出産・育児とイタリア移住で見えたこと | Hippocampus's Garden
2年前の記事で整理したデュアルキャリアの意思決定を、実際に運用した結果を振り返ります。出産、イタリア移住、交代制ワンオペ育児を経て見えた現実と学びをまとめました。
000
Shion Honda @shionhonda.bsky.social · 05/04/2026
I migrated my personal blog from Gatsby to Astro, but the main goal was to fix long-standing UI problems and simplify the site's structure. A New Look for This Blog | Hippocampus's Garden hippocampus-garden.com/gatsby_to_as...
hippocampus-garden.com
A New Look for This Blog | Hippocampus's Garden
I migrated this personal blog from Gatsby to Astro, but the main goal was to fix long-standing UI problems and simplify the site's structure.
000
Shion Honda @shionhonda.bsky.social · 09/03/2026
Tool Decathlon [Li+, 2026, ICLR] Toolathlon is a benchmark for language agents, which offers 32 apps and 604 tools to evaluate the capability to use diverse and complex tools. Some complex apps, like e-commerce, are container-based to ensure realistic simulations. arxiv.org/abs/2510.25726
000
Shion Honda @shionhonda.bsky.social · 24/01/2026
Manifold-Constrained Hyper-Connections [Xie+, 2026] HC expands Transformer's residual stream width to improve expressivity but breaks the identity mapping. mHC constrains residual mappings to be doubly stochastic to stabilize training and enable 27B-scale experiments. arxiv.org/abs/2512.248...
000
Shion Honda @shionhonda.bsky.social · 05/01/2026
2025年に読んだ本をまとめました。 読んだ本とともに2025年を振り返る | Hippocampus's Garden hippocampus-garden.com/books_2025/
hippocampus-garden.com
読んだ本とともに2025年を振り返る | Hippocampus's Garden
2025年に読んだ本を振り返ります。
000
Shion Honda @shionhonda.bsky.social · 31/12/2025
Wrapped up 2025 with a post: my top 5 AI papers of the year. DeepSeek-R1, RLVR limitations, agent benchmarks, long-task metrics, and why LMs hallucinate, plus pointers to related research and write-ups. Top 5 AI Papers of 2025 | Hippocampus’s Garden hippocampus-garden.com/ai_papers_20...
hippocampus-garden.com
Top 5 AI Papers of 2025 | Hippocampus's Garden
This article picks five notable AI papers from 2025 and summarizes their key ideas and limitations, including reinforcement learning for reasoning, agent benchmarks, long‑task metrics, and a statistic...
100
Shion Honda @shionhonda.bsky.social · 14/12/2025
Does RL incentivize reasoning capacity in LLM? [Yue+, 2025, NeurIPS] By evaluating pass@k at large k, the paper shows that RLVR improves sampling efficiency but doesn't introduce new reasoning patterns. Distillation expands the boundary by injecting the teacher distribution. arxiv.org/abs/2504.13837
020
Shion Honda @shionhonda.bsky.social · 29/11/2025
AIによって知識がコモディティ化していく時代において、人間の価値がどこで生まれるのかを、フロンティア知識と行動のプレミアムという観点から整理しました。 知識のコモディティ化と行動のプレミアム | Hippocampus's Garden hippocampus-garden.com/commodity_kn...
hippocampus-garden.com
知識のコモディティ化と行動のプレミアム | Hippocampus's Garden
AIによって知識がコモディティ化していく時代において、人間の価値がどこで生まれるのかを、フロンティア知識と行動のプレミアムという観点から整理しました。
000
Shion Honda @shionhonda.bsky.social · 23/11/2025
Weight-sparse transformers have interpretable circuits [Gao+, 2025] Training a Transformer with L0 norm fixed—so that most weights are zero—yields disentangled circuits. The role of each weight can be identified and visualized for simple tasks like closing quotation marks. arxiv.org/abs/2511.13653
000
Shion Honda @shionhonda.bsky.social · 22/11/2025
An Open-Ended Embodied Agent with LLMs [Wang+,2024, TMLR] Voyager is an embodied agent in Minecraft powered by GPT-4. It learns skills (executable code for performing complex actions), guided by a curriculum generated by LLM to encourage exploration. arxiv.org/abs/2305.16291 #NowReading
000
Shion Honda @shionhonda.bsky.social · 16/11/2025
Evaluating creative reasoning with Sudoku variants [Seely+, 2025] This paper proposes Sudoku-Bench, a memorization-proof benchmark using Sudoku variants to test LLMs’ insight and multi-step reasoning, with expert solving traces from YouTube included in the dataset. arxiv.org/abs/2505.16135
000
Shion Honda @shionhonda.bsky.social · 15/11/2025
Recursive Reasoning with Tiny Networks [Jolicoeur-Martineau, 2026] Tiny Recursion Model is a 7M-param network that performs internal recursive reasoning, eliminating the complexity of HRMs. With supervised training, it outperforms LLMs on Sudoku and ARC-AGI. arxiv.org/abs/2510.048... #NowReading
020
Shion Honda @shionhonda.bsky.social · 03/11/2025
Chasing leaderboard scores is a trap. We reproduced Tau-bench across different models and learned: - Hidden configs make apple-to-apple comparisons impossible - Reasoning boosts accuracy, but also latency That’s why you need an internal benchmark. medium.com/alan/benchma...
medium.com
Benchmarking AI Agents: Stop Trusting Headline Scores, Start Measuring Trade-offs
Don’t chase leaderboards. Run the benchmark yourself and map your score–latency–cost frontier to choose models for production.
000
Shion Honda @shionhonda.bsky.social · 30/10/2025
Wrote a short piece on why pass@k and pass^k tell different stories than mean and variance. If you enjoy thinking about evaluation quirks and probability, you might like it. Pass@k and Pass^k Tell Different Stories from Mean Success Rate | Hippocampus's Garden hippocampus-garden.com/pass_k/
hippocampus-garden.com
Pass@k and Pass^k Tell Different Stories from Mean Success Rate | Hippocampus's Garden
These metrics capture coverage and reliability.
000
Shion Honda @shionhonda.bsky.social · 19/10/2025
A Sober Look at LM Reasoning [Hochlehnert+, 2025, COLM] The authors evaluated 7B-class open models on math tasks and found that performance varied significantly with decoding parameters (random seeds, temperature etc). Many RL methods failed to reproduce reported gains. arxiv.org/abs/2504.070...
000
Shion Honda @shionhonda.bsky.social · 15/10/2025
What’s the difference between evaluating LLMs vs. AI agents? Why is it so hard to assess AI agents in real business settings? I’ve written my thoughts and a possible path forward. I plan to keep writing about this topic—please share your thoughts or feedback! medium.com/alan/benchma...
medium.com
Benchmarking AI Agents: The Challenge of Real-World Evaluation
AI agents need stateful benchmarks. Unlike LLMs, agents interact with databases and users. We explore why and how to evaluate them…
000
Shion Honda @shionhonda.bsky.social · 30/09/2025
Evaluating AI on real-world economically valuable tasks [Patwardhan+, 2025] OpenAI created a new benchmark, GDPval, asking experts from 9 industries that contribute to the US GDP. The current top model is Claude Opus 4.5, nearing the standards of industry experts. openai.com/index/gdpval/
000
Shion Honda @shionhonda.bsky.social · 27/09/2025
Meta Agents Research Environments [Andrews+, 2025] ARE is a platform that facilitates async interactions between agents and the environment (including real apps). The authors then upgraded the GAIA benchmark to Gaia2, which runs in a virtual mobile environment built on ARE. arxiv.org/abs/2509.17158
000
Shion Honda @shionhonda.bsky.social · 27/09/2025
Benchmark for General AI Assistants [Mialon+, 2023] GAIA evaluates AI's fundamental abilities, such as web search, image understanding, and file reading, rather than expert knowledge. While humans achieve 92%, GPT-4 does 15% (30% with plugins). arxiv.org/abs/2311.12983
000
Shion Honda @shionhonda.bsky.social · 20/09/2025
Some reasoning models disable temperature and top‑p. This short post explains how sampling works and why those knobs are off. Why You Can’t Set Temperature on GPT-5/o3 | Hippocampus's Garden hippocampus-garden.com/llm_temperat...
hippocampus-garden.com
Why You Can’t Set Temperature on GPT-5/o3 | Hippocampus's Garden
Some LLMs disable sampling knobs like temperature and top_p. Here’s why.
010
Shion Honda @shionhonda.bsky.social · 09/09/2025
Why Language Models Hallucinate [Kalai+, 2025] LM pre-training is essentially density estimation and doesn't help prevent hallucination. Post-training can alleviate the issue by rewarding acknowledging uncertainty over guessing, but current benchmarks are not designed so. arxiv.org/abs/2509.04664
010
Shion Honda @shionhonda.bsky.social · 24/08/2025
Stress Testing MCP-enabled Agents [Yin+, 2025] LiveMCP-101 is a benchmark of 101 real-world tasks that require the coordinated use of multiple MCP tools. As results could change across attempts, it also evaluates the resolution path. GPT-5 achieves a success rate <60%. arxiv.org/abs/2508.15760
000
Shion Honda @shionhonda.bsky.social · 14/08/2025
Evaluating Conversational Agents in a Dual-Control Environment [Barres+, 2025] To simulate a user who doesn't have perfect information, τ^2-Bench gives tools to the user agent. The new telecom domain is more challenging than the existing airline and retail sectors. arxiv.org/abs/2506.07982
020
Shion Honda @shionhonda.bsky.social · 04/08/2025
"Should I put this tool description in the prompt or pass it as a parameter?" I've been asked this question a few times, so I wrote a new post. It's not obvious that inside LLMs, everything becomes one token stream. Here's how it actually works 👇 hippocampus-garden.com/llm_serializ...
hippocampus-garden.com
How LLMs Actually Process Your Prompts, Tools, and Schemas | Hippocampus's Garden
A deep dive into how LLMs serialize prompts, output schemas, and tool descriptions into a token sequence, with examples from Llama 4's implementation.
000
Shion Honda @shionhonda.bsky.social · 03/08/2025
Subliminal Learning [Cloud+, 2025] This paper studies a phenomenon where LMs transmit behavioral traits (e.g., liking owls) via unrelated data (e.g., a sequence of numbers) during distillation. It doesn't happen when the teacher and student have different base models. arxiv.org/abs/2507.148...
000
Shion Honda @shionhonda.bsky.social · 02/08/2025
Open Agentic Intelligence [Kimi Team, 2025] Kimi K2 is an MoE with 32B/1T params, which achieves the SOTA performance among open-source non-thinking models. It's powerful in agentic tasks thanks to the RL in an agentic environment. MuonClip helps the stable training. arxiv.org/abs/2507.20534
000
Shion Honda @shionhonda.bsky.social · 20/07/2025
ACEBench [Chen+, 2025] ACEBench is a benchmark designed to assess LLM's tool calling capability. The dataset is generated by an LLM-based pipeline that includes incomplete instructions and real-world APIs. It compares the model’s tool call outputs with the ground truths. arxiv.org/abs/2501.128...
000
Shion Honda @shionhonda.bsky.social · 24/06/2025
We developed an AI agent that dynamically investigates insurance claims, automating 30% of support tickets. See how we built it leveraging the tool calling capability👇 Inside Alan’s Claim Agent: How an AI Investigates Claims just like a Human, only Faster medium.com/alan/inside-...
medium.com
Inside Alan’s Claim Agent: How an AI Investigates Claims just like a Human, only Faster
The story behind our AI agent that dynamically investigates insurance claims, automating 30% of support tickets it receives.
000
Shion Honda @shionhonda.bsky.social · 24/06/2025
Yet another real-world example of an AI agent that does the job.
000
Shion Honda @shionhonda.bsky.social · 24/05/2025
『コンピュータの構成と設計』(通称「パタヘネ本」)のまとめと感想です。 書評『コンピュータの構成と設計』 | Hippocampus's Garden hippocampus-garden.com/book_review_...
hippocampus-garden.com
書評『コンピュータの構成と設計』 | Hippocampus's Garden
『コンピュータの構成と設計』(通称「パタヘネ本」)のまとめと感想です。
000
Shion Honda @shionhonda.bsky.social · 27/04/2025
We often treat databases as a black box, but understanding what’s inside is crucial for becoming a good software engineer. I just finished reading Database Internals and published a review + summary! hippocampus-garden.com/book_review_...
hippocampus-garden.com
Book Review: Database Internals | Hippocampus's Garden
A deep dive into how databases work.
000
Shion Honda @shionhonda.bsky.social · 26/04/2025
Measuring AI Ability to Complete Long Tasks [Kwa+, 2025] "50%-task-completion time horizon" measures human experts' time to complete tasks that AI models can complete with a 50% success rate. The time horizon for software engineering tasks has been doubling every 7 months. arxiv.org/abs/2503.14499
000
Shion Honda @shionhonda.bsky.social · 13/04/2025
Developing is the best way to learn. I built an MCP server to connect Claude to Google Sheets — and wrote a guide so you can build your own too! How to Customize Claude with your own MCP Server | Hippocampus's Garden hippocampus-garden.com/claude_mcp/
hippocampus-garden.com
How to Customize Claude with your own MCP Server | Hippocampus's Garden
Learn how to extend Claude's capabilities by building your own Model Context Protocol server.
000
Shion Honda @shionhonda.bsky.social · 23/03/2025
Benchmark for Tool-Agent-User Interaction in Real-World [Yao+, 2024] τ-bench evaluates LLM agents' human interaction and rule adherence in real-world domains (retail and airline)—key aspects missing from existing evaluations. It also proposes pass^k to measure consistency. arxiv.org/abs/2406.12045
010
Shion Honda @shionhonda.bsky.social · 04/03/2025
New post on Alan's tech blog. I explained how DeepSeek R1 uses pure reinforcement learning to boost reasoning (no reward models needed). Plus, R1-inspired projects and quick insights from our tests! DeepSeek R1: Demystifying LLMs’ Reasoning Capabilities. medium.com/alan/deepsee...
medium.com
DeepSeek R1: Demystifying LLM’s Reasoning Capabilities
DeepSeek shocked the world by dropping an open-weight successor to OpenAI’s o1: R1. This post summarizes our learnings from the tech…
020
Reposted by Shion Honda
Alan engineering @alanengineering.bsky.social · 03/03/2025
In late January, DeepSeek shocked the world by dropping an open-weight successor to OpenAI's o1: R1. Their tech report discusses how to incentivize reasoning capability in LLMs. We share our learnings at:
medium.com
DeepSeek R1: Demystifying LLM’s Reasoning Capabilities
DeepSeek shocked the world by dropping an open-weight successor to OpenAI’s o1: R1. This post summarizes our learnings from the tech…
001
Shion Honda @shionhonda.bsky.social · 16/02/2025
Just finished "AI Engineering" by Chip Huyen—an excellent guide to building AI applications with foundation models. I love how it ties technical and product thinking together, something many technical books miss! hippocampus-garden.com/book_review_...
hippocampus-garden.com
Book Review: AI Engineering by Chip Huyen | Hippocampus's Garden
A detailed guide on how to build applications with foundation models.
000
Shion Honda @shionhonda.bsky.social · 09/02/2025
Simple test-time scaling [Muennighoff+, 2025] The authors reproduced the test-time scaling curve by fine-tuning Qwen2.5-32B with s1K, a set of reasoning traces generated by Gemini. They controlled the response length by appending “Wait” when the model tried to answer early. arxiv.org/abs/2501.193...
000
Shion Honda @shionhonda.bsky.social · 02/02/2025
Incentivizing Reasoning Capability in LLMs via RL [DeepSeek-AI, 2025] R1-Zero acquired o1-level reasoning capability by directly applying GRPO, which rewards correct answers, to the base model. The team also released the weights of R1 with SFT and distilled models. arxiv.org/abs/2501.129...
000
Shion Honda @shionhonda.bsky.social · 08/01/2025
📈 In 2024, we automated 20% of customer contacts using AI, while keeping customer satisfaction intact! Check out our latest blog for the behind-the-scenes👇 Bonus: our Marmot image generator is good at joking. Look at the iMac with Windows installed 😂 medium.com/alan/how-we-...
medium.com
How We Built Alan’s AI Assistant for Customer Support
At Alan, exceptional customer care isn’t just a service — it’s a core part of what sets us apart in the insurance industry.
000
Reposted by Shion Honda
Alan engineering @alanengineering.bsky.social · 08/01/2025
In 2024, we automated 20% of customer contacts with AI while maintaining the same customer satisfaction (the number is still growing!) 🚀 Learn about our journey:
medium.com
How We Built Alan’s AI Assistant for Customer Support
At Alan, exceptional customer care isn’t just a service — it’s a core part of what sets us apart in the insurance industry.
053
Shion Honda @shionhonda.bsky.social · 01/01/2025
年末年始は、自分のキャリアを見つめ直す良い機会。ということで、今後5年程度の未来を見据え、AI時代を生き抜くための方法について考えていたことをまとめました。 皆さんの生存戦略もぜひ教えてください! AI時代をどう生き抜くか | Hippocampus's Garden hippocampus-garden.com/ai_and_emplo...
hippocampus-garden.com
AI時代をどう生き抜くか | Hippocampus's Garden
今後5年程度の未来を見据え、AI時代を生き抜くための方法について自分の考えをまとめました。
010
Shion Honda @shionhonda.bsky.social · 29/12/2024
I've wrapped up 2024 with my top 10 favorite deep learning papers of the year! 🔍 Dive into the breakthroughs, insights, and ideas that shaped the field. Year in Review: Deep Learning Papers in 2024 | Hippocampus's Garden hippocampus-garden.com/deep_learnin...
hippocampus-garden.com
Year in Review: Deep Learning Papers in 2024 | Hippocampus's Garden
Reflecting on 2024's deep learning breakthroughs! Discover my top 10 favorite research papers that shaped the field this year.
010
Shion Honda @shionhonda.bsky.social · 27/12/2024
Language Modeling in a Sentence Representation Space [Barrault+, 2024] Large Concept Models predict sentence embeddings instead of tokens. They inherently generalize to languages that their embedding model can and are good at handling long context (e.g., summarization) arxiv.org/abs/2412.08821
000
Shion Honda @shionhonda.bsky.social · 22/12/2024
The Road Less Scheduled [Defazio+, 2024, NeurIPS] Schedule-Free achieves Pareto-optimal loss curves without any LR scheduling. The theory behind it unifies iterate averaging (theoretically good) and scheduling (empirically strong) , taking the best of both worlds. arxiv.org/abs/2405.15682
020
Shion Honda @shionhonda.bsky.social · 07/12/2024
Movie Gen [Polyak+, 2024] Movie Gen is a group of Transformer-based Flow Matching models that can collectively generate and edit videos with audio, setting a new SOTA on various tasks. Videos are patchfied by spatiotemporal autoencoders. arxiv.org/abs/2410.13720
100
Shion Honda @shionhonda.bsky.social · 24/11/2024
RLHF is not the only method for AI alignment. This post introduces modern algorithms like DPO, KTO, and DiscoPOP that offer simpler and more stable alternatives. Evolution of Preference Optimization Techniques | Hippocampus's Garden hippocampus-garden.com/preference_o...
hippocampus-garden.com
Evolution of Preference Optimization Techniques | Hippocampus's Garden
RLHF is not the only method for AI alignment. This article introduces modern algorithms like DPO and KTO that offer simpler and more stable alternatives.
061