Sign in

Michael R. Bock

@michaelrbock.com
42 followers 81 following 224 posts

co-founder of Column Tax // michaelrbock.com

PostsRepliesMedia
Michael R. Bock @michaelrbock.com · 13h
This was an interesting finding when running our eval: GPT-6.1 Sol isn't more accurate than GPT-6 Sol calculating tax returns, but it is cheaper and faster. At their best, both are tied for accuracy around 65%. But to hit ~50%: 6.1 Sol (low thinking): $0.17/return, 48s
120
Michael R. Bock @michaelrbock.com · 29/09/2026
This was surprising to me (but maybe shouldn't be if you're paying attention to Anthropic guidance): Opus 5.5 beats Sonnet 5.5 on COST On TaxCalcBench (testing AI's ability to calculate tax returns), Opus 5.5 at xhigh reasoning is both more accurate AND cheaper than Sonnet 5.5
110
Michael R. Bock @michaelrbock.com · 24/09/2026
The harder they think, the wider the gap. At high thinking, both got 46% of returns right. Sol: $0.24/return. Opus 5.5: $1.75. At max thinking: $0.34 vs. $3.45.
100
Michael R. Bock @michaelrbock.com · 24/09/2026
Both models are almost tied on accuracy, but not on time or cost: GPT-6 Sol: $0.34/return, ~2 min Opus 5.5: $3.45/return, ~7 min 10x the price and 3x the time for one fewer correct return.
100
Michael R. Bock @michaelrbock.com · 24/09/2026
Claude Opus 5.5 is the #2 model in the world at filing taxes according to TaxCalcBench. Unfortunately, it's also 10x more expensive than the new #1 model: GPT-6 Sol.
110
Michael R. Bock @michaelrbock.com · 01/09/2026
Not sure if this will be useful unless you're a CTO, Head of Engineering, or maybe Engineering Mgr, but figured I'd share: At some point (this has probably already happened), you won't be able to individually lead every project yourself.
100
Michael R. Bock @michaelrbock.com · 19/08/2026
This is something pretty interesting I'm watching right now: the race between U.S. vs. China for the best open model. Meta says it will release an open-weight version of Muse Spark 1.2 soon. China's Moonshot AI has already released the weights for Kimi K3.
100
Michael R. Bock @michaelrbock.com · 10/08/2026
Before starting Column Tax, I had never hired anyone. By the time we were acquired, I had hired 50 amazing folks. I used the same algorithm for every hire:
100
Michael R. Bock @michaelrbock.com · 29/07/2026
1/ Anthropic just released its 4th model in 2 months: Opus 5. It's half the cost of Fable 5. And surprisingly, for knowledge work tasks like tax filing: Opus performs even better than its more expensive counterpart. We tested Opus 5 on TaxCalcBench. Here are the results vs. Fable 5:
100
Michael R. Bock @michaelrbock.com · 13/07/2026
1/ Claude (Code, Cowork, Fable) has all the mindshare right now, but is OpenAI actually the best at knowledge work? Our eval for tax filing says: yes. GPT-5.6 Sol (with web search) just became the best model in the world on TaxCalcBench (v2).
100
Michael R. Bock @michaelrbock.com · 07/07/2026
5/ AI still fails to calculate tax returns accurately on its own. GPT-5.5 is the highest-scoring model. On its own, it only scores 24%. With web search tool use, that jumps to 54%. Good, but not good enough to trust it with a task that requires 100% correctness.
100
Michael R. Bock @michaelrbock.com · 07/07/2026
4/ - 50 high-quality test cases covering complex tax & financial situations - Realistic PDF inputs (e.g. W-2s, 1099s, K-1s, etc.) - Federal & state returns - Up-to-date, covering tax year 2025 The upshot?
100
Michael R. Bock @michaelrbock.com · 07/07/2026
3/ But the benchmark itself had some limitations: simple returns, federal-only, simplified (text-based) input data. TaxCalcBench v2 (2025) changes that:
100
Michael R. Bock @michaelrbock.com · 07/07/2026
2/ I'm excited to release TaxCalcBench v2, the biggest update to measuring AI's ability to calculate tax returns since June 2025. v1 of TaxCalcBench measured AI's ability to calculate 2024 returns, AI failed to complete the task: the highest-scoring model at the time (Gemini 2.5 Pro) only got ~30%.
100
Michael R. Bock @michaelrbock.com · 07/07/2026
1/ This AI benchmark will decide Intuit's stock price next year. Last year, I released TaxCalcBench, and it moved markets. Here's how:
100
Michael R. Bock @michaelrbock.com · 08/06/2026
1/ Here I am about to lead a Product Review meeting (and I'm not nervous at all). I know it's going to go well because I forced my reports to write a Problem/Solution Review doc.
100
Michael R. Bock @michaelrbock.com · 26/03/2026
I've been working on tax software for the past 5 years. This is the last year anyone will have to pay for TurboTax. You can try it yourself today: - add the Aiwyn Tax connector inside of Claude (link below) - give it access to your tax documents (W-2s, etc.) - ask Claude to prepare your tax return
110
Michael R. Bock @michaelrbock.com · 10/03/2026
2/At lower thinking budgets, Pro actually pulls ahead: Medium thinking: Pro 56.86% vs Standard 49.02% High thinking: Pro 58.82% vs Standard 56.86% Ultrathink: Pro 62.75% vs Standard 62.75% Pro is smarter per token of thought, but give the cheaper model enough thinking time and it catches up
100
Michael R. Bock @michaelrbock.com · 10/03/2026
1/ OpenAI just launched GPT-5.4 Pro, their premium model at 12x the API cost of standard GPT-5.4. $30/M input tokens, $180/M output vs. $2.50/$15. I ran TaxCalcBench on Pro. The result: exactly tied with standard GPT-5.4 12x the price, 0% improvement But the full story is more nuanced:
110
Michael R. Bock @michaelrbock.com · 09/03/2026
This is actually genius: a Chrome extension that allows you to drag across your Google calendar and then paste your free times as perfectly-formatted text!
000
Michael R. Bock @michaelrbock.com · 06/03/2026
2/ The thinking level gap is enormous: High: 56.86% Medium: 49.02% Low: 31.37% That's a 25-point spread between low and high. Low thinking GPT-5.4 would rank near the bottom of the leaderboard. High thinking puts it at #1.
100
Michael R. Bock @michaelrbock.com · 06/03/2026
1/ The rivalry between OpenAI & Anthropic continues: GPT 5.4 is now the best model in the world at filing taxes (better than Opus 4.6)! We Just ran TaxCalcBench on GPT-5.4. 56.86% of tax returns computed perfectly. That's #1 overall: the first model to break 55%, surpassing Claude Opus 4.6:
100
Michael R. Bock @michaelrbock.com · 03/03/2026
3/ An interesting wrinkle on thinking budgets: Ultrathink: 49.02% Medium: 49.02% High: 47.06% Lobotomized: 37.25% Low: 35.29% Medium thinking matches ultrathink exactly. More thinking doesn't always mean better, but some thinking is critical. The jump from low to medium is +14 points.
100
Michael R. Bock @michaelrbock.com · 03/03/2026
1/ We just ran TaxCalcBench on Gemini 3.1 Pro to test how it does filing taxes. 49.02% of tax returns computed perfectly. That's #2 overall, only 4 points behind Opus 4.6. And it now holds the best "correct by line" score of any model ever tested (88.54%). Updated leaderboard:
100
Michael R. Bock @michaelrbock.com · 20/02/2026
4/ Full updated rankings (using strict scoring where every line must be correct): Opus 4.6: 52.94% GPT-5 w/ Search: 41.67% GPT-5.2 Pro: 41.18% Sonnet 4.6: 37.25% <-- new Gemini 3 Pro: 36.27% Opus 4.5: 36.27% GPT-5.2: 33.82% Gemini 2.5 Pro: 32.35% 7 months ago, 32% was SOTA.
010
Michael R. Bock @michaelrbock.com · 20/02/2026
3/ Thinking budget matters enormously for tax. Same model, same prompt, different thinking levels: Sonnet 4.6 (ultrathink): 37.25% Sonnet 4.6 (no thinking): 19.61% Nearly 2x accuracy just from letting the model think longer.
100
Michael R. Bock @michaelrbock.com · 20/02/2026
1/ We just ran TaxCalcBench on Claude Sonnet 4.6. 37.25% of tax returns computed perfectly. That's a "mid-tier" model outscoring every single flagship model from 6 months ago. Updated leaderboard:
110
Michael R. Bock @michaelrbock.com · 19/02/2026
I started Column Tax in early 2021 and sold it to Aiwyn in late 2025. Starting a startup was the hardest thing I've ever done. But knowing certain things makes it easier. So I wrote down everything I learned through experience that I wish I had known at the start:
120
Michael R. Bock @michaelrbock.com · 18/02/2026
1/ One year ago, no AI model could calculate a single tax return correctly. Today, Claude Opus 4.6 gets 52.94% right. Here's the full timeline of how AI went from 0% to halfway to replacing TurboTax:
120
Michael R. Bock @michaelrbock.com · 11/02/2026
You Can Just Do Things: 8 months ago I decided to create the first-ever AI eval for Tax filing: TaxCalcBench. Today, it's the industry standard: Just in the past few months, companies have started citing TaxCalcBench in their marketing, announcements, and blog posts.
100
Michael R. Bock @michaelrbock.com · 10/02/2026
I need proof before I make claims, so I can now confidently say: Claude Opus 4.6 is an incredible model. No model has been able to do this before: Claude Opus 4.6 is now able to exactly correctly compute 52.94% of tax returns in the TaxCalcBench dataset, handily beating GPT-5 w/ Web Search.
120
Michael R. Bock @michaelrbock.com · 05/02/2026
What _can't_ Claude Opus 4.6 do? (I'll be testing if it can file taxes using TaxCalcBench today - stay tuned!)
010
Michael R. Bock @michaelrbock.com · 05/02/2026
One cold DM changed my life: Five years ago, @GavinNachbar & I had applied for the @southpkcommons Founder Fellowship and had an interview coming up. I (cold) DM'd someone I respected online and they helped us prepare for the interview.
100
Michael R. Bock @michaelrbock.com · 22/01/2026
It's amazing that Oura ring independently knows I recently sold the company I founded:
000
Michael R. Bock @michaelrbock.com · 20/01/2026
Here's what I was doing in 2016:
010
Michael R. Bock @michaelrbock.com · 07/01/2026
Claude Opus 4.5 (w/ Claude Code) is known as the best coding model today, but which model is the best at filing taxes? We, at Column Tax, tested the latest crop of frontier models and here's how they stacked up on TaxCalcBench:
120
Michael R. Bock @michaelrbock.com · 16/12/2025
1/ After 5 years, I’m proud to share that Column Tax has found a new home. We’ve been acquired by Aiwyn. I couldn’t be more sure this is the right move for our business, tech, and team. P.S. these pics are the moments we started & sold the company:
140
Michael R. Bock @michaelrbock.com · 03/11/2025
imagine falling for the most obvious spy of all time on bumble ??? (a friend sent me this screenshot, I'm married 😅)
010
Michael R. Bock @michaelrbock.com · 29/10/2025
How will we know that AI has really “made it”? The task that most exemplifies our ability to automate knowledge work is “doing your taxes”. At Column Tax we’re now within line of sight to fully automating taxes. We started the company at the perfect moment, with LLMs just on the horizon.
110
Michael R. Bock @michaelrbock.com · 23/10/2025
Positive review of my most popular blog post: "Hypothesis Sheets - how to navigate and exit the idea maze with a (good) startup idea". Glad to hear the founder whisper networks are still sharing this knowledge around.
110
Michael R. Bock @michaelrbock.com · 18/09/2025
3/ GPT-5 is impressive in many ways especially because it's knowledge cutoff is still September 2024 but it's not the leader in tax calculation today (even with maximal test time compute)
100
Michael R. Bock @michaelrbock.com · 18/09/2025
1/ GPT-5 is worse than Gemini 2.5 Pro at filing your taxes (but it's really close and they both can't do it yet) we proved it via our tax calculation benchmark:
110
Michael R. Bock @michaelrbock.com · 17/09/2025
I got married last month.🤵‍♂️👰‍♀️ Here's what it taught me about B2B2C tax software: Just kidding :) but I do really recommend getting married to the love of your life with all your friends & family around!
030
Michael R. Bock @michaelrbock.com · 13/08/2025
no one had even heard of git worktress before claude code
010
Michael R. Bock @michaelrbock.com · 03/08/2025
amazing ChatGPT Agent Mode use case: find & validate coupon codes without having to test them yourself
020
Michael R. Bock @michaelrbock.com · 23/07/2025
8/ Models are also inconsistent: using pass^k (a measure of reliability of a model across multiple runs on the same task), performance degrades with additional runs meaning models mess up in new & surprising ways when calculating tax returns.
100
Michael R. Bock @michaelrbock.com · 23/07/2025
7/ For some models, performance improves with increased inference-time compute (thinking budget tokens) but not for the best model (Gemini 2.5 Pro), suggesting alternative techniques/scaffolding/orchestration is required to get AI to do this tax calculation task.
100
Michael R. Bock @michaelrbock.com · 23/07/2025
6/ Models consistently: 1. Misuse tax tables 2. Make calculation errors For example, models will hallucinate line numbers on Forms or use incorrect eligibility limits.
100
Michael R. Bock @michaelrbock.com · 23/07/2025
5/ Takeaway: models can’t calculate tax returns reliably today. Even on this simplified data set and allowing the models to output to a simplified format, the best model only calculates 32.35% of returns correctly.
100
Michael R. Bock @michaelrbock.com · 23/07/2025
4/ TaxCalcBench is a dataset of 51 pairs of user inputs and the expected tax return output + a testing harness. We made the task easy for the models. We provide: - all of the data (e.g. W-2s) needed to file a return - the expected output in IRS XML format
100