Sign in

Michael R. Bock

@michaelrbock.com
42 followers 81 following 224 posts

co-founder of Column Tax // michaelrbock.com

PostsRepliesMedia
Michael R. Bock @michaelrbock.com · 01/10/2026
6 Sol (high thinking): $0.24/return, 84s Same accuracy, but 6.1 Sol is 29% cheaper & 1.75x faster.
010
Michael R. Bock @michaelrbock.com · 01/10/2026
This was an interesting finding when running our eval: GPT-6.1 Sol isn't more accurate than GPT-6 Sol calculating tax returns, but it is cheaper and faster. At their best, both are tied for accuracy around 65%. But to hit ~50%: 6.1 Sol (low thinking): $0.17/return, 48s
120
Michael R. Bock @michaelrbock.com · 29/09/2026
the guidance: x.com/edwinarbus/...
010
Michael R. Bock @michaelrbock.com · 29/09/2026
This was surprising to me (but maybe shouldn't be if you're paying attention to Anthropic guidance): Opus 5.5 beats Sonnet 5.5 on COST On TaxCalcBench (testing AI's ability to calculate tax returns), Opus 5.5 at xhigh reasoning is both more accurate AND cheaper than Sonnet 5.5
110
Michael R. Bock @michaelrbock.com · 24/09/2026
All results + model outputs are public here: github.com/column-tax/...
github.com
GitHub - column-tax/tax-calc-bench: Code & data for TaxCalcBench
Code & data for TaxCalcBench. Contribute to column-tax/tax-calc-bench development by creating an account on GitHub.
000
Michael R. Bock @michaelrbock.com · 24/09/2026
But Anthropic made the bigger generational leap: Opus 5 → 5.5: 42% → 62% (+10 returns) GPT-5.6 Sol → GPT-6 Sol: 62% → 64% (+1 return) And without web search, Opus 5.5 wins: 38% vs. 26%.
100
Michael R. Bock @michaelrbock.com · 24/09/2026
The harder they think, the wider the gap. At high thinking, both got 46% of returns right. Sol: $0.24/return. Opus 5.5: $1.75. At max thinking: $0.34 vs. $3.45.
100
Michael R. Bock @michaelrbock.com · 24/09/2026
Both models are almost tied on accuracy, but not on time or cost: GPT-6 Sol: $0.34/return, ~2 min Opus 5.5: $3.45/return, ~7 min 10x the price and 3x the time for one fewer correct return.
100
Michael R. Bock @michaelrbock.com · 24/09/2026
Claude Opus 5.5 is the #2 model in the world at filing taxes according to TaxCalcBench. Unfortunately, it's also 10x more expensive than the new #1 model: GPT-6 Sol.
110
Michael R. Bock @michaelrbock.com · 08/09/2026
For tax filing, we built TaxCalcBench so we can test new models within days, inspect every failure, and share the full results publicly. Whatever domain you work in, you need the same thing. Full results: github.com/column-tax/...
github.com
GitHub - column-tax/tax-calc-bench: Code & data for TaxCalcBench
Code & data for TaxCalcBench. Contribute to column-tax/tax-calc-bench development by creating an account on GitHub.
000
Michael R. Bock @michaelrbock.com · 08/09/2026
But they don’t tell you how it performs on your task, with your tools, at your reasoning budget. The model release cycle is now faster than most companies’ evaluation cycle. That is a problem.
100
Michael R. Bock @michaelrbock.com · 08/09/2026
- GPT-6 Astra w/ web search: 60% (20 of 50) The headline is Gemini. Compared to 3.7 Flash, 3.8 is almost double as good at tax return calculation. The upshot of all these releases? Release announcements tell you what a model is designed to do.
100
Michael R. Bock @michaelrbock.com · 08/09/2026
So we tested all of the new models on 50 hyper-realistic federal and state tax returns: - Gemini 3.8 Flash w/ web search: 52.5% (21 of 40 completed returns) - Claude Fable 5.1 w/ web search: 46% (23 of 50) - Meta Muse Spark 1.3 w/ web search: 26% (13 of 50)
100
Michael R. Bock @michaelrbock.com · 08/09/2026
- Wednesday: Google released Gemini 3.8 Flash - Wednesday: Meta released Muse Spark 1.3 - Friday: OpenAI announced GPT-6 Astra All four releases make some version of the same promise: better coding, better agents, and better long-horizon knowledge work.
100
Michael R. Bock @michaelrbock.com · 08/09/2026
Last week was pretty wild for AI: there were four major AI models launched in four days. All four are already on TaxCalcBench, including GPT-6 Astra. But the surprising result: Meta & Gemini are back in the race. Last week’s release calendar: - Tuesday: Anthropic released Claude Fable 5.1
110
Michael R. Bock @michaelrbock.com · 01/09/2026
In any case, here's the process I used, which worked really well - hope it's helpful: michaelrbock.com/execution/
michaelrbock.com
Execution Review: how to keep projects on-track
How to ensure projects are on-track, even when you can't lead them all.
000
Michael R. Bock @michaelrbock.com · 01/09/2026
This is honestly the most-fun part of the job: you're a manager now so you don't get to code or solve fun technical problems, but you do get to scratch the problem solving itch here: what's wrong with the project and how do you get it back on track?
100
Michael R. Bock @michaelrbock.com · 01/09/2026
At that point, you'll want to make sure that your DRIs (you have DRIs for every project, right??) are executing well, and step in to help when things are off-track.
100
Michael R. Bock @michaelrbock.com · 01/09/2026
Not sure if this will be useful unless you're a CTO, Head of Engineering, or maybe Engineering Mgr, but figured I'd share: At some point (this has probably already happened), you won't be able to individually lead every project yourself.
100
Michael R. Bock @michaelrbock.com · 19/08/2026
All 150 model outputs behind these three rows, plus their evaluation reports, are public in TaxCalcBench. Full results and model releases: github.com/column-tax/...
github.com
GitHub - column-tax/tax-calc-bench: Code & data for TaxCalcBench
Code & data for TaxCalcBench. Contribute to column-tax/tax-calc-bench development by creating an account on GitHub.
000
Michael R. Bock @michaelrbock.com · 19/08/2026
Plus, Kimi K3's weights are available today. Meta has announced Muse Spark 1.2's open-weight release, but it hasn't happened yet. Model comparisons need more than one score. The winner can change with the metric, the tools, and even what you mean by “open.”
100
Michael R. Bock @michaelrbock.com · 19/08/2026
- 8 of 50 perfect returns - 11 of 50 under lenient scoring - 68.28% line-level accuracy That last number is only 2.79 points above Kimi, despite Kimi running without web search. The upshot? There isn't one winner in the U.S. vs. China open-model race.
100
Michael R. Bock @michaelrbock.com · 19/08/2026
- Under lenient scoring, Kimi won: 6 correct returns to Meta's 4. - By-line, Kimi won by more than 10 percentage points: 65.49% to 55.35%. Meta produced one more perfect return. Kimi got much more of the tax forms right overall. Then we gave Muse Spark web search. It jumped to:
100
Michael R. Bock @michaelrbock.com · 19/08/2026
So we put the two models through TaxCalcBench: 50 hyper-realistic U.S. federal and state tax returns. Who won? It depends on what you measure. Without web search: - Meta Muse Spark 1.2 calculated 4 of 50 returns perfectly. Kimi K3 calculated 3.
100
Michael R. Bock @michaelrbock.com · 19/08/2026
This is something pretty interesting I'm watching right now: the race between U.S. vs. China for the best open model. Meta says it will release an open-weight version of Muse Spark 1.2 soon. China's Moonshot AI has already released the weights for Kimi K3.
100
Michael R. Bock @michaelrbock.com · 10/08/2026
This blog post includes the algorithm, examples, and more on every part of the hiring lifecycle in the age of AI: sourcing, interviewing, references, comp negotiation, and closing. Read it here: michaelrbock.com/hiring/
michaelrbock.com
The 4 step algorithm to making your first hires
The algorithm I used for every hire I made at Column Tax.
000
Michael R. Bock @michaelrbock.com · 10/08/2026
Before I made my first-ever hires, I was lucky enough to learn this 4-step algorithm that I've personally used for every hire I've made since. I also forced every hiring manager who worked for me to use it too.
100
Michael R. Bock @michaelrbock.com · 10/08/2026
As a founder or manager, hiring is likely your *most important job*. As AI continues to automate more of our busywork, the quality of the *people* you hire (and their taste) will make the biggest difference to your company/team's progress.
110
Michael R. Bock @michaelrbock.com · 10/08/2026
Before starting Column Tax, I had never hired anyone. By the time we were acquired, I had hired 50 amazing folks. I used the same algorithm for every hire:
100
Michael R. Bock @michaelrbock.com · 29/07/2026
5/ github.com/column-tax/...
github.com
GitHub - column-tax/tax-calc-bench: Code & data for TaxCalcBench
Code & data for TaxCalcBench. Contribute to column-tax/tax-calc-bench development by creating an account on GitHub.
000
Michael R. Bock @michaelrbock.com · 29/07/2026
4/ So the headline isn’t simply “Opus is better than Fable.” It’s that Opus earns its lead when both models are pushed to their maximum reasoning setting. At every lower setting, Fable matched or beat Opus on strict accuracy. All of the results and model outputs are public in TaxCalcBench:
200
Michael R. Bock @michaelrbock.com · 29/07/2026
3/ - Lowest thinking: Fable won, 18% to 10% - Low: tie, 16% to 16% - Medium: tie, 24% to 24% - High: Fable won, 32% to 28% - Ultrathink: Opus won, 40% to 34% At ultrathink, the gap was even wider under lenient scoring: 56% for Opus versus 44% for Fable.
100
Michael R. Bock @michaelrbock.com · 29/07/2026
2/ - Opus 5 w/ web search: 40% tax returns computed correctly (strict) - Fable 5 w/ web search: 35% But only when we turned thinking all the way up. We ran both models with web search on the same 50 realistic federal and state tax returns at five thinking levels:
100
Michael R. Bock @michaelrbock.com · 29/07/2026
1/ Anthropic just released its 4th model in 2 months: Opus 5. It's half the cost of Fable 5. And surprisingly, for knowledge work tasks like tax filing: Opus performs even better than its more expensive counterpart. We tested Opus 5 on TaxCalcBench. Here are the results vs. Fable 5:
100
Michael R. Bock @michaelrbock.com · 13/07/2026
4/ You have to benchmark your specific task. For tax filing: we built a custom eval (TaxCalcBench) that's become the industry standard. If you do taxes, you can follow along. If you do another knowledge work task, you should build your own benchmark!
000
Michael R. Bock @michaelrbock.com · 13/07/2026
3/ Compare that to Claude Fable 5 which is only at 34%. Claude Code/Cowork may have been first to market, but starting with GPT-5.5, the race to see which company/model is best at knowledge work has been neck-and-neck. What does this mean for your work?
100
Michael R. Bock @michaelrbock.com · 13/07/2026
2/ GPT-5.6 Sol (with web search tool use) just scored 58%, beating the previous-best model, GPT-5.5 (54%). That means for 29 out of 50 hyper-realistic and complex federal & state tax returns, GPT-5.6 Sol calculated the perfect output.
100
Michael R. Bock @michaelrbock.com · 13/07/2026
1/ Claude (Code, Cowork, Fable) has all the mindshare right now, but is OpenAI actually the best at knowledge work? Our eval for tax filing says: yes. GPT-5.6 Sol (with web search) just became the best model in the world on TaxCalcBench (v2).
100
Michael R. Bock @michaelrbock.com · 07/07/2026
6/ Lots has happened in the last 11 months: new model releases, copycat benchmarks, and lots of accounting startup launches. But one thing remains: TaxCalcBench is the industry standard for measuring AI's ability to file taxes. Check out the new leaderboard: github.com/column-tax/...
github.com
GitHub - column-tax/tax-calc-bench: Code & data for TaxCalcBench
Code & data for TaxCalcBench. Contribute to column-tax/tax-calc-bench development by creating an account on GitHub.
010
Michael R. Bock @michaelrbock.com · 07/07/2026
5/ AI still fails to calculate tax returns accurately on its own. GPT-5.5 is the highest-scoring model. On its own, it only scores 24%. With web search tool use, that jumps to 54%. Good, but not good enough to trust it with a task that requires 100% correctness.
100
Michael R. Bock @michaelrbock.com · 07/07/2026
4/ - 50 high-quality test cases covering complex tax & financial situations - Realistic PDF inputs (e.g. W-2s, 1099s, K-1s, etc.) - Federal & state returns - Up-to-date, covering tax year 2025 The upshot?
100
Michael R. Bock @michaelrbock.com · 07/07/2026
3/ But the benchmark itself had some limitations: simple returns, federal-only, simplified (text-based) input data. TaxCalcBench v2 (2025) changes that:
100
Michael R. Bock @michaelrbock.com · 07/07/2026
2/ I'm excited to release TaxCalcBench v2, the biggest update to measuring AI's ability to calculate tax returns since June 2025. v1 of TaxCalcBench measured AI's ability to calculate 2024 returns, AI failed to complete the task: the highest-scoring model at the time (Gemini 2.5 Pro) only got ~30%.
100
Michael R. Bock @michaelrbock.com · 07/07/2026
1/ This AI benchmark will decide Intuit's stock price next year. Last year, I released TaxCalcBench, and it moved markets. Here's how:
100
Michael R. Bock @michaelrbock.com · 08/06/2026
5/ Blog post (with Google Docs templates you can copy and examples!) here: michaelrbock.com/reviews/
michaelrbock.com
Problem/Solution Reviews: how to ensure you (& your team) are building the right thing
If you're in a race, you would never want to run faster in the wrong direction (in fact, you'd rather not move at all!)
000
Michael R. Bock @michaelrbock.com · 08/06/2026
4/ I reached out to folks I knew and was lucky to get an amazing template. It helped me ensure we kept building the right things, even as I was managing more. The template was so helpful that I wanted to memorialize it for other founders and Product & Engineering leaders:
100
Michael R. Bock @michaelrbock.com · 08/06/2026
3/ At some point in my company's lifecycle, we scaled the number of projects we were working on to the point that I could no longer lead every single one. As co-founder and head of Product & Engineering, this was a big transition.
100
Michael R. Bock @michaelrbock.com · 08/06/2026
2/ The best AI Product & Engineering leaders are using Problem/Solution Reviews to ensure their team is actually building the right product instead of just building software for the sake of it.
100
Michael R. Bock @michaelrbock.com · 08/06/2026
1/ Here I am about to lead a Product Review meeting (and I'm not nervous at all). I know it's going to go well because I forced my reports to write a Problem/Solution Review doc.
100
Michael R. Bock @michaelrbock.com · 26/03/2026
link to the connector: claude.com/connectors/...
claude.com
Aiwyn Tax | Claude
Prepare your federal & state tax return 100% accurately
110