Sign in

Takashi Ishida

@tksiia.bsky.social
31 followers 9 following 18 posts

🤖🗼 takashiishida.github.io

PostsRepliesMedia
Takashi Ishida @tksiia.bsky.social · 06/07/2026
I'm in Seoul for ICML 2026!🇰🇷 If you’re interested in evals, coding and long-horizon agents, or issues such as contamination, reward hacking, and model cheating, please stop by our posters (details below) or say hi if you see me around, I'd love to chat!
100
Takashi Ishida @tksiia.bsky.social · 26/06/2026
Excited to share CoffeeBench!!☕️☕️☕️ We evaluate LLM agents in a 90-day B2B coffee supply-chain economy spanning farmers, roasters, and retailers, where autonomous firms negotiate, manage inventory, set prices, handle invoices, and manage cash flow. arxiv.org/abs/2606.16613 github.com/sakanaai/cof...
2214
Reposted by Takashi Ishida
skythanawat.bsky.social @skythanawat.bsky.social · 23/06/2026
Coding agents are evaluated with unit tests: more tests passed = better model. But if tests or feedback are accessible, models may learn to game them. We introduce CapCode to detect suspiciously high scores, and CapReward to discourage them during RL. 🧵1/10
221
Takashi Ishida @tksiia.bsky.social · 21/04/2026
I will be at ICLR 2026 this week! @iclr-conf.bsky.social I would love to chat with people with similar research interests. Please check out our posters at Session 6, Pavilion 3 & Pavilion 4, on Saturday, April 25. See you there!
010
Takashi Ishida @tksiia.bsky.social · 21/04/2026
If you are at ICLR 2026 this week and are interested in model evaluation and financial benchmarks, please stop by our poster at Session 6, Pavilion 3, on Saturday, April 25! 👋🇧🇷
010
Reposted by Takashi Ishida
hardmaru @hardmaru.bsky.social · 20/04/2026
I am very proud of our team for releasing EDINET-Bench, and it is fantastic to see a Japanese financial dataset recognized at #ICLR2026 this week. We need more diverse, non-English datasets to evaluate models in the real world. Paper: openreview.net/forum?id=Dxn...
back arrowGo to ICLR 2026 Conference homepage
EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements

Large Language Models (LLMs) have made remarkable progress, surpassing human performance on several benchmarks in domains such as mathematics and coding. A key driver of this progress has been the development of benchmark datasets. In contrast, the financial domain poses higher entry barriers due to its demand for specialized expertise, and benchmarks remain relatively scarce compared to those in mathematics or coding. We introduce EDINET-Bench, an open-source Japanese financial benchmark designed to evaluate LLMs on challenging tasks such as accounting fraud detection, earnings forecasting, and industry classification. EDINET-Bench is constructed from ten years of annual reports filed by Japanese companies. These tasks require models to process entire annual reports and integrate information across multiple tables and textual sections, demanding expert-level reasoning that is challenging even for human professionals. Our experiments show that even state-of-the-art LLMs struggle in this domain, performing only marginally better than logistic regression in binary classification tasks such as fraud detection and earnings forecasting. Our results show that simply providing reports to LLMs in a straightforward setting is not enough. This highlights the need for benchmark frameworks that better reflect the environments in which financial professionals operate, with richer scaffolding such as realistic simulations and task-specific reasoning support to enable more effective problem solving. We make our dataset and code publicly available to support future research.
1182
Reposted by Takashi Ishida
Johannes Ackermann @johannesack.bsky.social · 26/02/2026
Tired of KL penalties constraining your model? But don't want your policy to just hack the reward? Try Gradient Regularization! We show it beats a KL penalty in RLHF, RLVR and LLM-as-a-Judge! 🧵1/7
121
Takashi Ishida @tksiia.bsky.social · 23/02/2026
📣We are excited to launch the CapBencher toolkit today! CapBencher caps the best achievable accuracy on purpose. If an LLM scores above the cap, it’s a red flag for leakage, contamination, or leaderboard gaming🚩 If you are creating a new benchmark, you might find it useful👉
110
Takashi Ishida @tksiia.bsky.social · 26/01/2026
Happy to share that our papers were accepted to ICLR 2026!🇧🇷 Big thanks to my co-authors! Scalable oversight: arxiv.org/abs/2510.22500 EDINET-Bench: arxiv.org/abs/2506.08762 Optimal classification error estimation: arxiv.org/abs/2505.20761
arxiv.org
Scalable Oversight via Partitioned Human Supervision
As artificial intelligence (AI) systems approach and surpass expert human performance across a broad range of tasks, obtaining high-quality human supervision for evaluation and training becomes increa...
010
Takashi Ishida @tksiia.bsky.social · 28/11/2025
We're excited to announce the launch of Google Developer Group AI for Science Japan!🎉 If you're interested, we’d love to have you join our community. GDG AI for Science Japan gdg.community.dev/gdg-ai-for-s...
gdg.community.dev
GDG AI for Science - Japan | Google Developer Groups
140
Reposted by Takashi Ishida
Johannes Ackermann @johannesack.bsky.social · 29/07/2025
Reward models do not have the capacity to fully capture human preferences. If they can't represent human preferences, how can we hope to use them to align a language model? In our #COLM2025 "Off-Policy Corrected Reward Modeling for RLHF", we investigate this issue 🧵
121
Takashi Ishida @tksiia.bsky.social · 29/09/2025
Released bibfixer 🎉 A tiny AI tool that cleans & standardizes your BibTeX files using LLMs + web search. No more tedious edits like fixing capitalization (ai -> AI), swapping arXiv for the conference version, or expanding "and others" into full author lists. Let bibfixer do the grunt work for you!
github.com
GitHub - takashiishida/bibfixer: A Python tool that automatically cleans, completes, and standardizes BibTeX entries using LLMs and web search.
A Python tool that automatically cleans, completes, and standardizes BibTeX entries using LLMs and web search. - takashiishida/bibfixer
010
Reposted by Takashi Ishida
hardmaru @hardmaru.bsky.social · 09/06/2025
EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements Paper: pub.sakana.ai/edinet-bench/ We just released a Japanese financial benchmark designed to evaluate the performance of AI Agents on challenging financial tasks like accounting fraud detection.
github.com
GitHub - SakanaAI/EDINET-Bench: Evaluating the performance of LLMs on Japanese challenging financial tasks.
Evaluating the performance of LLMs on Japanese challenging financial tasks. - SakanaAI/EDINET-Bench
192
Takashi Ishida @tksiia.bsky.social · 09/06/2025
Excited to announce EDINET-Bench, a financial LLM benchmark built from 40k annual reports in Japan! It features accounting fraud detection, earnings forecasting, industry classification, and includes our tool edinet2dataset as a foundation for designing new tasks. Hope researchers find it useful!
011
Reposted by Takashi Ishida
Conference on Language Modeling @colmweb.org · 27/05/2025
Our discussion period just started. Authors, please read our instructions carefully. We require responses by June 2. But, what you really want to hear about is stats .... right? -> 🧵
2175