Using this and other methods, we create BonaFide, a dataset of 3k labeled CoTs, which we use to evaluate faithfulness metrics, finding that most perform near chance!
We also show that many metrics are ill-suited for real-time deployment, taking >1k seconds to run per example.