Only ~half of interactions were rated usable. Gaps persist across dialect and gender groups. And QA-based evaluation (roughly: Does key information survive translation?) predicts usability far better than standard metrics. Dataset released CC-BY 4.0.