This is why 3 in 4 GenAI projects fail (and this was a survey of execs, so they failed in such a visible way that even the most shielded people in the organization notice that things got worse).
Eventually the outputs will flow far enough downstream to face a real test that isn't exec vibes.