Language models for MT are good at generating large candidate pools that contain good translations; they're less good at assigning the highest score to the best translation.
This is where reranking comes in: rescoring with COMET, noisy channel decoding, minimum Bayes risk, etc.