M2.5 didn't quant particularly well (noticeable benchmark degradation at Q4, so likely much worse at realistic long context work), and M2.7 is even worse (same quanters had trouble even creating their quants). Try the full quality hosted model first if you haven't used it yet.