Not sure what your use case is, but I assume it is relatively simple. For agentic, long-context tasks the benchmarks tell us that American open models are not even in the ballpark. Consider the Artificial Analysis scores against the closest Chinese models (non-reasoning, similar sizes):