I don't see how the "needle in the haystack" benchmark is an indicator of how models handle long context with *relevant* information (& why it's so popular). All it tests for is how well they can ignore noise and spot outliers, while most real use cases are about effectively using *all* context.