Sign in

Dinos Papakostas

@din0s.me
541 followers 121 following 77 posts

🧠 neural IR 📋 model evals 🏋️ lifting weights Incredibly optimistic about the future! 𝕏: din0s_ / din0s.me

PostsRepliesMedia
Dinos Papakostas @din0s.me · 07/12/2024
wait what happened, why do i have apple intelligence on my mac... 👀
030
Dinos Papakostas @din0s.me · 05/12/2024
on a similar note, this is another banger paper you should read
000
Dinos Papakostas @din0s.me · 04/12/2024
spent the day reading the OpenScholar and Tülu 3 papers in detail. AI2 x UW has been on fire lately, this is what real “Open AI” looks like 😉. big thanks to both author teams for such thorough and well-crafted work.
140
Dinos Papakostas @din0s.me · 28/11/2024
0100
Dinos Papakostas @din0s.me · 28/11/2024
i guess elon was right about anti-censorship huh? imagine getting permabanned for posting a hf dataset
7142
Dinos Papakostas @din0s.me · 27/11/2024
qwenbros we are winning
050
Dinos Papakostas @din0s.me · 26/11/2024
it's ok we can share this hat
a black hat with the text "gpu poor." written in the front
220
Dinos Papakostas @din0s.me · 25/11/2024
seems like a year old tweet import
110
Dinos Papakostas @din0s.me · 22/11/2024
0141
Dinos Papakostas @din0s.me · 22/11/2024
0224
Dinos Papakostas @din0s.me · 18/11/2024
I don't see how the "needle in the haystack" benchmark is an indicator of how models handle long context with *relevant* information (& why it's so popular). All it tests for is how well they can ignore noise and spot outliers, while most real use cases are about effectively using *all* context.
Benchmark results of the Qwen2.5-Turbo model on the "Needle in the Haystack" problem. The model gets 100% retrieval accuracy regardless of document depth or context length.
050