Sign in

Marcin Junczys-Dowmunt (Marian NMT)

@marian-nmt.bsky.social
166 followers 134 following 16 posts

NLP. NMT. Main author of Marian NMT. Research Scientist at Microsoft Translator. marian-nmt.github.io

PostsRepliesMedia
Marcin Junczys-Dowmunt (Marian NMT) @marian-nmt.bsky.social · 05/02/2025
Still no bookmarks?
000
Marcin Junczys-Dowmunt (Marian NMT) @marian-nmt.bsky.social · 28/01/2025
Hi, the Microsoft Translator research team is looking for an intern for the summer. If you a PhD student in Machine Translation, Natural Language Processing, or related, check it out: aka.ms/mtintern
aka.ms
Search Jobs | Microsoft Careers
054
Reposted by Marcin Junczys-Dowmunt (Marian NMT)
Clem Delangue 🤗 @clem.hf.co · 16/12/2024
Just 10 days after o1's public debut, we’re thrilled to unveil the open-source version of the technique behind its success: scaling test-time compute By giving models more "time to think," Llama 1B outperforms Llama 8B in math—beating a model 8x its size. The full recipe is open-source!
48319
Marcin Junczys-Dowmunt (Marian NMT) @marian-nmt.bsky.social · 16/12/2024
Rant: Apparently every vector-based sentence alignment tool insists on having an unusable file-based API.
000
Reposted by Marcin Junczys-Dowmunt (Marian NMT)
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 16/12/2024
Wrote up some notes on Microsoft's new Phi-4 LLM. They trained it on a LOT of synthetic data, and the details of how and why they did that are really interesting. simonwillison.net/2024/Dec/15/phi-4…
simonwillison.net
Phi-4 Technical Report
Phi-4 is the latest LLM from Microsoft Research. It has 14B parameters and claims to be a big leap forward in the overall Phi series. From [Introducing Phi-4: Microsoft’s Newest …
0146
Reposted by Marcin Junczys-Dowmunt (Marian NMT)
porcoesphino.bsky.social @porcoesphino.bsky.social · 14/12/2024
It's messier, but I think this one slaps the point home a bit stronger by adding the giant squid footage. I think unique weather, like lighting sprites, would make the point just as well.
Chart of time vs:
- number of cameras (exponentially increasing ),
- giant squid footage (exponentially increasing ),
- bigfoot footage (small and not increasing), and
- good quality UFO footage (small and not increasing)
2903
Reposted by Marcin Junczys-Dowmunt (Marian NMT)
Yoav Goldberg @yoavgo.bsky.social · 13/12/2024
the anthropomorphizing in this LLM scheming paper is through the roof and the interpretations are wild, but still a cute set of experiments and a fun skim, showing some interesting behaviors. arxiv.org/abs/2412.04984
arxiv.org
Frontier Models are Capable of In-context Scheming
Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives - ...
4373
Reposted by Marcin Junczys-Dowmunt (Marian NMT)
Artidoro Pagnoni @artidoro.bsky.social · 13/12/2024
🚀 Introducing the Byte Latent Transformer (BLT) – A LLM architecture that scales better than Llama 3 using patches instead of tokens 🤯 Paper 📄 dl.fbaipublicfiles.com/blt/BLT__Pat... Code 🛠️ github.com/facebookrese...
56015
Marcin Junczys-Dowmunt (Marian NMT) @marian-nmt.bsky.social · 13/12/2024
So... no edit button, huh?
220
Marcin Junczys-Dowmunt (Marian NMT) @marian-nmt.bsky.social · 13/12/2024
This place needs bookmarks.
000
Marcin Junczys-Dowmunt (Marian NMT) @marian-nmt.bsky.social · 11/12/2024
Hi!
1101