Sign in

Jørgen Lund

@jaalu.bsky.social
47 followers 85 following 64 posts

Industry Ph.D. student in ML, DIPS AS, UiT The Arctic University of Norway github.com/jaalu | he/him

PostsRepliesMedia
Jørgen Lund @jaalu.bsky.social · 19/05/2026
New article out in BMJ Open @bmj.com - making a tool for shared decision making between spine surgeons and patients with XGBoost regression predicting daily function and pain after spine surgery, and a nearest-neighbor lookup to show outcomes from similar cases bmjopen.bmj.com/content/16/5...
bmjopen.bmj.com
121
Jørgen Lund @jaalu.bsky.social · 26/02/2026
I think this is an example of the more important cybersecurity threat from LLMs -- they lower the bar for attacks which _were_ possible previously, but were practically too time-consuming to conduct against random people and projects
000
Reposted by Jørgen Lund
Northern Lights Deep Learning Conference 2027 @nldlconference.bsky.social · 24/02/2026
The Northern Lights Deep Learning ( #NLDL ) Conference returns in 2027! 🤖 ❄️ 🗓️ When: Jan. 12th – 14th 2027 (Main conference), Jan. 11th – 15th 2027 (NLDL Winter School) 📍 Where: UiT- The Arctic University of Norway, Tromsø, Norway More information about NLDL in the comments below 👇 #NLDL2027
121
Jørgen Lund @jaalu.bsky.social · 18/02/2026
One interesting thing - Haiku and Devstral both complied with a "make the user type in 'I have read the documentation' before making this exact change" prompt, the Copilot (Raptor-mini) model acknowledged it but then explained how to make the change
020
Reposted by Jørgen Lund
Alexander Doria @dorialexander.bsky.social · 01/02/2026
It took me weeks, but finally it's there: an overlong blogpost on synthetic pretraining. vintagedata.org/blog/posts/s...
38420
Reposted by Jørgen Lund
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 29/01/2026
this lecture series is really useful for people who took the roundabout route to programming: missing.csail.mit.edu
missing.csail.mit.edu
The Missing Semester of Your CS Education
Master powerful tools that will make you a more productive computer scientist and programmer.
0437
Reposted by Jørgen Lund
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 01/12/2025
Unit tests in research code feel like something that's slowing you down until they wind up saving you six months of suffering
4749
Reposted by Jørgen Lund
mr. TIM @timkellogg.me · 20/11/2025
i read through ~5 pages of the Olmo 3 tech report.. whoah this is the best and most detailed summary of the current state of SOTA LLM training nanochat is good for understanding LLM training, this tech report catches you up to SOTA methods
1444
Reposted by Jørgen Lund
Ai2 @ai2.bsky.social · 20/11/2025
Announcing Olmo 3, a leading fully open LM suite built for reasoning, chat, & tool use, and an open model flow—not just the final weights, but the entire training journey. Best fully open 32B reasoning model & best 32B base model. 🧵
16917
Reposted by Jørgen Lund
Northern Lights Deep Learning Conference 2027 @nldlconference.bsky.social · 14/11/2025
Secure your NLDL 2026 registration before the fees increase on December 1st! 🤖 ❄️ Registration is open until January 1st 2026, but we recommend registering early to avoid expensive hotel prices More info in the comments 👇
133
Jørgen Lund @jaalu.bsky.social · 13/11/2025
Today I learnt that OpenSSL really does not like it if you try to pass in an X.509 certificate which only consists of the word "Blah"
000
Jørgen Lund @jaalu.bsky.social · 13/11/2025
Playing around with the PleIAs "smallest viable model" Monad, and realizing that with 4-bit quantization (storing 56 M parameters in ~27 MB) and a SuperDisk drive (to use the FD32MB format), you could turn it into a chat model that fits on a standard 3.5 inch diskette
2190
Reposted by Jørgen Lund
Northern Lights Deep Learning Conference 2027 @nldlconference.bsky.social · 05/11/2025
We are excited to have Mihaela van der Schaar and Anders Boyd as Winter School speakers at the Northern Lights Deep Learning Conference 2026! Read more about van der Schaar's and Boyd's and other Winter School tutorials in the comments 👇
133
Jørgen Lund @jaalu.bsky.social · 27/10/2025
Having used the Mac for a bit, I realize most of my time is spent in the same software I was using on Windows (Obsidian, Zotero, Marimo), but Homebrew is quite nice, Ghostty is a good terminal, and UTM is a good QEMU/virtualization frontend for the Windows things I do need
000
Reposted by Jørgen Lund
David Marx @digthatdata.bsky.social · 25/10/2025
> RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets Paper: arxiv.org/abs/2502.09615 Web: www.liuisabella.com/RigAnything/ Code: github.com/Isabella98Li... Model: huggingface.co/Isabellaliu/...
092
Reposted by Jørgen Lund
Sung Kim @sungkim.bsky.social · 24/10/2025
Neural audio codecs: how to get audio into LLMs The plan: sandwich a language model in an audio encoder/decoder pair (=neural audio codec), allowing it to predict audio continuations. kyutai.org/next/codec-e...
0193
Reposted by Jørgen Lund
Sung Kim @sungkim.bsky.social · 20/10/2025
Deepseek's DeepSeek-OCR Model: huggingface.co/deepseek-ai/... Paper: github.com/deepseek-ai/... Repo: github.com/deepseek-ai/...
0152
Reposted by Jørgen Lund
mr. TIM @timkellogg.me · 15/10/2025
Is 32B-4bit equal to 16B-8bit? Depends on the task * math: precision matters * knowledge: effective param count is more important * 4B-8bit threshold — for bigger prefer quant, smaller prefer more params * parallel TTC only works above 4B-8bit arxiv.org/abs/2510.10964
A scatter plot titled “AIME25 — Total Memory vs. Accuracy (Qwen3)” compares model accuracy (%) against total memory usage (weights + KV cache, in GB) for various Qwen3 model sizes and quantization levels.

Axes:
	•	X-axis: Total Memory (Weight + KV Cache) [GB] (log scale, ranging roughly from 1 to 100)
	•	Y-axis: Accuracy (%), ranging from 0 to 75

Legend:
	•	Colors: model sizes —
	•	0.6B (yellow)
	•	1.7B (orange)
	•	4B (salmon)
	•	8B (pink)
	•	14B (purple)
	•	32B (blue)
	•	Shapes: precision levels —
	•	Circle: 16-bit
	•	Triangle: 8-bit
	•	Square: 4-bit
	•	Marker size: context length —
	•	Small: 2k tokens
	•	Large: 30k tokens

Main trend:
Larger models (rightward and darker colors) achieve higher accuracy but require significantly more memory. Smaller models (left, yellow/orange) stay below 30% accuracy. Compression (8-bit or 4-bit) lowers memory usage but can reduce accuracy slightly.

Inset zoom (upper center):
A close-up box highlights the 8B (8-bit) and 14B (4-bit) models showing their proximity in accuracy despite differing memory footprints.

Overall, the chart demonstrates scaling behavior for Qwen3 models—accuracy grows with total memory and model size, with diminishing returns beyond the 14B range.
3318
Reposted by Jørgen Lund
LaurieWired @lauriewired.bsky.social · 14/10/2025
GPU computing before CUDA was *weird*.
 Memory primitives were graphics shaped, not computer science shaped.
 Want to do math on an array? Store it as an RGBA texture. 
Fragment Shader for processing. *Paint* the result in a big rectangle.
6697
Reposted by Jørgen Lund
Simon Willison @simonwillison.net · 14/10/2025
nanochat by Andrej Karpathy is neat - 8,000 lines of code (mostly Python, a tiny bit of Rust) that can train an LLM on $100 of rented cloud compute which can then be served with a web chat UI on a much smaller machine simonwillison.net/2025/Oct/13/...
simonwillison.net
nanochat
Really interesting new project from Andrej Karpathy, described at length in this discussion post. It provides a full ChatGPT-style LLM, including training, inference and a web Ui, that can be …
420421
Reposted by Jørgen Lund
Northern Lights Deep Learning Conference 2027 @nldlconference.bsky.social · 10/10/2025
Only 7 days left to submit your abstract for the Northern Lights Deep Learning Conference 2026! 🤖 ❄️ 📅 Abstract submission deadline: October 17th 2025 More information about submission guidelines on nldl.org
022
Reposted by Jørgen Lund
Sung Kim @sungkim.bsky.social · 10/10/2025
Defying Transformers: Searching for "Fixed Points" of Pretrained LLMs by Jiacheng Liu He wondered what CAN'T be transformed by Transformers? So, he wrote a fun blog post on finding "fixed points" of your LLMs. If you prompt it with a fixed point token,
2232
Reposted by Jørgen Lund
thebes @vgel.me · 05/10/2025
new blog post! why do LLMs freak out over the seahorse emoji? i put llama-3.3-70b through its paces with the logit lens to find out, and explain what the logit lens (everyone's favorite underrated interpretability tool) is in the process. link in reply!
821048
Reposted by Jørgen Lund
Dmytro Mishkin @ducha-aiki.bsky.social · 01/10/2025
How Diffusion Models Memorize Juyeop Kim, Songkuk Kim, Jong-Seok Lee tl;dr: classifier-free-guidance is to blame arxiv.org/abs/2509.25705
0102
Reposted by Jørgen Lund
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
Very happy this paper got accepted to NeurIPS 2025 as a Spotlight! 😁 Main takeaway: In mechanistic interpretability, we need assumptions about how DNNs encode concepts in their representations (eg, the linear representation hypothesis). Without them, we can claim any DNN implements any algorithm!
0254
Reposted by Jørgen Lund
Naomi Saphra @nsaphra.bsky.social · 01/10/2025
really neat clear explainer for the new on “centralizing flows” to theoretically model learning dynamics
centralflows.github.io
Understanding Optimization in Deep Learning with Central Flows
14410
Jørgen Lund @jaalu.bsky.social · 02/10/2025
Half-serious question: Could one tackle every "language models produce X because Y is in the training data" hypothesis in one go by training a large retrieval transformer, and doing side-by-side evaluations with/without a filter for Y on the retriever? #MLSky
100
Reposted by Jørgen Lund
Alexander Doria @dorialexander.bsky.social · 27/09/2025
And new paper out: Pleias 1.0: the First Family of Language Models Trained on Fully Open Data How we train an open everything model on a new pretraining environment with releasable data (Common Corpus) with an open source framework (Nanotron from HuggingFace). www.sciencedirect.com/science/arti...
817951
Jørgen Lund @jaalu.bsky.social · 18/09/2025
Stepping outside my lane for a bit, NPM really needs trusted publishing/provenance, but I think a per-package "will use secure-context-only features" flag would go a long way too, there are few good reasons for a package to suddenly _start_ using eval(), Fetch, or executing arbitrary commands
stepsecurity.io
ctrl/tinycolor and 40+ NPM Packages Compromised - StepSecurity
The popular @ctrl/tinycolor package with over 2 million weekly downloads has been compromised alongside 40+ other NPM packages in a sophisticated supply chain attack dubbed
000
Jørgen Lund @jaalu.bsky.social · 16/09/2025
As someone interested in cryptography and privacy-preserving measures, it's cool to see differential privacy being applied to ML practically like this - this seems to basically stop memorization of passages from the training set, even allowing for approximate matches (up to 10% edit distance)
010
Reposted by Jørgen Lund
Gus @gusthema.bsky.social · 04/09/2025
We've just released an amazing Embedding model: EmbeddingGemma, the new best-in-class open embedding model! 🚀 🏆 Top multilingual model on MTEB (<500M) 💾 Runs on <200MB RAM ⚙️ Customizable output for on-device use 🧩 Integrated with your favorite tools developers.googleblog.com/en/introduci...
developers.googleblog.com
Introducing EmbeddingGemma: The Best-in-Class Open Model for On-Device Embeddings- Google Developers Blog
Discover EmbeddingGemma, Google's new on-device embedding model designed for efficient on-device AI, enabling features like RAG and semantic search.
26612
Reposted by Jørgen Lund
EPFL School of Computer and Communication Sciences @icepfl.bsky.social · 02/09/2025
EPFL, ETH Zurich & CSCS just released Apertus, Switzerland’s first fully open-source large language model. Trained on 15T tokens in 1,000+ languages, it’s built for transparency, responsibility & the public good. Read more: actu.epfl.ch/news/apertus...
15429
Jørgen Lund @jaalu.bsky.social · 01/09/2025
As someone who hasn't used MacOS in a minute (since... Sierra?), which MacOS utilities do people on #MLSky recommend? (Homebrew is a given, but not sure which terminal emulators people prefer now, for instance)
110
Reposted by Jørgen Lund
mr. TIM @timkellogg.me · 31/08/2025
Limits of vector search a new GDM paper shows that embeddings can’t represent combinations of concepts well e.g. Dave likes blue trucks AND Ford trucks even k=2 sub-predicates make SOTA embedding models fall apart www.alphaxiv.org/pdf/2508.21038
alphaxiv.org
On the Theoretical Limitations of Embedding-Based Retrieval | alphaXiv
View recent discussion. Abstract: Vector embeddings have been tasked with an ever-increasing set of retrieval tasks over the years, with a nascent rise in using them for reasoning, instruction-followi...
28322
Reposted by Jørgen Lund
SFI Visual Intelligence @sfi-vi.bsky.social · 01/09/2025
NLDL 2026 paper submission deadline extended 🤖 ❄️ Due to a large amount of requests, we wish to inform you that the paper submission deadline is extended until September 14th, 23:59 CEST More info about submissions at nldl.org
022
Jørgen Lund @jaalu.bsky.social · 29/08/2025
Interesting argument that the architectural blocks of language models are almost always surjective, implying that any given output text has at least one corresponding input which can generate it (They do note that *finding* that input isn't necessarily easy, and assume you control input embeddings)
100
Jørgen Lund @jaalu.bsky.social · 20/08/2025
Wonderful project which packages ELIZA and other retro chatbots in an OpenAI-compatible API, because we all need a code assistant that can ask "how does the NullReferenceException make you feel?"
120
Jørgen Lund @jaalu.bsky.social · 20/08/2025
Nvidia have released a new set of hybrid Mamba-Transformer language models, but interestingly they've also published the majority of the data used to pretrain them: research.nvidia.com/labs/adlr/NV...
research.nvidia.com
NVIDIA Nemotron Nano 2 and the Nemotron Pretraining Dataset v1
NVIDA Nemotron Nano 2 is a new hybrid Mamba-Transformer reasoning model that achieves on-par or better accuracies compared to comparably sized leading open models at up to 6x higher throughput. Nemotr...
000
Jørgen Lund @jaalu.bsky.social · 19/08/2025
Fun breakdown of the process of implementing a toy TPU in Verilog, starting "from scratch" www.tinytpu.com
tinytpu.com
Tiny TPU
An attempt to understand and build a TPU—by complete novices.
055
Jørgen Lund @jaalu.bsky.social · 07/08/2025
Interesting from DeepMind: Live Music Models arxiv.org/pdf/2508.04651 - low-latency music generation models which run at real time and respond to user input, letting you perform style transfer on your own music on the fly
100
Jørgen Lund @jaalu.bsky.social · 06/08/2025
Shower thought: if you only had pen and paper (no calculator), what kind of neural networks could you feasibly still compute predictions with? I think one could make something like the earliest LeNets work, with ReLU activation and some form of weight quantization to simplify things
110
Jørgen Lund @jaalu.bsky.social · 06/08/2025
Today's small but neat trick: setting up a simple web server in Powershell to serve a Marimo notebook to localhost, giving you a light Python + Numpy + Matplotlib setup on any Windows box without any installation or network requirements
010