Sign in

Martin Görner

@martin-gorner.bsky.social
440 followers 763 following 778 posts

AI/ML engineer. Previously at Google: Product Manager for Keras and TensorFlow and developer advocate on TPUs. Passionate about democratizing Machine Learning.

PostsRepliesMedia
Martin Görner @martin-gorner.bsky.social · 20/04/2026
New Voyager SDK for Axelera chips is out: community.axelera.ai/product-upda...
community.axelera.ai
Voyager SDK: New Pipeline Builder and More | Community
The latest version of Voyager® Software Development Kit (SDK) is here and this release touches nearly every layer of the stack. Whether you're deploying on new hardware, building custom inference pipe...
040
Martin Görner @martin-gorner.bsky.social · 12/01/2026
Expect future LLMs to adopt DroPE in their pre-training, and no longer be limited to their pre-trained context lengths.
020
Martin Görner @martin-gorner.bsky.social · 12/01/2026
SelfExtend (arxiv.org/abs/2401.01325), a no-fine-tuning context extension method I covered in this video youtu.be/TV7qCk1dBWA, is not even mentioned but it has been evaluated in a previous eval paper (arxiv.org/abs/2409.12181) and found to be rubbish (top-right).
100
Martin Görner @martin-gorner.bsky.social · 12/01/2026
This beats all other context extension methods, with or without fine-tuning, by a considerable margin.
100
Martin Görner @martin-gorner.bsky.social · 12/01/2026
Dropping Positional Embeddings, yes, just discarding them towards the end of LLM pre-training, unlocks context generalization in LLMs way beyond their pre-trained context length.
100
Martin Görner @martin-gorner.bsky.social · 21/10/2025
Announcing our next-gen chip: axelera.ai/news/axelera... • 628 TOPS • in-memory compute (IMC) matrix multipliers <- this is Axelera's tech edge • 16 Risc-V vector cores for that will handle pre- and post-processing directly on chip.
axelera.ai
Axelera Announces Europa AIPU, Setting New Industry Benchmark for AI Accelerator Performance, Power Efficiency and Affordability
Axelera® today announced Europa™, an AI processor unit (AIPU) that sets a new performance/price standard for multi-user generative AI and computer vision applications.
010
Martin Görner @martin-gorner.bsky.social · 16/07/2025
Full report here: www.hottech.com/industry-cov...
hottech.com
Evaluating AI Inference Accelerators For Machine Vision Applications — Hot Tech
In a head-to-head battle of AI accelerators, the results are in — and Axelera AI didn’t just win, it ran laps around the competition.
010
Martin Görner @martin-gorner.bsky.social · 16/07/2025
and check out the full report, which has data about more modern models like YOLO8L. For that model, compared to the best NVIDIA card tested, Axelera's Metis is: - 230% faster - 330% more power efficient and also about 3x cheaper
110
Martin Görner @martin-gorner.bsky.social · 16/07/2025
The secret sauce works! www.forbes.com/sites/daveal...
forbes.com
Axelera AI Accelerators Smoke Competitors In Machine Vision Research Study
Domain-specific accelerators are proving they can compete, and in some cases lead, in the metrics that matter most for real-world deployments.
130
Martin Görner @martin-gorner.bsky.social · 30/05/2025
4-chip Metis accelerator PCIe card coming soon: store.axelera.ai/collections/...
store.axelera.ai
PCIe AI accelerator card. Powered by 4 quad-core Metis AIPUs | Axelera AI Store
Axelera AI’s PCIe card, powered by 4 Metis AIPU, offers the highest performance inference acceleration on the market, combining ease of use, power efficiency, and scalability. Key Benefits: The highes...
010
Martin Görner @martin-gorner.bsky.social · 30/05/2025
PCIe and M.2 Metis boards available now: - PCIe: axelera.ai/ai-accelerat... - M.2: axelera.ai/ai-accelerat...
axelera.ai
Metis PCIe AI Inference Acceleration Card | Axelera AI
Looking for powerful & energy-efficient AI acceleration hardware that doesn't break budgets? Discover our PCIe AI inference accelerator card.
110
Martin Görner @martin-gorner.bsky.social · 30/05/2025
50+ models pre-configured in the model zoo are ready to run: github.com/axelera-ai-h...
github.com
100
Martin Görner @martin-gorner.bsky.social · 30/05/2025
Blog post by A-Tang Fan and Doug Watt about Axelera.ai's Voyager SDK: community.axelera.ai/product-upda...
community.axelera.ai
Simplifying Model and Pipeline Deployment with the Voyager SDK | Community
Axelera AI’s A-Tang Fan and Doug Watt explain how the Voyager SDK simplifies the complex task of deploying AI-powered video pipelines on edge devices. This blog explores how its model compiler, model ...
121
Martin Görner @martin-gorner.bsky.social · 07/05/2025
I am impressed and humbled by what the Axelera team was able to bring to market, on only three years, with the Metis chip and Voyager SDK. And it's just a beginning. We have an exciting roadmap ahead! axelera.ai
axelera.ai
Axelera AI - Extreme Performance, Excellent Efficiency. Accelerating Inference at the Edge.
Bring data insights to the edge, increasing the performance of your solutions with a cost-effective and efficient inference chip. Axelera’s AI processing unit is designed to seamlessly integrate into ...
010
Martin Görner @martin-gorner.bsky.social · 07/05/2025
The explosion of new AI models and capabilities, in advanced vision, speech recognition, language models, reasoning etc, needs a novel, energy-efficient approach to AI acceleration to deliver truly magical AI experiences, at the edge and in the datacenter.
110
Martin Görner @martin-gorner.bsky.social · 07/05/2025
I'm delighted to share that I joined the Axelera team this week to deliver the next generation AI compute platform. axelera.ai
230
Martin Görner @martin-gorner.bsky.social · 11/03/2025
You'll see it in this form in Karpathy's original "Pong from pixel" post karpathy.github.io/2016/05/31/rl/ as well as my "RL without a PhD" video from a while ago youtu.be/t1A3NTttvBA.... They also explore a few basic reward assignment strategies. Have fun, don't despair and RL!
youtube.com
TensorFlow and deep reinforcement learning, without a PhD (Google I/O '18)
On the forefront of deep learning research is a technique called reinforcement learning, which bridges the gap between academic deep learning problems and wa...
000
Martin Görner @martin-gorner.bsky.social · 11/03/2025
It is usually written in vector form, using the cross-entropy function. This time, I use 𝛑̅(sᵢₖ) for the *vector* of all move probabilities predicted from game state sᵢₖ, while 𝒙̅ᵢₖ is the one-hot encoded *vector* representing the move actually played in game i move k.
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
One more thing: In modern autograd libraries like PyTorch or JAX, the RL gradient can be computed from the following “pseudo-loss”. Don’t try to find the meaning of this function, it does not have any. It’s just a function that has the gradient we want.
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
So in conclusion, math tells us that Reinforcement Learning is possible, even in multi-turn games where you cannot differentiate across multiple moves. But math tells us nothing about how to do it in practice. Which is why it is hard.
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
What "rewards"? Well the "good" ones, that encourage the "correct" moves! This is pretty much like a delicious recipe saying you should mix "great" ingredients in the "correct" proportions 😭. NOT HELPFUL AT ALL 🤬 !!!
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
Now the bad news: what this equation really means is that the gradient we are looking for is the weighted sum of the gradients of our policy network over many individual games and moves, weighted by an unspecified set of “rewards”.
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
We wanted to maximize the expected reward and managed to approximate the gradients form successive runs of the game, as played by our policy network. We can run backprop after all ... at least in theory 🙁.
110
Martin Görner @martin-gorner.bsky.social · 11/03/2025
... string of zeros followed by the final reward - but we may have finer-grained rewarding strategies. The 1/N constant was folded into the rewards. The good news: yay, This is computable 🎉🥳🎊! The expression only involves our policy network and our rewards.
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
For a more practical application, let's unroll the log-probabilities into individual moves using eq. (1) and rearrange a little. We use the fact that the log of a product is a sum of logs. I have also split the game reward into separate game steps rewards rᵢₖ - worst case a ...
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
But look, this is a sum of a probability × some value. That's an expectation! Which means that instead of computing it directly, we can approximate it from multiple games gᵢ :
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
And now we can start approximating like crazy - and abandon any pretense of doing exact math 😅. First, we use our policy network 𝛑 to approximate the move probabilities.
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
Combining the last two equations we get:
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
We now use a mathematical cheap trick based on the fact that the derivative of log(x) is 1/x. With gradients, this cheap trick reads:
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
To maximize the expectation (3) we compute its gradient. The notation ∇ is the "gradient", or list of partial derivatives relatively to parameters θ. Differentiation is a linear operation so we can enter it into the sum Σ. Also, rewards do not depend on θ so:
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
For example, for a single dice, the possible values are 1, 2, 3, 4, 5, 6, the probability of each outcome is p=⅙ which gives us an expectation of 3.5. And you get roughly the same number by rolling the dice many times and averaging the result. It's the "law of large numbers".
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
In probabilities, the “expectation” of a random variable X is the weighted sum of all possible outcomes xₖ, weighted by their probabilities p(xₖ). The really nice thing about expectations is that you can approximate them: just repeat the experiment and average the outcomes:
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
Introducing the Reinforcement Learning ninja 🥷 hack: we can play many games and try and maximize the "expected reward". Yes, bear with me, this will end up being computable!
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
We have to sample and play multiple moves before a reward is known. With the reward, we would like to backprop through all the moves and adjust the parameters θ of the policy network but we cannot. This process is not differentiable end-to-end 😫.
110
Martin Görner @martin-gorner.bsky.social · 11/03/2025
To play, we sample a move from the predicted probabilities (i.e. roll the dice to pick a move, but with skewed probabilities as predicted by the network). And this is a problem because sampling is not a differentiable operation.
120
Martin Görner @martin-gorner.bsky.social · 11/03/2025
In the "policy gradients" approach to RL, a neural network called "policy network" is used to predict the next move. The network sees the game state and returns next move probabilities. We call 𝛑(xₖ) the probability it predicts for move xₖ. θ is the set of weights of the net.
110
Martin Görner @martin-gorner.bsky.social · 11/03/2025
1) Deterministic play, i.e. no random monsters in Mario, i.e. p(sₖ₊₁|xₖ,sₖ)=1 2) The next move can be predicted from the current game state alone: p(xₖ|sₖ, xₖ₋₁,sₖ₋₁, xₖ₋₂,sₖ₋₂,...)=p(xₖ|sₖ). Then use the cond. prob. formula p(A|B)=p(A,B)/p(B) to get eq (1).
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
Under two reasonable assumptions, the probability of a given game p(g) is simply the product of the individual move probabilities in their respective game states (conditional probabilities). This is eq. (1). The assumptions are:
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
More formally, a "game" g is a sequence of game states s and moves x. Let's call r(g), the reward for the game. The probability of the game can be computed (I'll explain how shortly) from individual move probabilities which can in turn be predicted by a neural network.
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
You play until the outcome is clear, then assign a reward or penalty. With an LLM, this works for math or programming questions when the final answer can be verified. A slight refinement is when you can assign rewards to intermediate steps in the game (ORM vs. PRM, see pic⇩)
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
As often in machine learning, the math is both fun and, in the end, pretty useless 😝. Let's set the stage: wether playing pong (single-player against a deterministic algorithm) or predicting the next token or sentence in a chain-of-thought LLM, the idea is the same:
100
Martin Görner @martin-gorner.bsky.social · 11/03/2025
Reinforcement Learning (RL) just landed a stellar breakthrough with reasoning language models. Yet, RL has a distinctly bad reputation. See “To RL or not to RL” (www.reddit.com/r/MachineLe...) on reddit. I'd like to revisit the basic math of RL to see why. Let's enter the dungeon!
111
Martin Görner @martin-gorner.bsky.social · 20/02/2025
My conclusion: to go beyond the LoRA standard with 10x fewer params, I like the simplicity of Transformers²'s SVF. And if you need more trainable weights, SVFT is an easy extension. Both use all singular values (full rank, no s.v. pruning) and are still cheap 😁. Happy tuning!
130
Martin Görner @martin-gorner.bsky.social · 20/02/2025
There are still more PEFT techniques - DoRA (arxiv.org/abs/2402.09353) which splits weights into magnitudes and directions than tunes those - AdaLoRA (arxiv.org/abs/2303.10512) with a complex mechanism for finding the best tuning rank for a given budget of trainable weights. ...
arxiv.org
AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Fine-tuning large pre-trained language models on downstream tasks has become an important paradigm in NLP. However, common practice fine-tunes all of the parameters in a pre-trained model, which...
110
Martin Görner @martin-gorner.bsky.social · 20/02/2025
Quick sanity check: how diverse are the principal values of a pre-trained LLM? Colab here, checking this on Gemma2 9B: colab.research.google.com/drive/1Igzf... Bottom line: 99% of them are in the 0.1 - 1.1 range. I'm not sure partitioning them into "large" and "small" makes that much sense...
110
Martin Görner @martin-gorner.bsky.social · 20/02/2025
- that truncating the bottom principal values from the SVD still offers a good approximation of the weights matrices - that the fine-tuning data distribution if close to the pre-training one Both questionable IMHO, so I won't detail the math. Some results:
110
Martin Görner @martin-gorner.bsky.social · 20/02/2025
Finally, I'd like to mention LoRA-XS (arxiv.org/abs/2405.17604). Very similar to PiSSA but slightly different mechanism. Also good results with significantly fewer params than LoRA. The paper offers a mathematical explanation of why this setup is "ideal' under two conditions:
110
Martin Görner @martin-gorner.bsky.social · 20/02/2025
The PiSSA paper also has an interesting finding: full fine-tuning is prone to over-fitting. You might get better results in the absolute with a low-rank fine-tuning technique.
110
Martin Görner @martin-gorner.bsky.social · 20/02/2025
Surprisingly, MiLoRA seems to have the upper hand, at least when tuning on math datasets which are probably fairly aligned with the original pre-training. Arguably, PiSSA should be better for bending the behavior of the LLM further from its pre-training.
120
Martin Görner @martin-gorner.bsky.social · 20/02/2025
From the paper: "PiSSA is designed to approximate full finetuning by adapting the principal singular components, which are believed to capture the essence of the weight matrices. In contrast, MiLoRA aims to adapt to new tasks while maximally retaining the base model’s knowledge."
110