Sign in

Petr Baudis (pasky)

@bluepasky.bsky.social
185 followers 707 following 85 posts

Rossum.ai. A variety of OSS & AI things in the past (Git, glibc, pre-AlphaGo Pachi, OpenTTD, ...). Your computer might be running some of my code (sorry).

PostsRepliesMedia
Petr Baudis (pasky) @bluepasky.bsky.social · 08/03/2026
tbf opus was done even faster and probably a better summary
000
Petr Baudis (pasky) @bluepasky.bsky.social · 08/03/2026
> bought a minipc > rgb stripe - can i turn it off? > found a windows binary blob, welp > fire up pi > gpt-5.4: "reverse engineer this exe, i'd like to control the LEDs from Linux" (that's the full prompt) > 10min. of `objdump -d` later: "I dug into it." > i can control my rgb stripe in linux > huh
100
Petr Baudis (pasky) @bluepasky.bsky.social · 24/05/2025
Is this about SotA AI in general or comparative to Gemini and Claude?
000
Petr Baudis (pasky) @bluepasky.bsky.social · 06/01/2025
More details: github.com/huggingface/...
github.com
GitHub - huggingface/lighteval: Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends - huggingface/lighteval
050
Reposted by Petr Baudis (pasky)
Petr Baudis (pasky) @bluepasky.bsky.social · 06/01/2025
Definitely believe them regarding technical capabilities of the models. (Ok maybe add a 6-12m buffer.) Where they are imho over indexing is maximum realistic speed of adaptation of the real world. Adoption even of the most amazing stuff will take time, and need a lot of infra.
031
Petr Baudis (pasky) @bluepasky.bsky.social · 06/01/2025
Definitely believe them regarding technical capabilities of the models. (Ok maybe add a 6-12m buffer.) Where they are imho over indexing is maximum realistic speed of adaptation of the real world. Adoption even of the most amazing stuff will take time, and need a lot of infra.
031
Reposted by Petr Baudis (pasky)
Yoav Goldberg @yoavgo.bsky.social · 05/01/2025
i was annoyed at having many chrome tabs with PDF papers having uninformative titles, so i created a small chrome extension to fix it. i'm using it for a while now, works well. today i put it on github. enjoy. github.com/yoavg/pdf-ta...
59822
Reposted by Petr Baudis (pasky)
bijan (spooky version) @bijan.bsky.social · 05/01/2025
good news: despite all the bleak shit going on in the world it is once again Awesome Games Done Quick week. go watch some speedruns www.twitch.tv/gamesdonequick
twitch.tv
GamesDoneQuick - Twitch
Awesome Games Done Quick 2025 - Benefiting the Prevent Cancer Foundation - Ori and the Blind Forest
22911
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
I reposted the thread here! :) bsky.app/profile/xpas...
011
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
Will be back with more later - by losing MCTS we also lost the exploration policy, how to plug it back? This is a repost of a Twitter thread I made yesterday - my experiment on whether I can reach BSky DL audience. Twitter's LLM scene is very lively, I'd love to see more of that here. 'nite! 16/16
010
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
And the pseudocode algorithm for quick reference. 15/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
But this is the gist of the magic. And it results in reported 38x convergence speedup compared to MCTS & impressive benchmark gains. 14/n
110
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
PRIME of course also contains tons of important technical details. (PPO policy with alternative-normalized advantages over raw rewards. The initial finetuning LLM snapshot staing as a reference and considering only token logit changes to it, makes the math work. Formally proving it's >> MCTS..) 13/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
Unlike just using ORM approach alone, this introduces an accretive effect - information on what works is shared across training batches through the reward LLM, and as it learns, it produces better guidance and the convergence of the main LLM speeds up. 12/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
"Continuously" learned? The reward model LLM and the main LLM epochs are interleaved - the estimates are learned in parallel with finetuning the main model, a sort of expectation-minimization dance. 11/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
...and this extra LLM is then used as a Process RM assigning a reward to each token based on its continuously learned estimate of how much that token is helpful. 10/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
Well, how do we know how to reward each token then? Why, by finetuning an *extra* copy of your LLM internally to use as per-token reward model. This extra LLM copy is finetuned using the Outcome RM approach (so sparse rewards just encouraging tokens that lead to good final outcomes)... 9/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
5. Finally, PRIME (Implicit Rewards PRM)! The basic question is: Instead of MCTS-like evaluating each CoT step by N rollouts, could we just run a beam search of N rollouts of CoT from start to end? 8/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
The problem now is that you need to roll out the CoT ten times for each candidate - a Monte Carlo approach. This is not efficient as you are wasting a lot of time on stupid CoT step candidates and lost causes. 7/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
4. Enter MCTS-inspired approaches for automated PRM supervision. Given a few next possible CoT steps, which one is more helpful, can we tell automatically? Well, try rolling out the rest of the CoT ten times for each candidate, and see which one reaches the right answer most frequently! 6/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
3. So let's give a per-CoT-step reinforcement using a PRM (Process Reward Model). Like teaching humans: don't just look at the final result, tell them if their approach was good. Naive idea: just use per step human supervision for the steps. But that's obviously unsustainable, too little data. 5/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
The problem is that each rollout gives you only final outcome info, no sense if any particular CoT step actually helped move towards the result. Convergence is slow, so is ood generalization etc. 4/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
2. Basic approach is to use ORM (Outcome Reward Model) - try answering a query by rolling out a CoT, and check if it led to the right answer. This gives a positive/negative reinforcement to each token in the CoT (each token in the particular CoT gets the same reward). 3/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
1. We are RL tuning an LLM to produce good CoTs (chains of thought). (Good == leading step by step to correct answers to complex queries.) 2/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
Quick primer for non-wizards about the post-MCTS LLM reasoning future: How will LLMs learn to reason efficiently? No math in this thread, ~simple words only! Let's go through the "Process Reinforcement through IMplicit REwards" (PRIME) method. 1/n curvy-check-498.notion.site/Process-Rein...
130
Reposted by Petr Baudis (pasky)
Marios Richards @mariosrichards.bsky.social · 24/12/2024
Science/maths/programming have a tendency to depreciate the value of grinding - smart people don't grind! - yes, that project basically took me only 30 minutes. The 3 and half hours of dead ends I went down obviously don't count. Or the hour I spent installing the wrong package.
5576
Petr Baudis (pasky) @bluepasky.bsky.social · 14/12/2024
Yes
000
Petr Baudis (pasky) @bluepasky.bsky.social · 13/12/2024
Fun fact, AlphaGo etc. actually suck at being deliberative, they are all about "iterative intuition deepening". It doesn't even "cache" local sequence outcomes. Even when permitted only very tiny search tree, AlphaGo will be better than 99%+ serious human Go players.
000
Petr Baudis (pasky) @bluepasky.bsky.social · 13/12/2024
Like these ancient Go (Weiqi) players... "Fan was said to have played very quickly, while Shi very slowly. Fan would sometimes go on a picnic, sing songs or take a nap while Shi was thinking." senseis.xmp.net?FanXipingAnd...
senseis.xmp.net
Fan Xiping and Shi Ding'an at Sensei's Library
Sensei's Library, page: Fan Xiping and Shi Ding'an, keywords: Culture & History, People. SL is a large WikiWikiWeb about the game of Go (Baduk, Weiqi). It's a collaboration and community site. Everyon...
100
Petr Baudis (pasky) @bluepasky.bsky.social · 13/12/2024
"Is System 2 thinking even real, chat?" Humans vary widely between being very intuitive or deliberative. Human-level AGI could plausibly exist even purely as System 1, or with only very basic System 2 mixed in. Just git gud at superhuman intuition.
120
Reposted by Petr Baudis (pasky)
Petr Baudis (pasky) @bluepasky.bsky.social · 07/12/2024
What can RBMK teach us about building AGI?
001
Reposted by Petr Baudis (pasky)
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
Can we build something lasting in a world where technology is easily reproducible? How much time do we have before Artificial General Intelligence (AGI) is here, and how much time do we have afterwards? 2/n
111
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
Thanks to @reshmasohoni.bsky.social @gwern.bsky.social (who disagrees!) and many others for reading a draft of the essay.
000
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
I go in depth on building software products for the age of AGI in this essay I just published: rossum.ai/blog/buildin... 14/14
rossum.ai
Build Software Products for the Age of AGI
What’s the impact of Artificial General Intelligence on software products? Does it make sense to build software in an AGI world?
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
...are you doubting the 2027 AGI timeline? What AI capabilities are outside the focus of general LLM vendors? What moat can tools, platforms and "places" provide? Why will AGI take long to disrupt the technology, won't AGI just rebuild everything? 13/14
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
Building AGI-ready products is not magic or unimaginable sci-fi technology. It is something you can consider today and answer in your product strategy. Just do the hard thing, do it well, and have a plan. 12/14
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
The moat of a brand, built on institutional knowledge gained through sweat & experience and a consistent track record, will not go away with AGI. In fact, your tech and track record should come together so that even AGI will want to use it. 11/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
But your final test will be: Is AGI buying your product? Since: Even AGI vastly technically proficient will have a buy-vs-build dilemma. Why should it waste computronium on sidequests? The opportunity cost will be huge for some time. www.noahpinion.blog/p/plentiful-... 10/n
noahpinion.blog
Plentiful, high-paying jobs in the age of AI
Comparative advantage is very subtle, but incredibly powerful.
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
Even when AGI comes - the world’s economy relies on the collaboration of companies. They take years to adopt any new technology, even with strong market incentives. It may take a decade for AGI to make a big impact on the economy. @jasoncrawford.org blog.rootsofprogress.org/big-tech-tra... 9/n
blog.rootsofprogress.org
Big tech transitions are slow
With implications for AI
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
Building your institutional knowledge (painfully gained from rolling the product out for 1000s of customers, with all the real-world painful mess and nitty-gritty) is an endurance sport, but also something ChatGPT will not replace in a second. 8/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
The true moat is your brand. And that's not jus marketing buzz. In B2B, customers know that institutional knowledge is what really makes or breaks your solution in the real world, and look at your track record to prove you have it. 7/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
It's good to build tech that requires a lot of painstakingly detail-oriented work (e.g. scaleable database engines), or AI capabilities that general LLM vendors aren’t going to care about. But you should not treat it as a moat, but merely as a head start. 6/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
If a San Francisco AI lab can ship AGI at almost any moment, where can we find the safety we need when building a lasting software product? Not in a technological advantage, that is for sure. But it's still worth investing in technology! 5/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
The mid-point estimate of many AI researchers in the top industry labs is that we may get from the current level to AGI around 2027. This tells you that you should have a clear plan to deal with this scenario in the next decade. 4/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
And what then – will AGI want to use any software products? I believe the future will be still good for software products that get their moats right. 3/n
100
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
Can we build something lasting in a world where technology is easily reproducible? How much time do we have before Artificial General Intelligence (AGI) is here, and how much time do we have afterwards? 2/n
111
Petr Baudis (pasky) @bluepasky.bsky.social · 09/12/2024
The product strategy of most SaaS isn't ready for AGI. A disconnect is widening between how product builders think, and how Artificial Intelligence researchers project short-term future. But I believe it's a solvable problem (in any world where a solution exists). 1/n🧵
111
Reposted by Petr Baudis (pasky)
Petr Baudis (pasky) @bluepasky.bsky.social · 29/11/2024
THEY HAVE PLAYED US FOR ABSOLUTE FOOLS
121
Petr Baudis (pasky) @bluepasky.bsky.social · 08/12/2024
If I'm defining my self-worth by anything, it's positive effect of my actions. Also, I try to live as a stoic, finding happiness in what I can control: my thoughts and actions. However, positive effect of my actions isn't just in my control! Only their intent is. /off figuring it out
020
Petr Baudis (pasky) @bluepasky.bsky.social · 08/12/2024
"The first step on the road of not hating yourself is appreciating the impact you have on the life of others." www.youtube.com/watch?v=dp2HnJB2a2s
youtube.com
Haibane Renmei: Escaping Emotional Purgatory - Anime Alphabet
YouTube video by Trixie the Golden Witch
000