Sign in

Sai Prasanna

@saiprasanna.in
2.2K followers 689 following 287 posts

See(k)ing the surreal Causal World Models for Curious Robots @ University of Tübingen/Max Planck Institute for Intelligent Systems 🇩🇪 #reinforcementlearning #robotics #causality #meditation #vegan

PostsRepliesMedia
Sai Prasanna @saiprasanna.in · 10/09/2025
arxiv.org/abs/2203.091...
arxiv.org
On the Pitfalls of Heteroscedastic Uncertainty Estimation with Probabilistic Neural Networks
Capturing aleatoric uncertainty is a critical part of many machine learning systems. In deep learning, a common approach to this end is to train a neural network to estimate the parameters of a hetero...
020
Sai Prasanna @saiprasanna.in · 10/09/2025
Use Beta NLL for regression when you also predict standard deviations, a simple change to NLL that works reliably better.
140
Sai Prasanna @saiprasanna.in · 03/08/2025
If open-endedness has to be fundamentally subjectively measured, what are the factors of the agent makes it so if we fix humans as the final arbiter or evaluator. Does embodiment/action space etc of the agent matter for a human evaluator of open-endedness?
010
Sai Prasanna @saiprasanna.in · 25/06/2025
🤣 generalrobots.substack.com/p/a-brief-in...
generalrobots.substack.com
A Brief, Incomplete, and Mostly Wrong History of Robotics
(An homage to one of my favorite pieces on the internet: A Brief, Incomplete, and Mostly Wrong History of Programming Languages)
020
Sai Prasanna @saiprasanna.in · 27/03/2025
But this is from the vibes of Tübingen from 1.5 days of visit. I have lived in Freiburg for 3 years
000
Sai Prasanna @saiprasanna.in · 27/03/2025
Freiburg
110
Sai Prasanna @saiprasanna.in · 27/03/2025
Tübingen
110
Sai Prasanna @saiprasanna.in · 27/03/2025
Tübingen: Freiburg:: Introvert:Extrovert
140
Sai Prasanna @saiprasanna.in · 27/03/2025
Had a discussion with a fellow not-so-political Indian colleague doing a PhD in computer science in Europe. He is now thinking twice on his plan to go for an exchange at an US lab
0181
Sai Prasanna @saiprasanna.in · 15/03/2025
contraptions.venkateshrao.com/p/discworld-...
contraptions.venkateshrao.com
Discworld Rules
And LOTR is brain-rot for technologists
091
Reposted by Sai Prasanna
Venkatesh Rao 🔹 @vgr.bsky.social · 08/03/2025
This might be the most fun I’ve had writing an essay in a while. Felt some of that old going-nuts-with-an-idea energy flowing. open.substack.com/pub/contrapt...
open.substack.com
Discworld Rules
And LOTR is brain-rot for technologists
4569
Reposted by Sai Prasanna
Tom Silver @tomssilver.bsky.social · 02/03/2025
This week's #PaperILike is "Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming" (Bertsekas 2024). If you know 1 of {RL, controls} and want to understand the other, this is a good starting point. PDF: arxiv.org/abs/2406.00592
arxiv.org
Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming
In this paper we describe a new conceptual framework that connects approximate Dynamic Programming (DP), Model Predictive Control (MPC), and Reinforcement Learning (RL). This framework centers around ...
0438
Sai Prasanna @saiprasanna.in · 02/03/2025
Curious to know which show
110
Sai Prasanna @saiprasanna.in · 01/03/2025
One strategy I guess is to have good stream of good (BS filter) and diverse (topics, areas) inputs (books, research papers, what not) And not get bogged by the fact that I am too distracted to go deep into one input stream (book or podcast or article or paper) at a time
010
Sai Prasanna @saiprasanna.in · 01/03/2025
Do any of my fellow fox-brained folks (@vgr.bsky.social) have good strategies for aiding background processing? I think background processing feels more foxy thing intutively @visakanv.com (not sure if you identify as a fox in the fox hedgehog dichotomy though)
000
Sai Prasanna @saiprasanna.in · 01/03/2025
I guess the trick would be to do actions that makes the mind and emotional states to be fertile for the background processing to happen consistently!
110
Sai Prasanna @saiprasanna.in · 01/03/2025
I realized how I background process tonnes of information, from work/research and emotional stuff. And it works well, leads to good research ideas, wise processing of tough situations! But It's so hard to learn to trust this as conscious thinking for solving problems feels more under my "control"
210
Sai Prasanna @saiprasanna.in · 01/03/2025
Conditioning gap in latent space world models is due to how uncertainty can go into latent posterior distribution or the learnt prior (dynamics model) and not conditioning on the future would put the uncertainty incorrectly into dynamics model.
000
Sai Prasanna @saiprasanna.in · 01/03/2025
To re-think I think the problems could be orthogonal. Clever hans pertains to teacher forcing during training leading to easy solutions for lot of the timesteps skewing it to not learning the hard timestep which is most important for test-time.
100
Sai Prasanna @saiprasanna.in · 01/03/2025
(Shame that argmax.org/blog is down now!! They're a really nice less known research group in Volkswagen doing important stuff in world models.) Anyways, If these two problems are related, just establishing that would be an amazing paper!
argmax.org
100
Sai Prasanna @saiprasanna.in · 01/03/2025
Blog web.archive.org/web/20241108... paper arxiv.org/abs/2101.07046 Applied to world models for pomdps web.archive.org/web/20241009...
web.archive.org
A Tale of Gaps - argmax.org
With variational auto-encoders (VAEs), it has become popular to approximate Bayesian inference with neural networks. This scales Bayesian inference to large datasets and deep generative models at the ...
120
Sai Prasanna @saiprasanna.in · 01/03/2025
Conditioning gap: When you train a value encoder that computes an approximate posterior that's conditioned partially (say on past tokens), then the posterior has a worse lower bound than one also conditioned on everything (also future tokens).
100
Sai Prasanna @saiprasanna.in · 01/03/2025
It reminds me of another problem, and I'm not sure if it's equivalent or if it's some dual problem. It's called the conditioning gap in latent space inference.
100
Sai Prasanna @saiprasanna.in · 01/03/2025
The fix involves modelling forward and backward directions. I haven't grokked it fully, but I learnt about the above problem there. I find this two papers a really nice sequence of a fundamental problem and then a solution!
100
Sai Prasanna @saiprasanna.in · 01/03/2025
And there is a new paper that claims to fix this for transformer architecture!!! They call it "belief state transformer". Apparently it fixes lots of practical problems arising due to clever hans cheat! arxiv.org/abs/2410.23506
arxiv.org
The Belief State Transformer
We introduce the "Belief State Transformer", a next-token predictor that takes both a prefix and suffix as inputs, with a novel objective of predicting both the next token for the prefix and the previ...
110
Sai Prasanna @saiprasanna.in · 01/03/2025
Since teacher forcing makes the model learn easy cheat for most easy tokens, the learning dynamics make it hard to find the correct strategy for the first token.
100
Sai Prasanna @saiprasanna.in · 01/03/2025
But teacher forcing makes it easy to predict all tokens after the first branching token by paying attention only to previous token and remembering or attending to the edge with this. This strategy doesn't work for the first token where there are the start branches
100
Sai Prasanna @saiprasanna.in · 01/03/2025
The easiest coorect solution for the model is to look at the edge with the goal (since it's star graph there is only one edge) and work the way backwards to the start (in it's computation) and output the path one by one forward.
100
Sai Prasanna @saiprasanna.in · 01/03/2025
Imagine a task where you give a list of edges of a star graph, start and end node, and train a model with a teacher forcing you to predict the list of tokens in the path from the start to the end. (edge 1, edge 2 ...) (start, goal) (start, intermediate1, intermediate 2 .. .goal)
star graph illustrating the clever hans trick
100
Sai Prasanna @saiprasanna.in · 01/03/2025
This failure occurs in distribution, not OOD. And it apparently is general for any model learning next-token prediction regardless of recurrence (linear or otherwise) or attention!!!
100
Sai Prasanna @saiprasanna.in · 01/03/2025
This is orthogonal to the more well-known compounding error problem in auto-regression and distribution mismatch issue in teacher forcing.
100
Sai Prasanna @saiprasanna.in · 01/03/2025
TIL: "Clever Hans cheat" for next-token prediction. A subtle but interesting issue with next-token prediction. In the purely forward next token prediction objective, teacher forcing can lead to learning dynamics where the models don't even generalize "in-distribution"!! arxiv.org/abs/2403.06963
arxiv.org
The pitfalls of next-token prediction
Can a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple multi-token objective...
1112
Sai Prasanna @saiprasanna.in · 27/01/2025
Break the Monday Productivity ceiling with this super awesome 4 hour techno set on.soundcloud.com/hXTcWTTsYUNK...
on.soundcloud.com
Yetti Meissner @ Sisyphos Hammerhalle 09/08/14
🖤 BOOKING CONTACT chris@stilvortalent.de
030
Sai Prasanna @saiprasanna.in · 26/01/2025
Their aesthetic is soooo gooooood m.youtube.com/watch?v=hGQu...
m.youtube.com
Glass Beams - 'Mahal EP' (Full Live Performance)
YouTube video by Glass Beams
020
Sai Prasanna @saiprasanna.in · 26/01/2025
Yep and here's a cover by then which is soooo gooooood m.youtube.com/watch?v=w_3h...
m.youtube.com
Glass Beams - One Raga to a Disco Beat (A cover of 'Raga Bhairav' by Charanjit Singh)
YouTube video by Glass Beams
030
Sai Prasanna @saiprasanna.in · 24/01/2025
I am 50/50 about this, they can also make deceptive gain in sample efficiency for problems that can be solved by stringing human knowledge together, so we don't make actual algorithm gains in sample efficiency for problems not solvable by humans? Or maybe this just requires tougher benchmarks
000
Sai Prasanna @saiprasanna.in · 20/01/2025
Caffinate and hard techno to keep the pace going 🔥 open.spotify.com/track/5WLHfd...
open.spotify.com
HATRED
Gostwork, BANDEE · HATRED · Song · 2025
010
Sai Prasanna @saiprasanna.in · 20/01/2025
Monday kick starter open.spotify.com/track/6QXjBA...
open.spotify.com
Enimatek
Kore-G · Enimatek · Song · 2023
110
Sai Prasanna @saiprasanna.in · 30/12/2024
If I have a really good photo that could be potentially used in many contexts, what's the best place to make money with it? My friend has a really good eye for photos and we want to try a side venture selling some of her stuff
020
Reposted by Sai Prasanna
Venkatesh Rao 🔹 @vgr.bsky.social · 27/12/2024
RIP Manmohan Singh. Dude changed all our lives in 1991 for the better. His stint as turnaround finance minister was revolutionary even if his later stint as PM was rather hapless (for which Nehru dynasty is more to blame).
en.wikipedia.org
Manmohan Singh - Wikipedia
2212
Reposted by Sai Prasanna
Ronen Tamari @ronentk.me · 25/12/2024
Looks like a cool study. Lots to learn from ants about large scale coordination www.pnas.org/doi/10.1073/... "Our results exemplify how simple minds can easily enjoy scalability while complex brains require extensive communication to cooperate efficiently." h/t @petersuber.bsky.social
pnas.org
Comparing cooperative geometric puzzle solving in ants versus humans | PNAS
Biological ensembles use collective intelligence to tackle challenges together, but suboptimal coordination can undermine the effectiveness of grou...
2266
Sai Prasanna @saiprasanna.in · 26/12/2024
This album is going to be timeless open.spotify.com/album/32yQDx...
open.spotify.com
Mahal
Glass Beams · EP · 2024 · 5 songs
231
Sai Prasanna @saiprasanna.in · 18/12/2024
Wednesday Quirky mood open.spotify.com/track/3RBhQ7...
open.spotify.com
Doing The Beeston Bump
Leafcutter John · Yes! Come Parade With Us · Song · 2019
010
Sai Prasanna @saiprasanna.in · 18/12/2024
Mass effect 3 ending "choices" don't feel so bad or unrealistic seen in this light of ultimate end of augmenting ourselves with sophisticated tools and gradually tools that seem more like us in some ways. I picked the merge.
010
Sai Prasanna @saiprasanna.in · 18/12/2024
Ted Chiangs criticism that genAI as used currently reduces the decision landscape of humans checkouts in this case as well www.newyorker.com/culture/the-...
newyorker.com
Why A.I. Isn’t Going to Make Art
To create a novel or a painting, an artist makes choices that are fundamentally alien to artificial intelligence.
130
Sai Prasanna @saiprasanna.in · 18/12/2024
V/LM AR glasses always had this Rick and Morty death crystals vibe to me.
Morty from Rick and morty with a blue death crystal embedded on head and text says "I do as crystal guides"
120
Sai Prasanna @saiprasanna.in · 18/12/2024
Maybe its not either/or, it depends on the level at which these things can be customised? But the gap between direct neuro response augmentation/shift and indirect ways feel unsurmountable, atleast in the AR glasses type interface. Maybe direct neuro augmentation devices in the future changes it
210
Sai Prasanna @saiprasanna.in · 18/12/2024
For example, for people whom it takes effort and energy to read facial expressions and empathise emotionally instead of cognitively, one can see AR glasses with VLM support easily fixing the baseline ability. But it comes at a cost of further letting direct emotional sensitivity degrade
120
Sai Prasanna @saiprasanna.in · 18/12/2024
Does augmenting ourselves with V/LLMs to cognitive gaps make self actualization even more difficult on average? Stands stark in contrast with (more difficult/slower to show positive outcomr) augmentation strategies like meditation or psychedelics
150
Reposted by Sai Prasanna
Andreas Kirsch @blackhc.bsky.social · 17/12/2024
The slides for my lectures on (Bayesian) Active Learning, Information Theory, and Uncertainty are online now 🥳 They cover quite a bit from basic information theory to some recent papers: blackhc.github.io/balitu/ and I'll try to add proper course notes over time 🤗
317628