Sai Prasanna @saiprasanna.in · 10/09/2025arxiv.org/abs/2203.091...arxiv.orgOn the Pitfalls of Heteroscedastic Uncertainty Estimation with Probabilistic Neural NetworksCapturing aleatoric uncertainty is a critical part of many machine learning systems. In deep learning, a common approach to this end is to train a neural network to estimate the parameters of a hetero... 020
Sai Prasanna @saiprasanna.in · 10/09/2025Use Beta NLL for regression when you also predict standard deviations, a simple change to NLL that works reliably better. 140
Sai Prasanna @saiprasanna.in · 03/08/2025If open-endedness has to be fundamentally subjectively measured, what are the factors of the agent makes it so if we fix humans as the final arbiter or evaluator. Does embodiment/action space etc of the agent matter for a human evaluator of open-endedness? 010
Sai Prasanna @saiprasanna.in · 25/06/2025🤣 generalrobots.substack.com/p/a-brief-in...generalrobots.substack.comA Brief, Incomplete, and Mostly Wrong History of Robotics(An homage to one of my favorite pieces on the internet: A Brief, Incomplete, and Mostly Wrong History of Programming Languages) 020
Sai Prasanna @saiprasanna.in · 27/03/2025But this is from the vibes of Tübingen from 1.5 days of visit. I have lived in Freiburg for 3 years 000
Sai Prasanna @saiprasanna.in · 27/03/2025Had a discussion with a fellow not-so-political Indian colleague doing a PhD in computer science in Europe. He is now thinking twice on his plan to go for an exchange at an US lab 0181
Sai Prasanna @saiprasanna.in · 15/03/2025contraptions.venkateshrao.com/p/discworld-...contraptions.venkateshrao.comDiscworld RulesAnd LOTR is brain-rot for technologists 091
Reposted by Sai PrasannaVenkatesh Rao 🔹 @vgr.bsky.social · 08/03/2025This might be the most fun I’ve had writing an essay in a while. Felt some of that old going-nuts-with-an-idea energy flowing. open.substack.com/pub/contrapt...open.substack.comDiscworld RulesAnd LOTR is brain-rot for technologists 4569
Reposted by Sai PrasannaTom Silver @tomssilver.bsky.social · 02/03/2025This week's #PaperILike is "Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming" (Bertsekas 2024). If you know 1 of {RL, controls} and want to understand the other, this is a good starting point. PDF: arxiv.org/abs/2406.00592arxiv.orgModel Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic ProgrammingIn this paper we describe a new conceptual framework that connects approximate Dynamic Programming (DP), Model Predictive Control (MPC), and Reinforcement Learning (RL). This framework centers around ... 0438
Sai Prasanna @saiprasanna.in · 01/03/2025One strategy I guess is to have good stream of good (BS filter) and diverse (topics, areas) inputs (books, research papers, what not) And not get bogged by the fact that I am too distracted to go deep into one input stream (book or podcast or article or paper) at a time 010
Sai Prasanna @saiprasanna.in · 01/03/2025Do any of my fellow fox-brained folks (@vgr.bsky.social) have good strategies for aiding background processing? I think background processing feels more foxy thing intutively @visakanv.com (not sure if you identify as a fox in the fox hedgehog dichotomy though) 000
Sai Prasanna @saiprasanna.in · 01/03/2025I guess the trick would be to do actions that makes the mind and emotional states to be fertile for the background processing to happen consistently! 110
Sai Prasanna @saiprasanna.in · 01/03/2025I realized how I background process tonnes of information, from work/research and emotional stuff. And it works well, leads to good research ideas, wise processing of tough situations! But It's so hard to learn to trust this as conscious thinking for solving problems feels more under my "control" 210
Sai Prasanna @saiprasanna.in · 01/03/2025Conditioning gap in latent space world models is due to how uncertainty can go into latent posterior distribution or the learnt prior (dynamics model) and not conditioning on the future would put the uncertainty incorrectly into dynamics model. 000
Sai Prasanna @saiprasanna.in · 01/03/2025To re-think I think the problems could be orthogonal. Clever hans pertains to teacher forcing during training leading to easy solutions for lot of the timesteps skewing it to not learning the hard timestep which is most important for test-time. 100
Sai Prasanna @saiprasanna.in · 01/03/2025(Shame that argmax.org/blog is down now!! They're a really nice less known research group in Volkswagen doing important stuff in world models.) Anyways, If these two problems are related, just establishing that would be an amazing paper!argmax.org 100
Sai Prasanna @saiprasanna.in · 01/03/2025Blog web.archive.org/web/20241108... paper arxiv.org/abs/2101.07046 Applied to world models for pomdps web.archive.org/web/20241009...web.archive.orgA Tale of Gaps - argmax.orgWith variational auto-encoders (VAEs), it has become popular to approximate Bayesian inference with neural networks. This scales Bayesian inference to large datasets and deep generative models at the ... 120
Sai Prasanna @saiprasanna.in · 01/03/2025 Conditioning gap: When you train a value encoder that computes an approximate posterior that's conditioned partially (say on past tokens), then the posterior has a worse lower bound than one also conditioned on everything (also future tokens). 100
Sai Prasanna @saiprasanna.in · 01/03/2025It reminds me of another problem, and I'm not sure if it's equivalent or if it's some dual problem. It's called the conditioning gap in latent space inference. 100
Sai Prasanna @saiprasanna.in · 01/03/2025The fix involves modelling forward and backward directions. I haven't grokked it fully, but I learnt about the above problem there. I find this two papers a really nice sequence of a fundamental problem and then a solution! 100
Sai Prasanna @saiprasanna.in · 01/03/2025And there is a new paper that claims to fix this for transformer architecture!!! They call it "belief state transformer". Apparently it fixes lots of practical problems arising due to clever hans cheat! arxiv.org/abs/2410.23506arxiv.orgThe Belief State TransformerWe introduce the "Belief State Transformer", a next-token predictor that takes both a prefix and suffix as inputs, with a novel objective of predicting both the next token for the prefix and the previ... 110
Sai Prasanna @saiprasanna.in · 01/03/2025Since teacher forcing makes the model learn easy cheat for most easy tokens, the learning dynamics make it hard to find the correct strategy for the first token. 100
Sai Prasanna @saiprasanna.in · 01/03/2025But teacher forcing makes it easy to predict all tokens after the first branching token by paying attention only to previous token and remembering or attending to the edge with this. This strategy doesn't work for the first token where there are the start branches 100
Sai Prasanna @saiprasanna.in · 01/03/2025The easiest coorect solution for the model is to look at the edge with the goal (since it's star graph there is only one edge) and work the way backwards to the start (in it's computation) and output the path one by one forward. 100
Sai Prasanna @saiprasanna.in · 01/03/2025Imagine a task where you give a list of edges of a star graph, start and end node, and train a model with a teacher forcing you to predict the list of tokens in the path from the start to the end. (edge 1, edge 2 ...) (start, goal) (start, intermediate1, intermediate 2 .. .goal) 100
Sai Prasanna @saiprasanna.in · 01/03/2025 This failure occurs in distribution, not OOD. And it apparently is general for any model learning next-token prediction regardless of recurrence (linear or otherwise) or attention!!! 100
Sai Prasanna @saiprasanna.in · 01/03/2025This is orthogonal to the more well-known compounding error problem in auto-regression and distribution mismatch issue in teacher forcing. 100
Sai Prasanna @saiprasanna.in · 01/03/2025TIL: "Clever Hans cheat" for next-token prediction. A subtle but interesting issue with next-token prediction. In the purely forward next token prediction objective, teacher forcing can lead to learning dynamics where the models don't even generalize "in-distribution"!! arxiv.org/abs/2403.06963arxiv.orgThe pitfalls of next-token predictionCan a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple multi-token objective... 1112
Sai Prasanna @saiprasanna.in · 27/01/2025Break the Monday Productivity ceiling with this super awesome 4 hour techno set on.soundcloud.com/hXTcWTTsYUNK...on.soundcloud.comYetti Meissner @ Sisyphos Hammerhalle 09/08/14🖤 BOOKING CONTACT chris@stilvortalent.de 030
Sai Prasanna @saiprasanna.in · 26/01/2025Their aesthetic is soooo gooooood m.youtube.com/watch?v=hGQu...m.youtube.comGlass Beams - 'Mahal EP' (Full Live Performance)YouTube video by Glass Beams 020
Sai Prasanna @saiprasanna.in · 26/01/2025Yep and here's a cover by then which is soooo gooooood m.youtube.com/watch?v=w_3h...m.youtube.comGlass Beams - One Raga to a Disco Beat (A cover of 'Raga Bhairav' by Charanjit Singh)YouTube video by Glass Beams 030
Sai Prasanna @saiprasanna.in · 24/01/2025I am 50/50 about this, they can also make deceptive gain in sample efficiency for problems that can be solved by stringing human knowledge together, so we don't make actual algorithm gains in sample efficiency for problems not solvable by humans? Or maybe this just requires tougher benchmarks 000
Sai Prasanna @saiprasanna.in · 20/01/2025Caffinate and hard techno to keep the pace going 🔥 open.spotify.com/track/5WLHfd...open.spotify.comHATREDGostwork, BANDEE · HATRED · Song · 2025 010
Sai Prasanna @saiprasanna.in · 20/01/2025Monday kick starter open.spotify.com/track/6QXjBA...open.spotify.comEnimatekKore-G · Enimatek · Song · 2023 110
Sai Prasanna @saiprasanna.in · 30/12/2024If I have a really good photo that could be potentially used in many contexts, what's the best place to make money with it? My friend has a really good eye for photos and we want to try a side venture selling some of her stuff 020
Reposted by Sai PrasannaVenkatesh Rao 🔹 @vgr.bsky.social · 27/12/2024RIP Manmohan Singh. Dude changed all our lives in 1991 for the better. His stint as turnaround finance minister was revolutionary even if his later stint as PM was rather hapless (for which Nehru dynasty is more to blame).en.wikipedia.orgManmohan Singh - Wikipedia 2212
Reposted by Sai PrasannaRonen Tamari @ronentk.me · 25/12/2024Looks like a cool study. Lots to learn from ants about large scale coordination www.pnas.org/doi/10.1073/... "Our results exemplify how simple minds can easily enjoy scalability while complex brains require extensive communication to cooperate efficiently." h/t @petersuber.bsky.socialpnas.orgComparing cooperative geometric puzzle solving in ants versus humans | PNASBiological ensembles use collective intelligence to tackle challenges together, but suboptimal coordination can undermine the effectiveness of grou... 2266
Sai Prasanna @saiprasanna.in · 26/12/2024This album is going to be timeless open.spotify.com/album/32yQDx...open.spotify.comMahalGlass Beams · EP · 2024 · 5 songs 231
Sai Prasanna @saiprasanna.in · 18/12/2024Wednesday Quirky mood open.spotify.com/track/3RBhQ7...open.spotify.comDoing The Beeston BumpLeafcutter John · Yes! Come Parade With Us · Song · 2019 010
Sai Prasanna @saiprasanna.in · 18/12/2024 Mass effect 3 ending "choices" don't feel so bad or unrealistic seen in this light of ultimate end of augmenting ourselves with sophisticated tools and gradually tools that seem more like us in some ways. I picked the merge. 010
Sai Prasanna @saiprasanna.in · 18/12/2024Ted Chiangs criticism that genAI as used currently reduces the decision landscape of humans checkouts in this case as well www.newyorker.com/culture/the-...newyorker.comWhy A.I. Isn’t Going to Make ArtTo create a novel or a painting, an artist makes choices that are fundamentally alien to artificial intelligence. 130
Sai Prasanna @saiprasanna.in · 18/12/2024V/LM AR glasses always had this Rick and Morty death crystals vibe to me. 120
Sai Prasanna @saiprasanna.in · 18/12/2024Maybe its not either/or, it depends on the level at which these things can be customised? But the gap between direct neuro response augmentation/shift and indirect ways feel unsurmountable, atleast in the AR glasses type interface. Maybe direct neuro augmentation devices in the future changes it 210
Sai Prasanna @saiprasanna.in · 18/12/2024For example, for people whom it takes effort and energy to read facial expressions and empathise emotionally instead of cognitively, one can see AR glasses with VLM support easily fixing the baseline ability. But it comes at a cost of further letting direct emotional sensitivity degrade 120
Sai Prasanna @saiprasanna.in · 18/12/2024Does augmenting ourselves with V/LLMs to cognitive gaps make self actualization even more difficult on average? Stands stark in contrast with (more difficult/slower to show positive outcomr) augmentation strategies like meditation or psychedelics 150
Reposted by Sai PrasannaAndreas Kirsch @blackhc.bsky.social · 17/12/2024The slides for my lectures on (Bayesian) Active Learning, Information Theory, and Uncertainty are online now 🥳 They cover quite a bit from basic information theory to some recent papers: blackhc.github.io/balitu/ and I'll try to add proper course notes over time 🤗 317628