Reposted by @calccon.bsky.socialcalccon.bsky.social @calccon.bsky.social · 27/06/2025SETOL: SemiEmpirical Theory of (Deep) Learning The draft is just about ready Why weightwatcher--and the HTSR theory--work github.com/CalculatedCo...github.com 021
calccon.bsky.social @calccon.bsky.social · 27/06/2025SETOL: SemiEmpirical Theory of (Deep) Learning The draft is just about ready Why weightwatcher--and the HTSR theory--work github.com/CalculatedCo...github.com 021
calccon.bsky.social @calccon.bsky.social · 11/06/2025Where does HTSR and the weightwatcher theory come from ? T𝐡𝐞 𝗪𝐢𝐥𝐬𝐨𝐧 𝐄𝐱𝐚𝐜𝐭 𝐑𝐞𝐧𝐨𝐫𝐦𝐚𝐥𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐆𝐫𝐨𝐮𝐩 A new principle of learning that is not only fundamental to our understanding of AI 🧠 I have a draft of the theory monograph up on github, and it is just about ready lnkd.in/gBsZ-QKF 000
calccon.bsky.social @calccon.bsky.social · 02/06/2025🎉 🚀 𝐀𝐩𝐩𝐫𝐨𝐚𝐜𝐡𝐢𝐧𝐠 𝟐𝟎𝟎𝐊 𝐝𝐨𝐰𝐧𝐥𝐨𝐚𝐝𝐬 🥳 💯 𝐖𝐞𝐢𝐠𝐡𝐭𝐖𝐚𝐭𝐜𝐡𝐞𝐫: 𝐃𝐚𝐭𝐚-𝐅𝐫𝐞𝐞 𝐃𝐢𝐚𝐠𝐧𝐨𝐬𝐭𝐢𝐜𝐬 𝐟𝐨𝐫 𝐃𝐞𝐞𝐩 𝐋𝐞𝐚𝐫𝐧𝐢𝐧𝐠 WeightWatcher is based on theoretical research into 𝑾ℎ𝒚 𝑫𝒆𝒆𝒑 𝑳𝒆𝒂𝒓𝒏𝒊𝒏𝒈, using the new 𝐓𝐡𝐞𝐨𝐫𝐲 𝐨𝐟 𝐇𝐞𝐚𝐯𝐲-𝐓𝐚𝐢𝐥𝐞𝐝 𝐒𝐞𝐥𝐟-𝐑𝐞𝐠𝐮𝐥𝐚𝐫𝐢𝐳𝐚𝐭𝐢𝐨𝐧 (HTSR), published in JMLR, Nature Comm., and NeurIPS weightwatcher.ai 000
calccon.bsky.social @calccon.bsky.social · 13/02/2025Reminder: tomorrow at 10AM PST we will be finishing the table read of section 4.2 of the SETOL monograph Here's the video from last week www.youtube.com/watch?v=0WhB... and the latest version of the paper can be found in theory-paperyoutube.comSETOL Paper Table Read Section 4 Part 1YouTube video by Calculation Consulting 100
calccon.bsky.social @calccon.bsky.social · 22/01/2025SETOL: SemiEmpirical Theory of (Deep Learning) & the connection to Renormalization Group Turns out, AI models obey a fundamental law of physics when they are trained well. A big thanks to the ML Research Jam for giving me the opportunity to present. www.slideshare.net/slideshow/se...slideshare.netSETOL: SemiEmpirical Theory of (Deep Learning)SETOL: SemiEmpirical Theory of (Deep Learning) - Download as a PDF or view online for free 010
calccon.bsky.social @calccon.bsky.social · 21/01/2025"We have not used perturbation theory—we have used an axe on the Hamiltonian" Ken Wilson ( Nobel Prize, Physics 1982 ) 000
calccon.bsky.social @calccon.bsky.social · 20/01/2025The weightwatcher theory paper is just about ready. SETOL: SemiEmpirical Theory of (Deep) Learning It's been a passion project of mine for nearly 10 years. before submitting it, I'd like to have a few people read it carefully and comment 210
calccon.bsky.social @calccon.bsky.social · 19/01/2025How did come up with the idea to look for power law signatures in deep learning models ? Why does power law behavior matter in neural systems and deep learning? Here’s the story: 100
calccon.bsky.social @calccon.bsky.social · 17/01/2025The quality of a NN layer has an Effective Free Energy obtained through a volume-preserving change of measure analogous to taking a single step of the Wilson Exact Renormalization Group For an ideal layer (alpha=2), this can be experimentally validated with weightwatcher 000
calccon.bsky.social @calccon.bsky.social · 13/01/2025You can see the emerging signatures of the Wilson Exact Renormalization Group in the best trained layers of modern LLMs like Llama and Falcon. If you're an old physics supernerd like me, that's super cool. But it's also super useful for AI people. weightwatcher.ai 000
calccon.bsky.social @calccon.bsky.social · 25/12/2024 calculatedcontent.com/2024/12/24/w...calculatedcontent.comWeightWatcher, HTSR theory, and the Renormalization GroupThere is a deep connection between the open-source weightwatcher tool, which implements ideas from the theory of Heavy Tailed Self-Regularization (HTSR) of Deep Neural Networks (DNNs), and the Wils… 000
calccon.bsky.social @calccon.bsky.social · 19/12/2024A quick weightwatcher workup on the new Falcon3 base models. As predicted by theory, the average weightwatcher layer alpha systematically decrease with increasing model size. The exception is the 10B model, which is an upscaled model. 000
calccon.bsky.social @calccon.bsky.social · 10/12/2024The theory behind weightwatcher is essentially an application of the Wilson Renormalization Group. The PL exponent alpha=2 is the analogous critical exponent separating the good generalization and overfit (i.e., spin-glass) phases of the NN layer. 000
calccon.bsky.social @calccon.bsky.social · 07/12/2024Why alpha=2 is the ideal state of a NN layer ? In our upcoming monograph, A SemiEmpirical Theory of (Deep) Learning, we show that the HTSR metrics can be derived as an phenomenological Effective Hamiltonian, but one that is governed by a scale-invariant partition function, just like the Wilson RG 200
calccon.bsky.social @calccon.bsky.social · 30/11/2024What is a SemiEmpirical Theory. Second pass. Comments ? Suggestions ? 010
calccon.bsky.social @calccon.bsky.social · 28/11/2024The SETOL layer quality metric, derived from statistical mechanics, correlated perfectly with the HTSR alpha layer quality metric. Good sign 010
calccon.bsky.social @calccon.bsky.social · 27/11/2024The weightwatcher theory (SETOL) posits that the quality of a NN layer is given by a sum of the integrated R-transform R(z), over the power law tail of the ESD. When alpha=2, the Inverse Wishart model is a good model of the ESD, and the branch cut starts right at the tail. 100
calccon.bsky.social @calccon.bsky.social · 26/11/2024Breaking down silos: Researchers highlight Nobel-winning AI breakthroughs and call for interdisciplinary innovation techxplore.com/news/2024-11... #AI #Physics #NobelPrizetechxplore.com 010
calccon.bsky.social @calccon.bsky.social · 26/11/2024I recently worked up a bunch of examples of Instruction Fine-Tuned models using weightwatcher. I hope they are useful to you. weightwatcher.ai/models.html 000
calccon.bsky.social @calccon.bsky.social · 26/11/2024 Most statistical mechanics is quite simple. Mean theory plus perturbation (i.e., Replica calculation ) in a real-world system, however, the correlations are nontrivial so you have to treat them phenomenologically 010
calccon.bsky.social @calccon.bsky.social · 25/11/2024 The recent Physics and Chemistry Nobel Prizes, AI, and the convergence of knowledge fields www.cell.com/patterns/ful... 000
calccon.bsky.social @calccon.bsky.social · 24/11/2024 new, large open-weights model out of China which compares to, even beats, Llama3.1 405B. Meet Hunyuan-Large by Tencent: a 389B param MOE (52B active). It is the largest open-source Transformer-based MoE model in the industry paper: arxiv.org/pdf/2411.02265lnkd.inLinkedInThis link will take you to a page that’s not on LinkedIn 010
calccon.bsky.social @calccon.bsky.social · 23/11/2024weightwatcher.ai #LLMs #LargeLanguageModels #AI #DNN #NeuralNetworks #FineTuning #InstructionFineTuning #DeepLearning #DeepNeuralNetworks 000
calccon.bsky.social @calccon.bsky.social · 23/11/2024In our work on the HTSR Theory and the new SETOL SemiEmpirical Theory of (Deep) Learning, we have discovered that the generalization performance of NNs are governed by a kind of Universality and a related Conservation Principle. 200
calccon.bsky.social @calccon.bsky.social · 22/11/2024WeightWatcher (w|w) is an open-source, diagnostic tool for analyzing Deep Neural Networks (DNN), without needing access to training or even test data. Here are over a dozen new examples of popular Instruction Fine-Tuned models. weightwatcher.ai/models.html 010