Han Bao @han-b.bsky.social · 08/10/2026I'm proud of my writing style and love scientific writing itself (and non-scientific one as well!), but not sure how much I should devote my efforts on it in this era 150
Han Bao @han-b.bsky.social · 03/10/2026I'm pretty satisfied with the current environment because I can spend a lot of time on research and discussions. This week, I had in-person discussions with another student to revise a paper everyday quite intensively. Not sure how long this style is allowed to me though 😅 080
Han Bao @han-b.bsky.social · 03/10/2026Now this is accepted by #NeurIPS2026! As being a PI, this is the first paper for me to support a student to complete. I really love the process to discuss and polish an idea and paper with a student. This owes to Xianliang's productivity a lot. 072
Reposted by Han BaoarXiv @arxiv.bsky.social · 01/10/2026arXiv has updated our policy on rate limiting for all submitters. This update was made to fairly distribute moderator time & support the arXiv community of staff, volunteers, readers & authors. Please read our announcement to learn more: blog.arxiv.org/2026/10/01/updated-r… 216874
Han Bao @han-b.bsky.social · 26/09/2026Hier je suis revenu du workshop pour des jaunes chercheurs au Japon. J'étais invité comme un conférencier principal, mais j'avais un peu honte de raconter mon parcours et leur donner des conseils😅 De toute façon c'était un bon événement. 020
Han Bao @han-b.bsky.social · 13/09/2026Last week I introduced this work in a domestic workshop among mathematicians and physicists, and one astrophysicist gave me a new insight to interpret the self-stabilizing potential! This is a real virtue of having a workshop with people from the communities next to us. 050
Han Bao @han-b.bsky.social · 12/09/2026I'm basically on that side. Mathematicians can be "interpreters" of formalized proofs, even instead of working on proofs themselves. The community has put too much emphasis to who first proves formally, but we can gradually move on to this "interpreting" role. (though this claim is nothing new😅) 110
Han Bao @han-b.bsky.social · 12/09/2026Yeah it is really a textbook example of Goodhart's law. Giving a solution to an unsolved problem is a "proxy" toward our community's understanding of the underlying mathematical structure. Solving problems is (maybe sometimes but) not often the ultimate goal. www.economist.com/science-and-...economist.comTop mathematicians are outraged by OpenAI’s methods24 Fields Medal winners have written a letter of objection 161
Reposted by Han Baomuchonov @muchonov.bsky.social · 11/09/2026それまで人件費として支払われていたコストがAIの労働代替によって海外の寡占的なテック企業にサービスフィーとして吸い上げられ、各国の企業が雇用を通じて社会に提供していた「消費者と需要を生み出す機能」が毀損されてゆくという問題について、もっと真剣に議論されるべきだと思います。 そっちのほうがLLMはAIかどうかなんて神学論争よりずっと重要。 4532319
Han Bao @han-b.bsky.social · 09/09/2026I completely resonate with your viewpoints. I would add an extra problem: all of us provide our intellectual outputs to tech giants almost for free (or even with paying subscription fee😅). This is not only economically irrational but also can be regarded as personal intrusion and exploitation. 010
Han Bao @han-b.bsky.social · 06/09/2026This assumes that the loss Hessian is (kinda) ill-conditioned, which is often reasonable in NN training. Many recent papers reported the heterogeneity of Hessian spectra for deep learning. This assumption naturally leads to [stepsize * sharpness] << 1, under which we can derive ODE. 010
Han Bao @han-b.bsky.social · 06/09/2026there. After the careful reading, I got the following conclusion. Briefly speaking, the stepsize \eta is not "dimension"-less, with the dimension of 1/sharpness, and they assume [stepsize * sharpness] to be infinitesimal, which is dimension-less now. In more detail, look at Assumption 5 (see fig). 110
Han Bao @han-b.bsky.social · 06/09/2026I spent the whole weekend to understand the crux of self-stabilization (arxiv.org/abs/2209.15594). To me, the most interesting part of this paper is how they model the edge of stability by ODE---because the stepsize is no longer infinitesimal at EoS and it is apparently not trivial to derive an ODEarxiv.orgSelf-Stabilization: The Implicit Bias of Gradient Descent at the Edge of StabilityTraditional analyses of gradient descent show that when the largest eigenvalue of the Hessian, also known as the sharpness $S(θ)$, is bounded by $2/η$, training is "stable" and the training loss decre... 140
Han Bao @han-b.bsky.social · 17/08/2026Pendant les vacances, j'ai regardé un film «Sirât» sans savoir qu'est-ce que ce film. Les paysage du désert avec des sons énormes sont magnifiques! Heureusement je pouvais le regarder à la dernière séance (il est fini au début d'août au Japon). www.imdb.com/title/tt3229...imdb.com 010
Han Bao @han-b.bsky.social · 07/08/2026Can't agree more. Previously one guy told me that if you want your theory paper to get accepted, the best thing is to write 30+ pages appendix (to overwhelm!). That's insane. Unless truly "necessary" (I know it's debatable), we should compress/distill formal proofs. Then we get better intuition. 071
Reposted by Han BaoAaron Roth @aaroth.bsky.social · 07/08/2026People used to be able to impress and intimidate reviewers with complicated proofs. This will change. In the age of AI inscrutable proofs are cheap. It is understandable proofs that are valuable. Opaque complexity is now it is a sign of laziness or lack of insight. 1638
Han Bao @han-b.bsky.social · 06/08/2026What I love with this work is the beauty of the proof. The error dynamics can be written as in the figure. This is an elementary differential inequality and solvable with undergrad mathematics. It's geometrically and analytically transparent. 020
Han Bao @han-b.bsky.social · 06/08/2026GD is known to have max-margin bias, but its convergence rate is extremely slow. However, we can observe fairly decent convergence behaviors in practice as seen in the figure. I have been thinking this for a while and figured out early-stage weak convergence is possible! arxiv.org/abs/2608.04382 182
Han Bao @han-b.bsky.social · 03/08/2026I agree with the latter. I'm not confident enough with the former; maybe the desk rejection or single-round rejection without rebuttal (which are common in journals) should be installed to conferences, but I'm not very experienced with these journal cultures. 100
Han Bao @han-b.bsky.social · 02/08/2026And the new initial meta review system is also stressful than I expected. I supposed it'd reduce rebuttal workload for both author and reviewer sides, but actually just gave us an extra burden. Very few ppl care the initial meta review and ppl do endlessly long rebuttals until the reviewers gave up🤦♂️ 061
Han Bao @han-b.bsky.social · 02/08/2026Really stressful to see AI-driven discussions during the review period this year... as an AC, I really don't have an idea how to facilitate discussion among AI-ish paper's authors and AI-ish reviewers (and the worst thing is that we cannot suppose they are truly AI even if it's very likely) 0110
Han Bao @han-b.bsky.social · 24/07/2026This VSCode plugin really changes my life recently! For some reason the default Claude plug-in shows LaTeX only in texts. marketplace.visualstudio.com/items?itemNa...marketplace.visualstudio.comClaude Code LaTeX - Visual Studio MarketplaceExtension for Visual Studio Code - Adds LaTeX math rendering to Claude Code chat — inline $...$, display $$...$$, and \(...\) / \[...\] 040
Han Bao @han-b.bsky.social · 19/07/2026Makes sense. I wouldn't ban AI usage in paper writing and review, but I don't think it's possible to submit 100% AI generated contents if we want to be responsible for what we output. Yet, I'm not sure how it'll be in a decade as human code review isn't feasible already and does not make sense... 020
Reposted by Han BaoPierre Alquier @pierrealquier.bsky.social · 18/07/2026AI is an interesting mechanism: it leads some scientists to reveal what they would not admit willingly, that is, that they actually don't care at all about scientific rigour. AI is not the problem here, it just makes the problem visible... 141
Han Bao @han-b.bsky.social · 18/07/2026way too rapid proliferation of 100% AI-generated reviews🤦♂️ 131
Han Bao @han-b.bsky.social · 10/07/2026While LLM/diffusion are extremely popular, we still can find classical topics in the poster session. This paper proposes convexified heterogeneous OT (unlike nonconvex Gromow-Wasserstein). In essence, | E[d(x,X)|y] - E[d(y,Y)|x] | is regarded as the base cost for OT. arxiv.org/abs/2606.02047arxiv.orgConvex Distance Operator Transport: A Convex and Geometry-Preserving FormulationWe introduce Convex Distance Operator Transport (CDOT), the first convex optimal transport framework that aligns distributions across heterogeneous domains by jointly preserving feature correspondence... 030
Han Bao @han-b.bsky.social · 05/07/2026I'll arrive at Seoul at Monday noon. Would love to see old friends and connect with new people there! 030
Reposted by Han BaoAaron Roth @aaroth.bsky.social · 03/07/2026On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.calibration-tutorial.github.ioCalibration, Decisions, and Collaboration in Learning | ICML 2026An ICML 2026 tutorial on making probabilistic predictions trustworthy for downstream decision-making and collaboration. 133711
Reposted by Han BaoTransactions on Machine Learning Research @tmlrorg.bsky.social · 22/06/2026TMLR has been facing an significant uptick in the number of submissions since the start of 2026. This is placing an extreme burden on our amazing team of reviewers & action editors. To ease this burden, TMLR will be implementing submission quotas, effective July 1. 1/n medium.com/@TmlrOrg/ann... 23014
Han Bao @han-b.bsky.social · 21/06/2026With my super limited knowledge (also helped by Claude): In primal - convex function (or its epigraph): X -> R - program: [input] -> [output] In dual - support function: (X -> R) -> R - CPS-ed program: [continuation (or partial evaluation)] -> [output] 010
Han Bao @han-b.bsky.social · 21/06/2026I participated a workshop among ML vs. PL (programming lang) people. For a long time I have zero understanding of CPS transform (en.wikipedia.org/wiki/Continu...), but a PL person told me it's essentially the dual transform from a convex set to a support function (!) Convex analysis everywhere...en.wikipedia.orgContinuation-passing style - Wikipedia 161
Han Bao @han-b.bsky.social · 11/06/2026It is always fantastic to read old strong papers! I read this by Koltchinskii (30 years ago) today. In this paper, he gave a convex-analytic view of the quantile function (quantile = convex conjugate supremum!), and extend it elegantly to multivariate rvs. projecteuclid.org/journals/ann... 082
Han Bao @han-b.bsky.social · 03/06/2026Here's also the projected page by Xianliang: yinleung.com/denoise-ortho/yinleung.comDenoise First, Orthogonalize Later — Understanding Momentum in MuonA theoretical and empirical study of why placing momentum before orthogonalization makes Muon work, framed as a spectral filter on the gradient stream. 010
Han Bao @han-b.bsky.social · 03/06/2026My first PhD student Xianliang worked hard out this: In Muon, polar decomposition should always precede momentum, which significantly improves signal recovery. I'm excited to share this since few theory has been working on the benefit of momentum in Muon! arxiv.org/abs/2606.03899arxiv.orgDenoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral FilteringMuon has recently demonstrated strong empirical performance in large language model training, but the theoretical role of momentum in Muon remains unclear. Existing analyses of Muon either remove mome... 1182
Han Bao @han-b.bsky.social · 30/05/2026アジア人が英語喋ると無視される、みたいなステレオタイプも若目のフランス人にはそんなに当てはまらないので、生半可なフランス語で喋ると英語で返ってきて心折れます…😂 むしろお店でもBonjourと元気よく挨拶する方が大事かも(挨拶をしないと一発で不審者扱い) 130
Han Bao @han-b.bsky.social · 26/05/2026I didn't expect he was still alive and well... huge legend🙏 010
Han Bao @han-b.bsky.social · 06/05/2026#AISTATS2026 It was extremely great to first see someone whose paper I have closely read, those with whom I collaborated recently without having met in-person, got invited to a next workshop, etc. Even though the time is quite challenging before the deadline😅, I really enjoyed the conference! 050
Reposted by Han BaoGautam Kamath @gautamkamath.com · 29/04/2026My new policy: if someone asks me to read something, I ask them how they used AI in creating it, and what validation/processing they applied to the AI outputs. (I disclose the same.) Been burned by giving too much attention to (undisclosed) slop folks have sent me... 1312
Han Bao @han-b.bsky.social · 30/04/2026D'ailleurs, j'ai pu profiter de mon séjour à Paris cette fois-ci pendant l'escale avant d'aller au Maroc. J'ai vu quelques endroits que j'aime là-bas---surtout le Centre de Pompidou, même s'il est en rénovation---aprés 7 ans! Maintenant, c'est le moment de se concentrer intensément sur le travail... 030
Han Bao @han-b.bsky.social · 30/04/2026Me: I highly look forward to exploring Morocco🤩 Reality: NeurIPS😅 050
Han Bao @han-b.bsky.social · 14/04/2026Half a year ago I got a quotation for a GPU server, and I redid it this month, and noticed that the server price almost doubled🤯 030
Han Bao @han-b.bsky.social · 13/04/2026Absolutely! I hope classes provided human still makes sense in this era😅 010
Reposted by Han BaoG. Wolfer @gwolfer.bsky.social · 13/04/2026If you meet the eligibility requirements for the LOTUS Program and are interested in working with me, feel free to reach out. www.jst.go.jp/program/indi...jst.go.jpOpen call for applications | LOTUS ProgrammeThis page provides the information regarding the open call for applications for FY2025. Applications only be accepted from Japanese organizations. 012
Han Bao @han-b.bsky.social · 13/04/2026Tomorrow is the very first class of my lecture at ISM (I'm gonna introduce learning theory and convex analysis). It's extremely useful for myself as well to review bunch of facts and proofs, but I need to rush because I've prepared for only half a semester😅 170