Sign in

Csaba Szepesvari

@skiandsolve.bsky.social
1.3K followers 223 following 92 posts

⛷️ ML Theorist carving equations and mountain trails | 🚴‍♂️ Biker, Climber, Adventurer | 🧠 Reinforcement Learning: Always seeking higher peaks, steeper walls and better policies. ualberta.ca/~szepesva

PostsRepliesMedia
Reposted by Csaba Szepesvari
ARLET @arletworkshop.bsky.social · 28/07/2025
Stay tuned for updates by following us or the organizers:
Alberto Metelli, @antoine-mln.bsky.social , Dirk van der Hoeven, Felix Berkenkamp, Francesco Trovò, @gioramponi.bsky.social Marco Mussi, @skiandsolve.bsky.social , and @tillfreihaut.bsky.social
arlet-workshop.github.io
ARLET
A simple, whitespace theme for academics. Based on [*folio](https://github.com/bogoli/-folio) design.
051
Csaba Szepesvari @skiandsolve.bsky.social · 10/07/2025
..actually, not only standard notation, but also to be able to speak about the loss (=log-loss) used to train today's LLMs.
000
Csaba Szepesvari @skiandsolve.bsky.social · 10/07/2025
No, it is not information retrieval. It is deducing new things from old things. You can do this by running a blind breadth-first (unintelligent) search producing all proofs of all possible statements. Just don't want errors. But this is not retrieval. It is computation.
010
Csaba Szepesvari @skiandsolve.bsky.social · 10/07/2025
Of course approximations are useful. The paper is narrowly focused on deductive reasoning which seem to require the exactness we talk about. The point is that regardless of whether you use quantum mechanics or the Newtonian one, you don't want your derivations mistake-ridden.
010
Csaba Szepesvari @skiandsolve.bsky.social · 10/07/2025
Worst-case vs. average case: yes! But I would not necessarily connect these to minimax vs. Bayes.
000
Csaba Szepesvari @skiandsolve.bsky.social · 10/07/2025
Yeah, admittedly, not a focus point of the paper. How about if the model produces a single response, the loss is the zero-one loss. Then the model better choose the label with the highest probability label, which is OK. Point of having mu: Not much point, just matching standard notation..
100
Csaba Szepesvari @skiandsolve.bsky.social · 10/07/2025
I am curious about these examples.. (and yes, I can construct a few, too, but I want to add more)
000
Csaba Szepesvari @skiandsolve.bsky.social · 10/07/2025
No, this is not correct: Learning 1[A>B] interestingly has the same complexity (provably). This is because 1[A>B] is in the "orbit" of 1[A>=B]. So the symmetric learning who is being taught 1[A>B] need to figure out it is not taught 1[A>=B].
100
Csaba Szepesvari @skiandsolve.bsky.social · 09/07/2025
Maybe. I am asking for much less here from the machines. I am asking for them just to be correct (or stay silent). No intelligence, just good old fashioned computation.
100
Csaba Szepesvari @skiandsolve.bsky.social · 09/07/2025
the solution is found..
000
Csaba Szepesvari @skiandsolve.bsky.social · 09/07/2025
Yes, transformers do not have "working memory". Also, I don't believe in that using them in AR mode is powerful enough for challenging problems. In a way, without "working memory", external "loop", we say the model should solve problems by free association ad infinitum or at least until
110
Csaba Szepesvari @skiandsolve.bsky.social · 09/07/2025
On the paper: Interesting but indeed there is little in common. On the problem studied in the paper: Would not a slightly more general statistical framework solve your problem? Ie measure error differently than through the prediction loss (AR models: parameters, spectral measure, etc.).
000
Csaba Szepesvari @skiandsolve.bsky.social · 09/07/2025
Yeah, I don't see the exactness happening that much on its own through statistical learning. Neither experimentally, nor theoretically. We have an example for illustrating this: use the uniform distribution for good coverage, teach transformers to compare m-bit integers using GD. Need 2^m examples.
300
Csaba Szepesvari @skiandsolve.bsky.social · 09/07/2025
Yeah, we cite this and this was a paper that got me started on this project!
010
Csaba Szepesvari @skiandsolve.bsky.social · 08/07/2025
First position paper I ever wrote. "Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence" arxiv.org/abs/2506.23908 Background: I'd like LLMs to help me do math, but statistical learning seems inadequate to make this happen. What do you all think?
arxiv.org
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
Sound deductive reasoning -- the ability to derive new knowledge from existing facts and rules -- is an indisputably desirable aspect of general intelligence. Despite the major advances of AI systems ...
3519
Csaba Szepesvari @skiandsolve.bsky.social · 06/05/2025
Our seminars are back. If you missed Max's talk, it is on YouTube and today I will host Jeongyeol from UWM who will talk about the curious case of why latent MDPs though scary at first sight might be tractable! Link to the seminar homepage: sites.google.com/view/rltheor...
sites.google.com
RL theory seminars
Welcome Looking for an online seminar that presents the latest advances in reinforcement learning theory? You just found it! We aim to bring you a virtual seminar (approximately) every Tuesday at 6pm ...
0223
Csaba Szepesvari @skiandsolve.bsky.social · 04/04/2025
Glad to see someone remembers these:)
070
Csaba Szepesvari @skiandsolve.bsky.social · 20/03/2025
should be distinguished. The reason they should not is because they are indistinguishable. So at least those need to be collapsed. So yes, one can start with redundant models, where it will appear you could have epistemic uncertainty, but this is easy to rule out. 2/2
000
Csaba Szepesvari @skiandsolve.bsky.social · 20/03/2025
I guess with a worst-case hat on, we just all die:) In other words, indeed, the distinction is useful inasmuch as the modelling assumptions are valid. And there the mixture of two Diracs over 0 and 1 actually is a bad example, because that says that two models that are identical as distributions 1/x
100
Csaba Szepesvari @skiandsolve.bsky.social · 20/03/2025
I guess I stop here:) 5/5
000
Csaba Szepesvari @skiandsolve.bsky.social · 20/03/2025
Well, yes, to the degree that the model you use correctly reflects what's going on. Example with drug trials, randomized patient allocation. Result is effectiveness. Meaning of aleatoric and epistemic uncertainty should be clear and they help with explaining outcomes of the trial. 4/x
100
Csaba Szepesvari @skiandsolve.bsky.social · 20/03/2025
One observes 1, there is epistemic uncertainty (the model could be the first or the second). Of course, nothing is black and white like this ever. And we talk about models here. Models are.. made up.. Usual blurb about usefulness of models. Should you care about this distinction? 3/x
100
Csaba Szepesvari @skiandsolve.bsky.social · 20/03/2025
Epistemic uncertainty refers to whether given the data (and prior information), we can surely identify the data generating model. Example: Model class has two distributions; one has support {0,1}, the other has support {1}. One observes 0. There is no epistemic uncertainty. 2/X
100
Csaba Szepesvari @skiandsolve.bsky.social · 20/03/2025
I don't get this: In the context of this terminology, data comes from a model. Aleatoric uncertainty refers to the case when this model is a Dirac! In the second case, the model is a mixture of two Dirac's. This is not a Dirac. Hence, there is aleatoric uncertainty. 1/X
100
Reposted by Csaba Szepesvari
Prof Ben Britton @bmatb.expmicromech.com · 15/03/2025
This is a very significant development - more fellowships, harmonized and typically higher stipends, and international students can apply #CanPoli www.nserc-crsng.gc.ca/NewsDetail-D...
nserc-crsng.gc.ca
NSERC - Latest News - Launch of the new Harmonized Tri-agency Scholarship and Fellowship programs
As announced in Budget 2024, the scholarship and fellowship programs administered by the three federal research funding agencies – the Canadian Institutes of Health Research (CIHR), the Natural Sciences and Engineering Research Council (NSERC), and the Social Sciences and Humanities Research Council (SSHRC) – have been streamlined into a new harmonized talent program called the Canada Research Training Awards Suite (CRTAS) that will open for applications in summer 2025.
23716
Reposted by Csaba Szepesvari
arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 11/03/2025
Dylan J. Foster, Zakaria Mhammedi, Dhruv Rohatgi: Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration arxiv.org/abs/2503.07453 arxiv.org/pdf/2503.07453 arxiv.org/html/2503.07453
155
Csaba Szepesvari @skiandsolve.bsky.social · 11/03/2025
But also we are how we act! So it's up to us all to behave so as to make statement true.
000
Csaba Szepesvari @skiandsolve.bsky.social · 09/03/2025
Who says mountain car is a toy problem? www.reddit.com/r/nonononoye...
reddit.com
From the nonononoyes community on Reddit: He was there for a while
Explore this post and more from the nonononoyes community
070
Csaba Szepesvari @skiandsolve.bsky.social · 07/03/2025
Yes, another gem from Rich!
010
Csaba Szepesvari @skiandsolve.bsky.social · 06/03/2025
www.youtube.com/watch?v=9_Pe... An interview with Rich. The humility of Rich is truly inspiring: "There are no authorities in science". I wish people would listen and live by this.
youtube.com
TURING AWARD WINNER Richard S. Sutton in Conversation with Cam Linke | No Authorities in Science
YouTube video by Amii
24013
Csaba Szepesvari @skiandsolve.bsky.social · 06/03/2025
That's all good: Bubbles join when they get up high into the blue sky:)
010
Csaba Szepesvari @skiandsolve.bsky.social · 06/03/2025
LOL, the dude in the long white labcoat missed my poster at NeurIPS, how could he do that? LOL
110
Csaba Szepesvari @skiandsolve.bsky.social · 23/02/2025
New journal that looks pretty good for your mathy type people caring about ML, AI and alike: mdlijournal.org
mdlijournal.org
Mathematics of Data, Learning, and Intelligence
070
Reposted by Csaba Szepesvari
Ryan Williams @rrwilliams.bsky.social · 21/02/2025
New paper: Simulating Time With Square-Root Space people.csail.mit.edu/rrw/time-vs-... It's still hard for me to believe it myself, but I seem to have shown that TIME[t] is contained in SPACE[sqrt{t log t}]. To appear in STOC. Comments are very welcome!
people.csail.mit.edu
1726375
Csaba Szepesvari @skiandsolve.bsky.social · 15/02/2025
Oh, it turns out you have to wait a little longer for *RLDM* to come to Edmonton! But!!! RLC will happen to be in Edmonton this year, so all is good, you RL folks are all covered!
060
Csaba Szepesvari @skiandsolve.bsky.social · 12/02/2025
Yet another opportunity to visit Edmonton in the summer. If you come, ask me to take you on the river valley trails, or to go climbing together😄
370
Csaba Szepesvari @skiandsolve.bsky.social · 12/02/2025
I also sense that you may be thinking that uniform error bounds are meant for error estimation. My sense is that they were meant for what they are, answering specific learnability questions. The literature on cross validation estimation is interesting though..
000
Csaba Szepesvari @skiandsolve.bsky.social · 12/02/2025
Fine, I bite: Maybe because classical results sufficed (chapter 8 in the classic text of Devroye, Györfi and Lugosi) and there is nothing much else to add (except for the nasty cases like policy evaluation on batch data where there is no lack of engagement).
100
Reposted by Csaba Szepesvari
Miro Dudik @mdudik.bsky.social · 16/01/2025
📣My team at Microsoft Research New York is hiring a senior researcher in AI, both broadly in AI/ML, and in some specific areas including science of deep learning and modular transfer learning. Apply by February 7, 2025 on the link below. jobs.careers.microsoft.com/global/en/jo...
jobs.careers.microsoft.com
Search Jobs | Microsoft Careers
0238
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
..whether the interval does indeed contain the unknown parameter.
000
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
Maybe just say that the theoretical guarantees are concerned with how frequently the random effects will produce and incorrect interval and in particular they don't guarantee that the interval will be correct every time or that the user will be able to tell given the data 1/2
110
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
From the associated blog post: "The theoretical guarantees only hold before data is collected." I don't like sentences like this. Guarantees hold regardless of what we do. @beenwrekt.bsky.social, just realized this is your stuff. Am I just too grumpy today? I certainly appreciate the text otherwise
110
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
We will get disinterested. At least some part of the community. You don't do math just for the sake of knowing what's true or false but you do it for the sake of knowing why..
130
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
I'm my book that wouldn't be a randomizing algorithm. Classic stat texts talk about this randomization to get exact coverage but imho that's a bit of 'smoke in the eyes' thing.
020
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
Yes, it seems like a good idea to first consider the problem when you not only design the interval but you also want the procedure that tells you where the interval contains the unknown parameter. And then show this can only be solved in uninteresting ways. Can't turn uncertainty to certainty.
010
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
So they would want to have not only an interval but an indicator which with no exceptions tells you where the interval contains the data. And that you can't have that is hard to swallow. 2/2
120
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
I guess I get it know: I guess what the author means is that people are unhappy about that you never know whether your confidence interval on the data you use it with actually contains the true parameter or not. 1/2
120
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
Also, I wonder what could be a non ex ante guarantee for a problem like confidence interval design.. (wondering about the word "only" in the excerpt).
210
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
Specifically, randomization may appear to help with getting some tradeoff exactly right, but the problem is that randomization can be misused badly; the criteria based on expectations allows such misuse. And imho this is important.
110
Csaba Szepesvari @skiandsolve.bsky.social · 13/01/2025
Maybe a better wording would have been: "even if we have no means of verifying success, ..." And I still don't know why we should be excited about randomization in interval design other than to address some weird optimality criteria related to lack of convexity.
110