Sign in

Leonardo Cotta

@cottascience.bsky.social
1K followers 280 following 71 posts

floptimistic from BH🔺🇧🇷 cottascience.github.io

PostsRepliesMedia
Leonardo Cotta @cottascience.bsky.social · 02/10/2026
the use of "engineering" was very precise, it's exactly what these tribes avoid doing ;)
000
Leonardo Cotta @cottascience.bsky.social · 29/09/2026
In terms of hiring, I feel like departments should finally give more weight to teaching - quality of material and lectures produced, etc. as this will be so key moving forward. Getting people that want/like to teach over paper hunters (true not only for TCS).
110
Leonardo Cotta @cottascience.bsky.social · 28/09/2026
Yep. Also, the specific bio advances claimed are bs and a PR stunt - as opposed to the maths stuff you linked here. But the original thread is a disservice, it lists the wrong reasons. The "finding" itself is bs, there's crispr-like seqs everywhere in bacteria, finding one is not relevant per se
011
Leonardo Cotta @cottascience.bsky.social · 27/09/2026
disclaimer: I'm no luddite. I feel like I can generate beautiful images and code with enough guidance, or an specific workflow. for writing? nothing seems to work. maybe the turing test should've been writing a book with good taste.
000
Leonardo Cotta @cottascience.bsky.social · 27/09/2026
maybe it's a skill issue of mine, but I still really really hate LLMs for long-form writing. I've tried many different ways: asking for a v0 and fixing myself. Writing a v0 and asking to improve, doing it piece by piece, one-shot, skills md file. everything still feels like garbage and no taste.
110
Leonardo Cotta @cottascience.bsky.social · 26/09/2026
100%. it's part of the anthropomorphization trend, where these things are treated as oracles rather than tools.
100
Leonardo Cotta @cottascience.bsky.social · 24/09/2026
scaling laws appeared as "learning curve models" long ago in the literature (90's), and finally people are going back to it. I love the chinchilla paper but that's not the whole story/method arxiv.org/abs/2509.19189
arxiv.org
Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules
Scaling laws have emerged as a unifying lens for understanding and guiding the training of large language models (LLMs). However, existing studies predominantly focus on the final-step loss, leaving o...
020
Leonardo Cotta @cottascience.bsky.social · 23/09/2026
this looks dope. I love the permutation p-value interpretation
110
Leonardo Cotta @cottascience.bsky.social · 09/09/2026
In the same way that mathematicians argue that their field has never been about the answers, but about the questions and their proofs, computer scientists should repeat Leslie Lamport's mantra: coding is not programming.
010
Leonardo Cotta @cottascience.bsky.social · 08/09/2026
ps: please do share the catalogue if possible! I think the ML community would appreciate it a lot!
110
Leonardo Cotta @cottascience.bsky.social · 08/09/2026
right, I totally believe that. It's quite often the case that it's a data problem: either small N or small D - more modalities/views are needed. Either way, I think understanding why, under what conditions, is more important than the failure itself. Eg we often just don't have data to make it work!
110
Leonardo Cotta @cottascience.bsky.social · 07/09/2026
tbf I think 90% of the time the problem is more with eval (metrics/tasks) than modelling, eg you don't even need a linear probe, random guesses (controls) are often sufficiently (many quotes) "powerful" ;)
100
Leonardo Cotta @cottascience.bsky.social · 06/09/2026
not sure I fully agree on (2). Coding develops and shares our understanding in the same way of math proofs. I agree there are many instances where (2) can be acceptable, but I find them similar to (1), i.e. non-critical code. Of course this is assuming you are one-shot'ing it, not co-developing.
010
Leonardo Cotta @cottascience.bsky.social · 05/04/2026
I'm trying to write about the history of scaling laws, and my go-to reference in the ML community is [1]. If anyone has good suggestions in asymptotic stats, I'm curious to read and help make the connections. [1] proceedings.neurips.cc/paper_files/...
proceedings.neurips.cc
Learning Curves: Asymptotic Values and Rate of Convergence
071
Leonardo Cotta @cottascience.bsky.social · 04/04/2026
I've finally deleted my twitter account, but as much as I love bksy's idea, it doesn't seem to be a good replacement. In terms of keeping up with science, linkedin and reddit have unbelievably proven more effective for me. Is there anything I'm missing here? A feed to follow, a better way to use it?
010
Leonardo Cotta @cottascience.bsky.social · 16/11/2025
I can only imagine how crazy it must be to be a PhD student submitting to ML conferences now. The process has always been noisy, but at this point it's selecting for either obfuscation or shallow ideas. You either intimidate the reviewer, or you write a blog post in latex.
020
Reposted by Leonardo Cotta
The Matter Lab @thematterlab.bsky.social · 21/08/2025
We're excited to present our latest article in Nature Machine Intelligence - Boosting the predictive power of protein representations with a corpus of text annotations. Link: www.nature.com/articles/s42... [1/4]
1125
Leonardo Cotta @cottascience.bsky.social · 12/08/2025
I’d add data/task understanding as a separate mid layer. Most papers I know break in the transition of high to mid.
110
Leonardo Cotta @cottascience.bsky.social · 09/08/2025
the goat of brazilian music w/ the best of (current) american music www.youtube.com/watch?v=jFUh...
youtube.com
Milton Nascimento & esperanza spalding: Tiny Desk (Home) Concert
YouTube video by NPR Music
020
Leonardo Cotta @cottascience.bsky.social · 31/07/2025
This is why I personally love TMLR. If it's correct and well-written let's publish. The interesting papers are the ones the community actively recognizes in their work, e.g. citing, extending, turning into products, etc. (independent process of publication).
010
Leonardo Cotta @cottascience.bsky.social · 31/07/2025
I agree with most of your thread, but classifying "uninteresting work" is quite hard nowadays. Papers became this "hype-seeking" game, where out of the 10 hyped papers of the month, at most 1 survives further investigation of the results. And even if we think we're immune to this, what is interest?
120
Leonardo Cotta @cottascience.bsky.social · 26/07/2025
I loved this new preprint by Lourie/Hu/ @kyunghyuncho.bsky.social . If you really wanna convince someone youre training a foundation model, or proposing better methodology, loss scaling laws aren't enough. It has to be tied w/ downstream performance. it shouldn't be vibes arxiv.org/abs/2507.00885
arxiv.org
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
Downstream scaling laws aim to predict task performance at larger scales from pretraining losses at smaller scales. Whether this prediction should be possible is unclear: some works demonstrate that t...
051
Leonardo Cotta @cottascience.bsky.social · 16/07/2025
We're at ICML, drop us a line if you're excited about this direction. 📄 Paper: arxiv.org/abs/2507.02083 💻 Code: github.com/h4duan/SciGym 🌍 Website: h4duan.github.io/scigym-bench... 🗂️ Dataset: huggingface.co/datasets/h4d...
010
Leonardo Cotta @cottascience.bsky.social · 16/07/2025
I'm very excited about our new work: SciGym. How can we scale scientific agents' evaluation? TLDR; Systems biologists have spent decades encoding biochemical networks (metabolic pathways, gene regulation, etc.) into machine-runnable systems. We can use these as "dry labs" to test AI agents!
120
Leonardo Cotta @cottascience.bsky.social · 30/06/2025
Also, I see ITCS more like a “out of the box”, “bold” idea or even new area, I don’t see the papers having simplicity as a goal, but just my experience.
100
Leonardo Cotta @cottascience.bsky.social · 30/06/2025
Mhm, I agree with the idealistic part, I certainly have seen the same. But I know quite a few papers that are aligned w the call, tbh this happens in any venue. I think the message and the openness to this kind of paper is important though
200
Leonardo Cotta @cottascience.bsky.social · 29/06/2025
I wish we had an ML equivalent of SOSA (Symposium On Simplicity in Algorithms). "simpler algorithms manifest a better understanding of the problem at hand; they are more likely to be implemented and trusted by practitioners; they are more easily taught" www.siam.org/conferences-....
130
Leonardo Cotta @cottascience.bsky.social · 14/06/2025
this is not my area, but if you think of it in terms of a randomized algorithm (BPP,PP), the hard part is usually the generation, at least for the algorithms we tend to design. e.g. Schwartz-Zippel Lemma. (Although in theory you can have the "hard part" in verification for any problem)
120
Leonardo Cotta @cottascience.bsky.social · 09/06/2025
It takes 1 terrible paper for knowledgeable people to stop reading all your papers, this risk is often not accounted for
110
Leonardo Cotta @cottascience.bsky.social · 08/06/2025
Maybe check Cat s22, it gives you the basics, eg whatsapp+gps and nothing else
020
Reposted by Leonardo Cotta
Quaid Morris @quaidmorris.bsky.social · 03/06/2025
Please check out our new approach to modeling somatic mutation signatures. DAMUTA has independent Damage and Misrepair signatures whose activities are more interpretable and more predictive of DNA repair defects, than COSMIC SBS signatures 🧬🖥️🧪 www.biorxiv.org/content/10.1...
biorxiv.org
Damage and Misrepair Signatures: Compact Representations of Pan-cancer Mutational Processes
Mutational signatures of single-base substitutions (SBSs) characterize somatic mutation processes which contribute to cancer development and progression. However, current mutational signatures do not ...
04117
Leonardo Cotta @cottascience.bsky.social · 31/05/2025
it just sounds like "see you three times" ;) it's like some people named "Sinho" that is often confused with portuguese/brazilians; but from what I heard it's a variation of Singh (not sure though)
110
Leonardo Cotta @cottascience.bsky.social · 18/04/2025
One simple way to reason about this: treatment assignment guarantees you have the right P(T|X). Self-selection changes P(X), a different quantity. Looking at your IPW estimator you can see that changing P(X) will bias regardless of P(T|X).
032
Leonardo Cotta @cottascience.bsky.social · 13/04/2025
I haven't been up to date with the model collapse literature, but it's crazy the amount of papers that consider the case where people only reuse data from the model distribution. This never happens, there's always some human curation or conditioning that yields some type of "real-world, new, data".
020
Leonardo Cotta @cottascience.bsky.social · 12/04/2025
this general idea of using an external world/causal model given by a human and using the LM only for inference is really cool ---it's also the insight behind our work in NATURAL. Do you guys think it's possible to write a more general software for the interface DAG->LLM_inference->estimate?
120
Leonardo Cotta @cottascience.bsky.social · 24/03/2025
This is my favourite "graph paper" of the last 1 or 2 years. We also need to start including non-NN baselines, e.g. fingerprints+catboost ---if the goal is real-world impact and not getting it published asap. I also recommend following @wpwalters.bsky.social's blog. arxiv.org/abs/2502.14546
arxiv.org
Position: Graph Learning Will Lose Relevance Due To Poor Benchmarks
While machine learning on graphs has demonstrated promise in drug design and molecular property prediction, significant benchmarking challenges hinder its further progress and relevance. Current bench...
060
Reposted by Leonardo Cotta
Derek Thompson @dkthomp.bsky.social · 27/02/2025
Unbelievable news. Pancreatic is one of the deadliest cancers. New paper shows personalized mRNA vaccines can induce durable T cells that attack pancreatic cancer, with 75% of patients cancer free at three years—far, far better than standard of care. www.nature.com/articles/s41...
13772111912
Leonardo Cotta @cottascience.bsky.social · 20/02/2025
Oh gotcha. I think it’s just super cheesy to quote feynman at this point haha but it’s a good philosophy to embrace
000
Leonardo Cotta @cottascience.bsky.social · 20/02/2025
In what contexts do you think it’s misused? Just curious, I’m a big fan and might be overusing it 😅
100
Reposted by Leonardo Cotta
Thomas Wolf @thomwolf.bsky.social · 19/02/2025
After 6+ months in the making and over a year of GPU compute, we're excited to release the "Ultra-Scale Playbook": hf.co/spaces/nanot... A book to learn all about 5D parallelism, ZeRO, CUDA kernels, how/why overlap compute & coms with theory, motivation, interactive plots and 4000+ experiments!
hf.co
The Ultra-Scale Playbook - a Hugging Face Space by nanotron
The ultimate guide to training LLM on large GPU Clusters
217952
Leonardo Cotta @cottascience.bsky.social · 19/02/2025
if you're feeling uninspired and getting nan's everywhere, you can give your codebase, describe the problem and ask for suggestions to try or debug. I think of it more as a debugger assistant than a code generator.
020
Leonardo Cotta @cottascience.bsky.social · 19/02/2025
I've always hated the "reasoning models" for code assistance since I think the most useful application of LLMs is really writing the boring helper functions and letting us focus on the hard work. However, I found o3 to be particularly useful when debugging ML code, e.g., 1/2
110
Leonardo Cotta @cottascience.bsky.social · 14/02/2025
if you remove one at a time you get reconstruction gnns 🙃 proceedings.neurips.cc/paper/2021/h...
proceedings.neurips.cc
Reconstruction for Powerful Graph Representations
010
Leonardo Cotta @cottascience.bsky.social · 14/02/2025
100%. Also, sometimes the use/task might be the same, but the user's notion of bias can vary. Eg people might expect group or individual notions of fairness.
000
Leonardo Cotta @cottascience.bsky.social · 26/01/2025
I wouldn’t call open-source democratic in this case. The models are free but the inference compute isn’t. Maybe democratic in the sense of a liberal democracy, but not in terms of accessibility. Agree w the rest though!
110
Leonardo Cotta @cottascience.bsky.social · 25/01/2025
The whole DeepSeek-R1 thing just highlights computer science's main feature: you can do A LOT with a small team and some (limited) resources. This is how we've been able to scale innovation and why free software is important.
060
Leonardo Cotta @cottascience.bsky.social · 24/01/2025
This is an amazing resource (of resources) for machine learners
020
Reposted by Leonardo Cotta
Sara Magliacane hiring PhDs in Saarland @smaglia.bsky.social · 23/01/2025
Sad after #AISTATS2025 and #ICLR2025 notifications? As we say in Italy, when a door closes, a bigger one opens ;) If you have a fantastic paper on #uncertainty #AI #ML #causality #statML #probabilisticmodels #reasoning #impreciseprobabilities etc, consider submitting to #UAI2025 🇧🇷 deadline 10 Feb 💥
24213
Reposted by Leonardo Cotta
Bruno Ribeiro (at #NeurIPS2024) @brunofmr.bsky.social · 09/01/2025
Slides of my presentation "Mathematical Foundations of Graph Foundation Models" yesterday at the AMS Session of the #JMM2025. The accompanying paper is coming soon. www.cs.purdue.edu/homes/ribeir...
042
Leonardo Cotta @cottascience.bsky.social · 30/12/2024
Learning Rust ~properly~ during my break and wow -- absolutely worth it! While we're all chasing GPU optimization, there's something magical about crafting efficient CPU-based apps. Clean and fast data processing can change our lives ;)
110