Sign in

Sergei Vassilvitskii

@vsergei.bsky.social
288 followers 135 following 4 posts

Algorithms, predictions, privacy. theory.stanford.edu/~sergei

PostsRepliesMedia
Sergei Vassilvitskii @vsergei.bsky.social · 14/02/2025
These two assumptions are enough to prove convergence bounds! More generally, this view presents a theoretical framework that unifies existing elements of synthetic data approaches, facilitating reasoning about when they might succeed or fail.
000
Sergei Vassilvitskii @vsergei.bsky.social · 14/02/2025
Instead of a weak learner, we assume access to models that can perfectly model an input distribution, which we call strong learners. But instead of iid samples, we have access to only weak information about the target distribution, i.e. weak data.
100
Sergei Vassilvitskii @vsergei.bsky.social · 14/02/2025
Synthetic Data is all the rage in LLM training, but why does it work? In arxiv.org/abs/2502.08924 we show how to analyze this question through the lens of boosting. Unlike boosting, however, our assumptions on the data and the learning method are inverted.
arxiv.org
Escaping Collapse: The Strength of Weak Data for Large Language Model Training
Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown that without proper...
183
Sergei Vassilvitskii @vsergei.bsky.social · 21/12/2024
Shameless plug on dp synthetic data: research.google/blog/protect...
research.google
Protecting users with differentially private synthetic training data
020