Flavio Calmon @fcalmon.bsky.social · 20/02/2025New paper on discretion in AI “alignment” — check out @maartenbuyl.bsky.social’s thread below! 060
Reposted by Flavio CalmonMaarten Buyl @maartenbuyl.bsky.social · 19/02/20259/n Full paper here: 🔗 arxiv.org/abs/2502.10441. Huge thanks to my amazing team of co-authors: @hadikh.bsky.social, @lucasmpaes.bsky.social, @claudiomv.bsky.social, @caiocvm.bsky.social, and @fcalmon.bsky.social. Done at @harvard.eduarxiv.orgAI Alignment at Your DiscretionIn AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are `better' or `safer.' We refer to this latitude as alignment discretion.... 031
Reposted by Flavio CalmonMaarten Buyl @maartenbuyl.bsky.social · 19/02/2025AI is built to “be helpful” or “avoid harm”, but which principles should it prioritize and when? We call this alignment discretion. As Asimov's stories show: balancing such principles for AI behavior is tricky. In fact, we find that AI has its own set of priorities. (comic by @xkcd.com)🧵👇 253
Reposted by Flavio CalmonBogdan Kulynych @bogdankulynych.bsky.social · 10/12/2024The standard practice in differential privacy of targeting ε at small δ is extremely lossy for interpreting the level of privacy protection. For many real-world algorithms (e.g., for DP-SGD), we can do much better! We show how in the #NeurIPS2024 paper: arxiv.org/abs/2407.02191 Short summary👇arxiv.orgAttack-Aware Noise Calibration for Differential PrivacyDifferential privacy (DP) is a widely used approach for mitigating privacy risks when training machine learning models on sensitive data. DP mechanisms add noise during training to limit the risk of i... 193
Reposted by Flavio CalmonBogdan Kulynych @bogdankulynych.bsky.social · 10/12/2024This is joint work with Felipe Gomez, Georgios Kaissis, @fcalmon.bsky.social, and @carmelatroncoso.bsky.social Happy to chat about it online, and in 🇨🇦+🇺🇸 next two weeks: - At the #NeurIPS2024 Friday Dec. 13 evening poster session. - Will also present in more detail on Tuesday Dec. 17 at Harvard. 112