Sign in

Differential Privacy Papers

@dppapers.bsky.social
558 followers 0 following 1.7K posts

🤖 new arXiv preprints mentioning "differential privacy" or "differentially private" in the title/abstract - unrelated quantum/FL papers + updates from differentialprivacy.org [Under construction.]

PostsRepliesMedia
Differential Privacy Papers @dppapers.bsky.social · 8h
BRACE: Differential Privacy for Dense Associative Memory with LSR Energy Chang Qu, Zhaoyang Shi arxiv.org/abs/2610.11218
BRACE: Differential Privacy for Dense Associative Memory with LSR Energy

Chang Qu, Zhaoyang Shi

http://arxiv.org/abs/2610.11218

Dense associative memory (DAM) provides an energy-based framework for memory retrieval with close connections to attention mechanisms in modern artificial intelligence. Despite growing interest in differential privacy for AI, the privacy of DAM retrieval dynamics remains relatively unexplored. In this paper, we develop a differential privacy framework for log-sum-ReLU (LSR) dense associative memory, whose finite-support retrieval dynamics pose distinctive challenges for privacy-preserving computation. We propose the Boundary-Responsive Adaptive Correction Evolution (BRACE) algorithm, a differentially private retrieval mechanism for LSR-DAM that adaptively corrects boundary-sensitive perturbations to control their cumulative effect over the retrieval trajectory. In theory, we prove that our method is minimax optimal by deriving dimension-independent terminal and full-trajectory retrieval error rates, with optimal dependence on the inverse temperature and, in the growing-horizon regime, the retrieval horizon. We further establish central limit theorems that enable uncertainty quantification for private retrieval by characterizing its asymptotic distribution and the additional variability introduced by privacy. Numerical experiments compare our proposed method with baseline differential privacy approaches and evaluate its retrieval accuracy. Together, our results provide a theoretical foundation for optimal privacy-preserving retrieval and uncertainty quantification in energy-based associative memory systems.
000
Differential Privacy Papers @dppapers.bsky.social · 8h
Soft Voting for Policy-Aware Private Data Synthesis Yingge Hu, Gautham Ramesh Babu, Mostafa Milani arxiv.org/abs/2610.11285
Soft Voting for Policy-Aware Private Data Synthesis

Yingge Hu, Gautham Ramesh Babu, Mostafa Milani

http://arxiv.org/abs/2610.11285

Blowfish privacy relaxes differential privacy (DP) by protecting only the attribute-value substitutions a data owner specifies as edges of a policy graph. A sparser policy can reduce the noise required by a mechanism, but only when the released statistic changes less across protected substitutions than across arbitrary DP neighbors. We study this question for evolutionary, nearest-neighbor DP synthesizers such as Private Evolution (PE) and its tabular instantiation Tab-PE, which score private records against a candidate population and release a noisy vote histogram. Their hard vote is constant inside each candidate's decision region and jumps at its boundary. Its policy-specific sensitivity therefore equals the full worst-case value whenever at least one protected substitution crosses a boundary, regardless of how short that substitution is. Because every round we examined contained such a substitution, the policy graph gave no reduction in noise. We propose BF-Soft, a temperature-smoothed soft vote whose response changes gradually with distance. Its sensitivity has a tight closed-form bound in the policy graph's reach and the temperature, independent of the number of candidates, and the bound can be computed once before synthesis. It also predicts from the policy alone when policy-aware smoothing cannot substantially reduce noise: protecting a flat categorical or binary attribute drives the reach to its maximum. On real and synthetic datasets under narrow numeric policies, BF-Soft reduces error relative to hard voting at strong privacy budgets, while the advantage reverses at weaker budgets. A public-data pilot predicts when soft voting is beneficial without spending private budget.
010
Differential Privacy Papers @dppapers.bsky.social · 8h
Minimax Gaussian Mechanisms for Continual Machine Unlearning Qi Kuang, Yin Xia arxiv.org/abs/2610.11628
Minimax Gaussian Mechanisms for Continual Machine Unlearning

Qi Kuang, Yin Xia

http://arxiv.org/abs/2610.11628

Machine unlearning updates a trained model after records are deleted, aiming to match exact retraining without repeating the full training procedure. We develop Gaussian mechanisms for Newton updates under sequential deletion requests. Using Gaussian differential privacy (GDP) and its adaptive composition rule, we show that the full sequence of released models is statistically difficult to distinguish from matched exact retraining. To calibrate these mechanisms for empirical risk minimization, we derive upper bounds on the error of the Newton approximation relative to exact retraining and on how this error changes after each deletion batch. Independent Gaussian noise is calibrated using bounds on the full residual at each release, whereas Gaussian random walk noise uses smaller bounds on residual increments. These bounds yield allocations minimizing the worst-case maximum noise variance across releases under the resulting GDP certification constraints. With count-based bounds, the random walk asymptotically matches the worst-case variance of a single release at deletion cap $M$, while independent noise incurs an additional factor of order $M$. Set-based bounds can reduce the noise variances by using gradients and Hessians of the deleted records. For singleton deletion, we further show that count-based independent noise, count-based random walk noise, and set-based independent noise are minimax among fixed Gaussian covariances under their respective residual or increment bounds. With set-based bounds, allowing variances to adapt to deleted records can improve on every fixed covariance by a factor of order $(\log M)^2$ on some data sequences. The residual and noise bounds also yield parameter and predictive consistency relative to exact retraining, uniformly over deletion policies. Simulations and a credit default data analysis evaluate bounds, noise varia
010
Differential Privacy Papers @dppapers.bsky.social · 8h
Local Sensitivity in Exponential Selection: Failure Modes and Valid Calibrations Dung Nguyen, Anil Vullikanti arxiv.org/abs/2610.11870
Local Sensitivity in Exponential Selection: Failure Modes and Valid Calibrations

Dung Nguyen, Anil Vullikanti

http://arxiv.org/abs/2610.11870

Selection is a task that chooses one element from a finite public candidate range to maximize a data-dependent score. In differential privacy (DP), the exponential mechanism (EM) samples a candidate at a temperature calibrated to the global sensitivity. In this paper, we study when dataset-dependent sensitivity can safely replace global sensitivity in private selection.
  We propose three valid approaches. First, a private, high-probability upper bound on local sensitivity yields approximate DP, and the method extends to finite higher-order sensitivity hierarchies. Second, our Propose-Test-Release (PTR) variant privately searches a finite public grid for a temperature scale rather than fixing it in advance. Third, smooth sensitivity supports several designs. A candidate-independent smooth geometric construction produces a sensitivity envelope that is admissible under the local dampening framework, which privacy is guaranteed for any admissible envelope. Additionally, a separate logarithmic transformation utilizes smooth sensitivity to produce a smoothed candidate score function with advantages: having controlled global sensitivity, and preserving the maximizers of the original utility score, i.e., candidates maximizing the utility. Both of the designs yield range-independent pure DP. Besides that, we also give two approximate DP private selectors using smooth sensitivity: a direct EM with smooth sensitivity calibrated to the candidate range and privacy parameters that matches a theoretical lower bound up to some constant factor, and one using a privatized smooth upper scale by analyzing the logarithmic transform of the smoothness. For every proposed mechanism, we derive a high-probability regret bound under its stated conditions.
000
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
Reward-Driven Learning under Prompt-Level Differential Privacy Jiachen Zhao, Antonia Januszewicz, Taeho Jung arxiv.org/abs/2610.07212
Reward-Driven Learning under Prompt-Level Differential Privacy

Jiachen Zhao, Antonia Januszewicz, Taeho Jung

http://arxiv.org/abs/2610.07212

Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be (ε,δ)-differentially private with respect to the presence of any one training problem. Taking the group of responses to one prompt as the privacy record, our method aggregates their gradients, clips the prompt's contribution once, adds Gaussian noise, and composes the privacy loss across updates, so the budget depends on neither the number of responses per prompt nor the clipping norm; to our knowledge this is the first differential privacy guarantee for RLVR training. We train Qwen2.5-1.5B-Instruct with LoRA at a per-run budget of ε=8 and compare, on the same prompts and at the same budget, a control that removes only the reward signal and two private supervised fine-tuning recipes. The reward signal improves accuracy over the control by 2.65 points on MATH and 3.24 on GSM8K, in every seed; the improvement survives a format-robust scorer, at 1.3 points on MATH, and is not explained by response length. At the same budget the private model outperforms both supervised recipes on MATH and GSM8K by 2.3 to 3.8 points, retains 85--90% of the gain of non-private GRPO on these tasks, and on MATH the noise of an eightfold tighter budget costs at most 1.2 points. The reward effect also carries to CommonsenseQA, an exploratory non-mathematical task. Verifier feedback thus remains a usable learning signal under prompt-level privacy.
000
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
The Cost of Differential Privacy in Linear-Quadratic Dynamic Games Chih-Yuan Chiu, Matthew Hale arxiv.org/abs/2610.07238
The Cost of Differential Privacy in Linear-Quadratic Dynamic Games

Chih-Yuan Chiu, Matthew Hale

http://arxiv.org/abs/2610.07238

Multi-agent coordination often requires strategic agents to share sensitive information about their states or objectives, creating a tension between performance and privacy. Our paper studies this tradeoff in stochastic linear-quadratic (LQ) dynamic games with heterogeneous agent objectives. In our framework, agents share noise-perturbed state and reference information with a cloud computer that computes feedback Nash equilibrium strategies, with the injected noise calibrated to provide differential privacy. We derive an analytical expression for each agent's infinite-horizon steady-state cost of privacy relative to the non-private game. Then, we prove that when agents' objectives are sufficiently aligned, the injection of privacy noise necessarily incurs a positive performance cost. In contrast, we characterize a class of games with sufficiently misaligned objectives across agents for which an agent's cost of privacy can be strictly negative. Thus, counterintuitively, noise can simultaneously protect privacy and improve the equilibrium performance of an agent when the objectives of interacting agents are sufficiently misaligned. Finally, we present numerical experiments which corroborate our theoretical contributions.
000
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
Sparse Kernel Mechanisms for Locally Differentially Private Discrete Channels Amirreza Zamani, Parastoo Sadeghi, Mikael Skoglund arxiv.org/abs/2610.07993
Sparse Kernel Mechanisms for Locally Differentially Private Discrete Channels

Amirreza Zamani, Parastoo Sadeghi, Mikael Skoglund

http://arxiv.org/abs/2610.07993

We study sparse locally private discrete mechanisms generated by a nonnegative kernel and an input-dependent admissible output support. The support set is intentionally small relative to the ambient alphabet and acts as a mechanism-design parameter. This formulation covers metric-ball supports, sparse exponential mechanisms, sparse staircase mechanisms, sparse randomized response, sparse additive kernels such as Skellam perturbations, and Poisson-binomial-inspired bounded-support channels. We give a general exact characterization of pure and approximate local differential privacy for this sparse-kernel class. Pure local differential privacy forces all supports to coincide, so genuinely input-dependent sparse supports are incompatible with pure privacy. In the approximate regime, the privacy defect decomposes exactly into support leakage and overlap excess loss. We instantiate the formula for several discrete mechanism families. For sparse staircase mechanisms, a shell-ratio condition eliminates overlap excess loss. For graph-metric supports, overlap of metric balls is necessary for nontrivial privacy. For sparse randomized response, the support-leakage term has a closed form in terms of support mismatch. For generic radius-truncated additive kernels and, in particular, sparse Skellam mechanisms, we obtain exact finite-support formulas in terms of the Skellam CDF and an exact Bessel-ratio overlap condition. Numerical evaluations illustrate how support radius and kernel shape separately affect support leakage and overlap excess loss.
000
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
One-Shot Private Confidence Regions via Resampling Shourya Pandey, Purnamrita Sarkar, Po-Ling Loh, Debepsita Mukherjee arxiv.org/abs/2610.08460
One-Shot Private Confidence Regions via Resampling

Shourya Pandey, Purnamrita Sarkar, Po-Ling Loh, Debepsita Mukherjee

http://arxiv.org/abs/2610.08460

We propose a simple framework for constructing differentially private confidence regions \textit{in one shot}, i.e., by adding noise only to the final resampling quantile instead of privatizing the estimator computed on each resample. The cost of privacy of our procedure is only logarithmic in the number of resamples $B$ under with-replacement ($m$-out-of-$n$) sampling and independent of $B$ under without replacement sampling (subsampling), avoiding the $\sqrt{B}$ factor that arises in previous works. We provide nonasymptotic Gaussian Differential Privacy (GDP) and utility guarantees for both subsampling and $m$-out-of-$n$ resampling, covering mean-like estimators with small global sensitivity as well as estimators admitting efficiently computable smooth sensitivity bounds, including quantiles and degenerate U-statistics. This allows us to also obtain private confidence regions for degenerate U-statistics where the private error is much smaller than the non-private error. In all, we provide a toolbox for widely applicable DP uncertainty quantification procedures under popular resampling strategies while avoiding the computational and privacy costs of privatizing many intermediate resample statistics.
000
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
LDPGraph: Locally Differentially Private Graph Synthesis by Exploiting Neighborhood Structure Jiawei Dong, Zhikun Zhang, Quan Yuan, Zhe Liu, Yunjun Gao arxiv.org/abs/2610.09642
LDPGraph: Locally Differentially Private Graph Synthesis by Exploiting Neighborhood Structure

Jiawei Dong, Zhikun Zhang, Quan Yuan, Zhe Liu, Yunjun Gao

http://arxiv.org/abs/2610.09642

The widespread application of graph data inevitably brings significant privacy risks, as its unprotected use can lead to the leakage of sensitive information. These risks are particularly acute in the setting with an untrusted curator, where the data remains decentralized and each user only holds the connections to their neighbors. To mitigate such privacy risks, we adopt local differential privacy (LDP) to collect users' private information and generate a synthetic graph. However, existing methods suffer from either excessive noise injection by perturbing the local adjacency lists or significant structural information loss due to the simplistic graph encoding process. To address these issues, we propose LDPGraph, an effective graph synthesis algorithm that takes one step further by exploiting neighborhood structures under LDP. To obtain neighborhood statistics beyond degrees, LDPGraph aggregates a noisy global view from perturbed adjacency lists and combines it with projected local connections to estimate node-level triangle counts. To correct the structural inconsistency caused by separately perturbing degree and triangle count, LDPGraph jointly refines them into feasible degree-triangle targets. To reconstruct a global graph from these estimated targets, LDPGraph adopts a triangle-first strategy that first preserves local clustering structures and then fulfills remaining degree requirements. Extensive experiments on four real-world datasets and multiple commonly used graph metrics validate the superiority of LDPGraph. Source code is available at https://github.com/ZJU-TrustAID/LDPGraph.
000
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
Quasi-Binarized Autoencoders: An Architecture-Independent Information Bottleneck for Medical Image Anomaly Detection Shouhei Hanaoka, Takahiro Nakao, Atsushi Takamatsu, Takeharu Yoshikawa, Osamu Abe arxiv.org/abs/2610.09670
Quasi-Binarized Autoencoders: An Architecture-Independent Information Bottleneck for Medical Image Anomaly Detection

Shouhei Hanaoka, Takahiro Nakao, Atsushi Takamatsu, Takeharu Yoshikawa, Osamu Abe

http://arxiv.org/abs/2610.09670

Unsupervised anomaly detection, which learns only from normal images, is a central task in medical image analysis and remains an open problem. Reconstruction-based methods pass an image through an encoder-decoder network trained on normal data and detect anomalies from the residual between the image and its reconstruction. This works only if the information passed from the encoder to the decoder is limited; otherwise the network learns an identity mapping and reconstructs anomalies too. This limit is usually imposed through architectural choices, tuned per dataset, that cannot be stated in bits. We introduce the quasi-binarizing (QB) layer, which squashes each latent element into [0, 1] and adds Laplace noise of scale 1/epsilon. Each element is then epsilon-locally differentially private, and the mutual information between an image and its reconstruction is bounded by a quantity that depends only on epsilon and the number of QB elements, whatever the encoder and decoder. Placing a QB layer on every encoder-decoder path, including all skip connections, we build QBAE, a seven-level attention U-Net with 32,768 QB elements. On the seven datasets of the MedIAnomaly benchmark, QBAE with one architecture and one configuration reaches a mean image-level AUROC of 0.828, the highest among methods that do not adapt to each dataset, and the best reported results on BraTS2021 (AUROC 0.911, pixel-level AP 0.838). The noise is kept at test time, so that every reconstruction satisfies the bound. Without input corruption, the bottleneck alone prevents identity collapse (mean AUROC 0.805 vs. 0.590). Code is available at https://github.com/hanaokalog/MedIAnomalyQB.
000
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
Efficient Provably Private Classification with a Tabular Foundation Model Talal Alrawajfeh, Cristiana Diaconu, Ossi Räisä, Sebastian Rodriguez Beltran, Yuan He, John Bronskill, Richard E. Turner, Antti Honkela arxiv.org/abs/2610.10068
Efficient Provably Private Classification with a Tabular Foundation Model

Talal Alrawajfeh, Cristiana Diaconu, Ossi Räisä, Sebastian Rodriguez Beltran, Yuan He, John Bronskill, Richard E. Turner, Antti Honkela

http://arxiv.org/abs/2610.10068

Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specific optimisation, suffers substantial utility loss under strong privacy, and is often difficult to apply correctly. Tabular foundation models adapt rapidly to new datasets, but existing models lack formal privacy guarantees, and are highly vulnerable to membership-inference attacks, limiting their use on sensitive data. Here we introduce PrivTab, an easy to use tabular foundation model for differentially private classification that embeds a privacy mechanism within its architecture. Pretrained on simulated datasets, PrivTab uses in-context learning to transform sensitive rows into compact, provably private summaries---effectively learning how to learn under privacy. PrivTab outperforms private linear and neural-network baselines under moderate-to-strong privacy, shows negligible membership leakage, maintains well-calibrated predictions under strong privacy, and reduces dataset fitting time by 10,000 times, requiring only a single forward pass. By combining formal privacy, speed, and easy of use, PrivTab brings recent advances in AI to applications where sensitive individual-level data have limited their adoption.
000
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
BetweenCut: Private Heavy-Node Classification with Doubly Logarithmic Error in Tree Height Ergute Bao, Graham Cormode, Xiaokui Xiao, Ting Yu arxiv.org/abs/2610.10075
BetweenCut: Private Heavy-Node Classification with Doubly Logarithmic Error in Tree Height

Ergute Bao, Graham Cormode, Xiaokui Xiao, Ting Yu

http://arxiv.org/abs/2610.10075

Finding heavy nodes in a tree---those whose counts exceed a given threshold---is a building block for analysis and learning over structured data. Achieving record-level differential privacy (DP) without sacrificing accuracy is challenging because each record contributes to counts along an entire root-to-leaf path, allowing privacy costs to accumulate across levels. Existing methods account for the multiple threshold comparisons for each record incur additive error margins of $Ω_{\varepsilon,δ}(\log h)$ or $Ω_{\varepsilon,δ}(\sqrt{\log h})$ for tree height $h$. We introduce \textsc{BetweenCut}, an $(\varepsilon,δ)$-DP algorithm with an additive error margin of $O_{\varepsilon,δ}(\log\log h)$, improving the existing bounds for deep trees. This error holds simultaneously for all nodes and is independent of the input database size.
010
Differential Privacy Papers @dppapers.bsky.social · 08/10/2026
Evaluating Sequence Assembly Strategies for Differentially Private Synthetic Time-Series Forecasting Guoxiong Long, Huizhen Huang, Qikun Cai, Tao Huang, Chen Hou arxiv.org/abs/2610.10222
Evaluating Sequence Assembly Strategies for Differentially Private Synthetic Time-Series Forecasting

Guoxiong Long, Huizhen Huang, Qikun Cai, Tao Huang, Chen Hou

http://arxiv.org/abs/2610.10222

Differentially private time-series generators commonly produce fixed-length synthetic windows, whereas downstream forecasting models often require long continuous training sequences. How these windows are assembled after generation can therefore alter the effective synthetic data presented to a forecaster, even when the trained generator remains unchanged. We study this post-generation sequence assembly process by systematically varying overlap rates and window-weighting schemes and evaluating the resulting sequences in terms of boundary continuity, statistical and temporal fidelity, and Train-on-Synthetic-Test-on-Real (TSTR) forecasting utility. Across four types of public datasets (ETTh1, ETTm1, Weather, and Appliances) and five forecasting models, the results reveal a clear forecaster-dependent assembly principle: downstream TSTR utility is jointly shaped by the forecaster, overlap rate, and window-weighting scheme, leading to distinct assembly preferences across forecasting models. Increased overlap generally improves boundary continuity, but improvements in continuity or individual fidelity diagnostics do not consistently reduce forecasting error, indicating that these diagnostics alone are insufficient for selecting assembly configurations. Complete five-forecaster assembly grids, together with matched Train-on-Real-Test-on-Real (TRTR) references, further characterize these regularities and quantify assembly-dependent utility relative to real-data training. We then validate the identified principles through additional analyses of robustness and generator variability.
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
Protecting Sensitive Data in Image Synthesis via PAC-Private Adaptation for Diffusion Models Boming Miao, Tao Zhang, Netanel Raviv, Murat Kantarcioglu, Bradley A. Malin, Yevgeniy Vorobeychik arxiv.org/abs/2610.04038
Protecting Sensitive Data in Image Synthesis via PAC-Private Adaptation for Diffusion Models

Boming Miao, Tao Zhang, Netanel Raviv, Murat Kantarcioglu, Bradley A. Malin, Yevgeniy Vorobeychik

http://arxiv.org/abs/2610.04038

Synthetic data are increasingly used as an alternative to sharing sensitive records. However, synthetic data generation does not guarantee privacy, as diffusion models trained or adapted on sensitive data remain susceptible to reconstruction attacks. Moreover, while approaches that use differential privacy (DP), such as DP-SGD, achieve provably private diffusion model training, the repeated gradient clipping and noise injection they require result in significant utility loss. An important limitation of DP-based privacy is that, although it has a provable relationship to reconstruction privacy (RP), that relationship is indirect. RP is defined in terms of limiting how much an adversary's posterior distribution over sensitive data differs from the prior, whereas DP provides guarantees by bounding the sensitivity of outputs to changes in individual records. This indirection is an important source of the utility loss. To address this, we propose a PAC-private diffusion model adaptation to achieve reconstruction privacy. Since PAC-privacy is defined directly with respect to posterior advantage over the prior, it directly implicates RP. To obtain scalable PAC privatization in high dimensions, we first learn a compact data-dependent diffusion model component using LoRA or Textual Inversion, and then calibrate anisotropic Gaussian noise from the covariance of repeated mechanism outputs. Unlike DP-SGD, our method perturbs the learned component only once after optimization, thereby avoiding privacy composition across gradient updates. We evaluate the framework on few-shot concept personalization and full-dataset image synthesis, and show that the proposed approach better preserves subject identity, generation quality, and downstream classification accura
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
Asymptotically Optimal Best Arm Identification with Fixed-Budget under Differential Privacy Keqin Chen, Jie Bian, Yulian Wu, Vincent Y. F. Tan arxiv.org/abs/2610.04600
Asymptotically Optimal Best Arm Identification with Fixed-Budget under Differential Privacy

Keqin Chen, Jie Bian, Yulian Wu, Vincent Y. F. Tan

http://arxiv.org/abs/2610.04600

Best arm identification under differential privacy is a pure-exploration problem in which both statistical efficiency and privacy protection must be achieved simultaneously. We study fixed-budget best arm identification for bandits under pure $ε$-differential privacy, where the learner must recommend an arm after a prescribed sampling budget while protecting the full transcript. We prove that the optimal exponential decay rate of the error probability is upper bounded by an instance-dependent privacy-aware transportation exponent that differs from the analogous quantity used to characterize the stopping time in fixed-confidence analysis by Jourdan and Azize [2025]. Guided by this exponent, we propose AO-Pri-BAI, an adaptive algorithm that maintains private running estimates through Laplace-tree mechanisms and learns a sampling design through a min--max interaction between hard alternatives and arm allocations. We prove that AO-Pri-BAI satisfies pure $ε$-differential privacy. We also establish that the exponent of the failure probability of AO-Pri-BAI matches the privacy-aware benchmark. Numerical studies show that even in the non-asymptotic setting, AO-Pri-BAI outperforms benchmark algorithms on various instances, complementing the theoretical analyses.
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
Communication-Efficient Differentially Private Gradient Tracking for Distributed Optimization via Local Updates Mihitha Maithripala, Chenyang Qiu, Zongli Lin arxiv.org/abs/2610.04854
Communication-Efficient Differentially Private Gradient Tracking for Distributed Optimization via Local Updates

Mihitha Maithripala, Chenyang Qiu, Zongli Lin

http://arxiv.org/abs/2610.04854

This paper studies privacy-preserving distributed optimization using gradient tracking with one communication-free local update between consecutive communication rounds. Since local computation alone does not protect gradient information, we perturb both the primal state and the gradient-tracking direction at communication iterations while keeping local updates noise-free. We establish an aggregate variable for tracking, characterize the limiting consensus point in the presence of added random noise, and prove convergence in mean under certain parameter conditions. We also establish infinite-horizon pure differential privacy for the complete communication transcript under an affine objective adjacency relation and geometrically decaying Laplace perturbations. Simulation results illustrate the advantages of the proposed method in terms of the convergence and optimization accuracy given a privacy budget.
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
Private Component-by-Component Learning Dvir Karni, Eliad Tsfadia arxiv.org/abs/2610.05102
Private Component-by-Component Learning

Dvir Karni, Eliad Tsfadia

http://arxiv.org/abs/2610.05102

We study differentially private learning problems in the realizable setting, where a hypothesis is specified by $k$ components. A direct iteration of private component learners is obstructed by a simple difficulty: an approximate choice of the next component may destroy exact realizability of the labeled sample, even when the next component is locally accurate. We restore realizability using the LabelBoost procedure of Beimel, Nissim, and Stemmer [SODA '15, Algorithmica '21] and recycle data through two alternating reservoirs. The resulting learner, for a target privacy $\varepsilon$, pays only $\widetilde O(\sqrt{k}/\varepsilon)$ overhead relative to the active sample requirement of a single component learning step at target accuracy $Θ(α/k)$. For learning $d$-dimensional halfspaces over a finite coordinate grid of size $L$, exact realizability makes the direct component-depth objective quasi-concave. Instantiating the framework with the IPConcave algorithm of Nissim, Tsfadia, and Yan [SODA '26] and with the quasi-concave optimizer of Cohen, Lyu, Nelson, Sarl'os, and Stemmer [STOC '23] yields a realizable sample complexity of $\widetilde{O}\left(\frac{1}{\varepsilon α}\cdot \min\{d^{2.5} \log^*L, \:\: d^{2.5} + d^{1.5} 2^{\log^*L}\}\right),$ which improves on the previously known bound of $\widetilde{O}\left(\frac{1}{\varepsilon α}\cdot\min\{\frac{1}α\cdot d^{5.5}\log^*L,\:\: d^{2.5}2^{\log^*L}\}\right).$ We also apply the framework to Boolean compositions: given proper private learners for classes $H_1,\ldots,H_k$, we obtain a proper private learner for $G(H_1,\ldots,H_k)$ for any fixed Boolean function $G:\{0,1\}^k\to\{0,1\}$. Compared with the closure theorem of Alon, Beimel, Moran, and Stemmer [COLT '20], this reduces the overhead on a common component sample bound from $\widetilde O(k/\varepsilon)$ to $\widetilde O(\sqrt{k}/\varepsilon)$.
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
Beyond Monolithic Perturbation: Heterogeneous Mechanism Design for Multi-Attribute Metric Differential Privacy Ruiyao Liu, Michael Oluwole, Chenxi Qiu arxiv.org/abs/2610.05561
Beyond Monolithic Perturbation: Heterogeneous Mechanism Design for Multi-Attribute Metric Differential Privacy

Ruiyao Liu, Michael Oluwole, Chenxi Qiu

http://arxiv.org/abs/2610.05561

Multi-attribute user records are inherently heterogeneous, often combining continuous, categorical, and binary attributes, and they frequently exhibit strong cross-attribute dependencies. Designing high-utility metric differential privacy (mDP) mechanisms for such records is challenging. Simple predefined mechanisms, such as distance-based noise, may be poorly aligned with task-specific utility loss, whereas fully optimization-based mechanisms can be computationally prohibitive for multi-attribute records.
  We propose Dependency-aware Heterogeneous Data Perturbation (DepHDP)}, a framework for multi-attribute mDP that combines dependency-aware attribute grouping with heterogeneous perturbation design. Rather than applying a one-size-fits-all mechanism, DepHDP selects an appropriate perturbation strategy for each attribute or attribute group, choosing between efficient predefined mechanisms, such as Laplace or Exponential mechanisms, and optimization-based designs. This selection is guided by both domain size and a predefined-noise adequacy criterion, which quantifies whether task-induced utility loss can be well explained by perturbation magnitude. To support scalable end-to-end optimization, DepHDP estimates group-level utility loss through sampling and lightweight surrogate modeling, and jointly optimizes privacy-budget allocation and group-wise mechanism design under a global $\ell_p$-metric mDP constraint. Across three case studies, DepHDP improves privacy--utility trade-offs over uniform baselines at lower computational cost than full-record OPT on evaluated domains.
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
Learning Sparse Support under Differential Privacy: Adaptive Algorithms and Minimax Limits Jia Gu, T. Tony Cai arxiv.org/abs/2610.05867
Learning Sparse Support under Differential Privacy: Adaptive Algorithms and Minimax Limits

Jia Gu, T. Tony Cai

http://arxiv.org/abs/2610.05867

We study exact support recovery under $(ε,δ)$-differential privacy in sparse high-dimensional linear regression. We introduce Saturated Propose-test-release, a general mechanism that privately releases the output of a discrete selector with probability one once its stability certificate reaches a finite threshold. Exploiting the coordinatewise geometry of the LASSO, we construct a computable support-stability score. The resulting computationally efficient Saturated LASSO satisfies worst-case $(ε,δ)$-differential privacy and achieves exact support recovery with high probability under explicit regularity and beta-min conditions. Maximizing sparsity-indexed certificates yields an adaptive procedure requiring no sparsity knowledge and having exactly the same finite-sample exact-recovery risk as the oracle fixed-sparsity procedure under common public tuning parameters. We also establish a minimax lower bound explicitly tracking $δ$: under its recovery conditions, Saturated LASSO is minimax optimal up to logarithmic factors in $n$ and $1/δ$ uniformly over $0 < δ\leq ε/16$; in the broad moderate-$δ$ regime, it further matches the lower-bound $δ$-dependence. A complementary information-theoretic construction with known sparsity attains the lower-bound rates up to constant factors under independent Gaussian design, at exponential computational cost. Simulations and a semi-synthetic study using public American Community Survey covariates illustrate the numerical performance of the proposed procedures.
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
DP-ES: Differentially Private Evolution Strategies for Prompt Optimization Ziniu Liu, Aiping Li, Yue Han, Han Yu, Junjian Zhang, Dong Zhu, Changjian Li, Shiqiang Zhang arxiv.org/abs/2610.06236
DP-ES: Differentially Private Evolution Strategies for Prompt Optimization

Ziniu Liu, Aiping Li, Yue Han, Han Yu, Junjian Zhang, Dong Zhu, Changjian Li, Shiqiang Zhang

http://arxiv.org/abs/2610.06236

Token-level differentially private (DP) prompt optimization methods such as DP-OPT can become unstable under tight privacy budgets: on GSM8K, DP-OPT obtains $49.5\pm28.5\%$ across 30 runs, and a logged search trajectory reveals prompt-template drift and noise-sensitive irreversible choices. We diagnose these as structural consequences of greedy token-by-token construction over privately aggregated counts. We then propose DP-ES (Differentially Private Evolution Strategies), a structurally cleaner alternative that maintains a population of full prompts, mutates them via LLM calls that never access the private dataset, and spends privacy only on sampled-Gaussian evaluation; deterministic or Gumbel-smoothed selection is post-processing. Under a conservative $(\varepsilon\leq1.0,δ=10^{-5})$ guarantee, DP-ES achieves 88.1% on GSM8K (+38.6 pp over DP-OPT, approximately 9 times lower standard deviation), 99.7% on MedQA, 73.5% on BANKING77, and 86.8% on Alpaca. It is also 2.5 times faster in wall-clock time and uses 3.3 times fewer logged private-data call groups than DP-OPT. Selection and population ablations, implementation-level noise checks, and a 200-profile exact-match memorization stress test complement the formal guarantee. Scope: Our experiments establish optimization robustness under DP noise, especially where prompt structure is critical; end-to-end validation on genuinely sensitive, non-saturated deployment data remains future work.
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
Differentially Private Mixing of Public Datasets Improves Private Learning Yufei Chen, Tejumade Afonja, Anvith Thudi, Nicolas Papernot arxiv.org/abs/2610.06636
Differentially Private Mixing of Public Datasets Improves Private Learning

Yufei Chen, Tejumade Afonja, Anvith Thudi, Nicolas Papernot

http://arxiv.org/abs/2610.06636

Many machine learning applications involve sensitive data and therefore require training under differential privacy (DP). However, DP training often degrades model utility. In some cases, first pre-training the model on "public" data before finetuning with DP on the sensitive data can reduce the drop in utility. However, the success of this depends on how relevant the selected public dataset is to the sensitive data. We introduce the first pipeline that privately learns the mixture of several public datasets to pretrain on for a given sensitive downstream task. Our key insight is that we can privately find the best mixture of multiple public datasets by privately learning a low-dimensional linear model. We tested our method on the NIH dataset for X-ray classification and the ENRON email dataset for language modeling. Applying our method to find tailored mixtures of X-ray datasets to pretrain on for diseases in the NIH ChestX-ray14 dataset, we improved macro AUC by up to 0.037 across privacy budgets compared to the baselines, with gains as large as +22.8% relative AUC on Cardiomegaly at $ε=1$. For DP training on the ENRON dataset, pre-training on our mixture of The Common Pile (a collection of public-domain text datasets) decreased test perplexity by 16% relative to the baseline mixtures.
000
Differential Privacy Papers @dppapers.bsky.social · 06/10/2026
Private online learning and prediction for Littlestone classes Amartya Sanyal arxiv.org/abs/2610.06822
Private online learning and prediction for Littlestone classes

Amartya Sanyal

http://arxiv.org/abs/2610.06822

We study mistake bounds for differentially private online learning and online prediction under oblivious realisable adversaries. Online learning requires the learner to release a hypothesis at each time step whereas in online prediction, the learner only needs to make predictions without releasing a hypothesis. Using a novel lower bound for private online learning and an upper bound for private prediction, we show that the sample complexity of these two problems are separated by a factor that grows with the time horizon for every class of finite Littlestone dimension $d$. First, we prove that every $\br{ε,δ}$-private online learner has a deterministic realisable stream of length $T$ on which the mistake bound is at least $\bE\bs{M_T}=\Om{\frac dε\log\br{ T}^{2/3}}$. In particular, this is the first non-trivial lower in the range $1/T<δ<1/\log T)$ left open in earlier works[SR22,DSS24,LWY24]. Second, we prove that for every class of of Littlestone dimension $d$, there exists an $(ε,δ)$-jointly private predictor with at most $2^{2^{cd^2}}ε^{-2}\log^2\br{2/\br{εδ}}$ expected mistakes, independently of $T$, for some absolute constant $c>0$. Thus, for every fixed class of finite Littlestone dimension when $δ=Θ\br{1/\log T}$, private learning requires $\Om{\br{\log T}^{2/3}}$ expected mistakes, whereas private prediction admits $\bigO{\br{\log\log T}^2}$.
000
Differential Privacy Papers @dppapers.bsky.social · 05/10/2026
Unifying Privacy Accounting: Information Equivalence and Information Loss Buxin Su, Qiaoshi Yang, Yiding Su, Chendi Wang arxiv.org/abs/2610.02414
Unifying Privacy Accounting: Information Equivalence and Information Loss

Buxin Su, Qiaoshi Yang, Yiding Su, Chendi Wang

http://arxiv.org/abs/2610.02414

Differential privacy (DP) admits several notions, but the choice among them may affect both privacy analysis and utility. In this paper, we consider four mainstream curve-based privacy notions within a unified information-theoretic framework. For a fixed ordered pair of output distributions, we establish information equivalence among the two directional privacy profiles of $(\varepsilon,δ)$-DP, the pair of hypothesis-testing trade-off functions, and the extended privacy-loss distribution. The exact Rényi differential privacy (RDP) curve joins this equivalence class whenever it is finite at some order greater than one. Under this mild condition, choosing among these notions changes only their semantic interpretation and computational requirements. In contrast, taking the maximum of the directional privacy profiles or compressing the RDP curve into a single zero-concentrated differential privacy (zCDP) parameter can lose information. We quantify the information loss between the exact RDP curve and its zCDP bound for standard noise mechanisms. This gap is zero for Gaussian noise but generally positive for Gaussian-mixture, Laplace, discrete Gaussian, and Poisson-subsampled Gaussian mechanisms. Moreover, this gap grows linearly with the number of independently composed mechanisms. Our information-theoretic perspective has practical consequences. At the same certified privacy level, retaining the full RDP curve rather than using zCDP reduces the required noise variance by up to $45\%$ for Gaussian-mixture noise in workloads comparable in size to the American Community Survey. For DP-SGD on Fashion-MNIST under Poisson subsampling, an RDP-based privacy accountant improves test accuracy by up to $8.73$ percentage points compared to a zCDP-based accountant when both are calibrated to the same $(\varepsilon,δ)$ guarantee.
000
Differential Privacy Papers @dppapers.bsky.social · 05/10/2026
HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Yan Zhang arxiv.org/abs/2610.02504
HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems

Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Yan Zhang

http://arxiv.org/abs/2610.02504

Balancing electricity demand and supply is increasingly difficult due to the inherent intermittency of renewable power generation and the stochastic power consumption. Grid operators require fine-grained, decision-relevant insights into household energy consumption to manage peak loads and design responsive tariffs, but increased transparency at this level raises significant privacy concerns. Traditional methods for explainable AI (XAI) can reveal sensitive information, while standard privacy techniques often reduce the usefulness of explanations. To address this issue, we introduce HXAI, a hierarchical framework that preserves privacy while enabling reasonable explainable analysis for grid-level demand management. HXAI consists of two main components: (1) a local model that generates fine-grained explanations within a secure, private environment, and (2) a zonal model that aggregates these explanations to support grid-level analysis while enforcing privacy through flexible privacy-budget management. We explicitly limit cumulative privacy exposure under repeated operator queries and show that the proposed framework preserves decision-relevant information without compromising household privacy. Experiments on both simulated and real-world energy datasets demonstrate that HXAI provides useful insights for zonal load management while ensuring that appliance-level consumption remains local and is never transmitted to grid operators. Our results show that preserving the semantic structure of explanations, rather than minimizing numerical error, is the key to XAI under differential privacy. This framework provides a way to achieve both privacy and explainability in energy management.
010
Differential Privacy Papers @dppapers.bsky.social · 05/10/2026
High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning Filip Kovačević, Edwige Cyffers, Stefano Sarao Mannelli, Marco Mondelli arxiv.org/abs/2610.02578
High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning

Filip Kovačević, Edwige Cyffers, Stefano Sarao Mannelli, Marco Mondelli

http://arxiv.org/abs/2610.02578

To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator. Our approach uses only summary statistics and it gives privacy guarantees either on labels only or jointly on features and labels, in terms of $ρ$-zero-concentrated differential privacy. The main technical contribution is a deterministic equivalent of the test error, which captures the interactions between sample size, covariance structure, model shift, regularization and privacy noise. Our theory allows to optimize hyperparameters (weights and ridge regularizers) and, more broadly, to decide when private external datasets are useful without accessing the data itself but only relying on population-level quantities. This provides a theoretically tractable foundation for private transfer learning, which we support via experiments on both synthetic and real-world datasets.
000
Differential Privacy Papers @dppapers.bsky.social · 05/10/2026
Differential Privacy of Gradient Descent on Perturbed Objectives Austin Watkins, Raman Arora arxiv.org/abs/2610.02716
Differential Privacy of Gradient Descent on Perturbed Objectives

Austin Watkins, Raman Arora

http://arxiv.org/abs/2610.02716

Objective perturbation adds a random linear term to a regularized empirical risk and releases the exact perturbed minimizer. We study the finite computation obtained by releasing the $N$-th iterate of deterministic gradient descent on $w\mapsto F(w;S)+\langle z,w\rangle$, where $z\sim\mathcal N(0,σ^2I_d)$ is drawn once before optimization. For strongly convex and smooth objectives with Lipschitz Hessian, we prove an explicit condition under which the map $z\mapsto w_N$ is a $C^1$-diffeomorphism on the bounded domains used in the privacy argument, with a quantitative lower bound on the smallest singular value of its Jacobian. This permits a direct change-of-variables analysis of the finite iterate. For generalized linear models, the resulting privacy-profile bound has no explicit ambient-dimension factor once the iteration condition holds, and its finite-iteration correction decreases geometrically. By letting the free truncation parameter grow slowly with $N$, we recover the corresponding exact-minimizer certificate in the limit. We also bound the expected excess empirical risk by $dσ^2/(2μ)$ plus a geometrically decreasing optimization term, and transfer the result to population risk without an additional multiplicative condition-number factor in the leading statistical terms.
010
Differential Privacy Papers @dppapers.bsky.social · 05/10/2026
ReLEAF: A Socio-Technical Framework Bridging Custodians and Researchers for Trustworthy Data Sharing Hibiki Ito, Chia-Yu Hsu, Hiroaki Ogata arxiv.org/abs/2610.02720
ReLEAF: A Socio-Technical Framework Bridging Custodians and Researchers for Trustworthy Data Sharing

Hibiki Ito, Chia-Yu Hsu, Hiroaki Ogata

http://arxiv.org/abs/2610.02720

Growing volumes of educational real-world data (ERWD) are being collected across learning platforms and institutional systems. Sharing these data within the Learning Analytics community offers substantial research opportunities, yet access remains limited by ethical, regulatory and governance constraints. Prior work has primarily focused on anonymisation techniques, but little attention has been paid to operational design of practical ERWD sharing, particularly how data custodians and researchers interact through privacy-preserving access mechanisms. To address this gap, we propose ReLEAF, a socio-technical framework that bridges data custodians and researchers by operationalising two-stage data sharing: 1) Differentially private synthetic data are shared for exploratory analysis, and 2) controlled real-data validation is performed on demand. Following a design-science research approach, we refine and formatively evaluate ReLEAF through three cycles involving 4 graduate students, 6 researchers, and 90 undergraduate students, respectively. Three design principles emerged through the cycles: P1) position privacy-preserving access mechanisms within the research workflow, P2) make the conditions for acceptable secondary use explicit and actionable, and P3) promote engagement with governance requirements rather than automate compliance decisions. Together, ReLEAF provides a concrete framework for trustworthy ERWD sharing, while the design principles offer transferable guidance for other data-sharing contexts.
000
Differential Privacy Papers @dppapers.bsky.social · 05/10/2026
Inner Momentum for Differentially Private Muon Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai arxiv.org/abs/2610.02738
Inner Momentum for Differentially Private Muon

Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai

http://arxiv.org/abs/2610.02738

Differentially private training clips each per-example gradient before adding noise. This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average. Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift. To curb this degradation, we propose averaging each sampled example's Muon gradient over the current model and a short history of recent models before clipping. The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F <= sigma_lambda sigma_G, bounding the clipping-induced distortion directly. We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization. In private GPT-2 fine-tuning on E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, and non-private diagnostics show 2-4% lower pre-noise polar error.
000
Differential Privacy Papers @dppapers.bsky.social · 05/10/2026
Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy Kaushik Bhargav Sivangi, Fani Deligianni arxiv.org/abs/2610.02943
Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy

Kaushik Bhargav Sivangi, Fani Deligianni

http://arxiv.org/abs/2610.02943

Multimodal 3D Human Pose Estimation (3D HPE) combines complementary information from RGB, LiDAR, and mmWave radar, but models trained on correlated observations from the same individuals, raise privacy risks overlooked by record level analysis. We present a unified framework for multimodal 3D HPE that couples kinematics-induced sensor fusion with subject level privacy auditing and private training. First, our multimodal model aligns modality specific joint representation, injects skeletal structure and adaptively aggregates complementary sensor evidence for accurate pose prediction. Second, we formulate a black-box subject membership inference attack for 3D HPE, complemented by an empirical pointwise maximal leakage analysis, which characterizes how individual attack score outcomes change inference about the membership outcome. Third, we instantiate user-level differential privacy via Action Temporal Stratification, a population weighted within-subject sampling strategy that enforces action and temporal coverage. We evaluate our framework on the MM-Fi dataset across three diverse experimental protocols. Source-code will be released upon acceptance.
000
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection Xudong Wang, Chris Ding, Tongxin Li, Jicong Fan arxiv.org/abs/2609.05253
GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection

Xudong Wang, Chris Ding, Tongxin Li, Jicong Fan

http://arxiv.org/abs/2609.05253

We introduce GLASS, a framework for graph-level anomaly detection (GLAD) that achieves robust cross-domain transferability through graph-language alignment on the unit hypersphere. GLASS builds a unified representation space by aligning a structure-aware graph encoder with an instruction-aware text embedding via a multi-slice soft cosine objective. Our framework serializes local, global, and semantic graph properties into a compact Graph Descriptor Prompt (GraphDP), creating a text bridge that enables domain-agnostic anomaly scoring. By enforcing multi-scale consistency through Matryoshka representation slices, the model captures anomalous deviations at multiple levels of granularity. We formulate anomaly detection as density estimation on the aligned hypersphere and introduce Spherical Multi-Modal Scoring (SMS), which instantiates von Mises-Fisher kernel density estimators in both graph and text embedding spaces. This probabilistic formulation recovers angular 1-nearest-neighbor scoring in the high-concentration limit, motivates the practical mean k-nearest-neighbor scorer, and provides a principled fusion of structural and semantic anomaly signals. The shared text embedding space further serves as a cross-domain bridge: by encoding a target domain's GraphDP without target-domain training data, GLASS performs zero-shot anomaly detection, and with only a handful of normal examples, few-shot adaptation via reference-set calibration. For privacy-sensitive deployment, we extend reference-set calibration with a bounded joint graph-text kernel summary that provides graph-record differential privacy while keeping the encoders fixed independently of the private target references. Across twelve benchmarks and three meta-domains, GLASS obtains the best average AUROC and rank compared with rece
000
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Energy Time-Series Imputation with Differentially Private Diffusion Models via Clipping-Aware Objective Conditioning Huizhen Huang, Yu Li, Tao Huang, Chen Hou arxiv.org/abs/2610.00209
Energy Time-Series Imputation with Differentially Private Diffusion Models via Clipping-Aware Objective Conditioning

Huizhen Huang, Yu Li, Tao Huang, Chen Hou

http://arxiv.org/abs/2610.00209

Reliable recovery of missing measurements is important for monitoring and analysis in energy time-series systems, where fine-grained measurements may contain sensitive temporal information. Diffusion models trained with differentially private stochastic gradient descent (DP-SGD) provide a promising framework for privacy-sensitive energy time-series imputation. Under cosine diffusion schedules, late timesteps correspond to low signal-to-noise ratio (SNR) conditions, where standard $\varepsilon$-prediction can induce large pre-clipping gradients. Such gradients are more likely to be clipped, reducing the retained optimization signal. The artificial intelligence (AI) contribution lies in formulating this objective--clipping interaction as an objective optimization problem under fixed-threshold DP-SGD and developing timestep-aware objective conditioning for diffusion-based energy time-series imputation. The method adopts $v$-prediction to mitigate late-timestep gradient amplification, uses static loss weighting as a uniform-scaling control, and introduces diffusion-schedule-aware dynamic weighting for stronger attenuation before clipping. For the engineering application, we evaluate the method on five real-world energy time-series datasets across random point missingness, contiguous block missingness, persistent outages, and multiple missing-data severities. Under matched DP-SGD settings, the proposed method consistently improves imputation utility over the $\varepsilon$-prediction baseline. Gradient diagnostics reveal lower upper-tail pre-clipping gradient norms, reduced clipping fractions, and stronger attenuation at late low-SNR timesteps, supporting the effectiveness of clipping-aware objective conditioning for energy time-series imputation.
000
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs Mohamed Shaaban, Mohamed Elmahallawy arxiv.org/abs/2610.00309
Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs

Mohamed Shaaban, Mohamed Elmahallawy

http://arxiv.org/abs/2610.00309

Large language models (LLMs) are increasingly deployed in privacy-critical domains (e.g., healthcare, finance, and government), but their propensity to memorize and disclose personally identifiable information (PII) poses serious security and compliance risks. Existing defenses typically force a trade-off between model utility, privacy protection, and access to fine-tuned private knowledge. We propose LoRA-Oriented Control via Keyed Entry Tokens (Locket), a practical framework that embeds fine-grained, policy-driven access control directly into LLM generation. Locket trains a set of lightweight LoRA (Low-Rank Adaptation) adapters, each encoding a distinct access policy (e.g., full reveal, partial redaction via PII masking, or reveal under a specified differential privacy level). A compact gating module is trained to associate a learned keyed entry token with exactly one LoRA adapter via sequence-level hard routing; the presence of a valid token acts as an authorization key that unlocks corresponding private knowledge, while an invalid or absent token triggers a privacy-preserving adapter that redacts or sanitizes sensitive content. This design ensures Locket remains fully compatible with off-the-shelf LLMs, supporting scalable deployment while satisfying regulatory and privacy requirements. We evaluate Locket across multiple datasets (Enron, ECHR, Yelp) and a diverse set of state-of-the-art LLMs, including Qwen3 (1.7B and 8B), Meta's Llama-3.2 (1B and 3B), and Google's Gemma-2-2B. Our extensive experiments demonstrate that, when the correct token is provided, Locket preserves perplexity comparable to fine-tuning on raw data (without any defense). Conversely, when the token is missing or invalid, it substantially reduces PII leakage while maintaining utility and perplexity on par with stron
010
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability Ceren Yıldırım, Kamer Kaya, Sinan Yıldırım, Erkay Savaş arxiv.org/abs/2610.01650
Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability

Ceren Yıldırım, Kamer Kaya, Sinan Yıldırım, Erkay Savaş

http://arxiv.org/abs/2610.01650

The increasing prevalence of decentralized data has led to a growing interest in federated learning, which enables collaborative model training without clients sharing their sensitive local data. However, FL alone does not sufficiently protect sensitive training data and is generally coupled with privacy-preserving techniques, such as differential privacy and homomorphic encryption. Although powerful, these techniques address separate concerns via different mechanisms, so relying on just one might prove insufficient or impractical for addressing challenges associated with federated learning. In this work, we propose a privacy-preserving federated learning framework that combines homomorphic encryption-based training with differential privacy-based model inspection and release. We adopt a Markov chain Monte Carlo-based Bayesian privacy estimation method to estimate the privacy of our proposed framework. Our results show that this method improves both model utility and estimated privacy over the baseline method that relies solely on differential privacy for training. In our experiments with the FEMNIST dataset, by the end of training, our method reaches a test loss of $1.09$, compared to $2.37$ for the differential privacy-only approach, while providing stronger estimated privacy protection, with the estimated posterior mean of the privacy parameter $ε$ of $4.32$, compared to $7.26$ for the differential privacy-only approach. We also show that intermittent model monitoring can preserve the encrypted training trajectory while, under our evaluated experimental setting, providing estimated privacy comparable to or stronger than the differential privacy-only approach.
011
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Detection and Resolution of Periodic Artifacts in OpenDP's Discrete Laplace Sampler Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta, Massimo Caccia arxiv.org/abs/2610.01907
Detection and Resolution of Periodic Artifacts in OpenDP's Discrete Laplace Sampler

Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta, Massimo Caccia

http://arxiv.org/abs/2610.01907

Differential privacy implementations rely on precise sampling from noise distributions to provide formal privacy guarantees. We report the discovery of systematic artifacts in OpenDP's discrete Laplace sampler that manifest as periodic distortions in the output distribution. Through systematic testing, we trace these artifacts to a faulty implementation in the rational arithmetic library used by the bernoulli_exp1 function, a low-level primitive that implements sampling from Bernoulli(e^(-x)) distributions. We present a diagnostic methodology that isolates the faulty component in the nested sampling hierarchy and propose an alternative implementation based on exact rational arithmetic that eliminates the artifacts. Statistical validation with 10^6 samples confirms that the corrected sampler produces outputs indistinguishable from the theoretical distribution at the tested precision level.
000
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Quantum Advantage for Two-Party Differential Privacy Daniel Alabi, Emil T. Khabiboulline arxiv.org/abs/2610.02113
Quantum Advantage for Two-Party Differential Privacy

Daniel Alabi, Emil T. Khabiboulline

http://arxiv.org/abs/2610.02113

We introduce information-theoretically private quantum protocols for two-party Hamming distance when both parties must output the same estimate. Classically, for input length $n$, information-theoretic protocols require $Ω(\sqrt{n})$ error under pure differential privacy and $Ω(\sqrt{n}/\log n)$ error under strong approximate differential privacy, whereas computational security permits $O(1)$ error. In Klauck's honest, nonpreemptive, message-preserving model, we give an $O(n)$-communication quantum protocol with pure $\varepsilon$ quantum differential privacy (QDP) and expected error at most $\frac{2}{\sinh \varepsilon}+γ$, for every $γ>0$. For approximate $(\varepsilon, δ)$ QDP, an exact hockey-stick divergence calculation yields strictly smaller error, while preserving the $O(1)$-versus-$Ω(\sqrt{n}/\log n)$ separation for $δ=o(1/n)$. Thus, quantum communication achieves $O(1)$ information-theoretic error, matching the accuracy available classically only under computational assumptions.
  The main construction uses a guarded coherent round trip and an equal-Gram rigidity principle that prevents an honest player from retaining input-dependent complementary information. We separate this model from weaker prescribed-channel privacy, which already admits an exact classical realization, and from fully retention-robust security, against which measurement-and-abort attacks remain possible. Therefore, we identify preservation of non-orthogonal quantum messages as a resource for privacy.
000
Differential Privacy Papers @dppapers.bsky.social · 01/10/2026
PrivCert: Certifying Statement Support under Differential Privacy Tsubasa Takahashi, Takumi Hiraoka arxiv.org/abs/2609.38934
PrivCert: Certifying Statement Support under Differential Privacy

Tsubasa Takahashi, Takumi Hiraoka

http://arxiv.org/abs/2609.38934

Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement carries about the underlying data. We identify this as an evidence gap: a private report may contain plausible claims without indicating whether they are strongly supported by the private dataset. We introduce PrivCert, a framework for privacy-preserving reporting that makes statement support explicit through privacy-preserving certificates and emit-or-abstain decisions. As a canonical instantiation, PrivCert-PF (Proposal-and-Filter) separates data-independent candidate discovery from private support certification, emitting only statements whose support passes a private evidence test. We provide theoretical grounding for this framework by characterizing the limits of implicit evidence under DP, deriving a sharp privacy--honesty frontier for single-statement certification, and establishing a worst-case cost for fine-grained multi-statement certification. Experiments on synthetic tasks and TAB, WildChat, and Yelp show that explicit certification maintains low unsupported emission, while free-text DP baselines frequently produce low-support claims under the same declared support semantics. We further show that the PrivCert contract can be realized with histogram, sparse-vector, and Gaussian mechanisms, and use DP synthetic data to illustrate an important boundary: support in a private proxy does not automatically certify support in the original data. Together, these results position privacy-preserving reporting as an evidence-design problem: not only how to generate private text, but what a private report can substantiate about its underlying data.
000
Differential Privacy Papers @dppapers.bsky.social · 01/10/2026
Certification-Based Differentially Private Learning Mihnea Ghitu, Matthew Wicker arxiv.org/abs/2609.39629
Certification-Based Differentially Private Learning

Mihnea Ghitu, Matthew Wicker

http://arxiv.org/abs/2609.39629

Differential privacy (DP) in machine learning is typically achieved by adding noise to model parameters (private learning) or to model outputs (private prediction). Recent work uses formal methods, namely abstract interpretation, to provide tighter privacy guarantees, but only for private prediction in classification settings. In this work, we investigate the use of formal methods as a general tool for tighter privacy analysis. First, we generalize the abstract gradient training (AGT) framework to private prediction in continuous, unbounded regression. Second, by reducing learning in parameterized models to a regression problem over the parameter space, we introduce Abstract Gradient Sampling (AGS), an algorithm that enables reachability-based analysis to provide guarantees for private learning. In both private prediction and private learning, we provide tightened privacy accounting for the AGT framework and a theoretical analysis demonstrating when our smooth sensitivity upper-bounds yield favourable privacy-utility trade-off. In practice, we validate that our regression bounds are tighter than global-sensitivity baselines on regression benchmarks, and, notably, yield the first finite privacy guarantees in settings where global prediction sensitivity is a priori unbounded. We also find that under matched conditions, our private learning algorithm can outperform standard private learners.
000
Differential Privacy Papers @dppapers.bsky.social · 01/10/2026
Privacy Foundations for Multi-Institutional Scientific Artificial Intelligence Olivera Kotevska, Sumit Jha, Aurélien Bellet, Rui Hu, Nathaniel D. Bastian, Rafael Ferreira da Silva, Ravi Madduri, Kibaek Kim arxiv.org/abs/2609.39787
Privacy Foundations for Multi-Institutional Scientific Artificial Intelligence

Olivera Kotevska, Sumit Jha, Aurélien Bellet, Rui Hu, Nathaniel D. Bastian, Rafael Ferreira da Silva, Ravi Madduri, Kibaek Kim

http://arxiv.org/abs/2609.39787

Scientific artificial intelligence (AI), spanning foundation models (FMs) to federated data-analysis pipelines, is becoming shared infrastructure across national laboratories, universities, hospitals, and industrial partners. This collaboration creates privacy risks whose natural unit is often an institution's participation, research strategy, or technical capability rather than a single record. Differential privacy (DP), federated learning (FL), secure computation, trusted execution, and provenance each protect parts of the stack, but their guarantees rarely compose across mixed-trust institutions, access tiers, and autonomous agents. This perspective recasts privacy for scientific AI as an assurance problem defined by six elements: protected asset, observer, channel, permitted disclosure, guarantee, and evidence. We demonstrate the framing through a claim register for a composite cross-institutional scenario and use it to assess the model lifecycle. Two of the resulting gaps are specific to leadership-class facilities: scheduler, allocation, and telemetry metadata expose an institution's resource posture, and instrument-attached control loops leak research strategy through timing and contention on shared accelerators. We identify six research priorities: institution-level guarantees, agent-communication privacy, cross-tier information flow, privacy-compatible reproducibility, leadership-scale accounting, and instrument side channels. The contribution is a common form for stating, comparing, and auditing claims whose guarantees otherwise remain fragmented across the scientific AI stack.
000
Differential Privacy Papers @dppapers.bsky.social · 01/10/2026
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD? Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine arxiv.org/abs/2609.40335
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine

http://arxiv.org/abs/2609.40335

Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.
010
Differential Privacy Papers @dppapers.bsky.social · 30/09/2026
Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising Om Shankar Tiwari, Navnit Shukla, Guanyu Wang, Akshay Jain arxiv.org/abs/2609.36153
Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising

Om Shankar Tiwari, Navnit Shukla, Guanyu Wang, Akshay Jain

http://arxiv.org/abs/2609.36153

B2B advertising targets a viewer's professional attributes (employer size and industry, function, seniority) and has obtained them by matching identities across sites. Safari and Firefox block third-party cookies, Google retired the Privacy Sandbox cohort APIs in 2025, and reverse-IP firmographics decay under remote work. We present SIF (Sealed Inference Frame), which infers coarse professional cohorts on the device and emits only a locally differentially private, taxonomy-coded label into the OpenRTB bid stream, with no cross-site identifier. It rests on a property of the web platform we make precise: a navigated cross-origin iframe is the only way third-party code obtains a policy it controls, so inference runs in WebAssembly even where the publisher's CSP forbids it, and a nested worker served with default-src 'none' gives the model no network. Even a malicious model leaks at most about 5 bits per site per week. Labels pass through a memoised k-ary randomised response keyed to the publisher's first-party identifier, which gives $\varepsilon$-local differential privacy, defeats averaging, and links requests no better than the identifier already sent. An org-conditional k-anonymity rule suppresses cells, more strictly on corporate networks than at home. Cohorts ride OpenRTB user.data in a LinkedIn-aligned taxonomy, and attribution uses LinkedIn's click-scoped li_fat_id without bridging identities. We report a crawl of CSP deployment on 7,969 top sites and 431 B2B publishers, Heavy-Ad budgets, closed-form privacy-utility trade-offs, a re-identification simulation, and an assessment of which attributes are predictable at all: company type and size are, seniority largely is not. On-device is a design property, not a consent exemption.
000
Differential Privacy Papers @dppapers.bsky.social · 30/09/2026
Understanding Private Evolution as Learning-Augmented Clustering Audra McMillan, Kunal Talwar, Felix Zhou arxiv.org/abs/2609.36678
Understanding Private Evolution as Learning-Augmented Clustering

Audra McMillan, Kunal Talwar, Felix Zhou

http://arxiv.org/abs/2609.36678

Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-case Wasserstein analyses would predict. We recast PE as generative model-augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then we can obtain much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then sample complexity depends on intrinsic, not ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm is competitive with standard baselines and can improve recall.
000
Differential Privacy Papers @dppapers.bsky.social · 30/09/2026
A Sharp Transition in Data Reconstruction under Differential Privacy Max Cairney-Leeming, Simone Bombari, Marco Mondelli arxiv.org/abs/2609.37344
A Sharp Transition in Data Reconstruction under Differential Privacy

Max Cairney-Leeming, Simone Bombari, Marco Mondelli

http://arxiv.org/abs/2609.37344

Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single $d$-dimensional training sample from a $ρ$-zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at $ρ\asymp d$ for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for $ρ\ll d$; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for $ρ\gg d$. Remarkably, the transition moves to $ρ\asymp s$ for data lying in an $s$-dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
TRAP: Understanding and Mitigating Privacy Memorization in Language Models Muhammed Ustaomeroglu, Ziyue Xu, Hanshen Xiao, Peter Cnudde, Guannan Qu, Holger R. Roth arxiv.org/abs/2609.32293
TRAP: Understanding and Mitigating Privacy Memorization in Language Models

Muhammed Ustaomeroglu, Ziyue Xu, Hanshen Xiao, Peter Cnudde, Guannan Qu, Holger R. Roth

http://arxiv.org/abs/2609.32293

Fine-tuning a language model on sensitive records can leave it able to reproduce them. We ask when this memorization arises and how to prevent it without knowing in advance which spans are sensitive. Our starting point is that most memorization scores and attacks share one statistical core: whether the model assigns a token more probability than some reference would. Taking as the reference a model trained on the complementary half of the same corpus gives the Target Reference Advantage (TRA), a per-token signal that separates what a model fit to a particular record from what it learned across records, and is cheap and differentiable. We then study what drives memorization during fine-tuning: it keeps growing well past the validation minimum, is larger on small datasets and at higher learning rates, and higher when the underlying task is harder. Early stopping removes much of it, but because it is chosen by aggregate validation loss it helps least for rare, hard-to-predict spans embedded in otherwise learnable text, which is exactly what sensitive information tends to be. We therefore introduce TRAP, a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference. On student essays with annotated personal information and clinical cases with patient identifiers, TRAP brings memorization near the level of an untrained model at little utility cost, where generic regularizers barely move and differential privacy gives up most of what fine-tuning bought.
020
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Differentially Private Approximation of the John Ellipsoid Bar Mahpud, Daniel Omer, Or Sheffet arxiv.org/abs/2609.32606
Differentially Private Approximation of the John Ellipsoid

Bar Mahpud, Daniel Omer, Or Sheffet

http://arxiv.org/abs/2609.32606

We study the problem of approximating the John ellipsoid (JE) of a given (centrally symmetric) polytope of $n$ constraints in a Euclidean space under differential privacy (DP). We give the first differentially private algorithm for this problem under the standard model, where neighboring datasets may differ arbitrarily in one a single constraint. Our work also extends to the complimentary problem of Minimum Enclosing Ellipsoid of $n$ points in the Euclidean space.
  Our approach is based on the recent non-private multiplicative-weights algorithm of~\cite{pmlr-v99-cohen19a}. First we introduce a non-private generalization of the Cohen et al algorithm, yielding a $(1+γ)$-approximation of the JE problem while violating at most $κn$ constraints in $O(\log(1/κ)/γ)$ iterations. This variant works by projecting the intermediate weights assigned to the constraints onto the set of $κ$-dense distributions, similarly to~\cite{bun2020efficientnoisetolerantprivatelearning}.
  We then design a $ρ$-zCDP variant of this algorithm by adding Gaussian noise to the weighted covariance matrix aggregated in each step of the algorithm. Under a mild goodness assumption on the data we can assert that the resulting noisy matrix is close to the true matrix, thereby achieving essentially the same guarantee as the non-private algorithm provided sufficiently many input points. Thus our method achieves an efficient DP poly-time algorithm under concrete sample complexity bounds.
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier arxiv.org/abs/2609.32835
FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors

Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier

http://arxiv.org/abs/2609.32835

As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.
001
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Geometry-Adaptive Mechanisms for Private Synthetic Data Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian arxiv.org/abs/2609.33363
Geometry-Adaptive Mechanisms for Private Synthetic Data

Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian

http://arxiv.org/abs/2609.33363

Generating differentially private synthetic data with meaningful Wasserstein utility guarantees is challenging in high dimensions. For datasets of size \(n\) on $[0,1]^d$ with $d\ge2$, existing pure \(\varepsilon\)-differentially private mechanisms achieve expected $1$-Wasserstein error of order $(\varepsilon n)^{-1/d}$, reflecting the curse of dimensionality. While this rate is optimal in the worst case, it can be overly pessimistic when the data are supported on a lower-dimensional set. We formalize this through a multiscale packing-growth dimension $k$, which captures the geometric complexity of the support via the growth of packing numbers across scales. We propose \emph{Adaptive Pruned-PMM}, a pure $\varepsilon$-differentially private mechanism that combines private depth selection with our pruned variant of the Private Measure Mechanism (PMM) of He et al.\ (2023). The mechanism supports deeper, geometry-adapted hierarchies with expected running time $O\!\left(d(n+d)\log(\varepsilon n)\right)$, which is near-linear in $n$ for fixed dimension and privacy budget. Under an external multiscale packing-growth condition with dimension $k$, we show that, for fixed positive privacy budgets and fixed geometry, the expected $1$-Wasserstein error is of order $(\varepsilon n)^{-1/k}$ for $k>1$ as $n$ grows. We also prove a lower bound under a corresponding internal packing-growth condition, showing that the exponent $1/k$ is sharp within this framework.
001
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits Idan Attias, Orin Levy, Alexander Ryabchenko, Yishay Mansour, Uri Stemmer arxiv.org/abs/2609.33888
Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits

Idan Attias, Orin Levy, Alexander Ryabchenko, Yishay Mansour, Uri Stemmer

http://arxiv.org/abs/2609.33888

Can vanilla policy optimization explore enough to achieve near-optimal regret in stochastic contextual bandits? We show that standard exponential policy updates driven by offline regression do so under realizability, without exploration bonuses or importance weighting. For $A$ actions, $T$ rounds, and a finite prediction class $F$, vanilla PO achieves $\widetilde O(\sqrt{AT\log(|F|)})$ regret with high probability. Our analysis reveals an implicit exploration mechanism of independent interest: gradual policy updates prevent actions from losing probability too quickly, allowing the regression oracle to learn their expected losses. We further develop a batched version using only $O(\log T)$ regression calls and policy switches, and show how private regression oracles yield differentially private contextual bandit algorithms without composition across batches. For a finite class, this gives pure $\varepsilon_{\rm priv}$-DP and regret $\widetilde O\left( \sqrt{AT \log(|F|/δ)}(1+\varepsilon_{\rm priv}^{-1/2}) \right)$. Finally, experiments across oracle-based contextual bandit algorithms, with and without privacy, demonstrate the practical effectiveness of policy optimization and the value of explicit exploration under stronger privacy constraints.
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Contraction of Rényi Divergences for Discrete Channels Adrien Vandenbroucque, Amedeo Roberto Esposito, Michael Gastpar arxiv.org/abs/2609.34570
Contraction of Rényi Divergences for Discrete Channels

Adrien Vandenbroucque, Amedeo Roberto Esposito, Michael Gastpar

http://arxiv.org/abs/2609.34570

We investigate Strong Data-Processing Inequality (SDPI) constants for Rényi Divergences on finite spaces. We study their dependence on the Rényi order $α$, proving that they are non-decreasing and that their scaling by $(α-1)$ is convex for $α\geq1$. We also identify several support restrictions on the probability measures involved in determining these constants. In particular, in the distribution-independent setting, measures supported on a common set of at most two points suffice to evaluate the Rényi-SDPI constant. For $α\in[0,1]$, we further prove equality with the $χ^2$-SDPI constant, while at order infinity we obtain a closed-form expression. In order to link contraction over product spaces to contraction along individual coordinates, we provide tensorisation bounds for arbitrary product channels analogous to those known for $\varphi$-Divergences. At finite orders, the Rényi-SDPI constants are bounded above and below through comparisons with the $χ^2$-Divergence and Hellinger Divergences, with sharpness established in multiple cases. At order infinity, we instead relate these constants to the contraction of Total Variation Distance. Finally, our findings are applied to local differential privacy (LDP) and the analysis of Markov chains. This yields sharp contraction guarantees for pure-LDP mechanisms, and connects Rényi-LDP to Rényi-SDPI constants. For Markov chains, we derive finite-time convergence bounds and exhibit a family of chains for which Rényi-SDPIs improve on classical $χ^2$-based bounds by arbitrarily large factors.
000
Differential Privacy Papers @dppapers.bsky.social · 28/09/2026
Revisiting Certified Defense with Differential Privacy on Vision Transformers Jun Yan, Weiquan Huang, Qixian Zhang, Yan Bai, Shutai Zhang arxiv.org/abs/2609.31310
Revisiting Certified Defense with Differential Privacy on Vision Transformers

Jun Yan, Weiquan Huang, Qixian Zhang, Yan Bai, Shutai Zhang

http://arxiv.org/abs/2609.31310

Certified defenses that incorporate differential privacy have proven effective on Convolutional Neural Networks (CNNs), furnishing rigorous robustness guarantees against norm-bounded adversaries. However, the certified robustness behavior of Pixel Differential Privacy (PixelDP) remains largely unexplored with the self-attention architecture now dominating the deep-learning landscape. Given that the Transformer has a profound impact on our daily applications from the digital world to the physical world, it is crucial to study certified robustness through differential-privacy-style stability. To fill this research gap, we revisit this construction in Vision Transformers and identify a failure mode that is largely hidden in the convolutional setting. When noise is injected after the patch embedding, the Laplace mechanism with the inherited grouped $\ell_1$ sensitivity bound collapses to chance-level accuracy across noise scales, whereas the Gaussian mechanism remains trainable. This contrast isolates the source of failure: not the injected noise itself, but the geometry of the sensitivity constraint. We show that the attenuation induced by the inherited $Δ_{1,1}$ projection increases with layer width and kernel size according to a random-matrix scale $C/(\sqrt{M}+\sqrt{N})$. Replacing the $\ell_1$-type constraint with a spectral-norm constraint eliminates the collapse across datasets and architectures, but creates a fundamental obstacle: the repaired models no longer satisfy the sensitivity condition required by the standard Laplace certificate. We resolve this mismatch by deriving a dimension-free $(\varepsilon,\ δ)$-privacy guarantee for the Laplace mechanism under $\ell_2$ sensitivity through concentration of the privacy loss.
010
Differential Privacy Papers @dppapers.bsky.social · 28/09/2026
Gap-free Differentially Private PCA for Gaussian Data Alina Ene, Huy L. Nguyen arxiv.org/abs/2609.31614
Gap-free Differentially Private PCA for Gaussian Data

Alina Ene, Huy L. Nguyen

http://arxiv.org/abs/2609.31614

We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
000