Sign in

Differential Privacy Papers

@dppapers.bsky.social
557 followers 0 following 1.7K posts

🤖 new arXiv preprints mentioning "differential privacy" or "differentially private" in the title/abstract - unrelated quantum/FL papers + updates from differentialprivacy.org [Under construction.]

PostsRepliesMedia
Differential Privacy Papers @dppapers.bsky.social · 13h
Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising Om Shankar Tiwari, Navnit Shukla, Guanyu Wang, Akshay Jain arxiv.org/abs/2609.36153
Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising

Om Shankar Tiwari, Navnit Shukla, Guanyu Wang, Akshay Jain

http://arxiv.org/abs/2609.36153

B2B advertising targets a viewer's professional attributes (employer size and industry, function, seniority) and has obtained them by matching identities across sites. Safari and Firefox block third-party cookies, Google retired the Privacy Sandbox cohort APIs in 2025, and reverse-IP firmographics decay under remote work. We present SIF (Sealed Inference Frame), which infers coarse professional cohorts on the device and emits only a locally differentially private, taxonomy-coded label into the OpenRTB bid stream, with no cross-site identifier. It rests on a property of the web platform we make precise: a navigated cross-origin iframe is the only way third-party code obtains a policy it controls, so inference runs in WebAssembly even where the publisher's CSP forbids it, and a nested worker served with default-src 'none' gives the model no network. Even a malicious model leaks at most about 5 bits per site per week. Labels pass through a memoised k-ary randomised response keyed to the publisher's first-party identifier, which gives $\varepsilon$-local differential privacy, defeats averaging, and links requests no better than the identifier already sent. An org-conditional k-anonymity rule suppresses cells, more strictly on corporate networks than at home. Cohorts ride OpenRTB user.data in a LinkedIn-aligned taxonomy, and attribution uses LinkedIn's click-scoped li_fat_id without bridging identities. We report a crawl of CSP deployment on 7,969 top sites and 431 B2B publishers, Heavy-Ad budgets, closed-form privacy-utility trade-offs, a re-identification simulation, and an assessment of which attributes are predictable at all: company type and size are, seniority largely is not. On-device is a design property, not a consent exemption.
000
Differential Privacy Papers @dppapers.bsky.social · 13h
Understanding Private Evolution as Learning-Augmented Clustering Audra McMillan, Kunal Talwar, Felix Zhou arxiv.org/abs/2609.36678
Understanding Private Evolution as Learning-Augmented Clustering

Audra McMillan, Kunal Talwar, Felix Zhou

http://arxiv.org/abs/2609.36678

Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-case Wasserstein analyses would predict. We recast PE as generative model-augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then we can obtain much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then sample complexity depends on intrinsic, not ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm is competitive with standard baselines and can improve recall.
000
Differential Privacy Papers @dppapers.bsky.social · 13h
A Sharp Transition in Data Reconstruction under Differential Privacy Max Cairney-Leeming, Simone Bombari, Marco Mondelli arxiv.org/abs/2609.37344
A Sharp Transition in Data Reconstruction under Differential Privacy

Max Cairney-Leeming, Simone Bombari, Marco Mondelli

http://arxiv.org/abs/2609.37344

Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single $d$-dimensional training sample from a $ρ$-zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at $ρ\asymp d$ for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for $ρ\ll d$; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for $ρ\gg d$. Remarkably, the transition moves to $ρ\asymp s$ for data lying in an $s$-dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
TRAP: Understanding and Mitigating Privacy Memorization in Language Models Muhammed Ustaomeroglu, Ziyue Xu, Hanshen Xiao, Peter Cnudde, Guannan Qu, Holger R. Roth arxiv.org/abs/2609.32293
TRAP: Understanding and Mitigating Privacy Memorization in Language Models

Muhammed Ustaomeroglu, Ziyue Xu, Hanshen Xiao, Peter Cnudde, Guannan Qu, Holger R. Roth

http://arxiv.org/abs/2609.32293

Fine-tuning a language model on sensitive records can leave it able to reproduce them. We ask when this memorization arises and how to prevent it without knowing in advance which spans are sensitive. Our starting point is that most memorization scores and attacks share one statistical core: whether the model assigns a token more probability than some reference would. Taking as the reference a model trained on the complementary half of the same corpus gives the Target Reference Advantage (TRA), a per-token signal that separates what a model fit to a particular record from what it learned across records, and is cheap and differentiable. We then study what drives memorization during fine-tuning: it keeps growing well past the validation minimum, is larger on small datasets and at higher learning rates, and higher when the underlying task is harder. Early stopping removes much of it, but because it is chosen by aggregate validation loss it helps least for rare, hard-to-predict spans embedded in otherwise learnable text, which is exactly what sensitive information tends to be. We therefore introduce TRAP, a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference. On student essays with annotated personal information and clinical cases with patient identifiers, TRAP brings memorization near the level of an untrained model at little utility cost, where generic regularizers barely move and differential privacy gives up most of what fine-tuning bought.
020
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Differentially Private Approximation of the John Ellipsoid Bar Mahpud, Daniel Omer, Or Sheffet arxiv.org/abs/2609.32606
Differentially Private Approximation of the John Ellipsoid

Bar Mahpud, Daniel Omer, Or Sheffet

http://arxiv.org/abs/2609.32606

We study the problem of approximating the John ellipsoid (JE) of a given (centrally symmetric) polytope of $n$ constraints in a Euclidean space under differential privacy (DP). We give the first differentially private algorithm for this problem under the standard model, where neighboring datasets may differ arbitrarily in one a single constraint. Our work also extends to the complimentary problem of Minimum Enclosing Ellipsoid of $n$ points in the Euclidean space.
  Our approach is based on the recent non-private multiplicative-weights algorithm of~\cite{pmlr-v99-cohen19a}. First we introduce a non-private generalization of the Cohen et al algorithm, yielding a $(1+γ)$-approximation of the JE problem while violating at most $κn$ constraints in $O(\log(1/κ)/γ)$ iterations. This variant works by projecting the intermediate weights assigned to the constraints onto the set of $κ$-dense distributions, similarly to~\cite{bun2020efficientnoisetolerantprivatelearning}.
  We then design a $ρ$-zCDP variant of this algorithm by adding Gaussian noise to the weighted covariance matrix aggregated in each step of the algorithm. Under a mild goodness assumption on the data we can assert that the resulting noisy matrix is close to the true matrix, thereby achieving essentially the same guarantee as the non-private algorithm provided sufficiently many input points. Thus our method achieves an efficient DP poly-time algorithm under concrete sample complexity bounds.
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier arxiv.org/abs/2609.32835
FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors

Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier

http://arxiv.org/abs/2609.32835

As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.
001
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Geometry-Adaptive Mechanisms for Private Synthetic Data Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian arxiv.org/abs/2609.33363
Geometry-Adaptive Mechanisms for Private Synthetic Data

Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian

http://arxiv.org/abs/2609.33363

Generating differentially private synthetic data with meaningful Wasserstein utility guarantees is challenging in high dimensions. For datasets of size \(n\) on $[0,1]^d$ with $d\ge2$, existing pure \(\varepsilon\)-differentially private mechanisms achieve expected $1$-Wasserstein error of order $(\varepsilon n)^{-1/d}$, reflecting the curse of dimensionality. While this rate is optimal in the worst case, it can be overly pessimistic when the data are supported on a lower-dimensional set. We formalize this through a multiscale packing-growth dimension $k$, which captures the geometric complexity of the support via the growth of packing numbers across scales. We propose \emph{Adaptive Pruned-PMM}, a pure $\varepsilon$-differentially private mechanism that combines private depth selection with our pruned variant of the Private Measure Mechanism (PMM) of He et al.\ (2023). The mechanism supports deeper, geometry-adapted hierarchies with expected running time $O\!\left(d(n+d)\log(\varepsilon n)\right)$, which is near-linear in $n$ for fixed dimension and privacy budget. Under an external multiscale packing-growth condition with dimension $k$, we show that, for fixed positive privacy budgets and fixed geometry, the expected $1$-Wasserstein error is of order $(\varepsilon n)^{-1/k}$ for $k>1$ as $n$ grows. We also prove a lower bound under a corresponding internal packing-growth condition, showing that the exponent $1/k$ is sharp within this framework.
001
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits Idan Attias, Orin Levy, Alexander Ryabchenko, Yishay Mansour, Uri Stemmer arxiv.org/abs/2609.33888
Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits

Idan Attias, Orin Levy, Alexander Ryabchenko, Yishay Mansour, Uri Stemmer

http://arxiv.org/abs/2609.33888

Can vanilla policy optimization explore enough to achieve near-optimal regret in stochastic contextual bandits? We show that standard exponential policy updates driven by offline regression do so under realizability, without exploration bonuses or importance weighting. For $A$ actions, $T$ rounds, and a finite prediction class $F$, vanilla PO achieves $\widetilde O(\sqrt{AT\log(|F|)})$ regret with high probability. Our analysis reveals an implicit exploration mechanism of independent interest: gradual policy updates prevent actions from losing probability too quickly, allowing the regression oracle to learn their expected losses. We further develop a batched version using only $O(\log T)$ regression calls and policy switches, and show how private regression oracles yield differentially private contextual bandit algorithms without composition across batches. For a finite class, this gives pure $\varepsilon_{\rm priv}$-DP and regret $\widetilde O\left( \sqrt{AT \log(|F|/δ)}(1+\varepsilon_{\rm priv}^{-1/2}) \right)$. Finally, experiments across oracle-based contextual bandit algorithms, with and without privacy, demonstrate the practical effectiveness of policy optimization and the value of explicit exploration under stronger privacy constraints.
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Contraction of Rényi Divergences for Discrete Channels Adrien Vandenbroucque, Amedeo Roberto Esposito, Michael Gastpar arxiv.org/abs/2609.34570
Contraction of Rényi Divergences for Discrete Channels

Adrien Vandenbroucque, Amedeo Roberto Esposito, Michael Gastpar

http://arxiv.org/abs/2609.34570

We investigate Strong Data-Processing Inequality (SDPI) constants for Rényi Divergences on finite spaces. We study their dependence on the Rényi order $α$, proving that they are non-decreasing and that their scaling by $(α-1)$ is convex for $α\geq1$. We also identify several support restrictions on the probability measures involved in determining these constants. In particular, in the distribution-independent setting, measures supported on a common set of at most two points suffice to evaluate the Rényi-SDPI constant. For $α\in[0,1]$, we further prove equality with the $χ^2$-SDPI constant, while at order infinity we obtain a closed-form expression. In order to link contraction over product spaces to contraction along individual coordinates, we provide tensorisation bounds for arbitrary product channels analogous to those known for $\varphi$-Divergences. At finite orders, the Rényi-SDPI constants are bounded above and below through comparisons with the $χ^2$-Divergence and Hellinger Divergences, with sharpness established in multiple cases. At order infinity, we instead relate these constants to the contraction of Total Variation Distance. Finally, our findings are applied to local differential privacy (LDP) and the analysis of Markov chains. This yields sharp contraction guarantees for pure-LDP mechanisms, and connects Rényi-LDP to Rényi-SDPI constants. For Markov chains, we derive finite-time convergence bounds and exhibit a family of chains for which Rényi-SDPIs improve on classical $χ^2$-based bounds by arbitrarily large factors.
000
Differential Privacy Papers @dppapers.bsky.social · 28/09/2026
Revisiting Certified Defense with Differential Privacy on Vision Transformers Jun Yan, Weiquan Huang, Qixian Zhang, Yan Bai, Shutai Zhang arxiv.org/abs/2609.31310
Revisiting Certified Defense with Differential Privacy on Vision Transformers

Jun Yan, Weiquan Huang, Qixian Zhang, Yan Bai, Shutai Zhang

http://arxiv.org/abs/2609.31310

Certified defenses that incorporate differential privacy have proven effective on Convolutional Neural Networks (CNNs), furnishing rigorous robustness guarantees against norm-bounded adversaries. However, the certified robustness behavior of Pixel Differential Privacy (PixelDP) remains largely unexplored with the self-attention architecture now dominating the deep-learning landscape. Given that the Transformer has a profound impact on our daily applications from the digital world to the physical world, it is crucial to study certified robustness through differential-privacy-style stability. To fill this research gap, we revisit this construction in Vision Transformers and identify a failure mode that is largely hidden in the convolutional setting. When noise is injected after the patch embedding, the Laplace mechanism with the inherited grouped $\ell_1$ sensitivity bound collapses to chance-level accuracy across noise scales, whereas the Gaussian mechanism remains trainable. This contrast isolates the source of failure: not the injected noise itself, but the geometry of the sensitivity constraint. We show that the attenuation induced by the inherited $Δ_{1,1}$ projection increases with layer width and kernel size according to a random-matrix scale $C/(\sqrt{M}+\sqrt{N})$. Replacing the $\ell_1$-type constraint with a spectral-norm constraint eliminates the collapse across datasets and architectures, but creates a fundamental obstacle: the repaired models no longer satisfy the sensitivity condition required by the standard Laplace certificate. We resolve this mismatch by deriving a dimension-free $(\varepsilon,\ δ)$-privacy guarantee for the Laplace mechanism under $\ell_2$ sensitivity through concentration of the privacy loss.
010
Differential Privacy Papers @dppapers.bsky.social · 28/09/2026
Gap-free Differentially Private PCA for Gaussian Data Alina Ene, Huy L. Nguyen arxiv.org/abs/2609.31614
Gap-free Differentially Private PCA for Gaussian Data

Alina Ene, Huy L. Nguyen

http://arxiv.org/abs/2609.31614

We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
000
Differential Privacy Papers @dppapers.bsky.social · 25/09/2026
DP-IPI: A Hybrid Differential Privacy Text Rewriting Mechanism for Indirect Personal Identifiers in Clinical Texts Ibrahim Baroud, Stephen Meisenbacher, Sebastian Möller, Florian Matthes, Roland Roller arxiv.org/abs/2609.29684
DP-IPI: A Hybrid Differential Privacy Text Rewriting Mechanism for Indirect Personal Identifiers in Clinical Texts

Ibrahim Baroud, Stephen Meisenbacher, Sebastian Möller, Florian Matthes, Roland Roller

http://arxiv.org/abs/2609.29684

Despite the strengths of modern anonymization and de-identification techniques, the risk of re-identification remains significant due to the indirect identifiers remaining in texts. To address this problem, recent works have applied text rewriting under Differential Privacy (DP) to prevent data linkage by perturbing texts via noise addition. Such methods privatize all tokens in a text indiscriminately, diminishing text quality and usability in critical domains such as in clinical settings. Focusing on indirect personal identifiers (IPIs), we introduce a utility-preserving DP text rewriting method that only privatizes spans containing IPIs. We show that our method effectively reduces re-identification risks in clinical texts while being producing more coherent and usable output texts, leading to higher privacy-utility trade-offs. In this, we demonstrate the effectiveness of hybrid text privatization, which leverages the promise of DP in an efficient, usable manner.
000
Differential Privacy Papers @dppapers.bsky.social · 25/09/2026
An Exposition of GPT Astra's Proof of Lower Bound on DP Continual Counting Jalaj Upadhyay arxiv.org/abs/2609.28528
An Exposition of GPT Astra's Proof of Lower Bound on DP Continual Counting

Jalaj Upadhyay

http://arxiv.org/abs/2609.28528

The goal of this note is to give a detailed proof, to the best of our understanding, of the recent presentation by Harrison and Leeman (arXiv:2609.17650v01 and arXiv:2609.17650v02) of the proof by Astra on the lower bound for differentially private continual counting. We believe a more natural and easy proof is possible and hope that this note will help in that effort.
  Prior to the initial preprint by Harrison and Leeman (arXiv:2609.17650v01), Bairaktari and Larsen (arXiv:2607.00876) gave an elegant proof to show a lower bound of $Ω(\log^{3/2}(n))$ for both pure and approximate-DP continual counting, and in personal communication had informed us that they have a proof of optimal $Ω(\log^{2}(n))$ for pure-differential private continual counting as well. They have subsequently published their $Ω(\log^{2}(n))$ bound, which is now a joint work of Bairaktari, Dahl, and Larsen (arXiv:2607.00876v3). Their new result is an elegant extension of their technique for approximate-differential privacy. Although the two proofs are technically different, the Astra argument uses related tree geometry introduced in Bairaktari and Larsen.
010
Differential Privacy Papers @dppapers.bsky.social · 25/09/2026
When Do Differentially Private Inputs Protect Graph Shift Operators? Andrew Campbell, Chenyue Zhang, Hang Liu, Victor Elvira, Anna Scaglione, Sean Peisert arxiv.org/abs/2609.28899
When Do Differentially Private Inputs Protect Graph Shift Operators?

Andrew Campbell, Chenyue Zhang, Hang Liu, Victor Elvira, Anna Scaglione, Sean Peisert

http://arxiv.org/abs/2609.28899

We study the differential privacy (DP) of a graph shift operator (GSO) when an analyst observes the output of a graph filter. In particular, we study the setting in which the input signals to the graph filter are drawn from a differentially private distribution. Unlike approaches that perturb the GSO or the filter output, we use the randomness already present in the inputs to protect the GSO. This yields an equivalent level of privacy protection to that of the perturbation methods without adding noise, and thus a better privacy-utility trade-off. We provide an explicit characterization of the privacy loss and its certificate in terms of the zeros of the graph filter. In doing so, we show that the log-likelihood ratio between the releases of two adjacent topologies is governed by the distances from each zero to the graph frequencies of the two GSOs. Then, by uniformly bounding the log-likelihood ratio over the adjacent topologies, we obtain an explicit $(\varepsilon,δ)$-DP guarantee for Gaussian inputs. We further show, via a Cramér--Rao bound, that the zero placement that limits the privacy loss also raises the floor on the adversary's reconstruction error. Finally, empirical validation is performed on a synthetic network of financial exposures, where the largest position a pair can conceal and the accuracy with which it can be sized are collinear across pairs. Both are set by the graph-frequency content of the pair, and the full network becomes recoverable only as the certified budget grows.
000
Differential Privacy Papers @dppapers.bsky.social · 25/09/2026
Forte: A sensitivity type system for imperative Rust Chiké Abuah arxiv.org/abs/2609.30254
Forte: A sensitivity type system for imperative Rust

Chiké Abuah

http://arxiv.org/abs/2609.30254

We introduce Forte, a sensitivity type system for Rust whose soundness rests on ownership. The graded sensitivity type systems, from Fuzz's linear grading to Solo's environment indices, are pure calculi: a claim about a value holds for the value's whole lifetime because nothing can mutate it. The imperative sensitivity analyses admit assignment to first-order variables and no references, so no question of aliasing arises in them. The programs that compute differentially private statistics in deployment are Rust, and they mutate through borrows. Forte closes this gap. Its central rules strongly update a sensitivity environment through an exclusive borrow, at a primitive call and across a checked function boundary; its soundness theorem is metric preservation over an operational semantics with a store, in which the exclusivity of &mut alone licenses framing across a mutating call, and two aliased borrows suffice to refute the theorem without it. Verus mechanizes the theorem, the function rule, and the refutation. Flux checks Forte as an ordinary library, with no fork of the compiler; a machine-checked theorem backs every deterministic primitive signature, and a correspondence theorem transports metric preservation to the programs the checker accepts. We evaluate Forte on mechanism kernels from OpenDP with genuine in-place mutation, matching the library's trusted stability maps with checked constants, covering the constructors that have no proof document, rejecting off-by-one diameters, tightened bounds, miscalibrated releases, and overspent budgets, and deriving one trusted constant as an inferred loop invariant.
001
Differential Privacy Papers @dppapers.bsky.social · 24/09/2026
What fidelity metrics miss: a structural check on synthetic educational data Hitoshi Inoue, Koichi Yasutake arxiv.org/abs/2609.27265
What fidelity metrics miss: a structural check on synthetic educational data

Hitoshi Inoue, Koichi Yasutake

http://arxiv.org/abs/2609.27265

Secondary use of educational records is increasingly mediated by platforms that share a differentially private synthetic version of a dataset and validate specific findings against the real data on request. The synthetic version is evaluated by comparing summary statistics of each variable, yet reported confirmation rates suggest that such comparisons do not predict which findings survive. We propose a structural check: the number of connected components of a weekly proximity graph over learners, tracked across a term. Across four annual cohorts of lower-secondary study-habit logs, the synthetic versions reproduced the level of this quantity and the shape of the weekly partition, but its variation across the term was between 2.6 and 4.9 times smaller than in the real data at a common working point, without exception, and those changes fell in different weeks: the synthetic cohorts single out the term's examination weeks and the real cohorts do not. We also show that a routine rule for setting the graph threshold makes naive comparisons between two datasets invalid, and illustrate this with an error of our own. The real curves are also distinguishable from marginal-preserving surrogates of themselves in all four cohorts, where three of the four synthetic ones are not, a comparison that needs no real data; these differences trace to what the generator was given.
000
Differential Privacy Papers @dppapers.bsky.social · 24/09/2026
Only Pay What You Must Spend: On-Demand Privacy Budget Payment for Differentially Private RAG Zhonghao Sun, Zhiliang Tian, Xinyue Fang, Shuo Ma, Juhua Zhang, Yiping Song, Dongsheng Li arxiv.org/abs/2609.27406
Only Pay What You Must Spend: On-Demand Privacy Budget Payment for Differentially Private RAG

Zhonghao Sun, Zhiliang Tian, Xinyue Fang, Shuo Ma, Juhua Zhang, Yiping Song, Dongsheng Li

http://arxiv.org/abs/2609.27406

Deploying large language models (LLMs) on sensitive data via Retrieval-Augmented Generation (RAG) introduces severe privacy risks. Recent studies apply Differential Privacy (DP) to LLMs with RAG for formal privacy guarantees. However, existing DP-RAG frameworks rapidly exhaust the privacy budget. Although recent efforts attempt to save the budget by narrowing the retrieval scope or sparsifying private generation, these methods themselves cumulatively consume the budget, whereas they could actually rely merely on public information or at a negligible one-time privacy cost. This mismatch fails to align budget expenditure with the model's actual reliance on private data, causing substantial waste on operations that require no private access. To address this, we propose SparsePay-RAG, adopting "only pay what you must spend" as its core principle. Using public information as a zero-privacy prior, it charges the privacy budget only for the private increment. Specifically, SparsePay-RAG narrows the retrieval scope via public topic-guided clustering, adaptively controls private access frequency without privacy cost through isotonic cross-layer trajectory fitting, and compresses per-access budget via DP contrastive decoding. Under strong privacy constraints, experiments show SparsePay-RAG achieves superior privacy-utility trade-offs over baselines.
010
Differential Privacy Papers @dppapers.bsky.social · 24/09/2026
Finite-Sample Binary Hypothesis Testing via Rényi Divergences: Strong Converse and Local Privacy Roberto Bruno, Adrien Vandenbroucque, Amedeo Roberto Esposito arxiv.org/abs/2609.27617
Finite-Sample Binary Hypothesis Testing via Rényi Divergences: Strong Converse and Local Privacy

Roberto Bruno, Adrien Vandenbroucque, Amedeo Roberto Esposito

http://arxiv.org/abs/2609.27617

We study asymmetric simple binary hypothesis testing between $H_0:P_0^{n}$ and $H_1:P_1^{n}$, based on $n$ independent and identically distributed observations. Leveraging a variational representation of Rényi divergence of order $α$, we derive our main result: a finite-sample converse with $α>1$. The bound uses both directions of the divergence $D_α(P_1\|P_0)$ and $D_α(P_0\|P_1)$, tensorises under product measures, and contains familiar data-processing converses as boundary cases. For comparison, we apply the same variational approach to general $f$-divergences and specialise it to total variation, $E_γ$, Hellinger, and Kullback Leibler divergences, thereby recovering familiar converses within a unified framework. Together with an achievability bound involving Rényi divergence with $α\in (0,1)$, the main converse recovers the phase transition of the optimal Type II error under the exponentially decaying Type I error constraint $\varepsilon_n=e^{-nr}$. Under regularity conditions, the optimal Type II error vanishes exponentially when $r<D(P_1\|P_0)$ and converges exponentially fast to one when $r>D(P_1\|P_0)$. We also derive sample-complexity bounds and extend both the converse and achievability analyses to locally differentially private observations, quantifying the cost of privacy and recovering the non-private achievability bound as the privacy constraint vanishes.
000
Differential Privacy Papers @dppapers.bsky.social · 24/09/2026
Contraction and Statistical Inference under Privacy for Uniformly Bounded Distributions Leonhard Grosse, Sara Saeidian, Tobias J. Oechtering, Mikael Skoglund arxiv.org/abs/2609.28297
Contraction and Statistical Inference under Privacy for Uniformly Bounded Distributions

Leonhard Grosse, Sara Saeidian, Tobias J. Oechtering, Mikael Skoglund

http://arxiv.org/abs/2609.28297

We investigate $c$-interior pointwise maximal leakage (PML) as a tool for contraction analyses and disclosure control. Based on the strong adversarial threat models from maximal leakage, $c$-interior PML generalizes local differential privacy (LDP) to data-generating distributions with densities uniformly bounded away from zero by $c>0$. Viewing $c$-interior PML as an algebraic constraint on a kernel yields more flexible (and often tighter) contraction analyses than standard LDP. We provide tight bounds on the Dobrushin coefficient, and bound the contraction coefficient of the Hockeystick-divergence. We further derive strong data processing inequalities on $f$-divergences under $c$-interior PML constraints when the input distributions to the divergence are restricted to be in the $c$-interior. These results extend beyond the regime of pure LDP to cover a larger class of kernels, including, e.g., arbitrary stochastic matrices. We apply the results to minimax theory and provide asymptotically optimal strategies under $c$-interior PML constraints for binary hypothesis testing and mean estimation. The results show that disclosure control with PML allows analysts to reason about systems in a more differentiated manner: For example, it allows us to quantify the privacy leakage of deterministic systems, and can give precise adversarial guarantees with respect to arbitrary distributional assumptions. Interestingly, a recurring theme in the disclosure analyses is that if the privacy problem is relatively regular (if the density bound $c$ is large), private inference can be possible without incurring any additional cost in terms of sample complexity.
000
Differential Privacy Papers @dppapers.bsky.social · 23/09/2026
Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy Anming Gu, Syamantak Kumar, Kevin Tian, Chutong Yang arxiv.org/abs/2609.26508
Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy

Anming Gu, Syamantak Kumar, Kevin Tian, Chutong Yang

http://arxiv.org/abs/2609.26508

Streaming principal component analysis (PCA) seeks to recover a leading spectral subspace in a single pass over a data stream. We give a new analysis of the ubiquitous Oja's algorithm [Oja82] for the most general, gap-free variant of this problem, where no eigengap assumptions are made on the underlying mean matrix, complemented by a nearly-matching lower bound. Prior works achieving near-optimal rates for streaming PCA either required gap assumptions [JJK+16, HNWW21], or were limited to rank-one updates [AZL17, Lia23]. Our proof only uses a second moment bound on the individual stochastic updates, bypassing the almost sure bounds needed by prior near-optimal analyses, and the analogous offline matrix Bernstein bound. We also extend our result to a Rayleigh quotient notion of approximate PCA, addressing an open question of [JJK+16]. As our main application, we give gap-free differentially private PCA guarantees for sub-Gaussian data, settling Conjecture 1.1 of [Bro26] up to logarithmic factors.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Differentially Private and Fairness-Audited Score Diffusion for Irregular Longitudinal Health Records Taimoor Ahmad arxiv.org/abs/2609.22401
Differentially Private and Fairness-Audited Score Diffusion for Irregular Longitudinal Health Records

Taimoor Ahmad

http://arxiv.org/abs/2609.22401

Sharing irregular longitudinal health records can accelerate model development, yet synthetic releases may leak participation, distort temporal dependence, suppress rare events, or reduce utility for underrepresented groups. We present TRUST LONGSYNTH, an auditable patient level private generator that combines bounded sufficient statistics, zCDP accounted Gaussian releases, conditional analytic score diffusion, block banded temporal covariance, separate missingness and gap models, and a protected event sampling floor with population weights.
  The method was evaluated on five independently generated, three cohort benchmarks containing 720 patients, fourteen irregular observation slots, six mixed variables, informative missingness, and a rare deterioration outcome. At epsilon = 12 and delta = 10 to the power of minus 5, TRUST LONGSYNTH achieved mean train synthetic test real AUPRC 0.342, Brier score 0.088, expected calibration error 0.082, correlation error 0.222, autocorrelation error 0.317, and membership attack AUROC 0.499.
  Relative to the private diagonal score baseline, AUPRC increased by 7.5 percent, while Brier, calibration, correlation, and autocorrelation errors decreased by 6.1 percent, 17.0 percent, 28.0 percent, and 30.1 percent, respectively. The method did not dominate every nonprivate or discrete baseline, and corrected paired tests were inconclusive with five seeds. Canary exposure was 1.8 percent, compared with 28.8 percent for DP Score in the same stress test.
  These findings support a transparent privacy utility fairness evaluation protocol, not clinical validity or unconditional release safety, and motivate governed external validation on real multi site records.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Locally Private Inference for Riemannian Stochastic Optimization Xiaotian Chang, Yangdi Jiang, Qirui Hu arxiv.org/abs/2609.22642
Locally Private Inference for Riemannian Stochastic Optimization

Xiaotian Chang, Yangdi Jiang, Qirui Hu

http://arxiv.org/abs/2609.22642

We develop inference for manifold-valued population minimizers when each observation belongs to a different participant and only locally private messages reach the analyst. The method releases randomized tangent gradients and combines them through Riemannian stochastic approximation and Polyak-Ruppert averaging. Directly inserting a private data surrogate into a nonlinear loss can shift its population target, whereas conditional centring of the released gradient preserves the first-order equation. We introduce symmetric-pair regression (SPR) to estimate the asymptotic variance from the same private messages used for point estimation, without holding out participants or requesting a second release. We prove the central limit theorem and consistency of the fully transcript-based sandwich covariance and intrinsic Wald region under local differential privacy. Simulations across various statistical problems and manifolds support the predicted decrease in estimation error and near-nominal coverage under moderate privacy. An application to NHANES anthropometric data illustrates private estimation of a leading body-size direction and its uncertainty.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Improved Private Sparse Covariance Estimation with Multiscale Threshold Tests Zihan Zhang arxiv.org/abs/2609.22783
Improved Private Sparse Covariance Estimation with Multiscale Threshold Tests

Zihan Zhang

http://arxiv.org/abs/2609.22783

We study differentially private covariance estimation in operator norm for mean-zero sub-Gaussian distributions with unknown covariance support and at most $k$ nonzero entries per row. We develop a multiscale random-threshold algorithm with sample complexity $\ot(k^2/α^2+k\sqrt d/(α\varepsilon))$ for $(\varepsilon,δ)$-differential privacy and error at most $ασ^2$, where $d$ is the dimension and $σ$ is a known sub-Gaussian scale. The bound improves the privacy-dependent term of the existing $\ot(k^2/α^2+k^{3/2}\sqrt d/(α\varepsilon))$ \citep{kumar2026curse} upper bound by a factor of $\sqrt k$, and matches the lower bound of $\widetildeΩ(k^2/α^2 + k\sqrt{d}/(α\varepsilon))$ in its applicable parameter regime.
  Our key technical ingredient is a direct operator-norm bound on the centered fluctuations of an ideal reconstruction, exploiting conditional independence rather than accumulating entrywise errors across each row. A multiscale allocation of threshold tests balances reconstruction variance against query sensitivity. Together, these ingredients sharpen the trade-off between approximation error and privacy protection, removing the additional $\sqrt{k}$ factor from the privacy-dependent sample complexity.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated Study Márton Pál Lipcsey-Magyar, Adrian Pekar arxiv.org/abs/2609.23521
Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated Study

Márton Pál Lipcsey-Magyar, Adrian Pekar

http://arxiv.org/abs/2609.23521

Residential traffic classification supports service management, but learning across homes must account for heterogeneous traffic and privacy constraints. Privacy-aware training may impose uneven costs across traffic categories. We study this tradeoff in simulated two-client federated learning using 1.62 million preprocessed gateway-collected flows across six categories. We compare a full-feature baseline, feature suppression (FS), and differentially private stochastic gradient descent (DP-SGD) under one fixed record-level privacy setting. FS-mild excludes four timing features from 16 model inputs; it provides no formal privacy guarantee. With size-proportional aggregation, FS-mild achieves higher combined macro-F1 and worst-group F1 (the minimum per-class F1 across homes) than DP-SGD in all five seeds at both model capacities under stratified and temporal splits. The tested DP-SGD configuration incurs pronounced minority-category losses, especially in the smaller home, but FS-mild does not uniformly improve on the full-feature baseline. On stratified-split models, loss-based and shadow-model membership probes show near-chance aggregate discrimination without a consistent ranking across probes; this does not establish equivalent privacy. These findings support FS as an input-minimization baseline, not a substitute for formal privacy.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Pattern-level Differential Privacy for High-utility Complex Event Processing He Gu, Thomas Plagemann, Vera Goebel, Maik Benndorf, Boris Koldehofe arxiv.org/abs/2609.23827
Pattern-level Differential Privacy for High-utility Complex Event Processing

He Gu, Thomas Plagemann, Vera Goebel, Maik Benndorf, Boris Koldehofe

http://arxiv.org/abs/2609.23827

Current privacy-preserving mechanisms (PPMs) in Complex Event Processing (CEP) systems are unnecessarily restrictive, reducing the utility of data received by data consumers. This article presents a novel approach to preserve privacy in CEP systems, improving the utility of detected event patterns by dynamically adapting the noise added to an unprotected data stream. We introduce a new guarantee named pattern-level differential privacy (DP), which enables us to apply and compare the strength of PPMs at the pattern level. We propose new pattern-level PPMs yielding pattern-level DP and analyze different trust settings of these PPMs and their requirements for context knowledge in the CEP system, e.g., the deployed queries. Our evaluation is based on three datasets (two real-world, one synthetic) and shows that the proposed PPMs increase data utility while preserving the same privacy level as the state-of-the-art PPMs. We use simulations to study the performance of our proposed PPMs in various practical scenarios. Furthermore, we demonstrate that computational complexity is not an obstacle to deployment.
000
Differential Privacy Papers @dppapers.bsky.social · 21/09/2026
Asymptotic Anytime-Valid Quantile Inference under Local Differential Privacy Leheng Cai, Qirui Hu, Shuyuan Wu arxiv.org/abs/2609.21338
Asymptotic Anytime-Valid Quantile Inference under Local Differential Privacy

Leheng Cai, Qirui Hu, Shuyuan Wu

http://arxiv.org/abs/2609.21338

Sequential quantile inference is difficult under local differential privacy because every record is randomized before reaching the analyst and the limiting quantile variance depends on an unknown density. We develop an online procedure that combines randomized response with dynamically chained parallel stochastic gradient descent (P-SGD). The resulting Polyak--Ruppert estimator admits a strong Gaussian approximation. A cross-chain quadratic statistic, computed entirely from private iterates, consistently estimates the limiting variance without a separate online density estimator. These results yield asymptotic confidence sequences and, under polynomial chain growth, asymptotic time-uniform coverage. Arm-wise constructions support locally private quantile best-arm identification, time-uniform simple-regret bounds, and sequential A/B tests of quantile treatment effects. Simulations and salary-data analyses illustrate the finite-sample behavior and practical use of the proposed methods.
000
Differential Privacy Papers @dppapers.bsky.social · 21/09/2026
Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees Shuo Huang, Gholamreza Haffari, Xingliang Yuan, Ting Yu, Lizhen Qu arxiv.org/abs/2609.21340
Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees

Shuo Huang, Gholamreza Haffari, Xingliang Yuan, Ting Yu, Lizhen Qu

http://arxiv.org/abs/2609.21340

Empirical identity leakage from released text is increasingly driven by attackers that combine large language models (LLMs) with auxiliary knowledge to link documents to individuals. Existing audits typically report success rates for specific attack pipelines but lack finite-sample statistical guarantees, while training-time protections such as differential privacy are difficult to translate into release-time decisions for individual natural-language documents. We introduce Conformal Privacy Auditing(CPA), a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries. CPA outputs a conformal ambiguity set of candidate identities that is guaranteed to contain the true identity with user-chosen confidence under exchangeability, together with an interpretable leakage proxy derived from set size. CPA supports both logit-access and sampling-only attackers, enabling audits of open-source models and proprietary API models in a unified framework. Across multiple release benchmarks and attacker configurations, CPA achieves calibrated coverage and reveals sharp shifts in certified identifiability as auxiliary knowledge, LLM augmentation, and release mechanisms vary, providing a statistically grounded basis for reporting and comparing release-time linkage risk across attacker configurations, datasets, and release mechanisms alike.
000
Differential Privacy Papers @dppapers.bsky.social · 18/09/2026
Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs Li Ge, Wenjie Qu, Weitao Feng, Yi Zeng, Jiaheng Zhang, Xiaofeng Wang, Wei Dong arxiv.org/abs/2609.20532
Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs

Li Ge, Wenjie Qu, Weitao Feng, Yi Zeng, Jiaheng Zhang, Xiaofeng Wang, Wei Dong

http://arxiv.org/abs/2609.20532

Wide adoption of machine learning has created growing policy and regulatory demand for protecting sensitive training data, with differential privacy (DP) emerging as a key mechanism. Yet a less-studied problem is how to certify the faithful execution of DP during training: an external verifier should be able to check that a released model was trained with proper DP protection, without accessing the private training data. Existing cryptographic approaches, such as zero-knowledge proofs, provide strong guarantees but often incur prohibitive overhead, in some cases by orders of magnitude. Trusted Execution Environments (TEEs) offer a more efficient alternative, but the multi-GPU TEE support needed for training and fine-tuning large language models remains limited to recent platforms and is absent or inefficient on legacy GPUs.
  To address this, we propose a practical framework for verifiable DP training using CPU-side TEEs together with untrusted GPUs. Our design addresses a fundamental efficiency-security tension: training entirely inside a CPU TEE is too slow, while unrestricted GPU offloading can allow malicious deviations from DP. We therefore offload expensive gradient computation to GPUs, while using the CPU TEE to efficiently verify the correct enforcement of DP on gradients through probabilistic checking. Our framework detects frequent full deviations from DP with high probability; for the utility-oriented forged-gradient attacks evaluated in this work, sparse deviations provide limited utility benefit and show no measurable additional membership leakage. Experiments further show that our approach nearly achieves a ``free lunch'': it incurs only modest overhead compared with standard GPU-based DP training, while effectively constraining malicious deviations from the
000
Differential Privacy Papers @dppapers.bsky.social · 18/09/2026
Empirical Analysis of Randomness Quality in Differential Privacy Mechanisms Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta, Massimo Caccia arxiv.org/abs/2609.20561
Empirical Analysis of Randomness Quality in Differential Privacy Mechanisms

Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta, Massimo Caccia

http://arxiv.org/abs/2609.20561

Differential Privacy (DP) relies on carefully calibrated random noise to protect individual privacy in statistical analyses. While theoretical work has analyzed DP under weakened randomness assumptions, the practical consequences of entropy degradation remain poorly understood. We present a systematic empirical investigation of how randomness quality affects differential privacy mechanisms using IBM's DiffPrivLib. We introduce progressively degraded entropy sources characterized by established test suites, starting from high-quality quantum True Random Number Generators (TRNGs) and cryptographically secure Pseudo-Random Number Generators (PRNGs) down to systematically manipulated sources with controlled entropy degradation. Through repeated experiments over one million queries on a reference database and complementary statistical tests, we directly analyze empirical Privacy Loss Random Variable distributions. Our results demonstrate that DP mechanisms reliably detect deviations when approximately 1 bit in every 8 to 16 is manipulated, with detection sensitivity varying significantly between bit-level biases and temporal correlations. We demonstrate that statistical detection of distributional anomalies does not necessarily correspond to actual privacy guarantee violations.
000
Differential Privacy Papers @dppapers.bsky.social · 17/09/2026
Intrinsic-Dimensional Wasserstein Guarantees for Private Synthetic Measures Yiyun He arxiv.org/abs/2609.17624
Intrinsic-Dimensional Wasserstein Guarantees for Private Synthetic Measures

Yiyun He

http://arxiv.org/abs/2609.17624

We study an $\varepsilon$-differentially private synthetic measure for $n$ points in $[0,1]^d$ by applying the existing PrivTree algorithm to construct an adaptive binary partition and then privately releasing its leaf masses. We consider the worst-case data model without any sampling or population-distribution assumption. The 1-Wasserstein error of the synthetic measure is $\widetilde O_d((\varepsilon n)^{-1/d})$ for $d\ge2$, which is optimal compared to the minimax lower bound up to a logarithmic factor.
  Moreover, for $d\ge3$ and $2<s\le d$, if the data set has covering number at most $Ar^{-s}$ over the relevant finite range of scales $r$, the expected error improves to $\widetilde O_{d,s}((\varepsilon n)^{-1/s})$. Thus the rate depends on a finite-scale intrinsic dimension rather than the ambient dimension, without requiring the recovery of a low-dimensional manifold. We also introduce a shifting technique to further avoid the exponential dependence of the constant on the ambient dimension $d$.
000
Differential Privacy Papers @dppapers.bsky.social · 17/09/2026
Tight Lower Bounds for Differentially Private Continual Counting Charlie Harrison, Ethan Leeman arxiv.org/abs/2609.17650
Tight Lower Bounds for Differentially Private Continual Counting

Charlie Harrison, Ethan Leeman

http://arxiv.org/abs/2609.17650

The Binary Tree Mechanism is a standard algorithm for differentially private continual counting, but its asymptotic optimality under pure differential privacy has remained unresolved since its introduction. We resolve this question. For fixed $0 < \varepsilon \le 1$, we prove asymptotically tight lower bounds of $Ω(\log^2 n)$ for worst-case expected $\ell_\infty$ error and $Ω(\log^3 n)$ for mean and maximum per-coordinate expected squared error. These bounds hold for arbitrary mechanisms, even when the entire stream is available in advance. The same lower bounds hold under approximate differential privacy whenever $δ\le n^{-c}$, for any fixed $c>0$. Our lower bounds match the Binary Tree Mechanism instantiated with Laplace noise, establishing its asymptotic optimality under both pure differential privacy and approximate differential privacy in the standard regime of $δ\ll1/n$. Our proof uses a single hard distribution with a bounded exponential score on a tree. A simple modification of the score allows the same framework to establish tight lower bounds for all three error measures.
020
Differential Privacy Papers @dppapers.bsky.social · 17/09/2026
QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing Shuaiqi Wang, Zinan Lin, Giulia Fanti arxiv.org/abs/2609.17995
QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing

Shuaiqi Wang, Zinan Lin, Giulia Fanti

http://arxiv.org/abs/2609.17995

Natural-language datasets support many downstream applications and research studies, but releasing text can reveal sensitive global properties of the underlying data source, such as the proportion of records associated with a particular gender, diagnosis, or political stance. Existing work has largely focused on property inference attacks that recover such global properties, while defenses for protecting these dataset-level secrets remain limited. Differential privacy, although effective for protecting individual records, provides only weak protection for aggregate properties. We propose Randomized Quantization for Text (QuanText), a training-free and large-language-model-agnostic data release mechanism that protects global secrets in textual datasets while preserving data utility. Given a dataset-level secret, such as the proportion of records with a particular diagnosis, and attributes whose utility should be preserved, such as topic and sentiment, QuanText perturbs both the secret distribution and the distributions of correlated attributes. It does so by constructing candidate release distributions over secret and non-secret attributes, randomly selecting a candidate sufficiently close to the private empirical distribution, and rewriting each private text sample to match the selected distribution using attribute-related snippets from the original text. QuanText is inspired by the Statistic Maximal Leakage (SML) framework, which bounds leakage about a secret function of a data distribution. Under idealized conditions, we show that QuanText satisfies an SML guarantee. Since these conditions may not hold exactly in practice, we also evaluate QuanText empirically on real-world datasets. Our results show that QuanText achieves a better empirical privacy-utility trade-off than competing data generation baselines.
000
Differential Privacy Papers @dppapers.bsky.social · 17/09/2026
Low-Rank Masking for Single-Server Matrix Multiplication Alejandro Cohen, Rafael G. L. D'Oliveira, Alex Sprintson arxiv.org/abs/2609.18876
Low-Rank Masking for Single-Server Matrix Multiplication

Alejandro Cohen, Rafael G. L. D'Oliveira, Alex Sprintson

http://arxiv.org/abs/2609.18876

We study the statistical privacy of outsourcing matrix multiplication over a finite field ${\mathbb F_q}$ to a single server using additive masks of rank at most $r$. For independent uniform $n\times n$ inputs, we show that uniform \emph{rank-ball masks} and products of independent uniform factors give maximal-correlation secrecy of at most $q^{-r}$ against the complete server view, with $O(n^2r)$ field operations for encoding and decoding. This secrecy captures how effectively the server is prevented from estimating functions of the inputs. We prove an asymptotically matching lower bound of this secrecy measure for $r=o(n)$, showing that both sampling methods are asymptotically optimal among input-independent additive masks of rank at most $r$, even when secret invertible transformations are allowed. We also characterize the posterior distribution for uniform rank-ball masks under arbitrary joint input distributions and prove approximate individual security for rows and columns under independent uniform inputs. Finally, we show that every input-independent additive mask of rank at most $r=o(n)$ requires $δ\to1$ in entry-level $(\varepsilon,δ)$-differential privacy for fixed field size $q$ and bounded $\varepsilon$.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
CANAL: Channel-Aware Noise Allocation for Differentially Private Feature Distillation in Medical Image Segmentation Armaghan Butt, Shuya Feng, Qing Tian arxiv.org/abs/2609.13271
CANAL: Channel-Aware Noise Allocation for Differentially Private Feature Distillation in Medical Image Segmentation

Armaghan Butt, Shuya Feng, Qing Tian

http://arxiv.org/abs/2609.13271

Medical image segmentation needs diverse training data, but hospitals hold complementary scans they cannot share for privacy and regulatory reasons. Knowledge distillation can bridge this gap by exporting learned feature representations instead of images, but those representations still encode patient-specific anatomy and remain vulnerable to membership-inference and feature-inversion attacks. Adding calibrated Gaussian noise restores a differential-privacy guarantee, yet three issues have been overlooked. First, prior DP feature-distillation pipelines re-sample noise at every student iteration, so each patient image is released many times and the privacy cost composes over those releases, growing by orders of magnitude. We present a sample-once-per-image release, realized by a single precomputation pass, under which each patient contributes one release. Second, uniform noise is wasteful because channels differ in task importance. Using task-gradient energy as the importance measure, we derive CANAL, a closed-form water-filling allocation that gives important channels proportionally less noise, and prove it strictly minimizes importance-weighted distortion at a fixed budget. Third, the clipping caps and importance scores that drive the allocation are themselves data-dependent, so releasing them in the clear silently breaks the guarantee. We give a DP-honest budget split that charges each to the privacy budget, so the reported epsilon is the true epsilon. Across three medical segmentation benchmarks spanning dermoscopy, colonoscopy, and ultrasound, CANAL retains more task-relevant signal than uniform noise at the same privacy budget.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Scalable Discrete-to-Continuous Channel Simulation for Compression and Privacy Joseph Rowan, Buu Phan, Ashish J. Khisti arxiv.org/abs/2609.12067
Scalable Discrete-to-Continuous Channel Simulation for Compression and Privacy

Joseph Rowan, Buu Phan, Ashish J. Khisti

http://arxiv.org/abs/2609.12067

Channel simulation has recently emerged as a useful component in machine learning systems where samples from a prescribed probability distribution are to be compressed. Yet, general channel simulation algorithms often suffer from high computational costs, random stopping times or, in the worst case, can require generating an infinite number of shared random samples. We introduce a scheme for both exact and approximate simulation of discrete-to-continuous channels which conversely uses a fixed number of random samples, and therefore has a runtime independent of the channel and the input. Unlike existing channel simulation schemes which generate a sequence of independent samples from a proposal distribution, our approach generates one sample, or alternatively a fixed number of samples, from each potential target distribution. We then apply a latent permutation to the samples before performing sample selection using an exponential race. Our scheme provides a flexible tradeoff between the number of generated samples and the compression rate. Using polar and multilevel coding, we scale our approach to handle long blocklengths in $O(n \log n)$ time in order to benefit from reduced per-symbol overhead. We conclude by demonstrating applications to variable-rate compression with stochastic VQ-VAEs and communication-efficient differentially private distributed mean estimation via exact simulation of the Gaussian mechanism.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
I Am No One: Style-Aware Paraphrasing for Text Anonymization Ahmed Sohair Khan, Estrid He, Monica Wachowicz, Elham Naghizade arxiv.org/abs/2609.12341
I Am No One: Style-Aware Paraphrasing for Text Anonymization

Ahmed Sohair Khan, Estrid He, Monica Wachowicz, Elham Naghizade

http://arxiv.org/abs/2609.12341

Authorship attribution models can re-identify users from seemingly anonymized text by exploiting stable stylistic fingerprints, even after explicit identifiers are removed, posing a growing privacy risk for text publishing and analytics. This risk extends to speech-derived text such as ASR transcripts of meetings and call-center conversations, where stylometric leakage can persist even after acoustic anonymization. Differential privacy-based anonymization often severely degrades text quality and utility. We propose a style-aware, prompt-driven anonymization approach that uses pretrained large language models to construct compact stylistic profiles from minimal samples and rewrite text to suppress identifiable style markers while preserving meaning. Across blog and review datasets, our approach reduces authorship attribution F1 by 60-70% while maintaining content quality and readability, substantially outperforming DP-based and non-DP baselines.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
A Differentially Private Federated Proximal Optimization Framework for Customer Churn Prediction in Heterogeneous Federated Telecom Networks Joydeb Kumar Sana, Subrata Chakraborty, M M Manjurul Islam arxiv.org/abs/2609.12470
A Differentially Private Federated Proximal Optimization Framework for Customer Churn Prediction in Heterogeneous Federated Telecom Networks

Joydeb Kumar Sana, Subrata Chakraborty, M M Manjurul Islam

http://arxiv.org/abs/2609.12470

Customer churn is one of the major issues in the telecommunication industry. To predict customer churn, conventional centralized machine learning approaches have been widely used. This centralized approach requires customer data to be stored in a central repository, which raises privacy concerns and may violate data protection regulations. Federated learning addresses this problem by allowing multiple telecom operators to collaboratively train a global model without transferring their raw customer data. However, real-world customer data are often heterogeneous (non-IID), which may negatively affect the performance of standard federated learning. Trained models can also suffer from privacy attacks. To address those issues, we propose a Differentially Private (DP) based Federated Proximal optimization (FedProx) framework. All experiments were performed on two publicly available telecom churn datasets. We trained Federated Averaging (FedAvg), DP-FedAvg, FedProx, and the proposed DP-FedProx framework. For baseline comparison, we also used several centralized and local models. To evaluate the models, we employed seven widely used evaluation metrics. The experimental results show that the FedProx based models consistently outperform the FedAvg based models. Compared with the best centralized model, the proposed DP-FedProx framework achieves competitive prediction performance with only a small reduction in accuracy while providing privacy guarantees. To explain our model, we conducted SHAP analysis which shows that DP-FedProx method priorities revenue group features. These results indicate that the proposed DP-FedProx framework provides a practical balance between prediction performance and data privacy protection.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Differential Privacy Meets Fixed Parameter Tractability: Algorithms and Lower Bounds Pritish Kamath, Ravi Kumar, Pasin Manurangsi arxiv.org/abs/2609.12508
Differential Privacy Meets Fixed Parameter Tractability: Algorithms and Lower Bounds

Pritish Kamath, Ravi Kumar, Pasin Manurangsi

http://arxiv.org/abs/2609.12508

We study combinatorial optimization problems under the constraint of $ε$-differential privacy ($ε$-DP). Given the strong lower bounds for explicitly outputting solutions, we work within the implicit representation framework of Gupta et al. (SODA 2010), where a private polynomial-time randomized "encoder" generates a representation of a solution, and a "decoder" uses this representation along with the input to extract a valid final solution.
  In this work, we generalize this framework by allowing the encoder to run in fixed-parameter tractable time. This circumvents approximation barriers inherent to polynomial-time algorithms and obtains improved guarantees for many fundamental combinatorial optimization problems.
  Finally, we establish the first representation-independent lower bounds for our framework. Assuming a non-uniform variant of the Gap Exponential Time Hypothesis, for sufficiently small $ε> 0$, we prove that no $ε$-DP encoder-decoder pair can achieve certain approximation guarantees, if the decoder runs in subexponential time. We further provide representation-dependent lower bounds that hold even for larger $ε$.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Shuffling is Not Enough: Breaking Permutation-Based Model Confidentiality in Hybrid FHE Inference Jiseung Kim, Hyung Tae Lee arxiv.org/abs/2609.12911
Shuffling is Not Enough: Breaking Permutation-Based Model Confidentiality in Hybrid FHE Inference

Jiseung Kim, Hyung Tae Lee

http://arxiv.org/abs/2609.12911

Hybrid fully homomorphic encryption~(FHE) inference improves the practicality of private inference by letting the server evaluate linear layers homomorphically while the client decrypts and applies nonlinearities. Recent schemes attempt to protect model confidentiality by returning noisy, output-permuted responses and appealing to shuffle-model differential privacy~(DP). We show that this protection fails in the correctness regime required by hybrid FHE systems. For a $d$-input linear layer, $d+1$ admissible queries suffice for exact recovery of a permutation-invariant layer summary, hence for perfect model distinguishability. We further show that input DP is orthogonal to model confidentiality and that the local-DP premise required for shuffle amplification cannot hold under correctness-bounded noise. We recover all linear layers of a \safhire{}-style ResNet-20 end-to-end from TFHE transcripts with zero error, using $d+1$ queries per layer for a total of $5{,}712$ direct queries. Under the same query model, we also confirm exact per-layer recovery on pretrained ImageNet-scale CNNs and ViT-B/16. The leaked spectra enable fingerprinting, lineage attribution, and improved logit-based extraction, while suppressing them destroys inference utility.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
High quantum local differential privacy breaks entanglement Sujeet Bhalerao, Theshani Nuradha, Felix Leditzky arxiv.org/abs/2609.13418
High quantum local differential privacy breaks entanglement

Sujeet Bhalerao, Theshani Nuradha, Felix Leditzky

http://arxiv.org/abs/2609.13418

Differential privacy provides a mathematical framework for guaranteeing privacy for sensitive data. In quantum information processing, the interaction of privacy constraints with quantum resources such as entanglement remains a question of interest. Given that the utility of many protocols, and often the presence of a quantum advantage, relies on quantum resources such as entanglement, it is crucial to understand when a privacy requirement for a quantum channel is compatible with the channel's ability to preserve entanglement. We study this question for quantum local differential privacy (QLDP). Our main result shows that every $\varepsilon$-QLDP channel with a $d$-dimensional input is entanglement-breaking whenever $\varepsilon\leq\log\frac{d}{d-1}$. We also prove an approximate version for $(\varepsilon,δ)$-QLDP, where channels in the same high-privacy regime are close in diamond norm to an entanglement-breaking channel. We further prove a composition result for a collection of private quantum channels having entangled inputs and global measurements in the high-privacy regime. Finally, we apply our results to private quantum learning theory. We prove that any learning protocol using arbitrary quantum memory on copies of the output of an entanglement-breaking channel can be simulated by a protocol that measures the corresponding unprocessed input copies one at a time while storing only classical information. Combining this result with our high-privacy entanglement-breaking theorem, we show that under sufficiently private local noise, a learning protocol with quantum memory for purity testing and bipartite product testing is subject to the sample complexity lower bounds for protocols with single-copy measurements on the noiseless tasks. We also obtain stronger sample complexity lower bounds when a single highly private chan
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Canaries in the Bank: Auditing User-Level Privacy in Private Evolution Sai Aparna Aketi, Enayat Ullah, Shripad Gade arxiv.org/abs/2609.13499
Canaries in the Bank: Auditing User-Level Privacy in Private Evolution

Sai Aparna Aketi, Enayat Ullah, Shripad Gade

http://arxiv.org/abs/2609.13499

Private Evolution (PE) generates high-fidelity synthetic data in federated settings without exposing users' raw data. It aggregates clipped user votes over a shared candidate bank into a differentially private histogram, with noise calibrated to the worst-case user contribution. However, it is unclear whether an adversary can realize this worst-case privacy loss while following the PE protocol. We introduce a protocol-aware empirical audit in which the server commits to a single shared candidate bank and replaces roughly 1% of its entries with probes derived from a known, non-private canary. We evaluate eight attacks, including an unchanged-bank baseline, exact copies, plausible paraphrases, and high-entropy synthetic nonces. Experiments on Yelp and Sentiment140 show that natural-text attacks remain substantially below the theoretical DP bound, while nonce-based attacks yield considerably stronger bounds and come closest to the mechanism's privacy ceiling. These results quantify the gap between formal worst-case privacy and leakage achievable through protocol-valid candidate-bank manipulation.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
A Graph-Based Framework for Extending Metric Differential Privacy Mechanisms Ruiyao Liu, Chenxi Qiu arxiv.org/abs/2609.14125
A Graph-Based Framework for Extending Metric Differential Privacy Mechanisms

Ruiyao Liu, Chenxi Qiu

http://arxiv.org/abs/2609.14125

Metric differential privacy (mDP) is well suited to structured secret domains, but directly constructing utility-aware mechanisms over large or fine-grained domains is often computationally prohibitive. We study extension-based mDP design, where a mechanism is first specified on a finite set of seed records and then extended to a larger target domain. To our knowledge, this is the first work to systematically formulate extension as a general design paradigm for mDP rather than a method-specific construction. We present a graph-based extension framework, identify three requirements for correctness, local mDP constraints, overlap consistency, and successor-level mDP preservation, and show that, under these conditions, the induced global mechanism is well defined and satisfies $ε$-mDP on the target domain. We further instantiate the framework with a tree-based extension algorithm for multi-resolution grids, where multi-dimensional extension is realized through one-dimensional interpolation and dimension-wise composition. Experiments on road-map datasets demonstrate that our approach achieves a strong utility-scalability trade-off while preserving exact mDP guarantees.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
PixCrypt: Fast Fine-Grained FHE with Range-Aware Caching Chao Wang, Shubing Yang, Xiaoyan Sun, Yan Bai, Jun Dai, Dongfang Zhao arxiv.org/abs/2609.14137
PixCrypt: Fast Fine-Grained FHE with Range-Aware Caching

Chao Wang, Shubing Yang, Xiaoyan Sun, Yan Bai, Jun Dai, Dongfang Zhao

http://arxiv.org/abs/2609.14137

Many analytics tasks require secure computation over encrypted data. In particular, fine-grained data such as pixel-level images require higher precision, as every pixel can directly affect outcomes in tasks like tumor segmentation and anomaly detection. While Multi-Party Computation (MPC) is interactive, Differential Privacy (DP) protects only aggregate values, and Partially Homomorphic Encryption (PHE) lacks multiplicative support, none of them can efficiently handle fine-grained data analytics. Fully Homomorphic Encryption (FHE) uniquely enables arbitrary operations on encrypted pixels but remains computationally expensive, posing significant challenges for both software and hardware accelerators. We present PixCrypt, a caching-based acceleration mechanism for fine-grained fully homomorphic encryption. PixCrypt replaces expensive fresh ciphertext generation with cache retrieval and coefficient-level operations across CKKS, BFV, and BGV, while randomized reconstruction ensures that ciphertexts do not repeat. Its linear noise growth reduces the need for bootstrapping and lowers NTT load, improving hardware accelerator efficiency. This design yields up to 35x faster fine-grained encryption and maintains IND-CPA (Indistinguishability under Chosen Plaintext Attack) security. Experiments on five real-world pixel-level image processing tasks show that PixCrypt significantly improves the practicality of FHE for privacy-preserving analytics.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Private Graph Property Testing Hendrik Fichtenberger, Abigail Gentle, Tamalika Mukherjee, Sayantan Sen arxiv.org/abs/2609.14394
Private Graph Property Testing

Hendrik Fichtenberger, Abigail Gentle, Tamalika Mukherjee, Sayantan Sen

http://arxiv.org/abs/2609.14394

Graph property testing asks whether a massive graph satisfies a given property, or is far from doing so, using only a sublinear number of queries to the graph. Since property testers typically inspect only a small, randomly sampled portion of the input, they appear naturally compatible with differential privacy and privacy amplification by subsampling. Despite this, few results link these two fields. We initiate a systematic study of differentially private graph property testing with the goal of designing efficient testers with formal privacy guarantees in the dense and bounded-degree graph models. We develop new privacy amplification theorems for several widely used graph-sampling procedures such as induced subgraph sampling, random walks and k-disc sampling. We then leverage these privacy amplification techniques to design a private canonical tester in the dense graph model, as well as private bipartiteness testers and subgraph freeness testers in the dense and bounded-degree graph models. Finally, using the new privacy amplification theorem for k-disc sampling, we prove that every property of hyperfinite graphs is privately testable. The resulting query complexities of our private testers are comparable to those of their non-private counterparts.
021
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
PIMENTO: A Privacy Framework for Querying Text Mushtari Sadia, Ang Chen, Amrita Roy Chowdhury arxiv.org/abs/2609.14745
PIMENTO: A Privacy Framework for Querying Text

Mushtari Sadia, Ang Chen, Amrita Roy Chowdhury

http://arxiv.org/abs/2609.14745

Currently, there are two state-of-the-art, complementary privacy guarantees: contextual integrity (CI) for what may flow, and differential privacy (DP) for what may be inferred. Yet neither maps cleanly onto natural language, leaving existing approaches unable to provide these guarantees for analytics over unstructured text. We address this gap with Pimento, a framework that takes three forms of natural language: text corpus, queries, and privacy policies; and grounds them into a relational database, creating a common substrate on which both guarantees can be enforced formally. With this design, we not only provide end to end privacy guarantees, but also improvement to utility through three key contributions: DP aware Text-to-SQL, which searches for correct queries requiring the least DP noise; CI aware Text-to-SQL, which compiles natural language policies into executable CI rules over the database; and a new privacy definition we call contextual differential privacy, which redefines the traditional DP neighborhood under CI, and yields a tighter smooth sensitivity bound. Across new benchmarks, Pimento selects the best query in 75.3% of cases (upto +45 points over baselines) and achieves zero leakage under correct policy grounding. To our knowledge, Pimento is the first framework to provide formal privacy guarantees for natural language analytics under CI, DP, and their composition.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Privacy Preserving Gossip Learning Erkan Bayram, Mohamed-Ali Belabbas, Tamer Başar arxiv.org/abs/2609.14778
Privacy Preserving Gossip Learning

Erkan Bayram, Mohamed-Ali Belabbas, Tamer Başar

http://arxiv.org/abs/2609.14778

We propose a decentralized privacy-preserving learning algorithm in which each agent holds a single private sample and a shared model. Samples are learned sequentially, and each update must preserve the endpoint mappings at previously learned samples while protecting private data. This gives each agent three roles: (i) a learner that updates the model parameters, (ii) a teacher whose sample is learned at the current iteration, and (iii) a protected agent whose sample has already been learned. We build on Tuning without Forgetting (TwF) method to preserve previously learned mappings and show that TwF provides an indistinguishability guarantee for the learner whenever the set of protected agents contains another sample with the same label. For the teacher, we formulate a minimax optimal control problem that models the differential privacy noise as a worst-case disturbance to prevent performance loss while maintaining the same level of privacy for the gradient. For the protected agents, we compute the projections locally and aggregate them using a private push-sum gossip protocol. We prove geometric convergence of the decentralized gossip algorithm and of the distributed projection for TwF.
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar arxiv.org/abs/2609.15039
SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar

http://arxiv.org/abs/2609.15039

User prompts provided to large language models (LLMs) may contain sensitive or private information that can be misused by remotely deployed models, such as through inadvertent memorization during retraining. One way to protect user prompts is to execute the LLM inside a trusted execution environment (TEE), with the guarantee that the service provider has no access to computations performed within or information exchanged with the TEE. However, current TEEs are primarily CPU-based and significantly slower than GPUs optimized for LLM inference. To circumvent this, Tramer and Boneh (2019) proposed Slalom, which splits neural network inference between a TEE and an untrusted GPU and encrypts intermediate inputs sent to the GPU. We extend this split-inference architecture to LLM inference and instead protect intermediate inputs using differential privacy. We show that masking intermediate representations is necessary by showing that a prompt-reconstruction attack can recover prompts from these representations with nearly 80% accuracy. Our main contribution is a global sensitivity analysis of key LLM functions, which bounds the required scale of differentially private noise. Unlike encryption, differential privacy avoids quantization, allowing the LLM to remain in the floating-point domain. We also derive an upper bound on floating-point error from masking and noise cancellation in the TEE as a function of the privacy parameter epsilon. We implement our architecture using Intel TDX and evaluate it with two LLMs: Llama-3.2-3B and Qwen3-4B. Our split execution is nearly twice as fast as fully CPU-based inference inside TDX and 5-15 seconds faster than encryption-based Slalom while achieving higher accuracy. Finally, we demonstrate that prompt reconstruction, ev
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Differentially Private Multicolor Discrepancy and Fair Division of Indivisible Goods Max Dupré la Tour arxiv.org/abs/2609.15372
Differentially Private Multicolor Discrepancy and Fair Division of Indivisible Goods

Max Dupré la Tour

http://arxiv.org/abs/2609.15372

We study the fair division of indivisible goods under pure differential privacy, continuing the line of work initiated by Manurangsi and Suksompong. For $n$ agents with nonnegative additive utilities over $m$ goods and a fixed privacy parameter, we give an entry-private algorithm that, with high probability, achieves consensus envy-freeness up to $O(\sqrt n+\log^3 m)$ goods. This substantially improves the dependence on $n$ over the previous $O(n\log m)$ guarantee for ordinary envy-freeness, while providing the stronger consensus guarantee. A key ingredient is a private algorithm for multicolor discrepancy, which may be of independent interest. Our algorithm may require exponential time.
  We also obtain substantially stronger guarantees under additional structure: when all item values belong to a public alphabet of size $D$, we give a polynomial-time entry-private algorithm achieving ordinary envy-freeness up to $O(\operatorname{polylog}(mD))$ goods with high probability.
  Finally, we prove an $Ω(\log n)$ lower bound on the number of goods that must be removed to achieve ordinary envy-freeness under entry privacy, for sufficiently many goods, even with binary utilities.
010
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Differentially Private Semantic Plans for Aggregate Insight Generation Behrooz Razeghi arxiv.org/abs/2609.16283
Differentially Private Semantic Plans for Aggregate Insight Generation

Behrooz Razeghi

http://arxiv.org/abs/2609.16283

\texttt{URANIA} provides end-to-end differential privacy (DP) for summaries of data-dependent clusters. However, its cluster--keyword release does not directly provide collection-wide aggregates for semantic concepts defined independently of the protected corpus. Records may express several concepts, records expressing the same concept may be assigned to different clusters, and cluster identities need not correspond across analyses. Consequently, cluster-level statistics do not directly provide comparable measurements of predefined concepts across collections or repeated analyses. We introduce \texttt{DP-SPIN}, a trusted-curator framework for aggregate measurement and summarization over semantic concepts fixed independently of the protected target records. Each record is mapped to a bounded sparse nonnegative vector over these concepts, whose sum forms a semantic sketch. A differentially private mechanism releases a semantic plan containing admitted concepts and noisy masses; normalized semantic-support values and support bins are obtained by post-processing. For user-level privacy, each user's aggregate contribution is clipped to a fixed bound. The language model receives only the plan and fixed decoding instructions, while a public verifier checks concept mentions, reported values, comparisons, and rank claims against the released plan. The final summary is differentially private by post-processing. We establish record- and user-level DP guarantees under add/drop and replacement adjacency. We evaluate \texttt{DP-SPIN} under record-level privacy on CFPB complaint narratives, Amazon All Beauty reviews, and Yelp restaurant reviews, and under user-level privacy on Amazon and Yelp. We compare \texttt{DP-SPIN} with non-private plan and summary references, DP keyword and category histogram baselines, and a \texttt{URANIA}-style baseline with a fixed p
000
Differential Privacy Papers @dppapers.bsky.social · 16/09/2026
Exact Asymptotic Efficiency under zCDP: Diameter-Constrained Information Geometry T. Tony Cai, Yicheng Li arxiv.org/abs/2609.16895
Exact Asymptotic Efficiency under zCDP: Diameter-Constrained Information Geometry

T. Tony Cai, Yicheng Li

http://arxiv.org/abs/2609.16895

Protecting individual privacy has become a central and urgent concern in modern data analysis, given the vast quantities of data now generated and processed. In this paper, we develop a systematic theory of exact asymptotic efficiency for regular parametric estimation under central zero-concentrated differential privacy. In the privacy regime, the governing object is a diameter-constrained information region: the set of information matrices generated by statistics with diameter at most one. For weighted quadratic loss, we show that the exact local minimax risk is an inverse information variational functional over this region.
  More generally, a mixed information region yields a unified efficiency constant across different regimes, covering the privacy regime and classical Fisher efficiency. A matching estimator releases the empirical mean of a nearly optimal bounded statistic with Gaussian noise and locally inverts its population moment map. Our theory differs from classical efficiency theory in its set-valued information geometry and loss-dependent efficient estimator.
  As examples, we provide closed-form constants and optimal procedures for various concrete models, including one-dimensional regular families, Gaussian means, categorical probability vectors, and regression models among others. The theory also transfers to Gaussian differential privacy through an exact parameter rescaling.
000