Sign in

Differential Privacy Papers

@dppapers.bsky.social
558 followers 0 following 1.7K posts

🤖 new arXiv preprints mentioning "differential privacy" or "differentially private" in the title/abstract - unrelated quantum/FL papers + updates from differentialprivacy.org [Under construction.]

PostsRepliesMedia
Differential Privacy Papers @dppapers.bsky.social · 4h
Unifying Privacy Accounting: Information Equivalence and Information Loss Buxin Su, Qiaoshi Yang, Yiding Su, Chendi Wang arxiv.org/abs/2610.02414
Unifying Privacy Accounting: Information Equivalence and Information Loss

Buxin Su, Qiaoshi Yang, Yiding Su, Chendi Wang

http://arxiv.org/abs/2610.02414

Differential privacy (DP) admits several notions, but the choice among them may affect both privacy analysis and utility. In this paper, we consider four mainstream curve-based privacy notions within a unified information-theoretic framework. For a fixed ordered pair of output distributions, we establish information equivalence among the two directional privacy profiles of $(\varepsilon,δ)$-DP, the pair of hypothesis-testing trade-off functions, and the extended privacy-loss distribution. The exact Rényi differential privacy (RDP) curve joins this equivalence class whenever it is finite at some order greater than one. Under this mild condition, choosing among these notions changes only their semantic interpretation and computational requirements. In contrast, taking the maximum of the directional privacy profiles or compressing the RDP curve into a single zero-concentrated differential privacy (zCDP) parameter can lose information. We quantify the information loss between the exact RDP curve and its zCDP bound for standard noise mechanisms. This gap is zero for Gaussian noise but generally positive for Gaussian-mixture, Laplace, discrete Gaussian, and Poisson-subsampled Gaussian mechanisms. Moreover, this gap grows linearly with the number of independently composed mechanisms. Our information-theoretic perspective has practical consequences. At the same certified privacy level, retaining the full RDP curve rather than using zCDP reduces the required noise variance by up to $45\%$ for Gaussian-mixture noise in workloads comparable in size to the American Community Survey. For DP-SGD on Fashion-MNIST under Poisson subsampling, an RDP-based privacy accountant improves test accuracy by up to $8.73$ percentage points compared to a zCDP-based accountant when both are calibrated to the same $(\varepsilon,δ)$ guarantee.
000
Differential Privacy Papers @dppapers.bsky.social · 4h
HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Yan Zhang arxiv.org/abs/2610.02504
HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems

Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Yan Zhang

http://arxiv.org/abs/2610.02504

Balancing electricity demand and supply is increasingly difficult due to the inherent intermittency of renewable power generation and the stochastic power consumption. Grid operators require fine-grained, decision-relevant insights into household energy consumption to manage peak loads and design responsive tariffs, but increased transparency at this level raises significant privacy concerns. Traditional methods for explainable AI (XAI) can reveal sensitive information, while standard privacy techniques often reduce the usefulness of explanations. To address this issue, we introduce HXAI, a hierarchical framework that preserves privacy while enabling reasonable explainable analysis for grid-level demand management. HXAI consists of two main components: (1) a local model that generates fine-grained explanations within a secure, private environment, and (2) a zonal model that aggregates these explanations to support grid-level analysis while enforcing privacy through flexible privacy-budget management. We explicitly limit cumulative privacy exposure under repeated operator queries and show that the proposed framework preserves decision-relevant information without compromising household privacy. Experiments on both simulated and real-world energy datasets demonstrate that HXAI provides useful insights for zonal load management while ensuring that appliance-level consumption remains local and is never transmitted to grid operators. Our results show that preserving the semantic structure of explanations, rather than minimizing numerical error, is the key to XAI under differential privacy. This framework provides a way to achieve both privacy and explainability in energy management.
010
Differential Privacy Papers @dppapers.bsky.social · 4h
High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning Filip Kovačević, Edwige Cyffers, Stefano Sarao Mannelli, Marco Mondelli arxiv.org/abs/2610.02578
High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning

Filip Kovačević, Edwige Cyffers, Stefano Sarao Mannelli, Marco Mondelli

http://arxiv.org/abs/2610.02578

To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator. Our approach uses only summary statistics and it gives privacy guarantees either on labels only or jointly on features and labels, in terms of $ρ$-zero-concentrated differential privacy. The main technical contribution is a deterministic equivalent of the test error, which captures the interactions between sample size, covariance structure, model shift, regularization and privacy noise. Our theory allows to optimize hyperparameters (weights and ridge regularizers) and, more broadly, to decide when private external datasets are useful without accessing the data itself but only relying on population-level quantities. This provides a theoretically tractable foundation for private transfer learning, which we support via experiments on both synthetic and real-world datasets.
000
Differential Privacy Papers @dppapers.bsky.social · 4h
Differential Privacy of Gradient Descent on Perturbed Objectives Austin Watkins, Raman Arora arxiv.org/abs/2610.02716
Differential Privacy of Gradient Descent on Perturbed Objectives

Austin Watkins, Raman Arora

http://arxiv.org/abs/2610.02716

Objective perturbation adds a random linear term to a regularized empirical risk and releases the exact perturbed minimizer. We study the finite computation obtained by releasing the $N$-th iterate of deterministic gradient descent on $w\mapsto F(w;S)+\langle z,w\rangle$, where $z\sim\mathcal N(0,σ^2I_d)$ is drawn once before optimization. For strongly convex and smooth objectives with Lipschitz Hessian, we prove an explicit condition under which the map $z\mapsto w_N$ is a $C^1$-diffeomorphism on the bounded domains used in the privacy argument, with a quantitative lower bound on the smallest singular value of its Jacobian. This permits a direct change-of-variables analysis of the finite iterate. For generalized linear models, the resulting privacy-profile bound has no explicit ambient-dimension factor once the iteration condition holds, and its finite-iteration correction decreases geometrically. By letting the free truncation parameter grow slowly with $N$, we recover the corresponding exact-minimizer certificate in the limit. We also bound the expected excess empirical risk by $dσ^2/(2μ)$ plus a geometrically decreasing optimization term, and transfer the result to population risk without an additional multiplicative condition-number factor in the leading statistical terms.
010
Differential Privacy Papers @dppapers.bsky.social · 4h
ReLEAF: A Socio-Technical Framework Bridging Custodians and Researchers for Trustworthy Data Sharing Hibiki Ito, Chia-Yu Hsu, Hiroaki Ogata arxiv.org/abs/2610.02720
ReLEAF: A Socio-Technical Framework Bridging Custodians and Researchers for Trustworthy Data Sharing

Hibiki Ito, Chia-Yu Hsu, Hiroaki Ogata

http://arxiv.org/abs/2610.02720

Growing volumes of educational real-world data (ERWD) are being collected across learning platforms and institutional systems. Sharing these data within the Learning Analytics community offers substantial research opportunities, yet access remains limited by ethical, regulatory and governance constraints. Prior work has primarily focused on anonymisation techniques, but little attention has been paid to operational design of practical ERWD sharing, particularly how data custodians and researchers interact through privacy-preserving access mechanisms. To address this gap, we propose ReLEAF, a socio-technical framework that bridges data custodians and researchers by operationalising two-stage data sharing: 1) Differentially private synthetic data are shared for exploratory analysis, and 2) controlled real-data validation is performed on demand. Following a design-science research approach, we refine and formatively evaluate ReLEAF through three cycles involving 4 graduate students, 6 researchers, and 90 undergraduate students, respectively. Three design principles emerged through the cycles: P1) position privacy-preserving access mechanisms within the research workflow, P2) make the conditions for acceptable secondary use explicit and actionable, and P3) promote engagement with governance requirements rather than automate compliance decisions. Together, ReLEAF provides a concrete framework for trustworthy ERWD sharing, while the design principles offer transferable guidance for other data-sharing contexts.
000
Differential Privacy Papers @dppapers.bsky.social · 4h
Inner Momentum for Differentially Private Muon Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai arxiv.org/abs/2610.02738
Inner Momentum for Differentially Private Muon

Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai

http://arxiv.org/abs/2610.02738

Differentially private training clips each per-example gradient before adding noise. This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average. Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift. To curb this degradation, we propose averaging each sampled example's Muon gradient over the current model and a short history of recent models before clipping. The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F <= sigma_lambda sigma_G, bounding the clipping-induced distortion directly. We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization. In private GPT-2 fine-tuning on E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, and non-private diagnostics show 2-4% lower pre-noise polar error.
000
Differential Privacy Papers @dppapers.bsky.social · 4h
Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy Kaushik Bhargav Sivangi, Fani Deligianni arxiv.org/abs/2610.02943
Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy

Kaushik Bhargav Sivangi, Fani Deligianni

http://arxiv.org/abs/2610.02943

Multimodal 3D Human Pose Estimation (3D HPE) combines complementary information from RGB, LiDAR, and mmWave radar, but models trained on correlated observations from the same individuals, raise privacy risks overlooked by record level analysis. We present a unified framework for multimodal 3D HPE that couples kinematics-induced sensor fusion with subject level privacy auditing and private training. First, our multimodal model aligns modality specific joint representation, injects skeletal structure and adaptively aggregates complementary sensor evidence for accurate pose prediction. Second, we formulate a black-box subject membership inference attack for 3D HPE, complemented by an empirical pointwise maximal leakage analysis, which characterizes how individual attack score outcomes change inference about the membership outcome. Third, we instantiate user-level differential privacy via Action Temporal Stratification, a population weighted within-subject sampling strategy that enforces action and temporal coverage. We evaluate our framework on the MM-Fi dataset across three diverse experimental protocols. Source-code will be released upon acceptance.
000
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection Xudong Wang, Chris Ding, Tongxin Li, Jicong Fan arxiv.org/abs/2609.05253
GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection

Xudong Wang, Chris Ding, Tongxin Li, Jicong Fan

http://arxiv.org/abs/2609.05253

We introduce GLASS, a framework for graph-level anomaly detection (GLAD) that achieves robust cross-domain transferability through graph-language alignment on the unit hypersphere. GLASS builds a unified representation space by aligning a structure-aware graph encoder with an instruction-aware text embedding via a multi-slice soft cosine objective. Our framework serializes local, global, and semantic graph properties into a compact Graph Descriptor Prompt (GraphDP), creating a text bridge that enables domain-agnostic anomaly scoring. By enforcing multi-scale consistency through Matryoshka representation slices, the model captures anomalous deviations at multiple levels of granularity. We formulate anomaly detection as density estimation on the aligned hypersphere and introduce Spherical Multi-Modal Scoring (SMS), which instantiates von Mises-Fisher kernel density estimators in both graph and text embedding spaces. This probabilistic formulation recovers angular 1-nearest-neighbor scoring in the high-concentration limit, motivates the practical mean k-nearest-neighbor scorer, and provides a principled fusion of structural and semantic anomaly signals. The shared text embedding space further serves as a cross-domain bridge: by encoding a target domain's GraphDP without target-domain training data, GLASS performs zero-shot anomaly detection, and with only a handful of normal examples, few-shot adaptation via reference-set calibration. For privacy-sensitive deployment, we extend reference-set calibration with a bounded joint graph-text kernel summary that provides graph-record differential privacy while keeping the encoders fixed independently of the private target references. Across twelve benchmarks and three meta-domains, GLASS obtains the best average AUROC and rank compared with rece
000
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Energy Time-Series Imputation with Differentially Private Diffusion Models via Clipping-Aware Objective Conditioning Huizhen Huang, Yu Li, Tao Huang, Chen Hou arxiv.org/abs/2610.00209
Energy Time-Series Imputation with Differentially Private Diffusion Models via Clipping-Aware Objective Conditioning

Huizhen Huang, Yu Li, Tao Huang, Chen Hou

http://arxiv.org/abs/2610.00209

Reliable recovery of missing measurements is important for monitoring and analysis in energy time-series systems, where fine-grained measurements may contain sensitive temporal information. Diffusion models trained with differentially private stochastic gradient descent (DP-SGD) provide a promising framework for privacy-sensitive energy time-series imputation. Under cosine diffusion schedules, late timesteps correspond to low signal-to-noise ratio (SNR) conditions, where standard $\varepsilon$-prediction can induce large pre-clipping gradients. Such gradients are more likely to be clipped, reducing the retained optimization signal. The artificial intelligence (AI) contribution lies in formulating this objective--clipping interaction as an objective optimization problem under fixed-threshold DP-SGD and developing timestep-aware objective conditioning for diffusion-based energy time-series imputation. The method adopts $v$-prediction to mitigate late-timestep gradient amplification, uses static loss weighting as a uniform-scaling control, and introduces diffusion-schedule-aware dynamic weighting for stronger attenuation before clipping. For the engineering application, we evaluate the method on five real-world energy time-series datasets across random point missingness, contiguous block missingness, persistent outages, and multiple missing-data severities. Under matched DP-SGD settings, the proposed method consistently improves imputation utility over the $\varepsilon$-prediction baseline. Gradient diagnostics reveal lower upper-tail pre-clipping gradient norms, reduced clipping fractions, and stronger attenuation at late low-SNR timesteps, supporting the effectiveness of clipping-aware objective conditioning for energy time-series imputation.
000
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs Mohamed Shaaban, Mohamed Elmahallawy arxiv.org/abs/2610.00309
Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs

Mohamed Shaaban, Mohamed Elmahallawy

http://arxiv.org/abs/2610.00309

Large language models (LLMs) are increasingly deployed in privacy-critical domains (e.g., healthcare, finance, and government), but their propensity to memorize and disclose personally identifiable information (PII) poses serious security and compliance risks. Existing defenses typically force a trade-off between model utility, privacy protection, and access to fine-tuned private knowledge. We propose LoRA-Oriented Control via Keyed Entry Tokens (Locket), a practical framework that embeds fine-grained, policy-driven access control directly into LLM generation. Locket trains a set of lightweight LoRA (Low-Rank Adaptation) adapters, each encoding a distinct access policy (e.g., full reveal, partial redaction via PII masking, or reveal under a specified differential privacy level). A compact gating module is trained to associate a learned keyed entry token with exactly one LoRA adapter via sequence-level hard routing; the presence of a valid token acts as an authorization key that unlocks corresponding private knowledge, while an invalid or absent token triggers a privacy-preserving adapter that redacts or sanitizes sensitive content. This design ensures Locket remains fully compatible with off-the-shelf LLMs, supporting scalable deployment while satisfying regulatory and privacy requirements. We evaluate Locket across multiple datasets (Enron, ECHR, Yelp) and a diverse set of state-of-the-art LLMs, including Qwen3 (1.7B and 8B), Meta's Llama-3.2 (1B and 3B), and Google's Gemma-2-2B. Our extensive experiments demonstrate that, when the correct token is provided, Locket preserves perplexity comparable to fine-tuning on raw data (without any defense). Conversely, when the token is missing or invalid, it substantially reduces PII leakage while maintaining utility and perplexity on par with stron
010
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability Ceren Yıldırım, Kamer Kaya, Sinan Yıldırım, Erkay Savaş arxiv.org/abs/2610.01650
Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability

Ceren Yıldırım, Kamer Kaya, Sinan Yıldırım, Erkay Savaş

http://arxiv.org/abs/2610.01650

The increasing prevalence of decentralized data has led to a growing interest in federated learning, which enables collaborative model training without clients sharing their sensitive local data. However, FL alone does not sufficiently protect sensitive training data and is generally coupled with privacy-preserving techniques, such as differential privacy and homomorphic encryption. Although powerful, these techniques address separate concerns via different mechanisms, so relying on just one might prove insufficient or impractical for addressing challenges associated with federated learning. In this work, we propose a privacy-preserving federated learning framework that combines homomorphic encryption-based training with differential privacy-based model inspection and release. We adopt a Markov chain Monte Carlo-based Bayesian privacy estimation method to estimate the privacy of our proposed framework. Our results show that this method improves both model utility and estimated privacy over the baseline method that relies solely on differential privacy for training. In our experiments with the FEMNIST dataset, by the end of training, our method reaches a test loss of $1.09$, compared to $2.37$ for the differential privacy-only approach, while providing stronger estimated privacy protection, with the estimated posterior mean of the privacy parameter $ε$ of $4.32$, compared to $7.26$ for the differential privacy-only approach. We also show that intermittent model monitoring can preserve the encrypted training trajectory while, under our evaluated experimental setting, providing estimated privacy comparable to or stronger than the differential privacy-only approach.
011
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Detection and Resolution of Periodic Artifacts in OpenDP's Discrete Laplace Sampler Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta, Massimo Caccia arxiv.org/abs/2610.01907
Detection and Resolution of Periodic Artifacts in OpenDP's Discrete Laplace Sampler

Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta, Massimo Caccia

http://arxiv.org/abs/2610.01907

Differential privacy implementations rely on precise sampling from noise distributions to provide formal privacy guarantees. We report the discovery of systematic artifacts in OpenDP's discrete Laplace sampler that manifest as periodic distortions in the output distribution. Through systematic testing, we trace these artifacts to a faulty implementation in the rational arithmetic library used by the bernoulli_exp1 function, a low-level primitive that implements sampling from Bernoulli(e^(-x)) distributions. We present a diagnostic methodology that isolates the faulty component in the nested sampling hierarchy and propose an alternative implementation based on exact rational arithmetic that eliminates the artifacts. Statistical validation with 10^6 samples confirms that the corrected sampler produces outputs indistinguishable from the theoretical distribution at the tested precision level.
000
Differential Privacy Papers @dppapers.bsky.social · 02/10/2026
Quantum Advantage for Two-Party Differential Privacy Daniel Alabi, Emil T. Khabiboulline arxiv.org/abs/2610.02113
Quantum Advantage for Two-Party Differential Privacy

Daniel Alabi, Emil T. Khabiboulline

http://arxiv.org/abs/2610.02113

We introduce information-theoretically private quantum protocols for two-party Hamming distance when both parties must output the same estimate. Classically, for input length $n$, information-theoretic protocols require $Ω(\sqrt{n})$ error under pure differential privacy and $Ω(\sqrt{n}/\log n)$ error under strong approximate differential privacy, whereas computational security permits $O(1)$ error. In Klauck's honest, nonpreemptive, message-preserving model, we give an $O(n)$-communication quantum protocol with pure $\varepsilon$ quantum differential privacy (QDP) and expected error at most $\frac{2}{\sinh \varepsilon}+γ$, for every $γ>0$. For approximate $(\varepsilon, δ)$ QDP, an exact hockey-stick divergence calculation yields strictly smaller error, while preserving the $O(1)$-versus-$Ω(\sqrt{n}/\log n)$ separation for $δ=o(1/n)$. Thus, quantum communication achieves $O(1)$ information-theoretic error, matching the accuracy available classically only under computational assumptions.
  The main construction uses a guarded coherent round trip and an equal-Gram rigidity principle that prevents an honest player from retaining input-dependent complementary information. We separate this model from weaker prescribed-channel privacy, which already admits an exact classical realization, and from fully retention-robust security, against which measurement-and-abort attacks remain possible. Therefore, we identify preservation of non-orthogonal quantum messages as a resource for privacy.
000
Differential Privacy Papers @dppapers.bsky.social · 01/10/2026
PrivCert: Certifying Statement Support under Differential Privacy Tsubasa Takahashi, Takumi Hiraoka arxiv.org/abs/2609.38934
PrivCert: Certifying Statement Support under Differential Privacy

Tsubasa Takahashi, Takumi Hiraoka

http://arxiv.org/abs/2609.38934

Differentially private (DP) text generation can protect individual records, but privacy alone does not specify what evidence a released statement carries about the underlying data. We identify this as an evidence gap: a private report may contain plausible claims without indicating whether they are strongly supported by the private dataset. We introduce PrivCert, a framework for privacy-preserving reporting that makes statement support explicit through privacy-preserving certificates and emit-or-abstain decisions. As a canonical instantiation, PrivCert-PF (Proposal-and-Filter) separates data-independent candidate discovery from private support certification, emitting only statements whose support passes a private evidence test. We provide theoretical grounding for this framework by characterizing the limits of implicit evidence under DP, deriving a sharp privacy--honesty frontier for single-statement certification, and establishing a worst-case cost for fine-grained multi-statement certification. Experiments on synthetic tasks and TAB, WildChat, and Yelp show that explicit certification maintains low unsupported emission, while free-text DP baselines frequently produce low-support claims under the same declared support semantics. We further show that the PrivCert contract can be realized with histogram, sparse-vector, and Gaussian mechanisms, and use DP synthetic data to illustrate an important boundary: support in a private proxy does not automatically certify support in the original data. Together, these results position privacy-preserving reporting as an evidence-design problem: not only how to generate private text, but what a private report can substantiate about its underlying data.
000
Differential Privacy Papers @dppapers.bsky.social · 01/10/2026
Certification-Based Differentially Private Learning Mihnea Ghitu, Matthew Wicker arxiv.org/abs/2609.39629
Certification-Based Differentially Private Learning

Mihnea Ghitu, Matthew Wicker

http://arxiv.org/abs/2609.39629

Differential privacy (DP) in machine learning is typically achieved by adding noise to model parameters (private learning) or to model outputs (private prediction). Recent work uses formal methods, namely abstract interpretation, to provide tighter privacy guarantees, but only for private prediction in classification settings. In this work, we investigate the use of formal methods as a general tool for tighter privacy analysis. First, we generalize the abstract gradient training (AGT) framework to private prediction in continuous, unbounded regression. Second, by reducing learning in parameterized models to a regression problem over the parameter space, we introduce Abstract Gradient Sampling (AGS), an algorithm that enables reachability-based analysis to provide guarantees for private learning. In both private prediction and private learning, we provide tightened privacy accounting for the AGT framework and a theoretical analysis demonstrating when our smooth sensitivity upper-bounds yield favourable privacy-utility trade-off. In practice, we validate that our regression bounds are tighter than global-sensitivity baselines on regression benchmarks, and, notably, yield the first finite privacy guarantees in settings where global prediction sensitivity is a priori unbounded. We also find that under matched conditions, our private learning algorithm can outperform standard private learners.
000
Differential Privacy Papers @dppapers.bsky.social · 01/10/2026
Privacy Foundations for Multi-Institutional Scientific Artificial Intelligence Olivera Kotevska, Sumit Jha, Aurélien Bellet, Rui Hu, Nathaniel D. Bastian, Rafael Ferreira da Silva, Ravi Madduri, Kibaek Kim arxiv.org/abs/2609.39787
Privacy Foundations for Multi-Institutional Scientific Artificial Intelligence

Olivera Kotevska, Sumit Jha, Aurélien Bellet, Rui Hu, Nathaniel D. Bastian, Rafael Ferreira da Silva, Ravi Madduri, Kibaek Kim

http://arxiv.org/abs/2609.39787

Scientific artificial intelligence (AI), spanning foundation models (FMs) to federated data-analysis pipelines, is becoming shared infrastructure across national laboratories, universities, hospitals, and industrial partners. This collaboration creates privacy risks whose natural unit is often an institution's participation, research strategy, or technical capability rather than a single record. Differential privacy (DP), federated learning (FL), secure computation, trusted execution, and provenance each protect parts of the stack, but their guarantees rarely compose across mixed-trust institutions, access tiers, and autonomous agents. This perspective recasts privacy for scientific AI as an assurance problem defined by six elements: protected asset, observer, channel, permitted disclosure, guarantee, and evidence. We demonstrate the framing through a claim register for a composite cross-institutional scenario and use it to assess the model lifecycle. Two of the resulting gaps are specific to leadership-class facilities: scheduler, allocation, and telemetry metadata expose an institution's resource posture, and instrument-attached control loops leak research strategy through timing and contention on shared accelerators. We identify six research priorities: institution-level guarantees, agent-communication privacy, cross-tier information flow, privacy-compatible reproducibility, leadership-scale accounting, and instrument side channels. The contribution is a common form for stating, comparing, and auditing claims whose guarantees otherwise remain fragmented across the scientific AI stack.
000
Differential Privacy Papers @dppapers.bsky.social · 01/10/2026
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD? Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine arxiv.org/abs/2609.40335
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine

http://arxiv.org/abs/2609.40335

Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.
010
Differential Privacy Papers @dppapers.bsky.social · 30/09/2026
Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising Om Shankar Tiwari, Navnit Shukla, Guanyu Wang, Akshay Jain arxiv.org/abs/2609.36153
Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising

Om Shankar Tiwari, Navnit Shukla, Guanyu Wang, Akshay Jain

http://arxiv.org/abs/2609.36153

B2B advertising targets a viewer's professional attributes (employer size and industry, function, seniority) and has obtained them by matching identities across sites. Safari and Firefox block third-party cookies, Google retired the Privacy Sandbox cohort APIs in 2025, and reverse-IP firmographics decay under remote work. We present SIF (Sealed Inference Frame), which infers coarse professional cohorts on the device and emits only a locally differentially private, taxonomy-coded label into the OpenRTB bid stream, with no cross-site identifier. It rests on a property of the web platform we make precise: a navigated cross-origin iframe is the only way third-party code obtains a policy it controls, so inference runs in WebAssembly even where the publisher's CSP forbids it, and a nested worker served with default-src 'none' gives the model no network. Even a malicious model leaks at most about 5 bits per site per week. Labels pass through a memoised k-ary randomised response keyed to the publisher's first-party identifier, which gives $\varepsilon$-local differential privacy, defeats averaging, and links requests no better than the identifier already sent. An org-conditional k-anonymity rule suppresses cells, more strictly on corporate networks than at home. Cohorts ride OpenRTB user.data in a LinkedIn-aligned taxonomy, and attribution uses LinkedIn's click-scoped li_fat_id without bridging identities. We report a crawl of CSP deployment on 7,969 top sites and 431 B2B publishers, Heavy-Ad budgets, closed-form privacy-utility trade-offs, a re-identification simulation, and an assessment of which attributes are predictable at all: company type and size are, seniority largely is not. On-device is a design property, not a consent exemption.
000
Differential Privacy Papers @dppapers.bsky.social · 30/09/2026
Understanding Private Evolution as Learning-Augmented Clustering Audra McMillan, Kunal Talwar, Felix Zhou arxiv.org/abs/2609.36678
Understanding Private Evolution as Learning-Augmented Clustering

Audra McMillan, Kunal Talwar, Felix Zhou

http://arxiv.org/abs/2609.36678

Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-case Wasserstein analyses would predict. We recast PE as generative model-augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then we can obtain much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then sample complexity depends on intrinsic, not ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm is competitive with standard baselines and can improve recall.
000
Differential Privacy Papers @dppapers.bsky.social · 30/09/2026
A Sharp Transition in Data Reconstruction under Differential Privacy Max Cairney-Leeming, Simone Bombari, Marco Mondelli arxiv.org/abs/2609.37344
A Sharp Transition in Data Reconstruction under Differential Privacy

Max Cairney-Leeming, Simone Bombari, Marco Mondelli

http://arxiv.org/abs/2609.37344

Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single $d$-dimensional training sample from a $ρ$-zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at $ρ\asymp d$ for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for $ρ\ll d$; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for $ρ\gg d$. Remarkably, the transition moves to $ρ\asymp s$ for data lying in an $s$-dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
TRAP: Understanding and Mitigating Privacy Memorization in Language Models Muhammed Ustaomeroglu, Ziyue Xu, Hanshen Xiao, Peter Cnudde, Guannan Qu, Holger R. Roth arxiv.org/abs/2609.32293
TRAP: Understanding and Mitigating Privacy Memorization in Language Models

Muhammed Ustaomeroglu, Ziyue Xu, Hanshen Xiao, Peter Cnudde, Guannan Qu, Holger R. Roth

http://arxiv.org/abs/2609.32293

Fine-tuning a language model on sensitive records can leave it able to reproduce them. We ask when this memorization arises and how to prevent it without knowing in advance which spans are sensitive. Our starting point is that most memorization scores and attacks share one statistical core: whether the model assigns a token more probability than some reference would. Taking as the reference a model trained on the complementary half of the same corpus gives the Target Reference Advantage (TRA), a per-token signal that separates what a model fit to a particular record from what it learned across records, and is cheap and differentiable. We then study what drives memorization during fine-tuning: it keeps growing well past the validation minimum, is larger on small datasets and at higher learning rates, and higher when the underlying task is harder. Early stopping removes much of it, but because it is chosen by aggregate validation loss it helps least for rare, hard-to-predict spans embedded in otherwise learnable text, which is exactly what sensitive information tends to be. We therefore introduce TRAP, a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference. On student essays with annotated personal information and clinical cases with patient identifiers, TRAP brings memorization near the level of an untrained model at little utility cost, where generic regularizers barely move and differential privacy gives up most of what fine-tuning bought.
020
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Differentially Private Approximation of the John Ellipsoid Bar Mahpud, Daniel Omer, Or Sheffet arxiv.org/abs/2609.32606
Differentially Private Approximation of the John Ellipsoid

Bar Mahpud, Daniel Omer, Or Sheffet

http://arxiv.org/abs/2609.32606

We study the problem of approximating the John ellipsoid (JE) of a given (centrally symmetric) polytope of $n$ constraints in a Euclidean space under differential privacy (DP). We give the first differentially private algorithm for this problem under the standard model, where neighboring datasets may differ arbitrarily in one a single constraint. Our work also extends to the complimentary problem of Minimum Enclosing Ellipsoid of $n$ points in the Euclidean space.
  Our approach is based on the recent non-private multiplicative-weights algorithm of~\cite{pmlr-v99-cohen19a}. First we introduce a non-private generalization of the Cohen et al algorithm, yielding a $(1+γ)$-approximation of the JE problem while violating at most $κn$ constraints in $O(\log(1/κ)/γ)$ iterations. This variant works by projecting the intermediate weights assigned to the constraints onto the set of $κ$-dense distributions, similarly to~\cite{bun2020efficientnoisetolerantprivatelearning}.
  We then design a $ρ$-zCDP variant of this algorithm by adding Gaussian noise to the weighted covariance matrix aggregated in each step of the algorithm. Under a mild goodness assumption on the data we can assert that the resulting noisy matrix is close to the true matrix, thereby achieving essentially the same guarantee as the non-private algorithm provided sufficiently many input points. Thus our method achieves an efficient DP poly-time algorithm under concrete sample complexity bounds.
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier arxiv.org/abs/2609.32835
FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors

Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier

http://arxiv.org/abs/2609.32835

As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.
001
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Geometry-Adaptive Mechanisms for Private Synthetic Data Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian arxiv.org/abs/2609.33363
Geometry-Adaptive Mechanisms for Private Synthetic Data

Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian

http://arxiv.org/abs/2609.33363

Generating differentially private synthetic data with meaningful Wasserstein utility guarantees is challenging in high dimensions. For datasets of size \(n\) on $[0,1]^d$ with $d\ge2$, existing pure \(\varepsilon\)-differentially private mechanisms achieve expected $1$-Wasserstein error of order $(\varepsilon n)^{-1/d}$, reflecting the curse of dimensionality. While this rate is optimal in the worst case, it can be overly pessimistic when the data are supported on a lower-dimensional set. We formalize this through a multiscale packing-growth dimension $k$, which captures the geometric complexity of the support via the growth of packing numbers across scales. We propose \emph{Adaptive Pruned-PMM}, a pure $\varepsilon$-differentially private mechanism that combines private depth selection with our pruned variant of the Private Measure Mechanism (PMM) of He et al.\ (2023). The mechanism supports deeper, geometry-adapted hierarchies with expected running time $O\!\left(d(n+d)\log(\varepsilon n)\right)$, which is near-linear in $n$ for fixed dimension and privacy budget. Under an external multiscale packing-growth condition with dimension $k$, we show that, for fixed positive privacy budgets and fixed geometry, the expected $1$-Wasserstein error is of order $(\varepsilon n)^{-1/k}$ for $k>1$ as $n$ grows. We also prove a lower bound under a corresponding internal packing-growth condition, showing that the exponent $1/k$ is sharp within this framework.
001
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits Idan Attias, Orin Levy, Alexander Ryabchenko, Yishay Mansour, Uri Stemmer arxiv.org/abs/2609.33888
Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits

Idan Attias, Orin Levy, Alexander Ryabchenko, Yishay Mansour, Uri Stemmer

http://arxiv.org/abs/2609.33888

Can vanilla policy optimization explore enough to achieve near-optimal regret in stochastic contextual bandits? We show that standard exponential policy updates driven by offline regression do so under realizability, without exploration bonuses or importance weighting. For $A$ actions, $T$ rounds, and a finite prediction class $F$, vanilla PO achieves $\widetilde O(\sqrt{AT\log(|F|)})$ regret with high probability. Our analysis reveals an implicit exploration mechanism of independent interest: gradual policy updates prevent actions from losing probability too quickly, allowing the regression oracle to learn their expected losses. We further develop a batched version using only $O(\log T)$ regression calls and policy switches, and show how private regression oracles yield differentially private contextual bandit algorithms without composition across batches. For a finite class, this gives pure $\varepsilon_{\rm priv}$-DP and regret $\widetilde O\left( \sqrt{AT \log(|F|/δ)}(1+\varepsilon_{\rm priv}^{-1/2}) \right)$. Finally, experiments across oracle-based contextual bandit algorithms, with and without privacy, demonstrate the practical effectiveness of policy optimization and the value of explicit exploration under stronger privacy constraints.
000
Differential Privacy Papers @dppapers.bsky.social · 29/09/2026
Contraction of Rényi Divergences for Discrete Channels Adrien Vandenbroucque, Amedeo Roberto Esposito, Michael Gastpar arxiv.org/abs/2609.34570
Contraction of Rényi Divergences for Discrete Channels

Adrien Vandenbroucque, Amedeo Roberto Esposito, Michael Gastpar

http://arxiv.org/abs/2609.34570

We investigate Strong Data-Processing Inequality (SDPI) constants for Rényi Divergences on finite spaces. We study their dependence on the Rényi order $α$, proving that they are non-decreasing and that their scaling by $(α-1)$ is convex for $α\geq1$. We also identify several support restrictions on the probability measures involved in determining these constants. In particular, in the distribution-independent setting, measures supported on a common set of at most two points suffice to evaluate the Rényi-SDPI constant. For $α\in[0,1]$, we further prove equality with the $χ^2$-SDPI constant, while at order infinity we obtain a closed-form expression. In order to link contraction over product spaces to contraction along individual coordinates, we provide tensorisation bounds for arbitrary product channels analogous to those known for $\varphi$-Divergences. At finite orders, the Rényi-SDPI constants are bounded above and below through comparisons with the $χ^2$-Divergence and Hellinger Divergences, with sharpness established in multiple cases. At order infinity, we instead relate these constants to the contraction of Total Variation Distance. Finally, our findings are applied to local differential privacy (LDP) and the analysis of Markov chains. This yields sharp contraction guarantees for pure-LDP mechanisms, and connects Rényi-LDP to Rényi-SDPI constants. For Markov chains, we derive finite-time convergence bounds and exhibit a family of chains for which Rényi-SDPIs improve on classical $χ^2$-based bounds by arbitrarily large factors.
000
Differential Privacy Papers @dppapers.bsky.social · 28/09/2026
Revisiting Certified Defense with Differential Privacy on Vision Transformers Jun Yan, Weiquan Huang, Qixian Zhang, Yan Bai, Shutai Zhang arxiv.org/abs/2609.31310
Revisiting Certified Defense with Differential Privacy on Vision Transformers

Jun Yan, Weiquan Huang, Qixian Zhang, Yan Bai, Shutai Zhang

http://arxiv.org/abs/2609.31310

Certified defenses that incorporate differential privacy have proven effective on Convolutional Neural Networks (CNNs), furnishing rigorous robustness guarantees against norm-bounded adversaries. However, the certified robustness behavior of Pixel Differential Privacy (PixelDP) remains largely unexplored with the self-attention architecture now dominating the deep-learning landscape. Given that the Transformer has a profound impact on our daily applications from the digital world to the physical world, it is crucial to study certified robustness through differential-privacy-style stability. To fill this research gap, we revisit this construction in Vision Transformers and identify a failure mode that is largely hidden in the convolutional setting. When noise is injected after the patch embedding, the Laplace mechanism with the inherited grouped $\ell_1$ sensitivity bound collapses to chance-level accuracy across noise scales, whereas the Gaussian mechanism remains trainable. This contrast isolates the source of failure: not the injected noise itself, but the geometry of the sensitivity constraint. We show that the attenuation induced by the inherited $Δ_{1,1}$ projection increases with layer width and kernel size according to a random-matrix scale $C/(\sqrt{M}+\sqrt{N})$. Replacing the $\ell_1$-type constraint with a spectral-norm constraint eliminates the collapse across datasets and architectures, but creates a fundamental obstacle: the repaired models no longer satisfy the sensitivity condition required by the standard Laplace certificate. We resolve this mismatch by deriving a dimension-free $(\varepsilon,\ δ)$-privacy guarantee for the Laplace mechanism under $\ell_2$ sensitivity through concentration of the privacy loss.
010
Differential Privacy Papers @dppapers.bsky.social · 28/09/2026
Gap-free Differentially Private PCA for Gaussian Data Alina Ene, Huy L. Nguyen arxiv.org/abs/2609.31614
Gap-free Differentially Private PCA for Gaussian Data

Alina Ene, Huy L. Nguyen

http://arxiv.org/abs/2609.31614

We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
000
Differential Privacy Papers @dppapers.bsky.social · 25/09/2026
DP-IPI: A Hybrid Differential Privacy Text Rewriting Mechanism for Indirect Personal Identifiers in Clinical Texts Ibrahim Baroud, Stephen Meisenbacher, Sebastian Möller, Florian Matthes, Roland Roller arxiv.org/abs/2609.29684
DP-IPI: A Hybrid Differential Privacy Text Rewriting Mechanism for Indirect Personal Identifiers in Clinical Texts

Ibrahim Baroud, Stephen Meisenbacher, Sebastian Möller, Florian Matthes, Roland Roller

http://arxiv.org/abs/2609.29684

Despite the strengths of modern anonymization and de-identification techniques, the risk of re-identification remains significant due to the indirect identifiers remaining in texts. To address this problem, recent works have applied text rewriting under Differential Privacy (DP) to prevent data linkage by perturbing texts via noise addition. Such methods privatize all tokens in a text indiscriminately, diminishing text quality and usability in critical domains such as in clinical settings. Focusing on indirect personal identifiers (IPIs), we introduce a utility-preserving DP text rewriting method that only privatizes spans containing IPIs. We show that our method effectively reduces re-identification risks in clinical texts while being producing more coherent and usable output texts, leading to higher privacy-utility trade-offs. In this, we demonstrate the effectiveness of hybrid text privatization, which leverages the promise of DP in an efficient, usable manner.
000
Differential Privacy Papers @dppapers.bsky.social · 25/09/2026
An Exposition of GPT Astra's Proof of Lower Bound on DP Continual Counting Jalaj Upadhyay arxiv.org/abs/2609.28528
An Exposition of GPT Astra's Proof of Lower Bound on DP Continual Counting

Jalaj Upadhyay

http://arxiv.org/abs/2609.28528

The goal of this note is to give a detailed proof, to the best of our understanding, of the recent presentation by Harrison and Leeman (arXiv:2609.17650v01 and arXiv:2609.17650v02) of the proof by Astra on the lower bound for differentially private continual counting. We believe a more natural and easy proof is possible and hope that this note will help in that effort.
  Prior to the initial preprint by Harrison and Leeman (arXiv:2609.17650v01), Bairaktari and Larsen (arXiv:2607.00876) gave an elegant proof to show a lower bound of $Ω(\log^{3/2}(n))$ for both pure and approximate-DP continual counting, and in personal communication had informed us that they have a proof of optimal $Ω(\log^{2}(n))$ for pure-differential private continual counting as well. They have subsequently published their $Ω(\log^{2}(n))$ bound, which is now a joint work of Bairaktari, Dahl, and Larsen (arXiv:2607.00876v3). Their new result is an elegant extension of their technique for approximate-differential privacy. Although the two proofs are technically different, the Astra argument uses related tree geometry introduced in Bairaktari and Larsen.
010
Differential Privacy Papers @dppapers.bsky.social · 25/09/2026
When Do Differentially Private Inputs Protect Graph Shift Operators? Andrew Campbell, Chenyue Zhang, Hang Liu, Victor Elvira, Anna Scaglione, Sean Peisert arxiv.org/abs/2609.28899
When Do Differentially Private Inputs Protect Graph Shift Operators?

Andrew Campbell, Chenyue Zhang, Hang Liu, Victor Elvira, Anna Scaglione, Sean Peisert

http://arxiv.org/abs/2609.28899

We study the differential privacy (DP) of a graph shift operator (GSO) when an analyst observes the output of a graph filter. In particular, we study the setting in which the input signals to the graph filter are drawn from a differentially private distribution. Unlike approaches that perturb the GSO or the filter output, we use the randomness already present in the inputs to protect the GSO. This yields an equivalent level of privacy protection to that of the perturbation methods without adding noise, and thus a better privacy-utility trade-off. We provide an explicit characterization of the privacy loss and its certificate in terms of the zeros of the graph filter. In doing so, we show that the log-likelihood ratio between the releases of two adjacent topologies is governed by the distances from each zero to the graph frequencies of the two GSOs. Then, by uniformly bounding the log-likelihood ratio over the adjacent topologies, we obtain an explicit $(\varepsilon,δ)$-DP guarantee for Gaussian inputs. We further show, via a Cramér--Rao bound, that the zero placement that limits the privacy loss also raises the floor on the adversary's reconstruction error. Finally, empirical validation is performed on a synthetic network of financial exposures, where the largest position a pair can conceal and the accuracy with which it can be sized are collinear across pairs. Both are set by the graph-frequency content of the pair, and the full network becomes recoverable only as the certified budget grows.
000
Differential Privacy Papers @dppapers.bsky.social · 25/09/2026
Forte: A sensitivity type system for imperative Rust Chiké Abuah arxiv.org/abs/2609.30254
Forte: A sensitivity type system for imperative Rust

Chiké Abuah

http://arxiv.org/abs/2609.30254

We introduce Forte, a sensitivity type system for Rust whose soundness rests on ownership. The graded sensitivity type systems, from Fuzz's linear grading to Solo's environment indices, are pure calculi: a claim about a value holds for the value's whole lifetime because nothing can mutate it. The imperative sensitivity analyses admit assignment to first-order variables and no references, so no question of aliasing arises in them. The programs that compute differentially private statistics in deployment are Rust, and they mutate through borrows. Forte closes this gap. Its central rules strongly update a sensitivity environment through an exclusive borrow, at a primitive call and across a checked function boundary; its soundness theorem is metric preservation over an operational semantics with a store, in which the exclusivity of &mut alone licenses framing across a mutating call, and two aliased borrows suffice to refute the theorem without it. Verus mechanizes the theorem, the function rule, and the refutation. Flux checks Forte as an ordinary library, with no fork of the compiler; a machine-checked theorem backs every deterministic primitive signature, and a correspondence theorem transports metric preservation to the programs the checker accepts. We evaluate Forte on mechanism kernels from OpenDP with genuine in-place mutation, matching the library's trusted stability maps with checked constants, covering the constructors that have no proof document, rejecting off-by-one diameters, tightened bounds, miscalibrated releases, and overspent budgets, and deriving one trusted constant as an inferred loop invariant.
001
Differential Privacy Papers @dppapers.bsky.social · 24/09/2026
What fidelity metrics miss: a structural check on synthetic educational data Hitoshi Inoue, Koichi Yasutake arxiv.org/abs/2609.27265
What fidelity metrics miss: a structural check on synthetic educational data

Hitoshi Inoue, Koichi Yasutake

http://arxiv.org/abs/2609.27265

Secondary use of educational records is increasingly mediated by platforms that share a differentially private synthetic version of a dataset and validate specific findings against the real data on request. The synthetic version is evaluated by comparing summary statistics of each variable, yet reported confirmation rates suggest that such comparisons do not predict which findings survive. We propose a structural check: the number of connected components of a weekly proximity graph over learners, tracked across a term. Across four annual cohorts of lower-secondary study-habit logs, the synthetic versions reproduced the level of this quantity and the shape of the weekly partition, but its variation across the term was between 2.6 and 4.9 times smaller than in the real data at a common working point, without exception, and those changes fell in different weeks: the synthetic cohorts single out the term's examination weeks and the real cohorts do not. We also show that a routine rule for setting the graph threshold makes naive comparisons between two datasets invalid, and illustrate this with an error of our own. The real curves are also distinguishable from marginal-preserving surrogates of themselves in all four cohorts, where three of the four synthetic ones are not, a comparison that needs no real data; these differences trace to what the generator was given.
000
Differential Privacy Papers @dppapers.bsky.social · 24/09/2026
Only Pay What You Must Spend: On-Demand Privacy Budget Payment for Differentially Private RAG Zhonghao Sun, Zhiliang Tian, Xinyue Fang, Shuo Ma, Juhua Zhang, Yiping Song, Dongsheng Li arxiv.org/abs/2609.27406
Only Pay What You Must Spend: On-Demand Privacy Budget Payment for Differentially Private RAG

Zhonghao Sun, Zhiliang Tian, Xinyue Fang, Shuo Ma, Juhua Zhang, Yiping Song, Dongsheng Li

http://arxiv.org/abs/2609.27406

Deploying large language models (LLMs) on sensitive data via Retrieval-Augmented Generation (RAG) introduces severe privacy risks. Recent studies apply Differential Privacy (DP) to LLMs with RAG for formal privacy guarantees. However, existing DP-RAG frameworks rapidly exhaust the privacy budget. Although recent efforts attempt to save the budget by narrowing the retrieval scope or sparsifying private generation, these methods themselves cumulatively consume the budget, whereas they could actually rely merely on public information or at a negligible one-time privacy cost. This mismatch fails to align budget expenditure with the model's actual reliance on private data, causing substantial waste on operations that require no private access. To address this, we propose SparsePay-RAG, adopting "only pay what you must spend" as its core principle. Using public information as a zero-privacy prior, it charges the privacy budget only for the private increment. Specifically, SparsePay-RAG narrows the retrieval scope via public topic-guided clustering, adaptively controls private access frequency without privacy cost through isotonic cross-layer trajectory fitting, and compresses per-access budget via DP contrastive decoding. Under strong privacy constraints, experiments show SparsePay-RAG achieves superior privacy-utility trade-offs over baselines.
010
Differential Privacy Papers @dppapers.bsky.social · 24/09/2026
Finite-Sample Binary Hypothesis Testing via Rényi Divergences: Strong Converse and Local Privacy Roberto Bruno, Adrien Vandenbroucque, Amedeo Roberto Esposito arxiv.org/abs/2609.27617
Finite-Sample Binary Hypothesis Testing via Rényi Divergences: Strong Converse and Local Privacy

Roberto Bruno, Adrien Vandenbroucque, Amedeo Roberto Esposito

http://arxiv.org/abs/2609.27617

We study asymmetric simple binary hypothesis testing between $H_0:P_0^{n}$ and $H_1:P_1^{n}$, based on $n$ independent and identically distributed observations. Leveraging a variational representation of Rényi divergence of order $α$, we derive our main result: a finite-sample converse with $α>1$. The bound uses both directions of the divergence $D_α(P_1\|P_0)$ and $D_α(P_0\|P_1)$, tensorises under product measures, and contains familiar data-processing converses as boundary cases. For comparison, we apply the same variational approach to general $f$-divergences and specialise it to total variation, $E_γ$, Hellinger, and Kullback Leibler divergences, thereby recovering familiar converses within a unified framework. Together with an achievability bound involving Rényi divergence with $α\in (0,1)$, the main converse recovers the phase transition of the optimal Type II error under the exponentially decaying Type I error constraint $\varepsilon_n=e^{-nr}$. Under regularity conditions, the optimal Type II error vanishes exponentially when $r<D(P_1\|P_0)$ and converges exponentially fast to one when $r>D(P_1\|P_0)$. We also derive sample-complexity bounds and extend both the converse and achievability analyses to locally differentially private observations, quantifying the cost of privacy and recovering the non-private achievability bound as the privacy constraint vanishes.
000
Differential Privacy Papers @dppapers.bsky.social · 24/09/2026
Contraction and Statistical Inference under Privacy for Uniformly Bounded Distributions Leonhard Grosse, Sara Saeidian, Tobias J. Oechtering, Mikael Skoglund arxiv.org/abs/2609.28297
Contraction and Statistical Inference under Privacy for Uniformly Bounded Distributions

Leonhard Grosse, Sara Saeidian, Tobias J. Oechtering, Mikael Skoglund

http://arxiv.org/abs/2609.28297

We investigate $c$-interior pointwise maximal leakage (PML) as a tool for contraction analyses and disclosure control. Based on the strong adversarial threat models from maximal leakage, $c$-interior PML generalizes local differential privacy (LDP) to data-generating distributions with densities uniformly bounded away from zero by $c>0$. Viewing $c$-interior PML as an algebraic constraint on a kernel yields more flexible (and often tighter) contraction analyses than standard LDP. We provide tight bounds on the Dobrushin coefficient, and bound the contraction coefficient of the Hockeystick-divergence. We further derive strong data processing inequalities on $f$-divergences under $c$-interior PML constraints when the input distributions to the divergence are restricted to be in the $c$-interior. These results extend beyond the regime of pure LDP to cover a larger class of kernels, including, e.g., arbitrary stochastic matrices. We apply the results to minimax theory and provide asymptotically optimal strategies under $c$-interior PML constraints for binary hypothesis testing and mean estimation. The results show that disclosure control with PML allows analysts to reason about systems in a more differentiated manner: For example, it allows us to quantify the privacy leakage of deterministic systems, and can give precise adversarial guarantees with respect to arbitrary distributional assumptions. Interestingly, a recurring theme in the disclosure analyses is that if the privacy problem is relatively regular (if the density bound $c$ is large), private inference can be possible without incurring any additional cost in terms of sample complexity.
000
Differential Privacy Papers @dppapers.bsky.social · 23/09/2026
Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy Anming Gu, Syamantak Kumar, Kevin Tian, Chutong Yang arxiv.org/abs/2609.26508
Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy

Anming Gu, Syamantak Kumar, Kevin Tian, Chutong Yang

http://arxiv.org/abs/2609.26508

Streaming principal component analysis (PCA) seeks to recover a leading spectral subspace in a single pass over a data stream. We give a new analysis of the ubiquitous Oja's algorithm [Oja82] for the most general, gap-free variant of this problem, where no eigengap assumptions are made on the underlying mean matrix, complemented by a nearly-matching lower bound. Prior works achieving near-optimal rates for streaming PCA either required gap assumptions [JJK+16, HNWW21], or were limited to rank-one updates [AZL17, Lia23]. Our proof only uses a second moment bound on the individual stochastic updates, bypassing the almost sure bounds needed by prior near-optimal analyses, and the analogous offline matrix Bernstein bound. We also extend our result to a Rayleigh quotient notion of approximate PCA, addressing an open question of [JJK+16]. As our main application, we give gap-free differentially private PCA guarantees for sub-Gaussian data, settling Conjecture 1.1 of [Bro26] up to logarithmic factors.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Differentially Private and Fairness-Audited Score Diffusion for Irregular Longitudinal Health Records Taimoor Ahmad arxiv.org/abs/2609.22401
Differentially Private and Fairness-Audited Score Diffusion for Irregular Longitudinal Health Records

Taimoor Ahmad

http://arxiv.org/abs/2609.22401

Sharing irregular longitudinal health records can accelerate model development, yet synthetic releases may leak participation, distort temporal dependence, suppress rare events, or reduce utility for underrepresented groups. We present TRUST LONGSYNTH, an auditable patient level private generator that combines bounded sufficient statistics, zCDP accounted Gaussian releases, conditional analytic score diffusion, block banded temporal covariance, separate missingness and gap models, and a protected event sampling floor with population weights.
  The method was evaluated on five independently generated, three cohort benchmarks containing 720 patients, fourteen irregular observation slots, six mixed variables, informative missingness, and a rare deterioration outcome. At epsilon = 12 and delta = 10 to the power of minus 5, TRUST LONGSYNTH achieved mean train synthetic test real AUPRC 0.342, Brier score 0.088, expected calibration error 0.082, correlation error 0.222, autocorrelation error 0.317, and membership attack AUROC 0.499.
  Relative to the private diagonal score baseline, AUPRC increased by 7.5 percent, while Brier, calibration, correlation, and autocorrelation errors decreased by 6.1 percent, 17.0 percent, 28.0 percent, and 30.1 percent, respectively. The method did not dominate every nonprivate or discrete baseline, and corrected paired tests were inconclusive with five seeds. Canary exposure was 1.8 percent, compared with 28.8 percent for DP Score in the same stress test.
  These findings support a transparent privacy utility fairness evaluation protocol, not clinical validity or unconditional release safety, and motivate governed external validation on real multi site records.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Locally Private Inference for Riemannian Stochastic Optimization Xiaotian Chang, Yangdi Jiang, Qirui Hu arxiv.org/abs/2609.22642
Locally Private Inference for Riemannian Stochastic Optimization

Xiaotian Chang, Yangdi Jiang, Qirui Hu

http://arxiv.org/abs/2609.22642

We develop inference for manifold-valued population minimizers when each observation belongs to a different participant and only locally private messages reach the analyst. The method releases randomized tangent gradients and combines them through Riemannian stochastic approximation and Polyak-Ruppert averaging. Directly inserting a private data surrogate into a nonlinear loss can shift its population target, whereas conditional centring of the released gradient preserves the first-order equation. We introduce symmetric-pair regression (SPR) to estimate the asymptotic variance from the same private messages used for point estimation, without holding out participants or requesting a second release. We prove the central limit theorem and consistency of the fully transcript-based sandwich covariance and intrinsic Wald region under local differential privacy. Simulations across various statistical problems and manifolds support the predicted decrease in estimation error and near-nominal coverage under moderate privacy. An application to NHANES anthropometric data illustrates private estimation of a leading body-size direction and its uncertainty.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Improved Private Sparse Covariance Estimation with Multiscale Threshold Tests Zihan Zhang arxiv.org/abs/2609.22783
Improved Private Sparse Covariance Estimation with Multiscale Threshold Tests

Zihan Zhang

http://arxiv.org/abs/2609.22783

We study differentially private covariance estimation in operator norm for mean-zero sub-Gaussian distributions with unknown covariance support and at most $k$ nonzero entries per row. We develop a multiscale random-threshold algorithm with sample complexity $\ot(k^2/α^2+k\sqrt d/(α\varepsilon))$ for $(\varepsilon,δ)$-differential privacy and error at most $ασ^2$, where $d$ is the dimension and $σ$ is a known sub-Gaussian scale. The bound improves the privacy-dependent term of the existing $\ot(k^2/α^2+k^{3/2}\sqrt d/(α\varepsilon))$ \citep{kumar2026curse} upper bound by a factor of $\sqrt k$, and matches the lower bound of $\widetildeΩ(k^2/α^2 + k\sqrt{d}/(α\varepsilon))$ in its applicable parameter regime.
  Our key technical ingredient is a direct operator-norm bound on the centered fluctuations of an ideal reconstruction, exploiting conditional independence rather than accumulating entrywise errors across each row. A multiscale allocation of threshold tests balances reconstruction variance against query sensitivity. Together, these ingredients sharpen the trade-off between approximation error and privacy protection, removing the additional $\sqrt{k}$ factor from the privacy-dependent sample complexity.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated Study Márton Pál Lipcsey-Magyar, Adrian Pekar arxiv.org/abs/2609.23521
Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated Study

Márton Pál Lipcsey-Magyar, Adrian Pekar

http://arxiv.org/abs/2609.23521

Residential traffic classification supports service management, but learning across homes must account for heterogeneous traffic and privacy constraints. Privacy-aware training may impose uneven costs across traffic categories. We study this tradeoff in simulated two-client federated learning using 1.62 million preprocessed gateway-collected flows across six categories. We compare a full-feature baseline, feature suppression (FS), and differentially private stochastic gradient descent (DP-SGD) under one fixed record-level privacy setting. FS-mild excludes four timing features from 16 model inputs; it provides no formal privacy guarantee. With size-proportional aggregation, FS-mild achieves higher combined macro-F1 and worst-group F1 (the minimum per-class F1 across homes) than DP-SGD in all five seeds at both model capacities under stratified and temporal splits. The tested DP-SGD configuration incurs pronounced minority-category losses, especially in the smaller home, but FS-mild does not uniformly improve on the full-feature baseline. On stratified-split models, loss-based and shadow-model membership probes show near-chance aggregate discrimination without a consistent ranking across probes; this does not establish equivalent privacy. These findings support FS as an input-minimization baseline, not a substitute for formal privacy.
000
Differential Privacy Papers @dppapers.bsky.social · 22/09/2026
Pattern-level Differential Privacy for High-utility Complex Event Processing He Gu, Thomas Plagemann, Vera Goebel, Maik Benndorf, Boris Koldehofe arxiv.org/abs/2609.23827
Pattern-level Differential Privacy for High-utility Complex Event Processing

He Gu, Thomas Plagemann, Vera Goebel, Maik Benndorf, Boris Koldehofe

http://arxiv.org/abs/2609.23827

Current privacy-preserving mechanisms (PPMs) in Complex Event Processing (CEP) systems are unnecessarily restrictive, reducing the utility of data received by data consumers. This article presents a novel approach to preserve privacy in CEP systems, improving the utility of detected event patterns by dynamically adapting the noise added to an unprotected data stream. We introduce a new guarantee named pattern-level differential privacy (DP), which enables us to apply and compare the strength of PPMs at the pattern level. We propose new pattern-level PPMs yielding pattern-level DP and analyze different trust settings of these PPMs and their requirements for context knowledge in the CEP system, e.g., the deployed queries. Our evaluation is based on three datasets (two real-world, one synthetic) and shows that the proposed PPMs increase data utility while preserving the same privacy level as the state-of-the-art PPMs. We use simulations to study the performance of our proposed PPMs in various practical scenarios. Furthermore, we demonstrate that computational complexity is not an obstacle to deployment.
000
Differential Privacy Papers @dppapers.bsky.social · 21/09/2026
Asymptotic Anytime-Valid Quantile Inference under Local Differential Privacy Leheng Cai, Qirui Hu, Shuyuan Wu arxiv.org/abs/2609.21338
Asymptotic Anytime-Valid Quantile Inference under Local Differential Privacy

Leheng Cai, Qirui Hu, Shuyuan Wu

http://arxiv.org/abs/2609.21338

Sequential quantile inference is difficult under local differential privacy because every record is randomized before reaching the analyst and the limiting quantile variance depends on an unknown density. We develop an online procedure that combines randomized response with dynamically chained parallel stochastic gradient descent (P-SGD). The resulting Polyak--Ruppert estimator admits a strong Gaussian approximation. A cross-chain quadratic statistic, computed entirely from private iterates, consistently estimates the limiting variance without a separate online density estimator. These results yield asymptotic confidence sequences and, under polynomial chain growth, asymptotic time-uniform coverage. Arm-wise constructions support locally private quantile best-arm identification, time-uniform simple-regret bounds, and sequential A/B tests of quantile treatment effects. Simulations and salary-data analyses illustrate the finite-sample behavior and practical use of the proposed methods.
000
Differential Privacy Papers @dppapers.bsky.social · 21/09/2026
Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees Shuo Huang, Gholamreza Haffari, Xingliang Yuan, Ting Yu, Lizhen Qu arxiv.org/abs/2609.21340
Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees

Shuo Huang, Gholamreza Haffari, Xingliang Yuan, Ting Yu, Lizhen Qu

http://arxiv.org/abs/2609.21340

Empirical identity leakage from released text is increasingly driven by attackers that combine large language models (LLMs) with auxiliary knowledge to link documents to individuals. Existing audits typically report success rates for specific attack pipelines but lack finite-sample statistical guarantees, while training-time protections such as differential privacy are difficult to translate into release-time decisions for individual natural-language documents. We introduce Conformal Privacy Auditing(CPA), a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries. CPA outputs a conformal ambiguity set of candidate identities that is guaranteed to contain the true identity with user-chosen confidence under exchangeability, together with an interpretable leakage proxy derived from set size. CPA supports both logit-access and sampling-only attackers, enabling audits of open-source models and proprietary API models in a unified framework. Across multiple release benchmarks and attacker configurations, CPA achieves calibrated coverage and reveals sharp shifts in certified identifiability as auxiliary knowledge, LLM augmentation, and release mechanisms vary, providing a statistically grounded basis for reporting and comparing release-time linkage risk across attacker configurations, datasets, and release mechanisms alike.
000
Differential Privacy Papers @dppapers.bsky.social · 18/09/2026
Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs Li Ge, Wenjie Qu, Weitao Feng, Yi Zeng, Jiaheng Zhang, Xiaofeng Wang, Wei Dong arxiv.org/abs/2609.20532
Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs

Li Ge, Wenjie Qu, Weitao Feng, Yi Zeng, Jiaheng Zhang, Xiaofeng Wang, Wei Dong

http://arxiv.org/abs/2609.20532

Wide adoption of machine learning has created growing policy and regulatory demand for protecting sensitive training data, with differential privacy (DP) emerging as a key mechanism. Yet a less-studied problem is how to certify the faithful execution of DP during training: an external verifier should be able to check that a released model was trained with proper DP protection, without accessing the private training data. Existing cryptographic approaches, such as zero-knowledge proofs, provide strong guarantees but often incur prohibitive overhead, in some cases by orders of magnitude. Trusted Execution Environments (TEEs) offer a more efficient alternative, but the multi-GPU TEE support needed for training and fine-tuning large language models remains limited to recent platforms and is absent or inefficient on legacy GPUs.
  To address this, we propose a practical framework for verifiable DP training using CPU-side TEEs together with untrusted GPUs. Our design addresses a fundamental efficiency-security tension: training entirely inside a CPU TEE is too slow, while unrestricted GPU offloading can allow malicious deviations from DP. We therefore offload expensive gradient computation to GPUs, while using the CPU TEE to efficiently verify the correct enforcement of DP on gradients through probabilistic checking. Our framework detects frequent full deviations from DP with high probability; for the utility-oriented forged-gradient attacks evaluated in this work, sparse deviations provide limited utility benefit and show no measurable additional membership leakage. Experiments further show that our approach nearly achieves a ``free lunch'': it incurs only modest overhead compared with standard GPU-based DP training, while effectively constraining malicious deviations from the
000
Differential Privacy Papers @dppapers.bsky.social · 18/09/2026
Empirical Analysis of Randomness Quality in Differential Privacy Mechanisms Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta, Massimo Caccia arxiv.org/abs/2609.20561
Empirical Analysis of Randomness Quality in Differential Privacy Mechanisms

Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta, Massimo Caccia

http://arxiv.org/abs/2609.20561

Differential Privacy (DP) relies on carefully calibrated random noise to protect individual privacy in statistical analyses. While theoretical work has analyzed DP under weakened randomness assumptions, the practical consequences of entropy degradation remain poorly understood. We present a systematic empirical investigation of how randomness quality affects differential privacy mechanisms using IBM's DiffPrivLib. We introduce progressively degraded entropy sources characterized by established test suites, starting from high-quality quantum True Random Number Generators (TRNGs) and cryptographically secure Pseudo-Random Number Generators (PRNGs) down to systematically manipulated sources with controlled entropy degradation. Through repeated experiments over one million queries on a reference database and complementary statistical tests, we directly analyze empirical Privacy Loss Random Variable distributions. Our results demonstrate that DP mechanisms reliably detect deviations when approximately 1 bit in every 8 to 16 is manipulated, with detection sensitivity varying significantly between bit-level biases and temporal correlations. We demonstrate that statistical detection of distributional anomalies does not necessarily correspond to actual privacy guarantee violations.
000
Differential Privacy Papers @dppapers.bsky.social · 17/09/2026
Intrinsic-Dimensional Wasserstein Guarantees for Private Synthetic Measures Yiyun He arxiv.org/abs/2609.17624
Intrinsic-Dimensional Wasserstein Guarantees for Private Synthetic Measures

Yiyun He

http://arxiv.org/abs/2609.17624

We study an $\varepsilon$-differentially private synthetic measure for $n$ points in $[0,1]^d$ by applying the existing PrivTree algorithm to construct an adaptive binary partition and then privately releasing its leaf masses. We consider the worst-case data model without any sampling or population-distribution assumption. The 1-Wasserstein error of the synthetic measure is $\widetilde O_d((\varepsilon n)^{-1/d})$ for $d\ge2$, which is optimal compared to the minimax lower bound up to a logarithmic factor.
  Moreover, for $d\ge3$ and $2<s\le d$, if the data set has covering number at most $Ar^{-s}$ over the relevant finite range of scales $r$, the expected error improves to $\widetilde O_{d,s}((\varepsilon n)^{-1/s})$. Thus the rate depends on a finite-scale intrinsic dimension rather than the ambient dimension, without requiring the recovery of a low-dimensional manifold. We also introduce a shifting technique to further avoid the exponential dependence of the constant on the ambient dimension $d$.
000
Differential Privacy Papers @dppapers.bsky.social · 17/09/2026
Tight Lower Bounds for Differentially Private Continual Counting Charlie Harrison, Ethan Leeman arxiv.org/abs/2609.17650
Tight Lower Bounds for Differentially Private Continual Counting

Charlie Harrison, Ethan Leeman

http://arxiv.org/abs/2609.17650

The Binary Tree Mechanism is a standard algorithm for differentially private continual counting, but its asymptotic optimality under pure differential privacy has remained unresolved since its introduction. We resolve this question. For fixed $0 < \varepsilon \le 1$, we prove asymptotically tight lower bounds of $Ω(\log^2 n)$ for worst-case expected $\ell_\infty$ error and $Ω(\log^3 n)$ for mean and maximum per-coordinate expected squared error. These bounds hold for arbitrary mechanisms, even when the entire stream is available in advance. The same lower bounds hold under approximate differential privacy whenever $δ\le n^{-c}$, for any fixed $c>0$. Our lower bounds match the Binary Tree Mechanism instantiated with Laplace noise, establishing its asymptotic optimality under both pure differential privacy and approximate differential privacy in the standard regime of $δ\ll1/n$. Our proof uses a single hard distribution with a bounded exponential score on a tree. A simple modification of the score allows the same framework to establish tight lower bounds for all three error measures.
020
Differential Privacy Papers @dppapers.bsky.social · 17/09/2026
QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing Shuaiqi Wang, Zinan Lin, Giulia Fanti arxiv.org/abs/2609.17995
QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing

Shuaiqi Wang, Zinan Lin, Giulia Fanti

http://arxiv.org/abs/2609.17995

Natural-language datasets support many downstream applications and research studies, but releasing text can reveal sensitive global properties of the underlying data source, such as the proportion of records associated with a particular gender, diagnosis, or political stance. Existing work has largely focused on property inference attacks that recover such global properties, while defenses for protecting these dataset-level secrets remain limited. Differential privacy, although effective for protecting individual records, provides only weak protection for aggregate properties. We propose Randomized Quantization for Text (QuanText), a training-free and large-language-model-agnostic data release mechanism that protects global secrets in textual datasets while preserving data utility. Given a dataset-level secret, such as the proportion of records with a particular diagnosis, and attributes whose utility should be preserved, such as topic and sentiment, QuanText perturbs both the secret distribution and the distributions of correlated attributes. It does so by constructing candidate release distributions over secret and non-secret attributes, randomly selecting a candidate sufficiently close to the private empirical distribution, and rewriting each private text sample to match the selected distribution using attribute-related snippets from the original text. QuanText is inspired by the Statistic Maximal Leakage (SML) framework, which bounds leakage about a secret function of a data distribution. Under idealized conditions, we show that QuanText satisfies an SML guarantee. Since these conditions may not hold exactly in practice, we also evaluate QuanText empirically on real-world datasets. Our results show that QuanText achieves a better empirical privacy-utility trade-off than competing data generation baselines.
000
Differential Privacy Papers @dppapers.bsky.social · 17/09/2026
Low-Rank Masking for Single-Server Matrix Multiplication Alejandro Cohen, Rafael G. L. D'Oliveira, Alex Sprintson arxiv.org/abs/2609.18876
Low-Rank Masking for Single-Server Matrix Multiplication

Alejandro Cohen, Rafael G. L. D'Oliveira, Alex Sprintson

http://arxiv.org/abs/2609.18876

We study the statistical privacy of outsourcing matrix multiplication over a finite field ${\mathbb F_q}$ to a single server using additive masks of rank at most $r$. For independent uniform $n\times n$ inputs, we show that uniform \emph{rank-ball masks} and products of independent uniform factors give maximal-correlation secrecy of at most $q^{-r}$ against the complete server view, with $O(n^2r)$ field operations for encoding and decoding. This secrecy captures how effectively the server is prevented from estimating functions of the inputs. We prove an asymptotically matching lower bound of this secrecy measure for $r=o(n)$, showing that both sampling methods are asymptotically optimal among input-independent additive masks of rank at most $r$, even when secret invertible transformations are allowed. We also characterize the posterior distribution for uniform rank-ball masks under arbitrary joint input distributions and prove approximate individual security for rows and columns under independent uniform inputs. Finally, we show that every input-independent additive mask of rank at most $r=o(n)$ requires $δ\to1$ in entry-level $(\varepsilon,δ)$-differential privacy for fixed field size $q$ and bounded $\varepsilon$.
000