Sign in

ArXiv Paperboy (Stat.ME+Econ.EM)

@paperposterbot.bsky.social
2.3K followers 2 following 27K posts

posts updates from arXiv rss feeds for methodology papers in Statistics and Econometrics. Also maintains an arxiv and posts random papers from it. maintainer: @apoorvalal.com source code: github.com/apoorvalal/bsky_paperbot

PostsRepliesMedia
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Asymptotic Null Distributions of Moran's $I$ and Assortativity in Large Networks By Braham, Rivest, Duchesne
This study investigates the asymptotic behavior of two dependence measures defined on networks, Moran's $I$ statistic and Newman's assortativity, under the null hypothesis that a Gaussian node attribute $Y$ is independent of the network structure. We demonstrate that the structure of the network directly affects the convergence rate to normality of these measures as the size of the network increases. We further establish that, in some instances, the mean values of these dependence measures under the null hypothesis remain non-negligible asymptotically and must therefore be explicitly accounted for when calculating the test statistics. Applications to a variety of simulated and real networks also reveal that the normal approximation performs well only when the network is not strongly heterogeneous. Network topology determines both the convergence rate to normality and whether the limiting distribution is Gaussian. In dense networks whose degree heterogeneity does not vanish, we further show that assortativity can fail to be a valid test statistic even though Moran's $I$ remains well behaved, whereas a dominating node invalidates both.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 High-Dimensional Statistical Inference for Sparse Support Vector Machines By Zeng, Huang
Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the $L_1$-penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Anytime-valid simulation-based hypothesis testing By Forr\'e, Brenner
For a given data distribution $(X_t)_{t \in \mathbb{N}} \sim Q$ i.i.d., we investigate the hypothesis testing problem: $H_0: Q = P_0$ vs. $H_1: Q = P_1$, for two different model probability distributions $P_0$ and $P_1$. In contrast to the standard setting, where analytic densities $p_0$ and $p_1$ are given, here, we consider the density-free setting, where we only have access to i.i.d. simulations $(Z^0_t)_{t \in \mathbb{N}} \sim P_0$ and $(Z^1_t)_{t \in \mathbb{N}} \sim P_1$. For this simulation-based hypothesis testing setting, we construct an e-test martingale, resulting in a sequential test with anytime-valid type-I error guarantees, approximate growth optimality, geometrically decaying type-II error bounds, and asymptotic power one. Most ingredients used in our constructions are variants of well known concepts. The value of this paper lies in the compact presentation of an effective, anytime-valid solution for the density-free simulation-based sequential hypothesis testing case.
010
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations By Pullu, Arslan, Balkac et al
Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.
010
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 A Functional Representation of Credit Behavior for Probability of Default Modeling By Brunholm, H{\o}jgaard, Nielsen et al
This paper proposes a framework for modeling probability of default via functional data analysis. By representing a series of credit variables as functions, we investigate whether intra-monthly information improves default predictability in linear models. We further show how a range of widely used variables, among them available funds, utilization rate, and overdraft, can all be derived from three quantities observed over time. Namely, 1) the type of each account, 2) the balance of the account, and 3) the size of the credit limit on the account. Retaining these quantities as continuous-time processes, rather than reducing them to monthly aggregated values, yields a continuous faithful representation of the borrower. When analyzing these three processes, we discovered that recurring events associated with the ordinal position among banking days created strong cyclical patterns. We therefore develop a relative time framework that aligns the recurring events across borrowers, ensuring that borrowers possess the same cyclical pattern, regardless of real time. We assess the framework using functional logistic regression. This approach accommodates the continuous representation while remaining closely related to a logistic regression model commonly used in credit risk practice, due to strict regulatory constraints. We show that, when equipped with an effective functional representation of transactional trajectories, the proposed model attains predictive performance on par with XGBoost while consistently outperforming logistic regression. Importantly, the model balances predictive performance of a machine learning model with the interpretability of linear default models, potentially enabling financial institutions to use the model, even under strict regulation.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects By Li, Ding
Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison. For leaderboard ranking, randomize-and-average selection is consistent under a mild separation condition, and a Williams square design can improve efficiency when item qualities are close. For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model's nonlinearity. Empirical results further support the validity of the proposed model-based inference beyond the first-order theory, including settings with higher-order sequence memory. We illustrate the approach in an application where AI judges compare two graphical model estimation methods.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Active Feature Acquisition for Cost-Efficient Temporal Prediction with Reduced Participant Burden By Science, Hill), Neuroscience et al
Accurate forecasting of pathological outcomes is a central problem in psychology. To do so, psychologists often collect intensive longitudinal data. However, in such studies, the desire to acquire a large number of variables for the sake of accurate prediction is often counteracted by the need to minimize participant burden. Acquiring more variables per occasion can yield better predictions, but having too many acquisitions increase the risk of non-response and attrition. Longitudinal Active Feature Acquisition (LAFA) is a principled approach to resolve this conundrum. Instead of requiring responses to every item at every acquisition occasion, LAFA produces a policy that seeks to optimally select dynamic subsets of items to be acquired at each timepoint while preserving our ability to forecast a specific outcome. However, existing LAFA methods are mostly based on Neural Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy. Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 HyperNSDE: Personalized Neural SDEs for Joint Static-Longitudinal Clinical Data Generation By Chassat, Guilloux
Synthetic patient data generation is a promising solution to the dual challenge of data scarcity and privacy constraints in healthcare machine learning. Realistic synthesis of patient-level clinical data requires jointly modeling heterogeneous static covariates, irregularly sampled longitudinal trajectories, and informative observation times - three tightly coupled components in practice yet rarely addressed together. We propose HyperNSDE, a continuous-time generative model that conditions a latent Neural SDE on static patient representations through a hypernetwork, allowing baseline characteristics to shape trajectory evolution beyond the initial condition without requiring a trajectory encoder, while stochastic latent dynamics capture realistic variability in generated paths. Observation times are modeled jointly through a latent-state-dependent intensity process, and training on irregular stochastic paths is stabilized via a deterministic-stochastic path decomposition with a non-adversarial signature-kernel objective. Experiments on simulated and real clinical datasets show improved observation-time fidelity and competitive performance, while matched-grid analyses reveal that forecasting and correlation metrics are affected by observation-grid regularity and trajectory smoothness.
010
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Assumption-lean logistic regression with missing covariates By Choudhury, Verchand, Samworth et al
Missing covariates are frequently encountered in supervised learning problems, and classical methods for estimation using such data use carefully chosen imputation schemes for missing data, or likelihood approximations that lead to nonconvex $M$-estimation problems. These methods and their relatives are suitable for scenarios in which the covariate distribution is known, and more broadly, have enjoyed tremendous success in linear models. But even in basic nonlinear problems such as logistic regression in moderate dimensions, such methods can experience drastic failure modes when the covariate distribution is unknown.
  Motivated by the need for reliable alternatives, we consider the problem of parameter estimation in logistic regression with missing covariates. Crucially, we operate in the assumption-lean setting where the covariate distribution is unknown (but bounded). We design a stochastic approximation method that is based on $Z$-estimation with a novel monotone operator, and establish that our algorithm is computationally efficient and achieves provable signal recovery at parametric rates under the hypothesis that covariates are missing completely at random. Our theory sharply characterizes the $\ell_2^2$ risk of the estimator in terms of the missingness profile, accommodating heterogeneous observation probabilities. Importantly, it shows that our method always outperforms the de facto ``complete-case'' estimator that ignores observations with any missing data. Even in the setting with homogeneous missingness (in which each covariate is observed independently with probability $q$), our bounds exhibit intricate and nonstandard dependence on $q$ that can yield significant improvements over using only complete cases. We complement our upper bounds with new information-theoretic lower bounds that show that this intricate dependence on $q$ is fundamental in a minimax sense.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Data Fusion for Errors-in-Variables By Zhao, Liu, Wang
We study errors-in-variables problems in which a target study contains only a single error-prone surrogate of an unobserved exposure, while an external source study provides repeated surrogate measurements from a different population. The measurement error distribution is allowed to depend on the observed error-free variables, and the error-free variable distribution itself may differ between studies. We introduce a conditional transportability assumption that enables the use of external repeated measurements under source-target heterogeneity. Together with additional replicate-error conditions, it identifies the target conditional measurement-error distribution. Building on this identification result, we develop a data-fusion estimator for a broad class of target functionals. The estimator combines conditional deconvolution, flexible nuisance estimation, and orthogonal correction that reduces first-order sensitivity to nuisance estimation. For the proposed estimator, we develop a unified spectral theory covering both diffuse-spectrum and finite atomic-spectrum target functionals, derive a general asymptotic expansion, and establish consistency and target-specific convergence-rate bounds. The resulting convergence-rate bounds depend jointly on the spectral properties of the measurement error, the latent exposure, and the target functional. For finite atomic-spectrum targets, we further establish joint Gaussian and bootstrap limits, yielding inference for smooth moment transformations under an additional centering condition. In the reported simulations, Fuse-EIV has small bias for the primary exposure-related coefficient. Applications to the National Health and Nutrition Examination Survey illustrate how accounting for population heterogeneity and error heteroscedasticity can change empirical conclusions.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Prediction-powered inference for time series across space By Rizvi, Burt, Srinivasan et al
The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data is observed over many spatial locations. For instance, crop yield might be observed over a large geographical area for recent years, but weather data (which is informative about crop yield) is available for a much longer period. The goal is to estimate, at each spatial location, the expected label (e.g., crop yield) in the future and provide a valid confidence interval for this value. The observed time period alone is too short for reliable estimates. Imputing missing labels with machine learning can cause substantial bias. Prediction-powered inference (PPI) can correct for this bias, but it relies on an i.i.d. assumption that breaks under our expected temporal dependencies. Heteroskedasticity and autocorrelation consistent (HAC) procedures account for temporal correlation, but have not been adapted to cases where some labels are imputed. We provide reliable point estimates and confidence intervals given: short labeled time series (across spatial locations), a longer unlabeled time series, and an imperfect predictor of labels given covariates. We show our method outperforms natural alternatives.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 When does conformal calibration need censoring weights? Cause-of-failure prediction sets under competing risks By Yang, Zhao
Split conformal prediction sets for competing-risks labels at a fixed horizon require calibration labels that right censoring can leave unobserved. Complete-case calibration guarantees coverage for the label-complete subpopulation, but its population coverage can deviate in either direction, even at the true class probabilities. We study how selection changes the score distribution near the population quantile. At the true score in our main simulation family, with independent draws, complete-case calibration covers 0.8723 at a nominal 0.900 when 22% of subjects are event-free at the horizon. In 4 of 48 further designs using per-draw normalisation, estimating cause incidences as one minus the exponential of the negative of the cause-specific Nelson-Aalen cumulative hazards, with the event-free probability as the clipped and renormalised remainder, puts complete-case coverage at least four standard errors below nominal. In these designs, weights from a correctly specified censoring model keep coverage near nominal. Extending the argument of Yi et al. (2025) to a cause label, we establish a finite-sample coverage lower bound with an explicit penalty for censoring-model error. Misspecifying the censoring model can also lead to under-coverage.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Measuring Gift Card Program Incrementality via Causal Data Fusion By Whitehouse, Betz, Zhang et al
Businesses regularly offer gift card programs to drive customer spending and increase engagement. A central question is how much incremental revenue these programs generate, and which channels drive it most efficiently. Measuring the incremental revenue associated with a gift card program is a challenging problem in causal inference, requiring a firm to infer how much each customer would have spent if they never received a gift card. Observational data on past customer purchasing behavior reveal possession of a gift card only when a customer makes a purchase, thus leaving a customer's treatment status systematically censored.
  In this paper, we develop a novel data fusion approach to overcome this missing data challenge. We identify and estimate incrementality by combining a large observational dataset with a smaller experimental dataset from a different population. Our approach relies on a mild transferability condition, which posits that the conditional relative treatment effect of gift card receipt on the decision to purchase is invariant across the two populations. We develop a flexible, machine learning-based estimator for the incremental revenue and establish its asymptotic normality.
  We apply our estimator across both first- and third-party channels through which Airbnb distributes gift cards, finding heterogeneity in incrementality across segments of the population. In particular, we find not only that third-party channels are more incremental than first-party ones, but also that "self-gifters" (i.e., customers likely to have purchased their own gift cards) are more incremental than the broader population.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Fridge: focused fine-tuning of ridge regression for personalize predictions By Hellton, Hjort
Statistical prediction methods typically require some form of fine-tuning of tuning parameter(s), with $K$-fold cross-validation as the canonical procedure. For ridge regression there exist numerous procedures, but common for all, including cross-validation, is that one single parameter is chosen for all future predictions. We propose instead to calculate a unique tuning parameter for each individual for which we wish to predict an outcome. This generates an individualized prediction by focusing on the vector of covariates of a specific individual. The focused ridge -- fridge -- procedure is introduced with a two-part contribution: 1) first we define an oracle tuning parameter minimizing the mean squared prediction error of a specific covariate vector, 2) then we propose to estimate this tuning parameter by using plug-in estimates of the regression coefficients and error variance parameter. The procedure is extended to logistic ridge regression by utilizing parametric bootstrap. For high-dimensional data, we propose to use ridge regression with cross-validation as the plug-in estimate, and simulations show that fridge gives smaller average prediction error than ridge with cross-validation for both simulated and real data. We illustrate the new concept for both linear and logistic regression models in two applications of personalized medicine: predicting individual risk and treatment response based on gene expression data. The method is implemented in the R package "fridge".
010
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Bayesian Machine Learning Methods For Large Scale Demand Estimation By Schmidt
This work studies how Bayesian machine learning methods can be used for large-scale demand estimation with many product categories. I compare two model classes, a latent factorization model and a mixed logit model and two Bayesian estimation approaches, Markov Chain Monte Carlo (MCMC) and Variational Inference (VI). The analysis combines a simulation study with an application to supermarket scanner data. The results show that the latent factorization model benefits from information across categories and improves its predictive performance as the dimensionality of the choice environment increases, whereas the mixed logit model does not exhibit the same pattern. MCMC delivers the highest predictive accuracy but is computationally intensive. VI achieves slightly lower predictive performance while substantially reducing runtime. In the empirical application, VI also outperforms the mixed logit benchmark. These findings highlight a trade-off between accuracy and computational feasibility in multi-category demand estimation.
020
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Elastic kernel Ridge regression, with applications in phonetics By Matteo, Stoecker, Pigoli et al
Predicting the shapes of entire curves requires accounting for nonlinear geometry and unknown alignments between curves. We develop elastic kernel ridge regression, a nonparametric method for planar curve responses with scalar, multivariate, or functional covariates. Using square-root velocity representations, we formulate penalized conditional Fr\'echet mean estimation in a shape space invariant to translation, rotation, scaling, and reparametrization. A vector-valued reproducing kernel Hilbert space provides a flexible nonlinear link through the spherical exponential map. We use an alternating algorithm to align each observed curve to its current fitted value, and updates the regression function. To facilitate this, we provide a new Euclidean quasi-Newton solver that exploits scale invariance of the reparametrization objective; this accelerates alignment while retaining accuracy in a numerical comparison. Simulations demonstrate the benefits of estimating alignment within the regression and respecting spherical geometry. Applied to vocal tract contours from a real-time magnetic resonance imaging recording, the method recovers missing frames and reconstructs shape trajectories from downsampled data. A speech inversion proof of concept on the same recording predicts tongue shapes from acoustic features, illustrating the method's potential. The method is implemented in the \texttt{R} package \texttt{sphereg2}.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Covariate-dependent Nonparametric $g$-modeling for regression via infinite Mixture-of-Expertizing class By Okazaki, Yano
Empirical Bayes $g$-modeling captures unit-level heterogeneity by estimating a latent prior distribution from observed data. In the existing formulations, however, the prior is shared by all units. In this paper, we develop a covariate-dependent g-modeling framework for regression in which the entire prior distribution of the regression coefficients is allowed to depend on covariates. We formulate the estimation of the prior as nonparametric maximum likelihood estimation (NPMLE) of the covariate-dependent prior, and show that the unrestricted problem is ill-posed. To resolve this, we introduce the infinite Mixture-of-Expertizing class of conditional priors, under which the NPMLE is precisely a softmax-gated Mixture of Experts (MoE) whose number of experts is not fixed in advance but is determined by the data. Building on a first-order optimality condition, we propose two exemplar-based estimation algorithms that select experts automatically, together with a post-hoc aggregation of experts for interpretation. On the theoretical side, we show that every conditional prior in the class is Lipschitz continuous in the covariates, and that aggregated softmax gates can approximate any continuous gate function. The effectiveness of the proposed NPMLE is shown through application to synthetic datasets and real datasets.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 3h
arXiv📈🤖 Bayesian heterogeneous copula mixtures with nonparametric margins: consistency, identifiability and tail asymmetry in physical fitness data By Liu
Finite mixtures of Clayton, Gumbel, Frank and Gaussian copulas can describe dependence that differs between the upper and lower tails. We study a two-stage Bayesian analysis in which ranks or kernel estimates replace the margins and the copula likelihood is evaluated at the resulting pseudo-observations. This pseudo-posterior is strongly consistent for the copula density whenever the log density admits a logarithmic boundary envelope. Both stages may use the same data, margins may be standardized within observed strata, and neither smoothness nor identifiability is required. The envelope holds for finite mixtures of Gaussian, Student, Clayton, Gumbel and Frank copulas in any fixed dimension; tail-dependence coefficients and conditional tail probabilities are therefore consistently estimated. We further prove that Clayton, Gumbel, Frank and Gaussian copulas are jointly finitely linearly independent, which makes mixture weights and components identifiable and consistently estimated. In simulations the pseudo-posterior matches multi-start maximum pseudo-likelihood in large samples. It is more stable in small samples and near independence (every component close to the independence copula), where tail coefficients are learned long before the weights and a marginal Metropolis sampler is up to twice as efficient as data augmentation. Two physical fitness datasets show mirror-image asymmetries. Among 8772 university students, sprint and jump performance are coupled mainly at the top. Among 5336 adults in a national health survey, low grip strength and low daily activity cluster together, increasingly with age, whereas high values do not. A Gaussian copula misses both patterns in held-out data.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Where Do Two Populations of Persistence Diagrams Differ? Calibrated Local Inference at a Fixed Budget By Bagchi, Bae, Mitra et al
Many two-sample tests for populations of persistence diagrams assess global differences without identifying the regions of the birth-death plane that contribute to them. We study simultaneous inference for local mean contrasts when the number of available diagrams is fixed. They are differences in expected weighted feature mass within $\ell_\infty$ neighborhoods at several centers and radii. We estimate these contrasts using additive landmark responses. A Gaussian multiplier bootstrap calibrates simultaneous confidence intervals while allowing unequal group covariances. The neighborhoods whose intervals exclude zero form a map with approximate family-wise error control, and selecting a subset of original intervals for display preserves their joint coverage guarantee. On the simultaneous coverage event, every reported neighborhood lies within twice its radius of the support of the mean-measure difference. A geometric result gives sufficient radius conditions for a displaced feature to produce a nonzero contrast. A comparison of sufficient detection thresholds quantifies the tradeoff between reducing the number of tested coordinates and reserving observations for an independent pilot. In simulations with 40 to 120 diagrams per class, the bands achieved 94%-98% simultaneous coverage under both the strict null and equal means with unequal covariances. In the latter setting, a permutation maximum and the pooled-t implementation of the two-stage persistence-image test of Moon and Lazar rejected in up to 32% and 26% of runs, respectively. In the fixed-budget simulations, spending a third of the observations on a pilot to choose landmarks or radii located changes less often than a prespecified grid at a single radius. On the MUTAG benchmark, the localized region concentrates on rings of fused-ring systems, an exploratory reading.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 The Missing Corner: When Sharp Causal Bounds Fail to Compose By Qi, Xiong
Sharp bounds on two potential-outcome means need not combine into sharp bounds on their difference. We characterize exactly when composition is guaranteed across environments sharing a joint potential-outcome law. For binary treatment and outcome, fix strict selection-probability bounds and interior treatment propensities. Subtracting the full model's sharp mean intervals gives sharp average-treatment-effect bounds for every compatible observed law if and only if the absolute selection bounds coincide across environments. Under this alignment, the complete mean region is rectangular; its endpoints and attaining causal models require $O(K)$ arithmetic operations for $K$ environments. For every mismatch, failure occurs on a relatively open set of positive observed tables at the prescribed propensities and persists under small joint drift. We give computable perturbation certificates and an exact family whose gap is first order in the mismatch. A separate finite-confounding family shows when the lost joint information changes an identified causal sign. These results distinguish local marginal sharpness from joint attainability after combining environments.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 The Noise Is the Signal: Correlated Sampling Error Is Rank-Informative for Proxy Metric Selection By Provenzano
North-star metrics such as customer lifetime value are often too slow and noisy to decide a short A/B test. Teams therefore rely on a proxy metric, commonly chosen by how closely its effects tracked the north star's across past experiments. Validating that choice, or any method for making it, is hard: the only benchmark is the noisy north star, and the number of available past experiments is limited. In addition, proxy and north-star effects are estimated on the same customers, so their sampling errors are correlated. Recent work at major experimentation platforms removes this shared error as contamination, improving estimates of the true-effect covariance. Choosing a proxy, however, is a ranking problem, and a better estimate need not give a better ranking. We measure agreement free of shared error by estimating the two effects on disjoint random halves of each experiment's customers. In an archive of 262 experiments and 69 candidate proxies, the shared error ranks the candidates in a similar order to this agreement (Spearman correlation 0.65): it carries information about proxy quality. The more of it a correction removes, the worse the ranking because removal discards part of the signal but leaves the main sources of ranking noise, the noisy north star and the limited number of experiments, untouched. Archive-calibrated simulations, in which the correct ranking is known, confirm this even when every correction receives the true sampling covariance. Held-out real experiments, evaluated on disjoint customer halves so that shared error cannot bias the comparison, closely reproduce the predicted ordering (Spearman correlation 0.93). Correction can still pay off with more experiments, but the number needed rises steeply with the north star's noise. We map this crossover and give platform teams three inexpensive checks for deciding from their own archive whether and how strongly to correct.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 TEASE: Targeted elastic spatial envelope By Krock
Envelope regression is a crucial part of multivariate linear modeling, leveraging separation into material and immaterial parts to provide parsimonious data reduction. Rekabdarkolaee et al. (2020) construct spatial envelope, a novel multivariate Gaussian process, extended by May et al. (2022) to linear coregionalization envelope. Our proposal (TEASE) combines spatial envelope and full-scale basis graphical lasso (LeDuc et al., 2025), leveraging targeted graphical elastic net (Kov\'acs et al., 2021) for immaterial precision matrices. Material part is a multivariate multiresolution Gaussian Markov random field (Kleiber et al., 2019; Caringi and Secchi, 2026). Enhanced envelope penalty (Kwon and Zou, 2025) regularizes regression coefficients. Demonstrations are on climate data.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 The Geometry of Existence and Uniqueness of Maximum Likelihood Estimation in Categorical Response Models By Sablica, Hornik, Rusch
Nonexistence of the maximum likelihood estimate (MLE) under separation is treated as a solved problem for binary logistic regression and as a scattered collection of model-specific results everywhere else. We show that it is one phenomenon with one criterion. In a latent polyhedral categorical response model, every observed outcome corresponds to a polyhedral event in latent variables whose faces shift linearly with the parameter. Random-utility choice, cumulative-link, ranking, multivariate binary and ordinal, sequential, adjacent-category logit, and fixed-score stereotype models belong to this family. The likelihood sees each observed factor only through the columns of its threshold map, the structure vectors. A finite MLE exists if and only if the pooled structure vector set has overlap. Sufficiency requires only continuity, necessity requires only strictly threshold-increasing probabilities, and neither likelihood concavity, exchangeability, nor full design rank is needed. Positive strictly log-concave latent densities and threshold-identifiable polyhedra then yield uniqueness on the estimable span. A single linear program determines whether overlap holds, and convex cone geometry measures the dimension of separation. An application using cumulative-link models for willingness to share health data identifies observations and model terms associated with nonexistence and illustrates how the diagnostics inform model revision and sensitivity analysis.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 What Should We Measure Next? Finding Identification Strategies by Refining Mechanisms By V\"a\"an\"anen, Cui, Cinelli et al
Canonical approaches in causal inference treat model specification as fixed, assuming that researchers directly translate all relevant domain knowledge into a causal model, which can then be used to deduce its logical implications. Yet, in practical applications, model specification is often an iterative process, and involves exploration and introspection: of the many aspects of the phenomenon one could investigate, which ones actually matter for the identification of the causal effect of interest? In this paper we study the problem of iterative identification in partially specified causal models. We focus on determining where observing variables that intercept a direct effect or a confounding path between two variables could enable identification in semi-Markovian models. We give necessary and sufficient graphical conditions for when observing such variables can render an unidentifiable query identifiable, together with an efficient algorithm for locating all such opportunities in a given causal diagram. Our results can help analysts better navigate the model space by drawing attention to the parts of their substantive knowledge that could result in a successful identification strategy.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Collaborative representations for targeted causal inference under outcome selection By Aguas
Estimating causal effects with missing outcomes requires learning outcome, exposure, and selection models. Flexible propensity learners can predict treatment or missingness well while worsening target estimation by emphasizing instruments, weak-overlap regions, or variation unrelated to outcome-regression bias. We propose a collaborative representation-based targeted learning method for recovered average treatment effects under outcome selection. The method learns low-dimensional exposure and selection representations from cross-fitted pseudo-outcomes that encode outcome-regression drift. A finite candidate library alternates targeted outcome updates with representation updates, and inner cross-validation selects representation complexity, collaboration strength, and regularization using a target-aware risk. We give high-level sufficient conditions for collaborative robustness, consistency, and asymptotic linearity after outer validation-fold targeting. Simulations show finite-sample bias reduction, with explicit bias-variance trade-offs across configurations and relative to competing estimators
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 The Topp-Leone XLindley Distribution: Properties, Estimation and Applications to Lifetime Data By M, Moharana, Maiti
In this article, we use a family of distributions developed by Topp-Leone to create a novel lifetime two-parameter distribution. The Topp-Leone XLindley distribution is the name used for it. We study several types of statistical and mathematical characteristics of this distribution, such as reliability functions, moments, moment-generating function, quantile function, and the Renyi entropy. For the purpose of estimating parameters under the proposed distribution model, this simulation study is conducted to evaluate the performance of the maximum likelihood estimates for the parameters of the Topp-Leone XLindley distribution. A comprehensive simulation investigation is carried out to evaluate the performance of the proposed methods. Furthermore, a practical data set of bank customer waiting times has been investigated for the purpose of illustration.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Scope-Restricted Backtracking Counterfactuals By Aguas, Jullum
Interventional counterfactuals evaluate statements such as "had $A$ been $a$" by replacing the assignment for $A$ while keeping the unit's exogenous background fixed. Backtracking counterfactuals reverse this logic: they preserve the structural assignments and instead vary background conditions across worlds. This semantics is natural for diagnostic explanation, but full backtracking may be too permissive because every exogenous coordinate can vary. We introduce scope-restricted backtracking counterfactuals, in which the modeler specifies which coordinates may vary while the remainder are shared across worlds. The scope determines the admissible carriers of the contrast, while a cross-world kernel governs the plausibility of their alternative values. We present model-based inference for general scope-restricted backtracking counterfactuals, derive sharp marginal bounds for a specified explanatory query, and establish frontier and compatibility conditions under which this explanatory query admits a modified-value policy representation. A linear-Gaussian example distinguishes local-noise, ancestor, irrelevant-noise, and infeasible explanations across scopes, while an application to the UCI Adult benchmark illustrates kernel specification, sensitivity, and how much of the explanation is carried by the cross-world coupling itself.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 State-space representation of early-life child mortality dynamics: linking event history analysis and orbit-based models By Lefotlha, Sherwell, Visaya
Understanding early-life child mortality requires methods that capture both statistically significant risk factors and the structure of individual life trajectories. Event History Analysis (EHA) estimates associations between covariates and mortality risk but may obscure rare or high-dimensional configurations of risk. Orbit Theory (OT) represents individuals as exact trajectories in a combinatorial state space, preserving full multivariate structure. We analyse longitudinal data across 12 variables for 31,081 children from the Agincourt Health and Demographic Surveillance System (South Africa, 1998-2008) using both methods. We show three things EHA alone cannot. First, maternal migration is structurally pivotal in the transition dynamics leading to child death despite being statistically non-significant in EHA. Second, maternal refugee status, while statistically significant, generates no trajectory dynamics and is correctly excluded from any structural reduction. Third, a three-variable subsystem (child status, mother status, mother migration) recovers EHA's principal findings within a reduced state space of at most 48 states while exposing small high-risk clusters invisible to regression-based analysis. Maternal death is the dominant EHA predictor; maternal migration organises 87% of the observable sequential transitions into child-death states. These are not contradictory: maternal death rarely appears as a preceding state because its role is terminal rather than sequential. Early-life mortality can be read at two levels simultaneously - statistically, through effect estimates, and structurally, through observed transitions.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Design-Driven Inference for Online Experiments: Jointly Optimal Evidence Collection, Testing, and Estimation By Yu, Wenbiao, Zhao et al
Online multifactorial experiments requires early stopping for practically negligible effects while controlling false stopping under continuous monitoring and retaining precise treatment-effect estimation. We develop a design-driven framework that chooses allocation and the evidence rule jointly, treating allocation as evidence collection and testing and estimation as complementary uses of the same evidence. For multifactor experiments with nuisance block effects and treatment-by-block interactions, we show that block-orthogonal designs remove nuisance contamination from the treatment score and form a complete class for worst-case evidence growth. Under a fixed information budget, an isotropic allocation and an explicit radial e-value jointly attain the minimax rate optimal for directionally unknown alternatives. The same allocation is universally optimal for maximum likelihood estimation of treatment effects, simultaneously achieving A-, D-, and E-optimality and establishing double optimality for testing and estimation. Batchwise replication yields an anytime-valid e-process, retains minimax rate optimality for cumulative conditional worst-case evidence growth, and preserves universal estimation optimality for the prespecified complete experiment. Simulations and a large language model prompt experiment illustrate gains in stopping efficiency and estimation precision.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Measuring Negotiation Traits with Collaborative Choice Data: A Two-Stage Latent Variable Model By Zhang, Chen, Hao et al
Consider a collaborative decision-making setting in which team members with different reward structures need to make a joint decision. Repeated collaborative decisions can reflect individual differences in team members' negotiation ability. However, extracting such information is challenging because each observed response arises from a collaborative process involving multiple participants and their reward structures. In this paper, we propose a two-stage probabilistic measurement model for structured collaborative choice tasks. In the first stage, known member-specific rewards and latent participant negotiation traits jointly determine team-option utilities, from which a latent team choice is generated. In the second stage, individual responses are modeled conditional on the latent team choice, with possible deviations that depend on participants' own rewards. Participant-level covariates can also be incorporated through a structural model. A simulation study shows satisfactory recovery of the participant traits and model parameters across different sample sizes and item sizes. An application to data from simulated collaborative negotiation tasks shows choice prediction accuracy well above the random-choice baseline and meaningful associations between the estimated traits and external criterion variables. The proposed framework provides a psychometric approach to measuring participant-level negotiation traits with collaborative choice data.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Vine Copula VAR:From Recursive Margins to Joint Forecast Inference By Ng, Tao
Joint-event forecasts often combine a dependence estimate based on past forecast errors with newly estimated marginal distributions. When each historical error retains the marginal fit available at its issue date, inference must account for an overlapping sequence of estimation errors. We derive their joint influence with the terminal forecast estimates in a stable Vine Copula VAR with normal innovation margins and a fixed, correctly specified Gaussian or positive Clayton vine. An intercept identity and the stable VAR filter reduce the historical correction to harmonically weighted innovation moments, while terminal slope uncertainty remains. The resulting covariance estimator gives asymptotically valid repeated-sample intervals for fixed one-sided event probabilities at the realized forecast state. In the Gaussian submodel, retaining issued transforms adds a positive semidefinite covariance term relative to refitting margins on the same observations. Monte Carlo simulations show that terminal-margin uncertainty is quantitatively more important than this additional term and that logit intervals improve lower-tail coverage in the designs studied. A real-time forecasting application to U.S. macroeconomic releases shows how marginal estimation contributes to uncertainty in predicted probabilities of joint contractions and identifies limitations of the stationary marginal model.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Debiased Machine Learning for Count Data: a Partially Linear Poisson Model Based on Neural Networks By Zhou, Shepherd, Xiang
We develop a novel debiased machine learning (DML) estimator to analyze count data, which are common in clinical or biomedical studies. Specifically, we estimate the effect of a binary treatment on count outcomes using a partially linear Poisson model. A key advantage of this model is that it represents the complex covariate structure through a nuisance function, which is estimated flexibly using neural networks. We use a Neyman-orthogonal score function to construct the estimator, reducing its sensitivity to errors in estimating the nuisance function. Cross-fitting is used to mitigate overfitting bias in neural networks. Under mild regularity conditions, this DML estimator is asymptotically normal with root-$n$ convergence. We derive a closed-form variance estimator and construct a Wald confidence interval for the treatment effect. Extensive simulations demonstrate that the proposed estimator reduces bias and root mean square error compared with the generalized linear model and the augmented inverse probability weighting (AIPW) estimator in most settings, while achieving a coverage probability near the nominal level. The proposed procedure is implemented in the R package $\texttt{PoissonDML}$. We apply our approach to a synthetic HIV cohort data to investigate the effect of treatment on AIDS-defining event counts of 7,141 analyzed patients. Our approach provides practical and valid estimation and inference for treatment effects under the partially linear Poisson model while combining modern machine learning methods to estimate complex covariate structures.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Large-scale linear hypothesis testing for high-dimensional online M-estimation By Ham, Yu, Jacobson
With the growing prevalence of streaming data, renewable estimation has become an important tool for updating statistical analyses without retaining historical raw observations. We study estimation and linear hypothesis testing in high-dimensional online M-estimation. Existing renewable procedures based on quadratic approximations can incur first-order approximation errors because the gradient of the unpenalized loss generally does not vanish at a preceding penalized estimator, posing a particular challenge for valid inference. To address this issue, we propose a corrected renewable loss that preserves historical first-order information and combine it with partial folded concave regularization and an iterative local linear approximation. We establish finite sample $\ell_1$- and $\ell_2$-error bounds, together with contraction bounds and strong oracle properties for both the constrained and unconstrained iterative estimators. Building on these results, we derive oracle Bahadur representations and establish asymptotically valid Wald and score tests for general linear hypotheses, allowing both the testing-oracle dimension and the number of restrictions to diverge. We further characterize their asymptotic power under local alternatives. The proposed procedure requires only recursively updated summaries while recovering the first-order inferential behavior of its pooled-data oracle counterpart. Simulation studies and real-data applications demonstrate the finite-sample performance of the proposed estimation and testing procedures.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Multiple Change Point Detection in the Mean and Covariance of High-Dimensional Matrix Time Series By Cen, Cho
This paper studies simultaneous change point detection in the first- and second-order structures of high-dimensional matrix-valued time series. Specifically, we extend the framework of main effect factor models to accommodate multiple change points so that shifts in the mean and covariance structures are characterized by changes in the trend-stationary main effects main effects and in the factor-driven common components, respectively. Under this framework, the two types of changes are not necessarily aligned and can be arbitrarily large without masking the effects of each other, and each of the detected change points is identified either with the row or column categorizations which improves their interpretability. We propose a unified procedure based on moving-sum statistics to detect and localize both types of changes. Under general regularity conditions that allow for temporal and cross-sectional dependence, we derive detection and localization guarantees for the proposed method. The finite-sample performance of the method is demonstrated through extensive experiments, and its practical usefulness is illustrated by a real data application to New York City Yellow Taxi trip records.
010
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 A perspective note on likelihood approximation and inference for complex simulation models using a chain of aggregated normalizing flows By Befekadu
We present a new perspective on the problem of likelihood approximation within the framework of simulation-based inference that promotes scalable and controllable simulation routines for large-scale data analysis, allows efficient parameter space exploration or smooth interpolation in high-dimensions and, thus, supports valid statistical treatments of hypothesis testings as well as uncertainty quantification. In particular, we consider a chain of $n$-aggregated normalizing flows for likelihood approximation scheme, where a set of upfront replicated observation datasets from the forward complex simulation model pass through the first set of bijective transformations, and then subsequently pass to the other sets of bijective transformations. Here, we assume that, for any $k \in \{1,\,2, \ldots, n\}$, the parameters corresponding to the first $k$ sets of bijective transformations are estimated sequentially, in some sense of optimality, for constructing flexible probability distributions, regardless of the remaining $(n-k)$ sets of bijective transformations. Moreover, our objects of interest are to highlight two complementary mathematical arguments that leverage an informatics-theoretic formalization, based-on empirical likelihood estimators under moment restrictions, and a sequential decision-making paradigm, with mixing distributions, for updating and aggregating the estimated parameters of the overall normalizing flows. As a by-product, the framework provides a reliable surrogate model, conditioned on the model parameters defining the forward computational simulation, that allows samples generation, with statistical powers, and facilitates computationally tractable scheme in the Bayesian paradigm for inference, hypothesis testings and uncertainty quantification.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Network Experiments with Edge Treatments and Node Outcomes By Kuriksha, Hung
We present a methodology for analyzing node-level outcomes while experimenting with edge-level treatments in a population connected by an undirected graph. Under our design, nodes are randomly assigned to test or control, and each edge inherits the treatment of its endpoints, with conflicts resolved by randomization. We use each node's assigned status as an instrument for its treatment exposure. We show that the Wald estimator is consistent for the global average treatment effect (GATE), even when the edge weights used to construct the exposure are misspecified. We formalize the assumptions needed both in terms of the linearity of potential outcomes and the sparsity of the graph, and prove the asymptotic normality of the Wald estimator. The estimator is straightforward to implement, requiring no assignment simulations over the graph. Our Monte Carlo study demonstrates the strong performance of our approach in the context of a social platform.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 4h
arXiv📈🤖 Existence and consistency of weighted maximum likelihood estimator for extreme quantile regression By Vidagbandji, Berred, Bertelle et al
Estimating extreme conditional quantiles faces two major challenges: understanding complex nonlinear relationships between variables and accurately extrapolating into the tails of the distribution, where data are sparse. To address both issues, we introduce a novel approach that combines the theoretical framework of block maxima with the predictive power of generalized random forests. The conditional distribution of maxima is modeled by a generalized extreme value (GEV) distribution whose parameters depend on the covariates and are estimated via a weighted maximum likelihood procedure, with weights derived from generalized random forests to capture complex high-dimensional structures. The conditional quantile estimator follows from the inversion of the estimated GEV distribution function. We establish theoretical results ensuring the existence and consistency of the proposed weighted maximum likelihood estimator. The method provides a robust and flexible framework for extreme quantile regression. The performance of the proposed approach is illustrated using simulated data, as well as an application to financial portfolio losses from stocks listed on the NYSE, AMEX, and NASDAQ.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 5h
arXiv📈🤖 Empirical-Bayes spectral partial pooling across related tasks By Mauri
Spectral methods are central to high-dimensional statistics and machine learning, underlying procedures for covariance estimation, matrix denoising, representation learning, clustering, and latent variable modeling. In this work, we focus primarily on spectral estimators for high-dimensional factor models, where leading singular vectors of the data matrix are used to estimate latent structure and covariance parameters. In many modern applications, however, data are collected across related but heterogeneous tasks, studies, domains, or populations. Applying such estimators separately to each task can lead to unstable estimates when sample sizes are limited relative to dimension, while completely pooling the data can obscure meaningful task-specific structure. We introduce Hierarchical Spectral Shrinkage (\texttt{HSS}), a scalable empirical-Bayes framework for partially pooling task-specific spectral estimators. The method regularizes the leading spectral directions of each task toward a data-adaptivelylearned common basis, while allowing the amount of shrinkage to vary across tasks and spectral components. The proposed estimator arises as the posterior mean in a surrogate Bayesian regression formulation of the empirical singular vectors. For factor models, \texttt{HSS} yields partially pooled estimators of the group-specific loading spaces and covariance matrices and can be combined with different spectral estimators by replacing their empirical singular vectors with hierarchically regularized counterparts. More broadly, the same construction provides a mechanism for extending spectral estimators to collections of related but heterogeneous datasets. We demonstrate substantial improvements in estimation accuracy and out-of-sample performance over separate, fully pooled, and shared-subspace approaches in synthetic experiments and multi-study gene-expression data.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 5h
arXiv📈🤖 lifelines-hc: Higher Criticism testing for sparse non-proportional hazard departures in Python By Kipnis, Galili, Yakhini
The log-rank test is the standard tool for two-sample survival comparison and has good power against proportional-hazards alternatives, but it loses power against sparse hazard departures, in which the hazard difference is concentrated in a small number of time intervals whose locations along the follow-up are unknown a priori (Kipnis, Galili and Yakhini, Biometrika 2026). Weighted log-rank tests -- Gehan-Wilcoxon, Tarone-Ware, Peto-Prentice, Fleming-Harrington -- each impose a pre-specified temporal emphasis. Combination procedures such as MaxCombo and the Yang-Prentice short-term/long-term hazard-ratio model relax that choice, but still scan only a small dictionary of global temporal shapes and remain insensitive to sparse departures occurring outside them.
  We introduce lifelines-hc, a Python package extending the lifelines survival library with the HCHG test: Higher Criticism applied to per-interval hypergeometric p-values. HCHG applies to right-censored two-sample data and detects sparse hazard departures at unknown locations without committing to any temporal pattern. Across four clinical case studies -- CheckMate 057 PFS (n=582), COMET-1 OS (n=1028), AZURE DFS (n=3359), and the Copenhagen Study Group for Liver Diseases cirrhosis trial (n=446, fully public individual patient data) -- HCHG achieves p <= 0.014 in every dataset, while the log-rank test is non-significant throughout (p >= 0.26). MaxCombo and the Yang-Prentice adaptive log-rank test detect the delayed-benefit crossing in CheckMate 057 (p < 0.001) but are non-significant on the remaining three (p >= 0.24), where the departure is sparse or multi-window rather than a smooth short-term/long-term hazard-ratio pattern.
  lifelines-hc is freely available under the MIT license at https://github.com/alonkipnis/lifelines-hc and via pip install lifelines-hc.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 5h
arXiv📈🤖 A Bayesian Multiscale Integrated Abundance Model for Estimating Latent Opioid Misuse Prevalence from Spatially Misaligned Data By Hepler, White, Cerda et al
Estimating small area prevalence of opioid misuse is critical for targeting public health interventions, yet direct measures are unavailable and related surveillance data are often reported on misaligned areal units. We propose a Bayesian Multiscale Integrated Abundance (MIA) model for estimating latent opioid misuse prevalence by jointly analyzing multiple indirect surveillance indicators observed on different geographic supports. The model extends existing integrated abundance models by representing all source geographies through a common set of atomic spatial units formed by the intersections of observed areal supports. Latent prevalence at the atomic level is modeled using a Fisher noncentral hypergeometric distribution, which preserves county-level prevalence totals while respecting local population constraints. To enable scalable inference, we develop a two-stage compositional Markov chain Monte Carlo algorithm that combines customized sampling strategies and parallel computing. Simulation studies show the proposed approach reduces bias and root mean squared error relative to common downscaling methods. We apply the model to Ohio data from 2010-2023, integrating state-level survey estimates, county-level counts of opioid overdose deaths and treatment admissions, and ZIP code tabulation area-level counts of emergency medical services naloxone administrations. Results reveal substantial within-county heterogeneity and identify localized areas of elevated opioid misuse prevalence that would be missed by county-level analyses.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 5h
arXiv📈🤖 Poisson empirical Bayes estimation of sums of random variables via minimum-distance methods By Favaro, Fortini, Jana
The estimation of sums of functions of observable and unobservable variables is a long-standing problem in statistics, with applications in many domains. We consider this problem in Poisson mixture models, where empirical Bayes provides a natural framework but nonparametric theory remains limited. We develop a nonparametric empirical Bayes methodology based on regular and coarsened minimum-distance estimation of the unknown mixing distribution. For a broad class of such sums, we establish large-sample guarantees showing that the resulting plug-in estimates asymptotically merge with the oracle Bayes estimate. In particular, when the mixing distribution has finite support, we obtain a nearly parametric convergence rate, up to a logarithmic factor. We then provide a finite-sample analysis of two representative sums: the total intensity among units whose observed count does not exceed a fixed threshold, and the number of units whose observed count exceeds their latent intensity. For the total-intensity sum, we establish a lower bound on the minimax regret and derive upper bounds for both regular and coarsened minimum-distance procedures under compact-support and subexponential assumptions on the mixing distribution. These upper bounds match the minimax rate up to logarithmic factors, with the coarsened procedure in the compact-support setting matching the rate exactly when the coarsening level is fixed. The above-average sum displays a markedly different behavior: a lower bound on the minimax regret shows that bounded regret is in general impossible to achieve. We identify additional stability conditions under which bounded regret can be recovered up to logarithmic factors, and show that these conditions are automatically satisfied when the mixing distribution has finite support. Numerical experiments on both synthetic and real data illustrate the performance of the proposed methodology.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 5h
arXiv📈🤖 Multi-Task Active Learning with Efficient Resource Allocation By Ye, Zhang, Qu
Many scientific studies allow costly auxiliary information to be collected during data labeling but not at deployment. Examples include diagnostic tests, laboratory assays, and expert evaluations. We study prediction under this deployment asymmetry, where auxiliary variables are selectively acquired during labeling under a budget constraint but systematically unavailable at prediction time, creating a missing-by-design problem that couples data acquisition, surrogate construction, and prediction. In this work, we introduce Active Learning with Cost-Adaptive Task Resource Allocations (ALCATRAs), a unified framework for selectively acquiring auxiliary information under resource constraints and leveraging that information to improve downstream prediction. ALCATRAs consists of two main components: a task-selection policy which strategically selects a sequence of cost-effective tasks for unlabeled data to perform, and a surrogate learning procedure which transfers knowledge from completed tasks to enhance model predictions. In theory, we show the effectiveness of surrogate models and sample-efficient task policies in improving the model's prediction error bound. Simulation studies and an application to the UCI heart disease cohort demonstrate improved sample efficiency of the proposed ALCATRAs framework relative to baselines under the studied settings.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 13h
arXiv📈🤖 Robust inference in inflated beta regression By Queiroz, Ferrari
The inflated beta regression model is widely used for modeling continuous proportions with values at the boundaries. Maximum likelihood estimation for these models is well-known for its sensitivity to outliers, which can severely distort inference and lead to misleading conclusions. We propose robust estimators that mitigate the lack of robustness in maximum likelihood-based inference while preserving the simplicity and interpretability of the inflated beta framework. Additionally, an algorithm is introduced to select tuning constants based on the data's robustness requirements. The proposed estimators' asymptotic and robustness properties are studied, and robust Wald-type tests are developed. Simulation studies and a real data application highlight the advantages and practical effectiveness of the proposed robust estimators.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 18h
arXiv📈🤖 Structural Classification of Locally Stationary Time Series Based on Second-order Characteristics By Qian, Ding, Li
Time series classification is crucial for numerous scientific and engineering applications. In this article, we present a numerically efficient, practically competitive, and theoretically rigorous classification method for distinguishing between two classes of locally stationary time series based on their time-domain, second-order characteristics. Our approach builds on the autoregressive approximation for locally stationary time series, combined with an ensemble aggregation and a distance-based threshold for classification. It imposes no requirement on the training sample size, and is shown to achieve zero misclassification error rate asymptotically when the underlying time series differ only mildly in their second-order characteristics. The new method is demonstrated to outperform a variety of state-of-the-art solutions, including wavelet-based, tree-based, convolution-based methods, as well as modern deep learning methods, through intensive numerical simulations and a real EEG data analysis for epilepsy classification.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 06/10/2026
arXiv📈🤖 Addressing the Between-Group Comparison Problem: Detecting Differences Between Correlation Matrix Populations due to Single-variable Perturbations for Resting State fMRI By Faran, Peer, Arzy et al
Resting-state fMRI has been known for decades as a promising method for evaluating cognitive and mental states, both in health and especially in disease, due to its ease of implementation as a short, standard MRI protocol. In clinical settings, a group of patients with a given disorder is typically compared to a group of healthy controls. This poses an inherent challenge of between-group comparison. We propose a new efficient model for characterizing changes to the temporal synchronization of brain activity measured using RS-fMRI between groups, summarized as individual correlation matrices. Our model posits that the between-group differences are the product of single-region effects describing the increase or decay of synchronization with the rest of the brain. This parsimonious model pools the correlation coefficients of each region with all others, and therefore can detect differences between groups even in small samples. Inference for this model accounts for the variability in individual correlation matrices, the within-group differences across individuals, and for the approximation error of the single-region model. This results in per-region estimates and confidence intervals for the parameters governing the difference between groups. In simulations, our model shows increased power to detect model-aligned alternatives compared with competing approaches. To demonstrate feasibility of the method in a clinical application, we use the model to analyze RS-fMRI correlation matrices in patients with transient global amnesia and healthy controls. Our model detects significant decreases in synchronization for the patient population in the amygdala after multiplicity correction as well as borderline decreases in memory-related brain regions that were not detected using mass-univariate tests without prior knowledge, suggesting its usefulness in the application of RS-fMRI in clinical settings.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 06/10/2026
arXiv📈🤖 Direct estimation and inference of differential Granger causality between two high-dimensional time series By Wang, Ma, Shojaie
Comparing dependence structures between two related multivariate time series is of fundamental interest in many scientific applications, where changes may occur in both directed temporal interactions and contemporaneous connectivity. Modeling each time series by a vector autoregressive (VAR) model, we propose a new framework for estimation and inference of differential Granger causality (DiffGC) and differential network (DiffNet) structures in high dimensions. The proposed method is based on a novel estimating equation derived from the Yule-Walker equations that directly links the difference between VAR transition matrices to the corresponding difference between precision matrices. Unlike separate estimation strategies that estimate the two VAR models individually and then take their difference, the proposed method directly targets the differential structures and requires sparsity only of the differences, thereby accommodating potentially dense individual networks, including hub structures. Building on the direct estimators, we develop selective inference procedures for both DiffNet and DiffGC parameters, providing valid post-selection inference while accounting for the data-driven screening process. Theoretically, we establish convergence rates and support recovery guarantees for the proposed estimators, derive asymptotic distributions for the selective inference targets, and obtain new consistency results for DiffNet estimation under weaker assumptions than existing methods. Simulation studies confirm the theoretical convergence rates and demonstrate accurate support recovery and robust inferential performance. An application to resting-state electroencephalography (EEG) data identifies substantial changes in both contemporaneous and Granger-causal connectivity across recording sessions.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 06/10/2026
arXiv📈🤖 Inverse Cross-spectral Neural Networks for Multivariate Time Series By Marinucci, Nino, D'Acunto et al
CoVariance Neural Networks and their extensions have emerged as effective tools for processing multivariate data, deriving graph shift operators directly from second-order statistics. These architectures, however, are designed for independent and identically distributed observations and do not fully capture the joint structure of temporal and cross-variable dependencies in multivariate time series. In this work, we introduce Inverse Cross-Spectral Neural Networks (iCSNNs), a class of graph neural networks for stationary multivariate time series whose shift operators are the inverse cross-spectral density (iCSD) matrices. These operators encode frequency-specific conditional relationships among variables, exploiting the decomposition provided by the spectral representation theorem. Leveraging spectral smoothness, frequencies are grouped into bands sharing a single iCSD operator, yielding a compact parametrisation that retains the frequency-dependent structure of the process. We further propose a joint learning procedure to estimate both the Fourier-domain dependence structure and the iCSNN parameters, adapting the iCSD operators to the downstream task. When tested on synthetic data, iCSNN outperforms baselines from different methodological families.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 06/10/2026
arXiv📈🤖 Valid Stopping in Adaptive Generator-Verifier Loops By Hegazy, Jordan, Dieuleveut
Numerous agentic workflows are based on a generator-verifier loop: a generator proposes candidates, a cheap verifier scores them, and the workflow terminates when a proposal is verified as good enough. The verifier typically proxies a more costly ground-truth oracle, and as the generator searches adaptively against it, false acceptances may accumulate. Proposals can pass the proxy but fail under the costlier ground-truth check. We study when to stop these loops while controlling the false discovery rate of the accepted proposals. Our construction introduces tools of independent interest in distribution-free statistical testing and conformal risk control, including analysis of $e$-values constructed through index betting and a novel conformal risk control procedure for non-monotone losses. We validate the approach in synthetic settings and on a protein-design benchmark.
000
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 06/10/2026
arXiv📈🤖 Measuring Trade Direction in a Prediction Market: Settlement Ground Truth and Trading-Cost Measurement on Polymarket By Dubach
Trade-sign errors can change measured trading costs even when classification accuracy is high. We validate the side of Polymarket's public trade prints against the taker leg of each print's on-chain settlement. On twelve selected days between April and August 2026, spanning both exchange generations, 92.1% to 100.0% of prints match a settled taker leg, with exact side and token agreement on all 24.5 million matched pairs. Mint-and-merge settlement makes that taker leg essential: pooling maker and taker legs changes the measured buy share. Signing the full cached tape before settlement selection yields 16.6 million prints signed by every rule; equal-day balanced accuracy is 0.942 for Lee-Ready, 0.768 for the tick test and 0.670 for retrospective bulk volume classification. On 16.4 million identical eligible fills, Lee-Ready raises effective spread by 0.429 cents per share under equal-fill weights; its realised-spread difference is -0.209 cents. Under share-volume weights, taker five-minute midpoint impact is 0.689 cents, while tick and bulk classifications give -0.374 and -0.119 cents. That aggregate sign reversal disappears when tied print rows are excluded. Distortion depends on error-weighted signed outcomes, sample selection and weighting. Receipt-time ordering and unobserved future-quote age limit these delivered-quote accounting quantities; they do not identify causal impact or private information. Supporting venue and collector analyses provide descriptive diagnostics.
010
ArXiv Paperboy (Stat.ME+Econ.EM) @paperposterbot.bsky.social · 06/10/2026
arXiv📈🤖 Finite-Sample Distribution Theory and Efficient Large-Scale Inference for Online Quantile Regression By Wei, Li, Wang et al
This paper studies online quantile regression for large-scale and streaming data using Stochastic SubGradient Descent (SSGD) with constant learning rates. Classical offline inference for quantile regression is computationally and memory intensive. Existing works of online inference for quantile regression provide only asymptotic guarantees and typically require sub-exponential tail conditions for distribution theory. To bridge these gaps, we introduce new techniques to prove a quenched central limit theorem (CLT) and finite-sample Gaussian approximation for SSGD under a finite-moment assumption. We further show that Ruppert-Polyak averaging with a constant learning rate has a non-vanishing bias and fails to satisfy CLT centering at the population target. Hence we propose suffix averaging to address this issue and establish its finite-sample Gaussian approximation. Based on these results, we provide an efficient online inference method for quantile regression that avoids covariance estimation. Numerical experiments show that our method achieves desirable empirical coverage rates and competitive performance compared to other inference methods. We also apply our approach to U.S. wage data to demonstrate its practical effectiveness.
000