Sign in

Cory McCartan

@corymccartan.com
698 followers 272 following 233 posts

Asst. Prof. of Statistics & Political Science at Penn State. I study stats methods, gerrymandering, & elections. Founder of UGSDW and proud alum of HGSU-UAW L. 5118. 🏳️‍🌈 corymccartan.com

PostsRepliesMedia
Cory McCartan @corymccartan.com · 15/09/2026
NYT poll of LV shows D+9 generic ballot ±3.2. Plugging that into my simple congressional swing model (tinyurl.com/cmchousemodel) gives 241D–194R, with 94% chance of D win. Effect of redistricting in that environment is R+4 seats, with TX and FL doing the heavy lifting, and NC slightly backfiring
010
Cory McCartan @corymccartan.com · 15/09/2026
From Alito dissent on mail ballots (L). How can this possibly be squared with the "major questions doctrine"!?!? (R from WV v. EPA)
010
Cory McCartan @corymccartan.com · 01/09/2026
He literally goes around with this slide. The same 'proof' would imply that no stationary AR(1) process with unbounded increments can exist. Error probability is not constant!
24210
Cory McCartan @corymccartan.com · 07/08/2026
Basic idea is to look at who doesn't get to elect their chosen candidate, but would have under a fair counterfactual plan. That difference = harm, which can be aggregated by & compared across different groups.
131
Cory McCartan @corymccartan.com · 03/08/2026
We also show random tree features have similar uncertainty quantification to full BART, despite not learning the tree structure
100
Cory McCartan @corymccartan.com · 03/08/2026
Enough theory—does this work in practice? We show random tree features are competitive with full BART, random forests, and xgboost, and lie along the Pareto frontier of accuracy & computation
100
Cory McCartan @corymccartan.com · 03/08/2026
Turns out (with some conditions) you only need n^(1/3) trees to achieve the optimal learning rate! Notably this rate is faster than existing rates for BART theory, which suffer from the curse of dimensionality* *not quite this simple, but basically
100
Cory McCartan @corymccartan.com · 03/08/2026
OK, but this is for infinitely many trees, right? What do we do in practice? We propose "random tree features," which are a BART model where the tree structure isn't learned
100
Cory McCartan @corymccartan.com · 03/08/2026
We can characterize the RKHS (function space) that BART's kernel corresponds to. It lies in between a weak class (just requiring 1 derivative; slow learning rates) and a smooth class (p derivatives; fast-ish rates), and has a nice minimax learning rate that depends only logarithmically on dimension!
100
Cory McCartan @corymccartan.com · 03/08/2026
The BART kernel is anisotropic, a.k.a. not rotation invariant, which means intuitively it prioritizes main effects over interactions (more about this below)
110
Cory McCartan @corymccartan.com · 03/08/2026
In fact, we show formally for the first time that as the # trees grows, BART converges to a Gaussian Process with a kernel we can describe in closed form! This means (a) no tree learning happens in the limit, and (b) we can study the kernel to learn things about the types of functions BART learns
100
Cory McCartan @corymccartan.com · 03/08/2026
We first try knocking out different parts of the BART model to see what matters most. Turns out, at least when you have many trees, it's (mostly) not being Bayesian, nor learning the tree structures, nor even the flexible part of trees per se
100
Cory McCartan @corymccartan.com · 03/08/2026
New WP w/@melodyyhuang.bsky.social studying the success of BART models, which regularly win causal inference competitions! We argue that BART should be thought of as a random features approximation to a limiting GP. This view helps understand BART & apply it in more places arxiv.org/abs/2607.28844
Seeing the Forest for the Trees: The Gaussian Process Limit of BART

Cory McCartan & Melody Huang

Abstract:
Bayesian Additive Regression Trees (BART) have shown state-of-the-art performance in both prediction and causal inference problems. Previous theoretical work has attempted to explain BART’s superior performance by establishing posterior contraction rates for standard BART models, but these rates depend strongly on the number of covariates. Here, we take a different approach and study the behavior of BART as the number of trees grows towards infinity. We show that in this regime, BART converges to a Gaussian process (GP) with a particular kernel. The kernel and its corresponding reproducing kernel Hilbert space (RKHS) have favorable inferential properties that help explain BART’s excellent performance. We introduce random tree features as an approximation to this limiting GP, and establish minimax-optimal learning rates for ridge
regression on these random features that depend only logarithmically on dimension. In addition to providing insight into the empirical success of BART, random tree features offer a computational benefit over traditional MCMC estimation. The random-features approximation also allows
practitioners to easily incorporate BART into any model which has a linear predictor, expanding the applicability and flexibility of BART.
192
Cory McCartan @corymccartan.com · 30/07/2026
Out today in the APSR: our paper studying which gerrymandering reforms work! Long story short, do what Michigan does! Slightly longer story: we project these different complex reforms onto a 1-dimensional "partisan leeway" axis, and do continuous-trt DiDiD there. doi.org/10.1017/S000...
First page of paper
140
Cory McCartan @corymccartan.com · 22/06/2026
Had a great time presenting Shiro Kuriwaki's and my new EI methods to EPSS in Belfast on Saturday! Our R package implementing all these methods (as well as partial ID bounds) is also now on CRAN! corymccartan.com/seine/
050
Cory McCartan @corymccartan.com · 11/06/2026
New R package on CRAN! ONNX is a runtime & file format for ML models. 'onnxr' lets you load & run models in about 2 lines of code! E.g. image detection running in ms from a pretrained model. Perfect for embeddings, smaller models, etc. And lots of .onnx available online! corymccartan.com/onnxr/
R code snippet:
library(onnxr)
model = onnx_model(model_path)
res = onnx_run(model, img, simplify = TRUE)Vermeer's "The Milkmaid" with person, bowls, dining table detected by ML model running via onnxr
010
Cory McCartan @corymccartan.com · 03/06/2026
Roberts court: "the mere fact of racially polarized voting is not relevant to proving racially polarized voting patterns"
The District Court also failed to follow our instruction in Callais that the mere fact that voters of differ-ent races vote for different parties is not relevant to proving racially polarized voting patterns. See id., at ___ (slip op., at 30).
032
Cory McCartan @corymccartan.com · 29/05/2026
Newly updated simple congressional model (tinyurl.com/cmchousemodel), now with LA map & recalibrated election model fit only to 2026-2024 data! Ds have lost 4 seats on avg to redistricting, but the loss grows to 5.6 seats in a D+6 environment. Ds 98% to win house in a D+6 environment, though
000
Cory McCartan @corymccartan.com · 29/04/2026
The NYT calls it a 'blow' to the VRA, but Thomas and Gorsuch are clear on what the effects of the decision will be: "an end" to the "misadventure" of "roughly proportional representation" for minority groups.
011
Cory McCartan @corymccartan.com · 29/04/2026
The egregious decision in Rucho now lets SCOTUS punch a massive loophole for VRA §2: partisan goals can justify disparate racial impacts. This is about districting, but the same logic would let states resurrect literacy tests, etc. as long as the stated goal was partisan discrimination.
plaintiffs’ illustrative maps must satisfy two conditions: Plaintiffs can-not use race as a districting criterion in drawing illustrative maps, and
illustrative maps must meet all the State’s legitimate districting ob-jectives, including traditional districting criteria and the State’s spec-ified political goals.
112
Cory McCartan @corymccartan.com · 22/04/2026
Time to plug my simple congressional model again, which I've updated with the new VA plan: tinyurl.com/cmchousemodel. Net effect of re-redistricting is D+1.5 seats in tied environment, growing to D+2 seats in a D+10 environment. Bit more competitive map, too.
021
Cory McCartan @corymccartan.com · 24/03/2026
New WP! Philip O'Sullivan, Kosuke Imai, and I generalize an earlier SMC sampler for redistricting plans, allowing it to scale better & be applied to multi-member districts, too, like the Dáil Éireann (Irish parliament) below. Look for a new 'redist' version soon(ish)! arxiv.org/abs/2603.22188
020
Cory McCartan @corymccartan.com · 27/02/2026
v0.2 of `bases` is up on CRAN! `bases` brings nonparametrics into your favorite modeling functions—using random Fourier features is as easy as sticking `b_rff()` into your formula. v0.2 brings graph Fourier & random convolutional features, `mgcv` integration, & more! corymccartan.com/bases/
Demo of random convolutional featureExample graph Fourier feature on a large lattice graphDemo of 'mgcv' integration
031
Cory McCartan @corymccartan.com · 13/02/2026
these people are sick
110
Cory McCartan @corymccartan.com · 13/01/2026
Last fall I shared new methods research with @shirokuriwaki.bsky.social on ecological inference—inferring individual relationships from aggregate data. Our new review WP frames past EI methods as linear models, and argues credible EI requires controlling for covariates arxiv.org/abs/2601.07668
The Role of Confounders and Linearity in Ecological Inference: A Reassessment

Abstract: Estimating conditional means using only the marginal means available from aggregate data is commonly known as the ecological inference problem (EI). We provide a reassessment of EI, including a new formalization of identification conditions and a demonstration of how these conditions fail to hold in common cases. The identification conditions reveal that, similar to causal inference, credible ecological inference requires controlling for confounders. The aggregation process itself creates additional structure to assist in estimation by restricting the conditional expectation function to be linear in the predictor variable. A linear model perspective also clarifies the differences between the EI methods commonly used in the literature, and when they lead to ecological fallacies. We provide an overview of new methodology which builds on both the identification and linearity results to flexibly control for confounders and yield improved ecological inferences. Finally, using datasets for common EI problems in which the ground truth is fortuitously observed, we show that, while covariates can help, all methods are prone to overestimating both racial polarization and nationalized partisan voting.
130
Cory McCartan @corymccartan.com · 03/12/2025
Interestingly, the TX gerrymander doesn't have any net impact in a D+12 environment Plots below show new effect of redistricting this cycle—cf left with the changes vs right if the 3-judge ruling holds
031
Cory McCartan @corymccartan.com · 03/12/2025
Davidson/Nashville swinging D+25 is wild Most swing districts next year won't look the way TN-07 does, with mainly rural counties that are shifting D by less
130
Cory McCartan @corymccartan.com · 03/12/2025
D+10-15 national environment would mean Dems end up ~100 seats over Rs in the House tinyurl.com/cmchousemodel
020
Cory McCartan @corymccartan.com · 18/11/2025
IF it holds, this would net Dems 1.6 seats, on average, due to mid-decade redistricting (D+1 seat in a Dem-favoring environment)
000
Cory McCartan @corymccartan.com · 17/11/2025
Net change in avg. Dem seats by state:
210
Cory McCartan @corymccartan.com · 17/11/2025
D+8.5 (±2) would be 245 Dem seats on average
table showing an average outcome of 245 to 190 under an 8.5 point Dem national environment (tinyurl.com/cmchousemodel)
062
Cory McCartan @corymccartan.com · 17/11/2025
Have updated my simple House model spreadsheet with currently enacted districting plans. Net effect is R+0.4 seats on average (!), with actually a _Dem_ advantage past a D+8 national environment. Copy, edit, & explore for yourself: tinyurl.com/cmchousemodel
230
Cory McCartan @corymccartan.com · 05/11/2025
NYC precinct map: Mamdani now vs the primary Purple = relative improvement vs primary Orange = relative loss vs primary (e.g. GOP voters) Takeaway? Mamdani improved significantly with Black voters since June! We see this in EI estimates as well: Mamdani likely won Black voters ~ 52/43 vs Cuomo
032
Cory McCartan @corymccartan.com · 05/11/2025
Precinct-level NYC data: Mamdani exceeded our voterfile-based expectations, Cuomo slightly exceeded them; Sliwa fell way behind, especially in places he was expected to do well!
100
Cory McCartan @corymccartan.com · 05/11/2025
In VA Gov precinct data, we are seeing a ~6pp shift on Spanberger vote share (y axis) versus 2024 president (x axis). Bit smaller in GOP precincts and bit larger in Dem precincts cf an R+5.5 shift (on vote share) from Biden '20 to McAullife in '21
020
Cory McCartan @corymccartan.com · 05/11/2025
Here at the CBS News data desk with @chriskenny.bsky.social and @simko.bsky.social! Looking at the VA numbers
0122
Cory McCartan @corymccartan.com · 04/11/2025
This election night I will be working the Data Desk at CBS News, focusing on the NYC mayoral race! Will try to post some things we are seeing in our precinct-level data and analyses, and maybe some cool maps like this one of Mamdani vs Harris support
130
Cory McCartan @corymccartan.com · 21/10/2025
As we showed in our paper, `seine` can strongly outperforms existing methods, which generally do not control for covariates (or do not do so efficiently)
100
Cory McCartan @corymccartan.com · 21/10/2025
Instead of a plot, you can also calculate a robustness value, which is a single-number summary of each estimand's sensitivity
100
Cory McCartan @corymccartan.com · 21/10/2025
The fun does not stop there! `seine` lets you immediately turn around and conduct a sensitivity analysis on your estimates. The `ei_sens()` function returns a data frame with different sensitivity parameters and biases. By default this can be plotted! Benchmarking to observed covariates works too!
100
Cory McCartan @corymccartan.com · 21/10/2025
The `ei_est()` function then does DML to estimate your main quantities of interest, and returns them in a tidy format. You can subset to produce subgroup estimates, and estimate linear contrasts as well—all with proper uncertainty quantification
100
Cory McCartan @corymccartan.com · 21/10/2025
Then you can quickly & easily fit a regression model and a Riesz representer to the EI specification These are implemented efficiently with the SVD, and the penalty is tuned automatically for the regression model!
100
Cory McCartan @corymccartan.com · 21/10/2025
You can then set up an EI specification which describes the problem and the covariates you will use.
100
Cory McCartan @corymccartan.com · 21/10/2025
`seine` first helps you preprocess your data into a format suitable for applying EI. Any messiness with proportions not adding to 1, etc. is handled.
100
Cory McCartan @corymccartan.com · 21/10/2025
A few weeks ago I shared a new WP on doing ecological inference—learning individual relationships from aggregate data, such as vote choice by race from precinct data. Excited now to introduce `seine`, our open-source R package for doing EI easily and efficiently! corymccartan.com/seine/
seine: Semiparametric Ecological Inference
Ecological inference (EI) is the statistical problem of learning individual-level associations from aggregate-level data. Without certain identifying assumptions and proper estimation methods, researchers can easily draw incorrect conclusions from aggregate data, finding patterns where none exist, missing important individual-level patterns, or even concluding an effect runs in the opposite direction than it actually does. This is known as the ecological fallacy.

The seine package allows researchers to perform modern ecological inference quickly, accurately, and transparently.

Double/debiased machine learning allows for controlling for confounding covariates, which increases the plausibility of identifying assumptions. Machine learning can be used to estimate the key regression model and avoid strong parametric assumptions made by existing EI methods.
Sensitivity analysis and benchmarking let researchers understand how violations of their assumptions will affect results.
A tidy interface makes the package modular and easy to use, and works well with pipe-based workflows.
Minimal dependencies and efficient estimation routines keep everything fast and lightweight.
173
Cory McCartan @corymccartan.com · 26/09/2025
The best news is that all of this is implemented in new software we've developed! I will write more about this next week. I've worked hard to make it ergonomic & efficient! corymccartan.com/seine/
110
Cory McCartan @corymccartan.com · 26/09/2025
This also leads naturally to a sensitivity analysis that lets researchers evaluate how violations of the key identifying assumption affect their inferences. (Heavily based off of the "Long Story Short" paper by Chernozhukov et al) We apply this to an air pollution example in the paper
100
Cory McCartan @corymccartan.com · 26/09/2025
How to do estimation with many covariates? Building on DML and Riesz regression, we propose a semiparametrically efficient series estimator that estimates nuisance functions within a restricted partially linear function space. You can get a root-n rate on your main estimand under ~weak conditions!
100
Cory McCartan @corymccartan.com · 26/09/2025
Take estimating vote choice by race, based on precinct data that have vote % and race % (but not both jointly). We (@shirokuriwaki.bsky.social and I) show that to do EI, you have to believe that each group's preferences are independent of the racial makeup of the precinct, given covariates
110
Cory McCartan @corymccartan.com · 26/09/2025
Very excited to share this week a new paper on EI that is two years in the making! What is EI? It's everywhere! EI is when you try to learn about individual relationships from aggregate data. We formalize identification & propose an efficient, assumption-lean estimator!
1111