Jack Fitzgerald @jackfitzgerald.bsky.social · 16/09/2026Yesterday, I had the honor of defending my dissertation and earning my PhD in Economics and Business Administration with cum laude distinction. In October, I will be joining the Meta-Research Innovation Center at Stanford as a Postdoctoral Scholar. I'm excited for the road ahead! 071
Jack Fitzgerald @jackfitzgerald.bsky.social · 03/07/2026In today's dose of dystopia, here's an ad I get hit with every time I use my preferred citation service 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 17/05/2026Ok what's the funniest part? I'm personally a fan of the unhinged 'best of luck' email signoff, though a close second is starting the email off with 'Hi Doctor, Hey.' 010
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026But controlling for contact probability erases most differences in business survey response, reducing maximum response rate differences between sectors from 32 to 9 p.p. Establishments registered to residential addresses get response rate deficits cut from 18 to 5 p.p.! 10/ 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026Most worrying: 70% of establishments in the LISA register are registered at a residential address. These establishments are 18 p.p. less likely to respond to LISA's business surveys than the average office, holding one of the lowest response rates among facility types. 9/ 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026Across different sectors and facility types, response rates to LISA's regional business surveys can vary by up to 32 percentage points. 8/ 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026Responsive establishments hire fewer people, who are less likely to work fulltime. This is consistent with efficient division of labor: the sorts of firms efficient enough to delegate fulltime and parttime tasks are more able to allocate administrative labor to answer surveys. 7/ 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026About 19% of Dutch establishments provided responses to LISA’s surveys in 2022. Nearly all of the remaining data is imputed from prior years. In my paper, I discuss LISA's imputation methods and propose alternatives, which I show empirically improve predictive performance. 6/ 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026This LISA data can be provided to researchers at certain Dutch institutions through the FIRMBACKBONE infrastructure, and I leverage this data to examine the differences between establishments that do and don’t respond to LISA’s regional business surveys. 4/ 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026But in the Netherlands, there is a 'business census'. Each year, the Landelijk Informatiesysteem voor Arbeidsplaatsen (LISA) runs regional business surveys and combines data from these surveys with administrative data on all establishments in the Netherlands. 3/ 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026Like all surveys, business surveys have the problem that respondents often aren't representative of the population. But in human surveys, you can deal with this by comparing your sample to, e.g., a census. Problem: in most settings, there's no 'business census' to compare to. 2/ 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026Fresh preprint! Using Dutch administrative data, I show that businesses who respond to business surveys are typically unrepresentative, and that this unrepresentativeness is driven by which businesses are contacted for those surveys. 1/ doi.org/10.31235/osf... 1205
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026Consequently, we find that log-like specifications in our replication sample are statistically significant 40-49% more frequently than in the general causal economics literature, and published test statistics are *really* likely to be just beyond 5% significance thresholds. 14/ 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026This happens because messing with unit scale (or the c in ln(Z+c)) allows you to overfit the data. In sample-split simulation data, the log-like specifications that yield the most spuriously significant results within-sample have the worst out-of-sample predictability. 12/ 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026We show this in simulation evidence: even with a placebo treatment and an outcome made of random noise, you get a >30% increase in rejection rates by mining over unit scalings. We also observe sweet spots in ~21% of our simulation draws. 11/ 111
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026We also discovered that in ln(Z+c) specifications, you can get sweet spots both in unit scale and in constant c. 9/ 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026We discovered that t-statistics in log-like specifications can be non-monotonic in unit scale, creating local optima in t-statistics that can briefly dip into rejection regions. This doesn’t just matter for point estimates: it affects studies’ entire conclusions. 8/ 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026Two of our robustness checks involved scaling variables up or down by a factor of 1000 before transformation. For 38% of estimates, *both* of these checks shrunk t-statistics. This pointed us to the existence of what we call ‘sweet spots’. 7/ 120
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026These specifications are *really* non-robust. Just removing the log-like transformation changes 36% of conclusions and significantly sign-flips 12% of estimates. Other checks change conclusions for 14-36% of estimates. 6/ 120
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026New preprint! We reanalyze 46 papers that use log-like specifications (ln(Z+1), inverse hyperbolic sine etc). We find widespread non-robustness, and we show through theory + simulation how these models drive spurious significance. 1/ doi.org/10.31222/osf... 1144
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/02/2026When your spam targets won't submit so you *demand submission* Wishing y'all luck today 130
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026The original paper explicitly highlights how much terrorism ‘declined’ in the UK after the COVID-19 pandemic. But the decline after 2020 is only that stark because all terrorism-related variables are imputed to 0, without disclosure, after the UK left the EU in 2020. 22/x 121
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026Trying to reproduce the paper’s main table using the variables actually disclosed in the paper yields estimates that are much smaller than those published in the paper. Almost everything loses statistical significance, and many coefficients are plainly inestimable. 15/x 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026None of these variables can be reconstructed by taking IHS transformations of the underlying rate variables they’re supposedly constructed from. Again, rates of zero somehow get mapped to many different positive values, as do values that are missing due to division by zero. 14/x 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026WEA's repository includes NewArrest, CharSin, ConvictionSin, and SentenceSin. Though these variables are supposedly constructed from rates that often require division by 0, there are no missing values in any of these variables. These are used to produce WEA’s main table. 13/x 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026This paper had 1 revision round before acceptance. A reviewer raised concerns about how the data transformation affected country-years with zero attacks. After a few extra sentences + citations about the IHS transformation, the reviewer was satisfied and accepted the paper. 9/x 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026The 305 post-2006 country-years that experience zero terror attacks also get assigned to 292 different positive values of DVSin. This is impossible, as the IHS of 0 is 0. This implies that the paper’s main outcome cannot possibly be constructed as described in the paper. 7/x 130
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026This main outcome is hard-coded in the dataset as ‘DVSin’. Since the dataset contains attack + population counts, I can directly compute per capita attack rates. DVSin is actually significantly *negatively* correlated with per capita attack rates + their IHS transformation. 6/x 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026The top panel shows terror attacks over time in each country. The bottom panel shows the paper’s main outcome, reported as the inverse hyperbolic sine (IHS) of per capita attack rates. Impossibly, the 9 countries with 0 attacks have positive IHS attack rates that change over time. 5/x 130
Jack Fitzgerald @jackfitzgerald.bsky.social · 02/09/2025Had a great time presenting my job market paper at the Lindau Nobel Meeting in Economic Sciences! 🔗 : osf.io/d7sqr_v1/ #LINOecon #EconSky 040
Jack Fitzgerald @jackfitzgerald.bsky.social · 13/03/2025What academic journal should I start? Wrong answers only 310
Jack Fitzgerald @jackfitzgerald.bsky.social · 07/03/2025I'll be waking up early (7 AM CET) on Tuesday, March 12 to present my job market paper at 5 PM Sydney time! If you're awake too, stop by to hear me talk about equivalence testing, replication-based methods research, and the robustness of null results in economics! 040
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024We also offer the tst() command in the eqtesting R package, the tsti command in Stata, and Jamovi code. You can visit the paper to find download instructions for all, + guidelines for implementation. We hope you find it useful! (8/9) osf.io/preprints/ps... 131
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024To make things easy, we offer the ShinyTST app, a point-and-click Shiny app that tells you which test/confidence interval is relevant, provides p-values, and visualizes test results given an estimate, standard error, and SESOI. 7/9 jack-fitzgerald.shinyapps.io/shinyTST/ 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024Practical significance conclusions about an estimate can be easily inferred from double-banded confidence intervals that combine the estimate’s (1 - α) CI (e.g., its 95% CI) with its (1 - 2α) CI (e.g., its 90% CI). 6/9 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024The three-sided testing (TST) framework combines two-sided minimum effects tests for inferiority/superiority with the two one-sided tests (TOST) equivalence testing procedure. TST can provide stat. sig. evidence that estimates are practically significant, or practically = 0. 4/9 120
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024Estimates can also be stat. sig. bounded outside of Δ (e.g., blue estimate). What should we conclude about estimates like these blue/orange estimates? Standard equivalence testing frameworks don't give us clear answers. We introduce researchers to a framework that does. 3/9 110
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024Equivalence testing lets us test whether estimates are stat. sig. bounded beneath practically negligible effect size Δ (e.g., pink estimate). But estimates can be both stat. sig. diff. from zero and stat. sig. bounded beneath Δ. 2/9 120
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024New paper for holiday reading! @isager.bsky.social and I provide an introduction to three-sided testing, a framework for testing estimates' practical significance. We offer a tutorial, Shiny app, + commands/code in #Rstats, #Jamovi, + #Stata. 1/9 osf.io/preprints/psyarxiv/8y925 #EconSky #PsychSky 85413
Jack Fitzgerald @jackfitzgerald.bsky.social · 09/12/2024In more recent news, I thoroughly enjoyed presenting The Need for Equivalence Testing in Economics at the Netherlands Reproducibility Network Symposium and Platform for Young Meta-Scientists Symposium, with great discussion from Tsz Keung Wong! 020
Jack Fitzgerald @jackfitzgerald.bsky.social · 09/12/2024Two weeks ago, I had a wonderful time presenting The Need for Equivalence Testing in Economics at the Leibniz Open Science Day! (pic: @prashantgarg.bsky.social) 020
Jack Fitzgerald @jackfitzgerald.bsky.social · 19/11/2024I had an excellent time presenting this paper to the Behavioural Insights for Business and Policy Network at the University of New South Wales. A huge thanks to @impartialspectator.bsky.social for hosting! 030
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024I’ve also spent an extensive amount of time yelling about how stat. insig. bias estimates are really bad evidence that there’s negligible/zero bias. For an in-depth discussion, see my job market paper. 17/19 🔗: jack-fitzgerald.github.io/files/The_Ne... 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024TE-irrelevant biases can badly misidentify TE-relevant biases, even up to the point of complete sign-flips. Trying to learn about hypothetical biases on TEs from hypothetical bias experiments that only vary stakes conditions can yield very misleading conclusions. 15/19 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024Here’s a simulated example where hypothetical stakes increase the outcome’s standard deviation, but decrease the TE’s standard error. Just because your outcome is more precisely measured doesn’t necessarily mean your TE will be! 13/19 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024This means that you can’t identify IHB in an experiment where you just randomize stakes conditions between groups and take differences in mean outcomes between those groups. If you try to infer IHBs from the CHBs estimated in these experiments, you can be badly misled. 11/19 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024Intuitively, that’s for two reasons. 1) You can’t identify an interaction effect if all you know is the avg marginal effect of one of the variables in the interaction. 2) You shouldn’t expect hypothetical stakes to impact all interventions’ TEs on an outcome in the exact same way. 10/19 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024I term the hypothetical bias relevant for TEs ‘interactive hypothetical bias (IHB)’, because it reflects the interaction effect between hypothetical stakes and the intervention of interest. CHB doesn’t identify this bias. 9/19 100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024In elicitation experiments, the only hypothetical bias we care about is the average marginal effect of hypothetical stakes on the outcome. I call this ‘classical hypothetical bias (CHB)’ because it’s the bias identified in most prior hypothetical bias studies. 7/19 100