Sign in

Jack Fitzgerald

@jackfitzgerald.bsky.social
1.5K followers 268 following 193 posts

Postdoctoral Scholar at Stanford University. Working on applied econometrics, replication, and metascience. jack-fitzgerald.github.io. Likes/reposts aren’t endorsements, views are my own.

PostsRepliesMedia
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/09/2026
Yesterday, I had the honor of defending my dissertation and earning my PhD in Economics and Business Administration with cum laude distinction. In October, I will be joining the Meta-Research Innovation Center at Stanford as a Postdoctoral Scholar. I'm excited for the road ahead!
Me, holding my PhD diploma.
071
Jack Fitzgerald @jackfitzgerald.bsky.social · 03/07/2026
In today's dose of dystopia, here's an ad I get hit with every time I use my preferred citation service
Ad for AI detector service reading: "Using AI for your paper? Professors Can Detect It. Check & Humanize Before Submitting. 99% Detection Accuracy. High Readability as Human. Real-time AI Detection. One-click humanization. Score human on all detectors."
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 17/05/2026
Ok what's the funniest part? I'm personally a fan of the unhinged 'best of luck' email signoff, though a close second is starting the email off with 'Hi Doctor, Hey.'
Subject line: Your Article Is Needed For Page Layout

Hi Doctor,

Hey. We are writing to you right away about the invitation we sent you in March 2026.

We haven't received your manuscript yet, even though we've talked about it before. We're very close to finishing the next issue of [redacted].

This is your last chance to say for sure that you will be there.

We don't want to close this issue without giving you one last chance to send in your most recent work because your academic work has always been very useful to the scientific community.

We really appreciate your participation, and sending in your work quickly will make our next publication much better.

The last day to send in your work is May 29, 2026. If you confirm right away, you can get more time.

We respectfully ask that you confirm by responding with one of the following choices so that we can effectively resolve the matter:

 

Yes - I'm interested. Kindly move on to the following steps.


             No - I won't be taking part right now.


We will accept articles like Original Research, Reviews, Case Reports/Case Studies, Opinions and Perspectives, Images/Short Communications, and other things like these.

A short cover letter should be included with your submission.

To finish the editorial schedule and avoid more follow-ups, we need your confirmation within the next 24 hours. We'll think you're not interested in this issue if you don't respond right away.

Your work will not only make the journal better, but it will also show how important you are to the scientific community. We really hope to include your important work.

Thank you for paying attention right away. We are waiting for you to confirm.

Best of luck,
[redacted]
010
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
But controlling for contact probability erases most differences in business survey response, reducing maximum response rate differences between sectors from 32 to 9 p.p. Establishments registered to residential addresses get response rate deficits cut from 18 to 5 p.p.! 10/
Figure 4 in the paper. Table version: Appendix Table A2.Figure 5 in the paper. Table version: Appendix Table A3.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Most worrying: 70% of establishments in the LISA register are registered at a residential address. These establishments are 18 p.p. less likely to respond to LISA's business surveys than the average office, holding one of the lowest response rates among facility types. 9/
Figure 1 in the paper.Figure 5 in the paper. Table version: Appendix Table A3.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Across different sectors and facility types, response rates to LISA's regional business surveys can vary by up to 32 percentage points. 8/
Figure 4 in the paper. Table version: Appendix Table A2.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Responsive establishments hire fewer people, who are less likely to work fulltime. This is consistent with efficient division of labor: the sorts of firms efficient enough to delegate fulltime and parttime tasks are more able to allocate administrative labor to answer surveys. 7/
Table 4 in the paper.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
About 19% of Dutch establishments provided responses to LISA’s surveys in 2022. Nearly all of the remaining data is imputed from prior years. In my paper, I discuss LISA's imputation methods and propose alternatives, which I show empirically improve predictive performance. 6/
Table 1 in the paper.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
This LISA data can be provided to researchers at certain Dutch institutions through the FIRMBACKBONE infrastructure, and I leverage this data to examine the differences between establishments that do and don’t respond to LISA’s regional business surveys. 4/
Source: https://firmbackbone.nl/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
But in the Netherlands, there is a 'business census'. Each year, the Landelijk Informatiesysteem voor Arbeidsplaatsen (LISA) runs regional business surveys and combines data from these surveys with administrative data on all establishments in the Netherlands. 3/
Source: https://lisa.bij12.nl/
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Like all surveys, business surveys have the problem that respondents often aren't representative of the population. But in human surveys, you can deal with this by comparing your sample to, e.g., a census. Problem: in most settings, there's no 'business census' to compare to. 2/
Source: sketchplanations.com/sampling-bias
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Fresh preprint! Using Dutch administrative data, I show that businesses who respond to business surveys are typically unrepresentative, and that this unrepresentativeness is driven by which businesses are contacted for those surveys. 1/ doi.org/10.31235/osf...
Paper title and abstract.
1205
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
Consequently, we find that log-like specifications in our replication sample are statistically significant 40-49% more frequently than in the general causal economics literature, and published test statistics are *really* likely to be just beyond 5% significance thresholds. 14/
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
This happens because messing with unit scale (or the c in ln(Z+c)) allows you to overfit the data. In sample-split simulation data, the log-like specifications that yield the most spuriously significant results within-sample have the worst out-of-sample predictability. 12/
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
We show this in simulation evidence: even with a placebo treatment and an outcome made of random noise, you get a >30% increase in rejection rates by mining over unit scalings. We also observe sweet spots in ~21% of our simulation draws. 11/
111
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
We also discovered that in ln(Z+c) specifications, you can get sweet spots both in unit scale and in constant c. 9/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
We discovered that t-statistics in log-like specifications can be non-monotonic in unit scale, creating local optima in t-statistics that can briefly dip into rejection regions. This doesn’t just matter for point estimates: it affects studies’ entire conclusions. 8/
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
Two of our robustness checks involved scaling variables up or down by a factor of 1000 before transformation. For 38% of estimates, *both* of these checks shrunk t-statistics. This pointed us to the existence of what we call ‘sweet spots’. 7/
120
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
These specifications are *really* non-robust. Just removing the log-like transformation changes 36% of conclusions and significantly sign-flips 12% of estimates. Other checks change conclusions for 14-36% of estimates. 6/
120
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
New preprint! We reanalyze 46 papers that use log-like specifications (ln(Z+1), inverse hyperbolic sine etc). We find widespread non-robustness, and we show through theory + simulation how these models drive spurious significance. 1/ doi.org/10.31222/osf...
1144
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/02/2026
When your spam targets won't submit so you *demand submission* Wishing y'all luck today
130
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026
The original paper explicitly highlights how much terrorism ‘declined’ in the UK after the COVID-19 pandemic. But the decline after 2020 is only that stark because all terrorism-related variables are imputed to 0, without disclosure, after the UK left the EU in 2020. 22/x
Line graphs displaying time series of the raw number of terrorist attacks over time in each of 28 EU member states. A red box highlights the time series for the United Kingdom. Based on WEA's data.
121
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026
Trying to reproduce the paper’s main table using the variables actually disclosed in the paper yields estimates that are much smaller than those published in the paper. Almost everything loses statistical significance, and many coefficients are plainly inestimable. 15/x
Image of panel (a) in Table 1 from the Matters Arising.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026
None of these variables can be reconstructed by taking IHS transformations of the underlying rate variables they’re supposedly constructed from. Again, rates of zero somehow get mapped to many different positive values, as do values that are missing due to division by zero. 14/x
4x2 array of scatterplots with OLS lines of best fit. In the first row of graphs, the y-axis represents values of NewArrest; the x-axis of the left (right) graph represents values of (IHS) arrest rates. In the second row of graphs, the y-axis represents values of CharSin; the x-axis of the left (right) graph represents values of (IHS) charge rates. In the third row of graphs, the y-axis represents values of ConvictionSin; the x-axis of the left (right) graph represents values of (IHS) convictionrate. In the fourth row of graphs, the y-axis represents values of SentenceSin; the x-axis of the left (right) graph represents values of (IHS) sentence. In the top two graphs, the OLS line of best fit is downward-sloping, whereas it is upward-sloping in the bottom six graphs. No line of best fit in the second column of graphs perfectly intersects all points.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026
WEA's repository includes NewArrest, CharSin, ConvictionSin, and SentenceSin. Though these variables are supposedly constructed from rates that often require division by 0, there are no missing values in any of these variables. These are used to produce WEA’s main table. 13/x
Image of the first few observations from WEA's replication dataset. Red boxes highlight variables ConvictionSin, SentenceSin, CharSin, and NewArrest. Source: https://zenodo.org/records/8196717.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026
This paper had 1 revision round before acceptance. A reviewer raised concerns about how the data transformation affected country-years with zero attacks. After a few extra sentences + citations about the IHS transformation, the reviewer was satisfied and accepted the paper. 9/x
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026
The 305 post-2006 country-years that experience zero terror attacks also get assigned to 292 different positive values of DVSin. This is impossible, as the IHS of 0 is 0. This implies that the paper’s main outcome cannot possibly be constructed as described in the paper. 7/x
Displays two scatterplots with OLS lines of best fit. In each graph, the y-axis shows values of DVsin. In the left scatterplot, the x-axis is the raw attack rate (attacks/population in hundreds of thousands), whereas in the right scatterplot, the x-axis is this attack rate transformed using the inverse hyperbolic sine function. In both graphs, a vertical line of dots is visible when x = 0, and the OLS line of best fit slopes downwards. Two red boxes highlight the vertical lines of dots where x  = 0 in each graph.
130
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026
This main outcome is hard-coded in the dataset as ‘DVSin’. Since the dataset contains attack + population counts, I can directly compute per capita attack rates. DVSin is actually significantly *negatively* correlated with per capita attack rates + their IHS transformation. 6/x
Displays two scatterplots with OLS lines of best fit. In each graph, the y-axis shows values of DVsin. In the left scatterplot, the x-axis is the raw attack rate (attacks/population in hundreds of thousands), whereas in the right scatterplot, the x-axis is this attack rate transformed using the inverse hyperbolic sine function. In both graphs, a vertical line of dots is visible when x = 0, and the OLS line of best fit slopes downwards.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 08/01/2026
The top panel shows terror attacks over time in each country. The bottom panel shows the paper’s main outcome, reported as the inverse hyperbolic sine (IHS) of per capita attack rates. Impossibly, the 9 countries with 0 attacks have positive IHS attack rates that change over time. 5/x
Time series of different measures of terrorist attacks over time in each of 28 EU member states. The top panel shows national time series for the raw count of terrorist attacks. The bottom panel shows national time series for 'DVsin'. Based on data from WEA.
130
Jack Fitzgerald @jackfitzgerald.bsky.social · 02/09/2025
Had a great time presenting my job market paper at the Lindau Nobel Meeting in Economic Sciences! 🔗 : osf.io/d7sqr_v1/ #LINOecon #EconSky
040
Jack Fitzgerald @jackfitzgerald.bsky.social · 13/03/2025
What academic journal should I start? Wrong answers only
Email inviting me, a PhD candidate, to be the editor-in-chief of a brand new journal of my choosing.
310
Jack Fitzgerald @jackfitzgerald.bsky.social · 07/03/2025
I'll be waking up early (7 AM CET) on Tuesday, March 12 to present my job market paper at 5 PM Sydney time! If you're awake too, stop by to hear me talk about equivalence testing, replication-based methods research, and the robustness of null results in economics!
Paper title: The Need for Equivalence Testing in Economics. Written by Jack Fitzgerald, Vrije Universiteit Amsterdam and Tinbergen Institute. Date: February 5, 2025. Abstract: Equivalence testing can provide statistically significant evidence that economic relationships are practically negligible. I demonstrate its necessity in a large-scale reanalysis of estimates defending 135 null claims made in 81 recent articles from top economics journals. 36-63% of estimates defending the average null claim fail lenient equivalence tests. In a prediction platform survey, researchers accurately predict that equivalence testing failure rates will significantly exceed levels which they deem acceptable. Obtaining equivalence testing failure rates that these researchers deem acceptable requires arguing that nearly 75% of published estimates in economics are practically equal to zero. These results imply that Type II error rates are unacceptably high throughout economics, and that many null findings in economics reflect low power rather than truly negligible relationships. I provide economists with guidelines and commands in Stata and R for conducting credible equivalence testing and practical significance testing in future research.
040
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024
We also offer the tst() command in the eqtesting R package, the tsti command in Stata, and Jamovi code. You can visit the paper to find download instructions for all, + guidelines for implementation. We hope you find it useful! (8/9) osf.io/preprints/ps...
Abstract of the paper; can be accessed at the URL in the post.
131
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024
To make things easy, we offer the ShinyTST app, a point-and-click Shiny app that tells you which test/confidence interval is relevant, provides p-values, and visualizes test results given an estimate, standard error, and SESOI. 7/9 jack-fitzgerald.shinyapps.io/shinyTST/
An estimate (18000), standard error (3000), smallest effect size of interest (10000), and significance level (0.05) have been supplied to the left-hand panel. A notice at the top indicates that results are asymptotically approximate. Blue text indicates the 90% confidence interval. Red text indicates the 95% confidence interval. Red text denotes that the superiority test is the relevant test (because the estimate is above the upper delta bound). Black text provides the test z-statistic, p-value, and practical significance conclusion. A graph below plots the estimate in black, a 90% confidence interval in blue, and a 95% confidence interval in red, all of which are above the upper delta bound of 10000.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024
Practical significance conclusions about an estimate can be easily inferred from double-banded confidence intervals that combine the estimate’s (1 - α) CI (e.g., its 95% CI) with its (1 - 2α) CI (e.g., its 90% CI). 6/9
In the inferiority region, a 95% confidence interval is used to assess whether the estimate is significantly bounded beneath the lower delta bound. In the inferiority region, a 90% confidence interval is used to assess whether the estimate is significantly bounded within the delta bounds, even if the 95% confidence interval crosses one of the delta bounds. In the superiority region, a 95% confidence interval is used to assess whether the estimate is significantly bounded above the upper delta bound. The bottom four estimates have confidence intervals crossing the delta bounds, implying that results are inconclusive.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024
The three-sided testing (TST) framework combines two-sided minimum effects tests for inferiority/superiority with the two one-sided tests (TOST) equivalence testing procedure. TST can provide stat. sig. evidence that estimates are practically significant, or practically = 0. 4/9
Panel A shows an inferiority test, where H0 states that the estimate is greater than the lower delta bound and HA states that the estimate is less the lower delta bound. Panel B shows a TOST procedure, where the null hypothesis is that the estimate is either above the upper delta bound or below the lower delta bound, and the alternative hypothesis is that the estimate is between the delta bounds. Panel C shows a superiority test, where the null hypothesis states that the estimate is less than the upper delta bound, and the alternative hypothesis states that the estimate is above the upper delta bound. Panel D shows a three-sided tests, where the inferiority region is bounded below the lower delta bound, the equivalence region is bounded between the delta bounds, and the superiority region is bounded above the upper delta bound.
120
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024
Estimates can also be stat. sig. bounded outside of Δ (e.g., blue estimate). What should we conclude about estimates like these blue/orange estimates? Standard equivalence testing frameworks don't give us clear answers. We introduce researchers to a framework that does. 3/9
Pink estimate is statistically significantly bounded between the delta bounds. Blue estimate is statistically significantly bounded above the upper delta bound. The confidence interval of the orange estimate intersects one of the delta bound, but does not cross zero.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024
Equivalence testing lets us test whether estimates are stat. sig. bounded beneath practically negligible effect size Δ (e.g., pink estimate). But estimates can be both stat. sig. diff. from zero and stat. sig. bounded beneath Δ. 2/9
Pink estimate is statistically significantly bounded between the delta bounds. Blue estimate is statistically significantly bounded above the upper delta bound. The confidence interval of the orange estimate intersects one of the delta bound, but does not cross zero.
120
Jack Fitzgerald @jackfitzgerald.bsky.social · 20/12/2024
New paper for holiday reading! @isager.bsky.social and I provide an introduction to three-sided testing, a framework for testing estimates' practical significance. We offer a tutorial, Shiny app, + commands/code in #Rstats, #Jamovi, + #Stata. 1/9 osf.io/preprints/psyarxiv/8y925 #EconSky #PsychSky
In the inferiority region, a 95% confidence interval is used to assess whether the estimate is significantly bounded beneath the lower delta bound. In the inferiority region, a 90% confidence interval is used to assess whether the estimate is significantly bounded within the delta bounds, even if the 95% confidence interval crosses one of the delta bounds. In the superiority region, a 95% confidence interval is used to assess whether the estimate is significantly bounded above the upper delta bound. The bottom four estimates have confidence intervals crossing the delta bounds, implying that results are inconclusive.
85413
Jack Fitzgerald @jackfitzgerald.bsky.social · 09/12/2024
In more recent news, I thoroughly enjoyed presenting The Need for Equivalence Testing in Economics at the Netherlands Reproducibility Network Symposium and Platform for Young Meta-Scientists Symposium, with great discussion from Tsz Keung Wong!
Presenting talk.
020
Jack Fitzgerald @jackfitzgerald.bsky.social · 09/12/2024
Two weeks ago, I had a wonderful time presenting The Need for Equivalence Testing in Economics at the Leibniz Open Science Day! (pic: @prashantgarg.bsky.social)
Presenting talk.
020
Jack Fitzgerald @jackfitzgerald.bsky.social · 02/12/2024
Does this count
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 19/11/2024
I had an excellent time presenting this paper to the Behavioural Insights for Business and Policy Network at the University of New South Wales. A huge thanks to @impartialspectator.bsky.social for hosting!
030
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024
I’ve also spent an extensive amount of time yelling about how stat. insig. bias estimates are really bad evidence that there’s negligible/zero bias. For an in-depth discussion, see my job market paper. 17/19 🔗: jack-fitzgerald.github.io/files/The_Ne...
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024
TE-irrelevant biases can badly misidentify TE-relevant biases, even up to the point of complete sign-flips. Trying to learn about hypothetical biases on TEs from hypothetical bias experiments that only vary stakes conditions can yield very misleading conclusions. 15/19
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024
Here’s a simulated example where hypothetical stakes increase the outcome’s standard deviation, but decrease the TE’s standard error. Just because your outcome is more precisely measured doesn’t necessarily mean your TE will be! 13/19
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024
This means that you can’t identify IHB in an experiment where you just randomize stakes conditions between groups and take differences in mean outcomes between those groups. If you try to infer IHBs from the CHBs estimated in these experiments, you can be badly misled. 11/19
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024
Intuitively, that’s for two reasons. 1) You can’t identify an interaction effect if all you know is the avg marginal effect of one of the variables in the interaction. 2) You shouldn’t expect hypothetical stakes to impact all interventions’ TEs on an outcome in the exact same way. 10/19
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024
I term the hypothetical bias relevant for TEs ‘interactive hypothetical bias (IHB)’, because it reflects the interaction effect between hypothetical stakes and the intervention of interest. CHB doesn’t identify this bias. 9/19
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/11/2024
In elicitation experiments, the only hypothetical bias we care about is the average marginal effect of hypothetical stakes on the outcome. I call this ‘classical hypothetical bias (CHB)’ because it’s the bias identified in most prior hypothetical bias studies. 7/19
100