Sign in

Jack Fitzgerald

@jackfitzgerald.bsky.social
1.5K followers 268 following 193 posts

Postdoctoral Scholar at Stanford University. Working on applied econometrics, replication, and metascience. jack-fitzgerald.github.io. Likes/reposts aren’t endorsements, views are my own.

PostsRepliesMedia
Reposted by Jack Fitzgerald
Institute for Replication @i4replication.bsky.social · 25/09/2026
Over the past two years, I4R partnered with Psychological Science (@psychscience.bsky.social) on a large reproducibility project. 67 independent teams reproduced 64 articles published in 2024–2025, about 44% of eligible papers. Paper: www.econstor.eu/bitstream/10... Here's what we learned 🧵
18958
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/09/2026
Yesterday, I had the honor of defending my dissertation and earning my PhD in Economics and Business Administration with cum laude distinction. In October, I will be joining the Meta-Research Innovation Center at Stanford as a Postdoctoral Scholar. I'm excited for the road ahead!
Me, holding my PhD diploma.
071
Reposted by Jack Fitzgerald
Social Science Prediction Platform (SSPP) @socscipredict.bsky.social · 31/08/2026
📣 Does involving journal editors in replication material requests impact the availability of such materials? @jackfitzgerald.bsky.social and coauthors invite your predictions! 🕐Time: 15 min 🗓️ Closes: Sept 15 📚 Field: Economics socialscienceprediction.org/predict/r/30...
socialscienceprediction.org
The Impact of Journal Involvement on Availability of Replication Materials
I subject the models defending 135 null claims in 81 articles recently published in top economics journals to equivalence testing, which assesses whether an estimate is significantly bounded within a ...
022
Jack Fitzgerald @jackfitzgerald.bsky.social · 03/08/2026
Not overcome, just speaking about a completely different context (standard NHST). The SESOI is actually an infeasible power target in practical significance tests. E.g., I have (virtually) zero power to detect whether the SESOI is greater than the SESOI.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 02/08/2026
However, we do offer some guidance in the supplementary materials on a 'safeguard power' approach; i.e., having a backup effect size to power to in case the effect size you observe is in a different TST region than the effect size you expect.
010
Jack Fitzgerald @jackfitzgerald.bsky.social · 02/08/2026
The effect size you expect to observe will be the same regardless of whether you're using standard NHST or TST. We shouldn't change the effect size we're powering to just because we're using a different testing framework.
200
Jack Fitzgerald @jackfitzgerald.bsky.social · 31/07/2026
Hey, other coauthor jumping in; this is necessarily inductive because the parameters of interest are random and need to be estimated/expected. If the true effect size were certainly known a priori, then we wouldn't need estimation or inferential testing to know whether it's inside/outside the SESOI!
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 03/07/2026
In today's dose of dystopia, here's an ad I get hit with every time I use my preferred citation service
Ad for AI detector service reading: "Using AI for your paper? Professors Can Detect It. Check & Humanize Before Submitting. 99% Detection Accuracy. High Readability as Human. Real-time AI Detection. One-click humanization. Score human on all detectors."
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
For details, please check out our updated paper! 10/ osf.io/preprints/me...
osf.io
OSF
000
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
IMPORTANT CAVEAT: In quasi-experiments, quantile regression requires extremely strong identification assumptions that are unlikely to hold in practice. We don’t recommend using our commands for that. But they may be useful to analyze RCTs or simple correlations/associations. 9/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
We provide the lzqreg command in Stata (available on SSC) and the lzrq package/function in R (available on CRAN). 8/ doi.org/10.32614/CRA...
doi.org
lzrq: Quantile Regression for Logarithmic Relationships with Non-Positive Outcome Values
Provides the lzrq() function for estimating logarithmic regression slopes in quantile regression models, permitting the outcome variable to take on non-positive values. lzrq() conducts regression afte...
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
You might not be able to find a Ψ large enough to meet that criterion, in which case your data doesn’t permit credible estimation using this method (🤷‍♂️). Our commands automate the search for such a Ψ, report results if it can be found, and suppress results if it cannot. 7/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
If you can find a Ψ* high enough that all median regression-predicted values are < -Ψ*, then the resulting coefficients are interpretable as logarithmic intensive-margin slopes between your covariates and Y’s median, and are identical for any Ψ >= Ψ*. 6/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
Generally, one can run quantile regression after transforming Y with the ‘calibrated extensive margin’ transformation (Chen & @jondr44.bsky.social 2024), setting positive values to ln(Y) and setting non-positive values to constant -Ψ. 5/ doi.org/10.1093/qje/...
doi.org
Logs with Zeros? Some Problems and Solutions*
Abstract. When studying an outcome Y that is weakly positive but can equal zero (e.g., earnings), researchers frequently estimate an average treatment effe
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
E.g., say you have a binary treatment. If the median of Y is positive in both treatment and control groups, then even if there are some non-positive values in the data, the logarithmic difference between those medians is defined, and can be recovered using median regression. 4/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
Idea: in mean regression, the point estimate is computed with every value in the outcome, so even one non-positive value messes up logarithmic slope estimation. But quantile regression doesn't use every value to compute the point estimate, just the conditional quantiles. 3/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
Most notably, we highlight a recent idea from Liu & Kaplan (2025) to use quantile regression to estimate intensive-margin logarithmic relationships, which is possible even with 0s/negatives in the data. 2/ drive.google.com/file/d/1F3dn...
drive.google.com
Liu_Kaplan_quantile_log.pdf
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/06/2026
An update on our log-like specifications paper! We’ve updated our recommendations and added new commands in Stata and R for one of our recommended alternatives. 1/ osf.io/preprints/me...
osf.io
OSF
196
Jack Fitzgerald @jackfitzgerald.bsky.social · 18/06/2026
Happy to have the CBS Replication Games (and my replication research more broadly) featured on the ODISSEI Podcast! Tune in if you're interested in hearing @drtomemery.bsky.social and I nerd out over replication!
062
Reposted by Jack Fitzgerald
ODISSEI @odissei.bsky.social · 17/06/2026
How can games change research culture? 🧩 On this month's ODISSEI podcast, @drtomemery.bsky.social talks with Jack Fitzgerald about replication, and why it needs to be more embedded in research. That's why the Replication Games were developed, Jack explains. 🎧 Tune in to learn more: edu.nl/eqhuc
012
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/05/2026
Our paper on the effect of large language models on replicators' success in conducting reproducibility + robustness checks is now published at PNAS. Before advocating that we let AI completely take over the replication process, please read our paper! doi.org/10.1073/pnas...
doi.org
PNAS
Proceedings of the National Academy of Sciences (PNAS), a peer reviewed journal of the National Academy of Sciences (NAS) - an authoritative source of high-impact, original research that broadly spans...
181
Jack Fitzgerald @jackfitzgerald.bsky.social · 17/05/2026
Ok what's the funniest part? I'm personally a fan of the unhinged 'best of luck' email signoff, though a close second is starting the email off with 'Hi Doctor, Hey.'
Subject line: Your Article Is Needed For Page Layout

Hi Doctor,

Hey. We are writing to you right away about the invitation we sent you in March 2026.

We haven't received your manuscript yet, even though we've talked about it before. We're very close to finishing the next issue of [redacted].

This is your last chance to say for sure that you will be there.

We don't want to close this issue without giving you one last chance to send in your most recent work because your academic work has always been very useful to the scientific community.

We really appreciate your participation, and sending in your work quickly will make our next publication much better.

The last day to send in your work is May 29, 2026. If you confirm right away, you can get more time.

We respectfully ask that you confirm by responding with one of the following choices so that we can effectively resolve the matter:

 

Yes - I'm interested. Kindly move on to the following steps.


             No - I won't be taking part right now.


We will accept articles like Original Research, Reviews, Case Reports/Case Studies, Opinions and Perspectives, Images/Short Communications, and other things like these.

A short cover letter should be included with your submission.

To finish the editorial schedule and avoid more follow-ups, we need your confirmation within the next 24 hours. We'll think you're not interested in this issue if you don't respond right away.

Your work will not only make the journal better, but it will also show how important you are to the scientific community. We really hope to include your important work.

Thank you for paying attention right away. We are waiting for you to confirm.

Best of luck,
[redacted]
010
Jack Fitzgerald @jackfitzgerald.bsky.social · 02/05/2026
The LISA data primarily covers data on employment and firm characteristics (what kind of facility is it in, what sector, etc). I've seen some use the data to track firm survival over time. Turnover is unfortunately not a part of the data.
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 01/05/2026
What kinds of later outcomes would you be interested in?
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
If you want to learn more, check out the preprint! 12/ doi.org/10.31235/osf...
doi.org
OSF
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
So most unrepresentativeness in business survey comes down to which businesses are contacted. That means there's a lot of potential to improve representativeness, generalizability, and/or response rates by understanding and modifying the design of business surveys. 11/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
But controlling for contact probability erases most differences in business survey response, reducing maximum response rate differences between sectors from 32 to 9 p.p. Establishments registered to residential addresses get response rate deficits cut from 18 to 5 p.p.! 10/
Figure 4 in the paper. Table version: Appendix Table A2.Figure 5 in the paper. Table version: Appendix Table A3.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Most worrying: 70% of establishments in the LISA register are registered at a residential address. These establishments are 18 p.p. less likely to respond to LISA's business surveys than the average office, holding one of the lowest response rates among facility types. 9/
Figure 1 in the paper.Figure 5 in the paper. Table version: Appendix Table A3.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Across different sectors and facility types, response rates to LISA's regional business surveys can vary by up to 32 percentage points. 8/
Figure 4 in the paper. Table version: Appendix Table A2.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Responsive establishments hire fewer people, who are less likely to work fulltime. This is consistent with efficient division of labor: the sorts of firms efficient enough to delegate fulltime and parttime tasks are more able to allocate administrative labor to answer surveys. 7/
Table 4 in the paper.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
About 19% of Dutch establishments provided responses to LISA’s surveys in 2022. Nearly all of the remaining data is imputed from prior years. In my paper, I discuss LISA's imputation methods and propose alternatives, which I show empirically improve predictive performance. 6/
Table 1 in the paper.
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Critically, conditional on a number of stratification/exception variables, LISA's business surveys are randomly distributed. So if you control for those variables, you can isolate which firms are most likely to respond to business surveys *conditional on being contacted*. 5/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
This LISA data can be provided to researchers at certain Dutch institutions through the FIRMBACKBONE infrastructure, and I leverage this data to examine the differences between establishments that do and don’t respond to LISA’s regional business surveys. 4/
Source: https://firmbackbone.nl/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
But in the Netherlands, there is a 'business census'. Each year, the Landelijk Informatiesysteem voor Arbeidsplaatsen (LISA) runs regional business surveys and combines data from these surveys with administrative data on all establishments in the Netherlands. 3/
Source: https://lisa.bij12.nl/
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Like all surveys, business surveys have the problem that respondents often aren't representative of the population. But in human surveys, you can deal with this by comparing your sample to, e.g., a census. Problem: in most settings, there's no 'business census' to compare to. 2/
Source: sketchplanations.com/sampling-bias
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 30/04/2026
Fresh preprint! Using Dutch administrative data, I show that businesses who respond to business surveys are typically unrepresentative, and that this unrepresentativeness is driven by which businesses are contacted for those surveys. 1/ doi.org/10.31235/osf...
Paper title and abstract.
1205
Reposted by Jack Fitzgerald
Peder M Isager @isager.bsky.social · 28/04/2026
My article "Three-Sided Testing to Establish Practical Significance" with @jackfitzgerald.bsky.social is now published in AMPPS! journals.sagepub.com/doi/10.1177/... Three-sided testing is an improved version of TOST that lets you test for equivalence, superiority and inferiority simultaneously.
journals.sagepub.com
Sage Journals: Discover world-class research
Subscription and open access journals from Sage, the world's leading independent academic publisher.
12815
Jack Fitzgerald @jackfitzgerald.bsky.social · 29/04/2026
Happy to see my paper with @isager.bsky.social on three-sided testing published at Advances in Methods and Practices in Psychological Science! If you're interested in practical significance testing, give it a read! doi.org/10.1177/2515...
doi.org
Three-Sided Testing to Establish Practical Significance: A Tutorial - Peder Mortvedt Isager, Jack Fitzgerald, 2026
Researchers may want to know whether an observed statistical relationship is either meaningfully negative, meaningfully positive, or small enough to be consider...
071
Jack Fitzgerald @jackfitzgerald.bsky.social · 22/03/2026
My NHB paper is literally about an article where problems with IHS transformations revealed data irregularities that (partly) resulted in that article's retraction. That NHB paper doesn't take any stance on IHS specifications, let alone a *pro* stance. 2/ doi.org/10.1038/s415...
doi.org
Imputations, inverse hyperbolic sines and impossible values - Nature Human Behaviour
Nature Human Behaviour - Imputations, inverse hyperbolic sines and impossible values
010
Jack Fitzgerald @jackfitzgerald.bsky.social · 22/03/2026
Given my stance on log-like specifications, I was surprised to learn that there's a 'news' article on my paper in Nature Human Behaviour, claiming that it actually advocates for the use of IHS specifications. This is categorically untrue. 1/ t.co/ZGjAmHZhLv
t.co
https://scienmag.com/imputations-and-inverse-hyperbolic-sines-unveiling-data-mysteries/
220
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
But for full details, nothing will beat reading the paper. Give it a look! 18/ doi.org/10.31222/osf...
doi.org
OSF
010
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
For the on-the-go researchers out there, we’ve made teaching slides to make the paper’s findings more digestible. 17/ jack-fitzgerald.github.io/files/Log-Li...
jack-fitzgerald.github.io
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
Huge shoutout to the rest of the research team who made this possible: @jopieboy.bsky.social, @fialalenka.bsky.social @essieconomist.bsky.social, and @davidvalenta.bsky.social. 16/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
We have a couple of recommendations on how to deal with the logs-with-zeros problem in the paper. But our biggest advice is this: 🛑stop🛑 using log-like specifications. They are actively polluting the literature with spuriously significant results. 15/
120
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
Consequently, we find that log-like specifications in our replication sample are statistically significant 40-49% more frequently than in the general causal economics literature, and published test statistics are *really* likely to be just beyond 5% significance thresholds. 14/
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
You don’t need p-hacking for this to cause problems. If either researchers file-drawer statistically insignificant results, or journals select statistically significant results, the most spuriously significant log-like specifications can be overrepresented in the literature. 13/
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
This happens because messing with unit scale (or the c in ln(Z+c)) allows you to overfit the data. In sample-split simulation data, the log-like specifications that yield the most spuriously significant results within-sample have the worst out-of-sample predictability. 12/
110
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
We show this in simulation evidence: even with a placebo treatment and an outcome made of random noise, you get a >30% increase in rejection rates by mining over unit scalings. We also observe sweet spots in ~21% of our simulation draws. 11/
111
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
This creates a multiple hypothesis testing problem. There’s no ‘right/wrong’ scale in which to measure a variable and no ‘right/wrong’ constant c to add to ln(Z+c). So you get an infinite number of tests that are equally theoretically valid, but most give different results. 10/
100
Jack Fitzgerald @jackfitzgerald.bsky.social · 16/03/2026
We also discovered that in ln(Z+c) specifications, you can get sweet spots both in unit scale and in constant c. 9/
100