lorenabarba.com
Beware cargo-cult reproducibility
Over just a few weeks, I received a string of very similar cold emails from seemingly unrelated senders. Each asks the same thing: would I perform a “bounded, independent reproduction” of a “frozen” computational package and return a signed attestation of exactly what I observed. Different names, different “institutes,” the same script. When I described the pattern on LinkedIn, my colleague Kyle Niemeyer replied within the hour that he’d been getting them too. So this is not a fluke aimed at one inbox; it’s a pattern.
The senders speak a fluent language of computational rigor: frozen manifests, expected outputs, tolerances, hashes, known limitations, and attestation templates. They carefully reassure me that they are “not asking me to endorse” anything. The full vocabulary of computational reproducibility, deployed with unusual polish, is itself a tell of something fishy going on.
## The anatomy of the ask
Looking closely, the requests share the same structure. The object to be reproduced is always chosen because it reproduces cleanly: a classical textbook calculation with a known analytical answer, or a public benchmark dataset anyone can download. Rerunning such a thing “independently” is nearly guaranteed to succeed, and succeeding proves nothing about the claim the sender actually cares about, likely some proprietary method, framework, or holdout result hiding behind a curtain of smoke. The reproduction is just a decoy. The rite is performed on a target selected precisely so that it cannot fail.
What is being solicited, then, is not scrutiny; it is my name on a document that can later be cited as validation. A signed attestation that “Prof. Barba independently reproduced the package” transfers nicely into an investor deck, a grant narrative, a patent file, or a landing page, where the careful boundaries of what I actually checked fall away and only the endorsement-shaped residue remains. The ask is engineered to feel low-cost and high-virtue: bounded, neutral, adversarial, “a failure would be equally informative.” Every phrase is borrowed from genuine practice, which is exactly what makes the counterfeit convincing.
And it is a counterfeit of something real. Independent reproduction is one of the most powerful instruments we have. Venues like the journal ReScience C exist precisely to publish careful, open, retrospective replications, and conference artifact-evaluation committees do serious, thankless work checking whether code truly produces the results a paper advertises. The cargo-cult version mimics the vocabulary of these efforts while inverting their purpose: real reproduction puts the claim at risk; this performs a ceremony that protects it.
## Why now, and why me
I propose a name for the phenomenon: _AI-manufactured reproducibility theater_. The uncanny sameness across unrelated senders is probably not coordination; it is the same class of LLM tooling pointed at the public norms of my field and aimed back at me. Ask a model who ought to independently verify a computational result, and it will surface the field’s most visible names. The targeting is an artifact of visibility, nothing more, which is also why Kyle and no doubt others with a reproducibility profile are hearing from the same quarter. What the models added is scale and polish, not novelty. They industrialized a performance that was already being staged.
This is uncomfortable, but it’s the reason I think this deserves more than a warning. With visibility come requests dressed in flattery and scholarly collaboration that quietly seek to borrow one’s reputation; I have fielded those for well over a decade. After a long commitment to reproducibility practice and advocacy, including as an expert contributor to the National Academies report _Reproducibility and Replicability in Science_ (2018), my honest assessment is that **very little has actually changed, apart from the performative**.
Reproducibility became cool at some point, and much of what followed was ceremony: badges, statements, checklists, artifact links, much of it decoration. Consider that the most basic and most widely awarded artifact badge certifies only that files were deposited in a persistent archive, not that anyone verified them, nor that they work. About thirty years ago Buckheit and Donoho, channeling Jon Claerbout, claimed that a paper about computational results is not the scholarship but the advertisement for it; the scholarship is the code and data that produced the figures. We have spent decades getting more elaborate at the advertising. The current wave did not invent cargo-cult reproducibility (alluding to Feynman’s 1974 Caltech commencement address, “Cargo Cult Science“), but it did accelerate it, and the cold emails in my inbox are the logical endpoint: outreach that seeks the _appearance_ of verification with no verification inside.
To colleagues, especially those with visibility: expect these messages, and read them for the gap between form and substance. Ask whether the proposed reproduction touches the claim or merely reproduces a decoy. Decline to lend your name to an attestation whose only real function is to be quoted as your endorsement. And if you do want to reproduce something, do it through a venue built to publish the results honestly, including negative ones.
> Evidence is not the same as the performance of evidence.