Sign in

Noah Haber

@whaleactually.com
2.9K followers 308 following 337 posts

econ, epi, stats, meta, causal inference mutant scientist, epistemic humility fairy godmother, chaos muppet. doing researchy metasciencey stuff at the Center for Open Science

PostsRepliesMedia
Noah Haber @whaleactually.com · 29/09/2026
Up on the COS.io blog now: "What is in the Registered Revisions Meta Trial Development Report" www.cos.io/blog/what-is... Yeah this is a post about a post about a report. But it's an unusual report worth a read, since it focuses on all the lessons learned, dead ends, straight up errors, etc.
cos.io
What is in the Registered Revisions Meta Trial Development and Pilot Report?
COS Principal Research Scientist Noah Haber shares how the SCORE project developed a collaborative reproducible manuscript workflow, walks through two different approaches, and introduces an open-sour...
032
Noah Haber @whaleactually.com · 17/09/2026
Can't beat Google Docs and similar for writing collaboration for larger teams and/or less code-oriented coauthors, so the placeholder-style workflow lives on for now. But having more good options and closing that gap is a huge win.
000
Noah Haber @whaleactually.com · 17/09/2026
Got a preview of the upcoming Quarto 2 collaborative editor a few months ago, and it's genuinely very very cool and huge step toward practical collaborative code-driven manuscript workflows. May become my default for solo-ish projects or where collaborators are at least a little code savvy.
cos.io
Toward Collaborative Reproducible Manuscripts
COS Principal Research Scientist Noah Haber shares how the SCORE project developed a collaborative reproducible manuscript workflow, walks through two different approaches, and introduces an open-sour...
1152
Noah Haber @whaleactually.com · 16/09/2026
I am stealing this exchange for an upcoming quick blog post @jeremyfreese.bsky.social @dingdingpeng.the100.ci
120
Noah Haber @whaleactually.com · 14/09/2026
For the curious: doi.org/10.1016/j.an... or free version on arXiv arxiv.org/abs/2004.04251 Comes complete with an R package to auto reveal assumptions hidden in your DAG to make you hate it/yourself
020
Noah Haber @whaleactually.com · 14/09/2026
Why yes I am 40 how did you know
110
Noah Haber @whaleactually.com · 14/09/2026
I have a paper that is about the fundamental and immutable severity of our causal inferential powers and our tendency to fool ourselves about them and it has a hidden Missy Elliott lyric and title that's a backronym from the Blondie comics. Sometimes, you just need the levity to get through.
150
Reposted by Noah Haber
Center for Open Science @cos.io · 09/09/2026
⬇️ The Registered Revisions Meta Trial Pilot Report documents the development & pilot of a new approach to peer review that asks authors to precommit to substantive revisions before conducting new analyses. It also details the multi-journal trial design, workflows, infrastructure, & lessons learned.
041
Noah Haber @whaleactually.com · 09/09/2026
I bet you in particular would find the meta trial design bit interesting. There's a whole giant class of multi-unit policy experiments that are thwarted by coordination, incentive, and design flexibility problems. Maybe the meta trial idea goes a long way toward enabling them?
000
Reposted by Noah Haber
Noah Haber @whaleactually.com · 08/09/2026
... and that's where Registered Revisions comes in The Registered Revisions Meta Trial Pilot Report is OUT NOW!! osf.io/preprints/me... But new(ish) policy involves developing a new(ish) multi-headed trial design, workflows, infrastructure, piloting, mistakes, and a super team. And it's all here.
22117
Noah Haber @whaleactually.com · 08/09/2026
IT'S OUT NOW!!!! osf.io/preprints/me... bsky.app/profile/did:...
osf.io
OSF
000
Noah Haber @whaleactually.com · 08/09/2026
So check it out! We think it might have genuine potential to really change the way we do both policy impact evaluation and peer review. I am STOKED to show eveyrone the process and not just the result (main results coming next year). We want your thoughts! Reply, email, takedown thread, whatever.
050
Noah Haber @whaleactually.com · 08/09/2026
And there's some bonuses in there. You want simulations???? Heck yeah you do. We needed some way to understand the whole Registered Revisions and policy process, plus bonuses like pre-writing our analysis code based on simulated data. How about a live demo webapp? Yup.
130
Noah Haber @whaleactually.com · 08/09/2026
Most importantly, what did the pilot journal editors (coauthors) think about the whole thing? What worked well, what was frustrating, and what went horribly wrong? We compiled our editorial team's thoughts so you can see for yourselves with (we think?) no sugar coating, good and bad alike.
130
Noah Haber @whaleactually.com · 08/09/2026
How about workflows and templates for integrating the Registered Revisions policy into standard editorial processes? Yeah we got that. You want infrastructure? You're covered. A whole webapp based system for participant management and tracking designed to work with normal editorial touchpoints.
130
Noah Haber @whaleactually.com · 08/09/2026
Part 2: the Pilot, with a crew of kickass journal editors and co designers. Unlike standard pilots confirming or adjusting a well-developed trial, we needed to build this plane as we were flying it. The "pathfinder" pilot was more co-development than test implementation. Everyone pitched in a TON.
130
Noah Haber @whaleactually.com · 08/09/2026
<record scratch> "You may be wondering how we got here". That's Part 1, and it starts with a curveball (no spoilers). Nearly everything in this project was developed by necessity (sometimes unexpected), from the policy itself to the meta trial design to the (non-obvious) outcomes selection.
130
Noah Haber @whaleactually.com · 08/09/2026
To estimate its impact, we are doing something out of the box: a "meta trial" Rather than one big trial with a bunch of journals or a hopeful meta-analysis of independent trials, we do a bit of both. One centralized design and infrastructure, but with flexible independent journal-based trials.
141
Noah Haber @whaleactually.com · 08/09/2026
Registered Revisions is like a mini registered report for when reviewers/editors imply some new analysis or data collection be performed during standard peer review. It's what @dingdingpeng.the100.ci describes here, but lightly formalized and with developed workflows bsky.app/profile/ding...
170
Noah Haber @whaleactually.com · 08/09/2026
This report is a little different than what you might be used to. You won't see results or implications (due next year). This is a deep dive into the design and development of the policy, methods, ideas, etc. And most importantly, what didn't go to plan. Which was a lot. bsky.app/profile/whal...
121
Noah Haber @whaleactually.com · 08/09/2026
... and that's where Registered Revisions comes in The Registered Revisions Meta Trial Pilot Report is OUT NOW!! osf.io/preprints/me... But new(ish) policy involves developing a new(ish) multi-headed trial design, workflows, infrastructure, piloting, mistakes, and a super team. And it's all here.
22117
Noah Haber @whaleactually.com · 08/09/2026
Most importantly, what did the pilot journal editors (coauthors) think about the whole thing? What worked well, what was frustrating, and what went horribly wrong? We compiled our editorioal team's thoughts so you can see for yourselves with (we hope) no sugar coating.
000
Noah Haber @whaleactually.com · 08/09/2026
<record scratch> "You may be wondering how we got here". That's Part 1, and it starts with a curveball (no spoilers). Nearly everything in this project was developed by necessity (sometimes unexpected), from the policy itself to the meta trial design to the (non-obvious) outcomes selection.
000
Noah Haber @whaleactually.com · 08/09/2026
This kind of sausage-making what-went-wrong and what-else-did-we-try process stuff is crucial to the sciencing. But it's rarely disclosed, and even rarer to be the main focus of an entire design and development paper. (please someone give me a better analogy than sausages i hate it)
121
Noah Haber @whaleactually.com · 08/09/2026
Still under moderation, but a teaser: Personal favorite bit of the report is getting to talk about what went wrong in a way that I don't often see in an academic-ish paper. Mistakes made, ideas scrapped, and even a particularly bad code bug that I personally am responsible for (and fixed).
110
Noah Haber @whaleactually.com · 02/09/2026
<refresh> <refresh> uuuuuugh come onnnnnn <refresh> * human mods do your thing as long as you need, i appreciate you
110
Noah Haber @whaleactually.com · 02/09/2026
One might call that policy some kind of "registered revision" perhaps, and mayhap some folks went out and made a mega weird randomized meta trial situation, and allegedly a giant report describing the devlopment and piloting of such a thing is getting released shortly, who can say...
1174
Noah Haber @whaleactually.com · 02/09/2026
Oh boy do I have something for you specifically coming in the next few days...
160
Reposted by Noah Haber
Paul Whaley @dangerwhale.bsky.social · 01/09/2026
Spoiler alert: Noah tried to do research with editors. I feel like someone should've warned him.
021
Noah Haber @whaleactually.com · 01/09/2026
If you love any of the following, stay tuned for something I am stoked to share in the next few days: * Weird study design * Things super not going to plan * Workflow and infrastructure * Entirely exploratory preregistration * Fun stats/inference issues * Simulations
1203
Noah Haber @whaleactually.com · 13/08/2026
I repeat: Mt. Washington, at a mere 6.3k ft / 1.9k meters, has the record for the Fastest. Recorded. Windspeed. On Earth. Ever. You can mess yourself up anywhere in the mountains if you want to, but Mt. Washington has a penchant for messing you up just because it's grumpy that day.
021
Noah Haber @whaleactually.com · 13/08/2026
Northeasterners understand that the shortest distance between two points is a straight line. The trails were built before switchbacks had even been invented.
210
Noah Haber @whaleactually.com · 13/08/2026
We used to have a rule that you never under any circumstances let Mt Washington know you are there to climb. You are there to take the gear for a walk. You might be exploring ideas for future climbing. But never ever are you there to actually climb rock or ice or anything up there. No sir.
110
Noah Haber @whaleactually.com · 13/08/2026
The west coast may have a lock on altitude training, but Mt Washington produces world class Marco Polo champions every winter.
100
Noah Haber @whaleactually.com · 13/08/2026
There are two kinds of people: those who see a serene landscape in this photo, and those who know. Source: sectionhiker.com/mt-washingto...
110
Noah Haber @whaleactually.com · 13/08/2026
Coming out of bluesky lurking as a former northeasterner/ north carolinian climber, now colorado-ite, to note: Mt. Washington is the true avatar of the northeast: mean, small, and ready to mess you up in a HURRY if you disrespect it. Also fastest officially recorded windspeed anywhere on earth.
1100
Noah Haber @whaleactually.com · 13/07/2026
Always going to be a tradeoff. For the SCORE, we landed waaay in the direction of using the most standard doc editors. For a smaller tech-savvy team github/quarto/overleaf etc makes more sense. Though I did see something recently that suggests we might be getting a better tradeoff frontier soon...
110
Noah Haber @whaleactually.com · 13/07/2026
Definitely a big risk w/ Google Docs, MS Word is a little better bet for long term reliability since the file format is functionally open now. The solution we went with converts Google Docs to MS Word (or skips that step and starts from MS Word) for that and a few other reasons.
110
Reposted by Noah Haber
Joe Bak-Coleman @jbakcoleman.bsky.social · 13/07/2026
Ah, I never used it. I had never been deeply interested in 1-shot reproducibility of full papers until realizing it made revising so much easier; every table and number is just a reference to artifact generated by the code. Data can change, errors can be found; just rerun it.
131
Reposted by Noah Haber
Center for Open Science @cos.io · 07/07/2026
⚙️ New on the COS blog: Principal Research Scientist Noah Haber (@whaleactually.com) shares how the SCORE project developed a collaborative reproducible manuscript workflow, walks through two different approaches, and introduces an open-source R package. www.cos.io/blog/toward-...
cos.io
Toward Collaborative Reproducible Manuscripts
COS Principal Research Scientist Noah Haber shares how the SCORE project developed a collaborative reproducible manuscript workflow, walks through two different approaches, and introduces an open-sour...
0159
Noah Haber @whaleactually.com · 07/07/2026
Of course! I LOVE Quarto, would love to see what y'all have cooking. Sending a DM now
010
Noah Haber @whaleactually.com · 07/07/2026
This blog post shows how we made a *collaborative* reproducible manuscript workflow for SCORE papers It was not only possible, but by far the most *practical* option. The post discusses workflows for reproducible manuscripts, a demo R package for making them, and worked examples. Enjoy!
A chart showing a workflow for markup-style workflows. Data enters in from the left, and a script includes analysis code and narrative text all mixed into the same market script, which produces the final populated manuscript.This image shows a workflow for placeholder-style reproducible manuscripts. In this case, data and analysis code are combined to produce analysis outputs. Separately, a manuscript text is written and collaborated on, containing placeholders for where analysis values will go. A script knits these two together, inserting analysis values where the placeholders are, producing the final populated manuscript.
140
Noah Haber @whaleactually.com · 07/07/2026
In a reproducible manuscript, the data, analysis code and narrative text are all knit together, making a full pipeline from data/code to final manuscript automated and reproducible. This idea has been around forever, but these workflows usually make collaborative writing difficult.
120
Noah Haber @whaleactually.com · 07/07/2026
For most scientific papers, coathors collaborate in the manuscript doc and someone manually enters analysis results/tables/figures into it. In this case, reproducibility effectively stops at the data and code. What this blog post presupposes is, what if it didn't? www.cos.io/blog/toward-...
cos.io
Toward Collaborative Reproducible Manuscripts
The first active phase of COS's Benchmarking LLM Agents on Scientific Tasks project has produced ReplicatorBench—a benchmark for evaluating LLM agents on research replication in social & behavioral sc...
2286
Noah Haber @whaleactually.com · 07/05/2026
When I started working in global health a little over 10 years ago, infectious disease (particularly policy) was said to be a declining field to be almost entirely replaced with chronic disease. That ... did not happen.
050
Noah Haber @whaleactually.com · 24/04/2026
I believe you are looking for "TED talk grifter"
130
Noah Haber @whaleactually.com · 13/04/2026
~100 downloads of the reproducible manuscript packages now, and so far no contacts about errors or not being able to get the package to run etc. Looks like things are working, or at least not obviously horribly broken. Neat.
010
Noah Haber @whaleactually.com · 11/04/2026
Wasn't there for the actual replications for SCORE, but can say a lot of pains were made to try to match the original designs and sample sizes, with repli sample sizes generally larger than the origs. Free metarxiv paper link below, appendix describes repli process in detail osf.io/preprints/me...
osf.io
OSF
010
Noah Haber @whaleactually.com · 10/04/2026
Yep, totally agreed, that can be super valuable, as can be critical reproduction approaches. FWIW, there are some reasons to try to replicate a "poor" design (e.g. separating statistical noise / publication bias from design, particularly if design flaws can be mitigated or bounded), but more niche
020
Noah Haber @whaleactually.com · 10/04/2026
What do you mean by "replicate" here? The papers frame "replication" as using same methods as the original, but with new data, so hard to demonstrate flaws in the method when repeating them. Are you thinking of a more critique oriented approach? Lots of ways to define/frame the word "replicate"
110