Sign in

Nic Fishman

@njw.fish
291 followers 410 following 88 posts

Using computers to read the entrails of modernity (statistics, optimization, machine learning). Currently: Stats PhD @Harvard Previously: CS/Soc @Stanford, Stat/ML @Oxford njw.fish

PostsRepliesMedia
Nic Fishman @njw.fish · 09/03/2026
For journals to maintain a 5% false-discovery rate after a 172x cost decline, the required number of passing robustness checks jumps from ~50 to ~7,000. That's a 140-fold increase in mandatory disclosure. Disclosure must scale linearly with the researcher's testing capacity.
220
Nic Fishman @njw.fish · 09/03/2026
Editors have two levers: 1. Tighten standards: lower the p-value cutoff 2. Force disclosure: require m passing specifications Our first result: switching to p<0.001 does not work. Stricter thresholds force researchers to search more, but she still only reports her best result.
130
Nic Fishman @njw.fish · 09/03/2026
Conventional significance testing with p < 0.05 and handful of robustness checks was designed for a world where each specification cost effort. That world is in the past. When testing is 172x cheaper and researchers search thousands of specifications they will find hits to report.
150
Nic Fishman @njw.fish · 06/03/2026
To show this works across modality we also have a sequence example: TCR repertoire forecasting. On longitudinal TCR-seq from COVID-19 patients, source+target conditioning with discrete flow matching cuts error by >60%, learning from patients observed at only a single timepoint.
100
Nic Fishman @njw.fish · 06/03/2026
We can do the same comparison in clonal dynamics in hematopoiesis: in lineage-traced scRNA-seq with ~6K clones, only ~2K are observed at multiple timepoints. Source+target conditioning leverages orphan clones to improve fate prediction across all transport mechanisms.
100
Nic Fishman @njw.fish · 06/03/2026
The semi-supervised task has many real world applications. In a mass cytometry drug screen across 10 patients and 11 drugs we compare source-conditioned models to a source-target models. For IID tasks SC wins, but STC generalizes better to held-out patients (the real task).
100
Nic Fishman @njw.fish · 06/03/2026
There are real world applications for this any-to-any task: Batch correction in scRNA-seq! On a 56-donor murine pancreas dataset, source+target DCT outperforms K-to-K baselines, scVI, and Harmony on held-out donors.
100
Nic Fishman @njw.fish · 06/03/2026
On synthetic Gaussians, the difference is stark. A simple K-to-K baseline memorizes training distributions and fails OOD (Voronoi pattern). Source+target DCT interpolates smoothly through embedding space, and the gap widens as K grows.
100
Nic Fishman @njw.fish · 06/03/2026
New paper: "Distribution-Conditioned Transport" Modern scientific datasets don't contain one population — they contain thousands. Clones, donors, patients, each its own distribution. DCT learns transport maps that generalize across them, including to distributions never seen during training.
131
Nic Fishman @njw.fish · 26/05/2025
8/ 🧬 Synthetic promoter design Sequences binned by expression → GDEs provide a powerful embedding space for regulatory sequence design. 🦠 SARS-CoV-2 spike protein dynamics GDEs embed monthly lineage distributions and recover smooth latent chronologies.
100
Nic Fishman @njw.fish · 26/05/2025
7/ 🧫 Morphological profiling (20M images) GDEs model phenotype distributions induced by perturbations and generalize to unseen conditions. 🧬 DNA methylation (253M reads) We learn tissue-specific methylation patterns directly from raw bisulfite reads — no alignment necessary.
110
Nic Fishman @njw.fish · 26/05/2025
6/ 🧬 scRNA-seq lineage tracing Each clone is a population of cells → a distribution over expression. GDEs predict clonal fate better than prior approaches. 🧪 CRISPR perturbation effects GDEs can improve zero-shot prediction of transcriptional response distributions.
100
Nic Fishman @njw.fish · 26/05/2025
5/ GDEs are built to scale: at inference, you can embed or generate from large samples without retraining. We show that GDE embeddings are asymptotically normal, grounding this with results from empirical process theory.
100
Nic Fishman @njw.fish · 26/05/2025
4/ With this simple setup, we find surprisingly elegant geometry: GDE latent distances track Wasserstein-2 (W₂) distances across modalities (shown for multinomial distributions) Latent interpolations recover optimal transport paths (shown for Gaussians and Gaussian mixtures)
110
Nic Fishman @njw.fish · 26/05/2025
🚨 New preprint 🚨 We introduce Generative Distribution Embeddings (GDEs) — a framework for learning representations of distributions, not just datapoints. GDEs enable multiscale modeling and come with elegant statistical theory and some miraculous geometric results! 🧵
5459
Nic Fishman @njw.fish · 12/12/2024
Another key insight: Not all fairness metrics are created equal. Our sensitivity analysis shows how different metrics respond to measurement biases - some are surprisingly fragile, others more robust. 3 / 5
A graph demonstrating the differential fragility problem.
130
Nic Fishman @njw.fish · 12/12/2024
Our research reveals a critical challenge: real-world datasets often have multiple measurement errors, and small measurement errors can completely change fariness analysis. We analyzed 14 benchmark datasets to understand these complex interactions. 2 / 5
A table showing the prevalence of different biases in FairML datasets.
130
Nic Fishman @njw.fish · 12/12/2024
Excited to share our new work on causal sensitivity analysis for fairness metrics at #NeurIPS2024! We've developed a causal sensitivity analysis framework to understand how underlying measurement biases (encoded by DAGs) impact machine learning fairness evaluations. 1 / 5
A pair of DAGs representing proxy bias and selection bias.
4245
Nic Fishman @njw.fish · 06/10/2023
Okay here’s my best shot at the core ideas:
210