Sign in

Eric Kernfeld

@ekernf01.bsky.social
1.4K followers 422 following 315 posts

Statistician and computational biologist; uw alum; jhu student. He/him. ekernf01.github.io

PostsRepliesMedia
Eric Kernfeld @ekernf01.bsky.social · 30/12/2025
Happy New Year. I have just donated $9,000 to malaria prevention. www.againstmalaria.com/MyNets.aspx?...
againstmalaria.com
Against Malaria: My nets
I've just donated to The Against Malaria Foundation, would you do your part in saving lives too?
010
Eric Kernfeld @ekernf01.bsky.social · 12/10/2025
This is so friggin cool. You don't have to rewrite all that FORTRAN to do autodiff on it. "Enzyme"
041
Eric Kernfeld @ekernf01.bsky.social · 19/09/2025
I love using the UKB RAP. As soon as I see that £, my brain goes "I shall keep-safe this file upon the clouds and it shall cost you nigh thruppence fortnightly." innit
010
Eric Kernfeld @ekernf01.bsky.social · 30/08/2025
For people that have tried to hire someone recently, did you encounter this problem? www.linkedin.com/posts/olga-v...
linkedin.com
Failed to weed out LLM-generated applications for bioinformatics role | Olga Sazonova, Ph.D. posted on the topic | LinkedIn
Well, that's it. Hiring is f*cked. I'm just opened a contract role for a bioinformatics project. I know it's a wild market at the moment - each open role gets flooded with resumes, making it hard to separate qualified candidates from the noise. I'm convinced LLMs are partially to blame. So I devised a process to select qualified, committed candidates and avoid generic applications. Instead of resumes and cover letters, I created an intake questionnaire requiring bespoke effort to reduce cookie-cutter applications. The questions focused on specific scenarios from past experience, and therefore shouldn't be directly outsourced to chatbots. I also accidentally made the application link hard to access, but decided not to fix it. After all, computational biology often requires hacking and clever work-arounds. Finally, I included an honor code asking people not to use LLMs. Initially, I felt good about my approach. I had 20 applications after two days, and the candidates were differentiating themselves: A few didn't answer all questions, a few didn't have the right background, and several candidates provided well-written, topical, plausible answers. It was only after I read a few of these "high signal" applications in a row that the alarm bells started ringing: - Multiple candidates highlighted the same methodological paper - Many cited required skills in the exact same order - A high percentage reported identical troubleshooting scenarios involving differential gene expression studies with lab-derived batch effects The nail in the coffin was repeated phrases like "my initial hypothesis was a bug in [popular tool]" and "the spurious result disappeared." Clearly, candidates were using LLMs. Instead of learning about fitness for the role, I was learning how LLMs answer my "clever" questions. So...I failed. My approach was naive, perhaps even hypocritical. After all, I use LLMs for technical documentation myself. Still, it's really disappointing. I'm no closer to a hiring process that identifies qualified, motivated, and trustworthy candidates for remote work. Now what, y'all? | 109 comments on LinkedIn
020
Eric Kernfeld @ekernf01.bsky.social · 16/08/2025
All parrots are stochastic.
000
Eric Kernfeld @ekernf01.bsky.social · 13/08/2025
Found it!!
Photo of Eric with mom, dad, and brother Paul. They are holding flowers and balloons to celebrate after Eric's defense.
160
Eric Kernfeld @ekernf01.bsky.social · 13/08/2025
Job Search Recap ekernf01.github.io/job_search_2...
Sankey diagram of Eric's job search efforts and opportunities. Source data: 

Months unemployed [1] Biking
Months unemployed [1] Stressing
Months unemployed [1] Applying 
Months unemployed [2] Networking
Months unemployed [0] Waiting to joke about "came in a fluffer"

Jobs applied to (no referral) [64] No follow up
Jobs applied to (with referral) [4] No follow up
Jobs applied to (with referral) [3] Interview
Jobs applied to (no referral) [1] Interview
Info interviews conducted [30] No follow up
Info interviews conducted [2] Interview
Credible unsolicited inquiries [2] Interview
Credible unsolicited inquiries [3] Location mismatch

Interview [2] Job offer
160
Eric Kernfeld @ekernf01.bsky.social · 13/08/2025
Thanks to everyone who has supported me in large and small ways during this uncertain interval.
040
Eric Kernfeld @ekernf01.bsky.social · 13/08/2025
4) I (constructively) peer-reviewed two papers; I wrote ten blog posts; I conducted about 30 informational interviews; and I applied to about 65 jobs. It's tough out there.
160
Eric Kernfeld @ekernf01.bsky.social · 13/08/2025
3) I rode my bike from Washington, DC to Pittsburgh and back, through rain, mud, stiff headwinds, dozens of downed trees, and a couple of sub-freezing nights.
Picture of Eric and his bike in front of tons of unstable rock at the mouth of the Paw-Paw Tunnel, part of the C&O Canal Towpath national park
150
Eric Kernfeld @ekernf01.bsky.social · 13/08/2025
2) I defended my PhD. Ask me if you want the slides and the recording! I somehow lost track of the cute photo I was going to show. :[ Will post in replies if I can find it.
260
Eric Kernfeld @ekernf01.bsky.social · 13/08/2025
1) I have accepted a job at Alden Scientific working to improve the world's best methods for predicting long-term health.
170
Eric Kernfeld @ekernf01.bsky.social · 13/08/2025
I am proud to announce a lot of personal progress from this spring and summer. 🧵
180
Eric Kernfeld @ekernf01.bsky.social · 28/07/2025
i hope you can keep your cool through this.
120
Eric Kernfeld @ekernf01.bsky.social · 28/07/2025
I definitely think everyone working on virtual cells would be happy with more training data. What do you mean when you say the cell should be smaller?
110
Eric Kernfeld @ekernf01.bsky.social · 28/07/2025
In October 2024, I twote that "something is deeply wrong" with what we now call virtual cell models. A lot has happened since then. How am I updating? New blog post: ekernf01.github.io/virtual-cell...
ekernf01.github.io
A recap of virtual cell releases circa June 2025
In October 2024, I twote that “something is deeply wrong” with what we now call virtual cell models. A lot has happened since then: modelers are advancing new architectures and mining new sources of i...
1122
Reposted by Eric Kernfeld
Elizabeth Wood, PhD @lizbwood.bsky.social · 22/07/2025
The biggest challenge for AI in biology isn't just models, it's the data used to train them. Standard biological data isn't built for AI. To unlock generative AI for drug discovery, we must rethink how we generate and capture data. 1/
Hardware/wetware codesigned data loop VISTA makes use of generative model sampling and synthesis "on chip" on-board by leveraging oligosynthesis setup shown here.
2309
Reposted by Eric Kernfeld
Sara Rouhanifard @srouhanifard.bsky.social · 17/07/2025
1/🧵 🚨 New paper published in RNA! Scientists often say anecdotally that RNA modifications are disrupted in immortalized cells — but no one’s really tested it. So we did: ψ-mapping in primary T cells vs. Jurkat cells using direct RNA-seq. 📄 tinyurl.com/TCellPsiRNA #nanopore #RNA #pseudouridine
141
Eric Kernfeld @ekernf01.bsky.social · 29/06/2025
Fuck, I'm so sorry.
010
Eric Kernfeld @ekernf01.bsky.social · 26/06/2025
😋
010
Eric Kernfeld @ekernf01.bsky.social · 26/06/2025
What it felt like to write this series:
Classic "midwit meme" with dummy expanding (a+b)^2 as a^2 + b^2, midwit (superimposed ekernf01 profile pic) crying and expanding (a+b)^2 as a^2 + 2ab + b^2, and wizard (superimposed Sasha Gusev profile pic) expanding E[(a+b)^2] as E[a^2] + E[b^2] because E[ab]=0.
010
Eric Kernfeld @ekernf01.bsky.social · 26/06/2025
Blogpost: ekernf01.github.io/vanilla Image credits:
Vanilla (Linkage Disequilibrium Score Regression) By B.navez - Own work, CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=436896
Strawberry (stratified LD score regression) By Ivar Leidus - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=114685374
Cranberry (cross-trait LD score regression) By Ɱ - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=42145142
Lime (Signed LD Score Regression) By Ivar Leidus - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=97307282
Melon (Mediated Expression Score Regression) By Filo gèn' - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=68127982
100
Eric Kernfeld @ekernf01.bsky.social · 26/06/2025
BLOG ALERT! 🧬 If you do RNA-seq, ATAC-seq, ChIP-seq, or modeling thereof, you may be overlooking LD score regression methods. These should be standard tools to study genes, variants, and regions en masse, but they are hard to understand. I wrote intros:
(flowery font "Corsiva" with bright colors)

The exotic FLAVORS of LD Score Regression

Vanilla (Linkage Disequilibrium Score Regression) distinguishes whether bias is from polygenicity or population structure using GWAS summary stats
Strawberry (stratified LD score regression) quantifies which genomic regions are enriched for disease risk using GWAS summary stats + functional annotations
Cranberry (cross-trait LD score regression) estimates genetic correlations between traits using summary stats from multiple GWAS
Lime (Signed LD Score Regression) tests whether predicted allele effects directionally align with per-allele disease risk using GWAS + deep learning
Melon (Mediated Expression Score Regression) quantifies how much disease risk is mediated through gene expression using GWAS + eQTL effects
160
Eric Kernfeld @ekernf01.bsky.social · 25/06/2025
Fast-followers are more lucrative than first-in-class drugs. centuryofbio.com/p/commoditiz... How, then, can companies protect investments in target discovery? Code obfuscators protect source code from reverse engineering. Could mechanism obfuscators protect drug target identities?
centuryofbio.com
On Modality Commoditization
And what might come next
000
Eric Kernfeld @ekernf01.bsky.social · 23/06/2025
Thank you. It seems this doesn't quite work to define the genetic association parameters it's supposed to define, so I will substitute another thing that makes sense to me.
010
Eric Kernfeld @ekernf01.bsky.social · 23/06/2025
Help me understand the definition of genetic covariance from the xt-ldsc paper. If beta is an argmax, then isn't cbeta also an argmax for any positive scalar c?
110
Eric Kernfeld @ekernf01.bsky.social · 22/06/2025
Is anyone making giant post-perturbation cDNA libraries and sticking them in the freezer while waiting for sequencing costs to drop below a certain point? editlife.substack.com/p/part-1-the...
editlife.substack.com
Part 1. The DNA Sequencing Arms Race: The Fight to Dethrone Illumina
Multi-part series exploring the tech and politics around short/long-read sequencing advances
110
Eric Kernfeld @ekernf01.bsky.social · 19/06/2025
Found it. link.springer.com/article/10.1...
link.springer.com
Clipper: p-value-free FDR control on high-throughput data from two conditions - Genome Biology
High-throughput biological data analysis commonly involves identifying features such as genes, genomic regions, and proteins, whose values differ between two conditions, from numerous features measure...
100
Eric Kernfeld @ekernf01.bsky.social · 19/06/2025
Interesting. I feel like this problem ought to have a general solution by now! Like a multivariate lmer or glmer package. But I am not up to date on the software ecosystem. Jessica Jingyi Li's group has an interesting and flexible FDR control method based on simulated data. Will link . . .
110
Eric Kernfeld @ekernf01.bsky.social · 19/06/2025
Terrific opportunity. If I were not committed to the east coast for personal reasons, I would apply to this immediately.
041
Eric Kernfeld @ekernf01.bsky.social · 19/06/2025
Could you write out more detail about the structure of the model you are looking for?
100
Eric Kernfeld @ekernf01.bsky.social · 19/06/2025
It seems like a big selling point of this method is the run-time on large datasets. Arguably necessary to keep up with new measurement tech.
110
Eric Kernfeld @ekernf01.bsky.social · 19/06/2025
I have not understood the method yet, but I am happy to see they structure the bootstrapping according to which cells come from which sample.
110
Eric Kernfeld @ekernf01.bsky.social · 19/06/2025
Yes, that's the paper! I didn't mean to obfuscate the source.
110
Eric Kernfeld @ekernf01.bsky.social · 19/06/2025
It's really cool to live in a time when a paper on differential expression testing will casually drop a new 160,000-cell perturb-seq experiment. www.ncbi.nlm.nih.gov/geo/query/ac...
ncbi.nlm.nih.gov
GEO Accession viewer
NCBI's Gene Expression Omnibus (GEO) is a public archive and resource for gene expression data.
180
Eric Kernfeld @ekernf01.bsky.social · 17/06/2025
If you see this, I want your answer to this question from @alz_zyd_ on twitter: "What is the largest scope/time project, of any kind, you've accomplished in your life so far?"
101
Reposted by Eric Kernfeld
EPIGENETIC HULK @epigenetichulk.bsky.social · 13/06/2025
EPIGENETIC HULK READY TO SMASH AGAIN! GET IN LOSERS!
1116750
Reposted by Eric Kernfeld
Eli Weinstein @eliweinstein.bsky.social · 29/05/2025
Thrilled to announce that I am joining DTU in Copenhagen in the fall, as an assistant professor of chemistry. My research group will focus on fundamental methodology in machine learning for molecules.
85811
Eric Kernfeld @ekernf01.bsky.social · 06/06/2025
never heard of it -- thank you for the rec!
010
Eric Kernfeld @ekernf01.bsky.social · 05/06/2025
Another case showing we must be wary of Gell-Mann amnesia in evaluating molecular bio applications of AI. rachel.fast.ai/posts/2025-0...
rachel.fast.ai
Rachel Thomas, PhD - Deep learning gets the glory, deep fact checking gets ignored
an AI researcher going back to school for immunology
051
Reposted by Eric Kernfeld
Han Yuan @hy395.bsky.social · 02/06/2025
1/ DNA sequence models like Borzoi predict gene expression and variant effects across tissues — but how can someone adapt the model to a custom experiment? @drkbio.bsky.social, Johannes Linder and I propose a solution via parameter-efficient fine-tuning (PEFT). www.biorxiv.org/content/10.1...
biorxiv.org
Parameter-Efficient Fine-Tuning of a Supervised Regulatory Sequence Model
DNA sequence deep learning models accurately predict epigenetic and transcriptional profiles, enabling analysis of gene regulation and genetic variant effects. While large-scale training models like E...
13312
Reposted by Eric Kernfeld
Jayati Sharma, PhD, ScM @jayatirsharma.bsky.social · 02/06/2025
Are you an early-career researcher working on the computational analysis of population-level human genetics data? We want to hear from you about if, how, and why you use population descriptors in your research! Fill out our short survey: forms.gle/SCiNUq71wgi5...
Do you work with human computational genetic/genomics data? Are you a trainee or early career scientist? Take our survey! And participate in our research study on how you use population descriptors in your work 

Survey link and QR Code are in the flyer
1106
Reposted by Eric Kernfeld
Valentine Svensson @nxn.se · 02/06/2025
The duplication crisis: the other replication crisis - www.worksinprogress.news/p/the-duplic...
worksinprogress.news
The duplication crisis: the other replication crisis
How bad publishing incentives hinder long-term thinking in computational biology research
032
Eric Kernfeld @ekernf01.bsky.social · 29/05/2025
new advance in classifying species-of-origin for long reads. Potentially very valuable for diagnostics!! But hot damn, look at that memory consumption.
030
Eric Kernfeld @ekernf01.bsky.social · 28/05/2025
Wake up, babe. New cat coat color mechanism just dropped. www.theguardian.com/us-news/2025...
theguardian.com
Scientists solve the mystery of ginger cats – helped by hundreds of cat owners
Discovery has implications for ‘all cells and tissues’ and research bridged gap between scientists and non-scientists
040
Eric Kernfeld @ekernf01.bsky.social · 28/05/2025
I'm just gonna leave this here for no reason at all. stopstick.com
stopstick.com
Home
Stop Stick® is Law enforcement's choice in pursuit termination and pursuit prevention. Stop Stick tire deflation devices end pursuits safely.
040
Eric Kernfeld @ekernf01.bsky.social · 27/05/2025
This is really interesting to look at. I expected more dark ink on the east coast, especially NYC. In Providence, is that tangly black coastline, or dark blue?
000
Eric Kernfeld @ekernf01.bsky.social · 25/05/2025
There's a lot going on with transcriptome transformers, virtual cells, and perturbation prediction. I updated my recap: ekernf01.github.io/perturbation...
table with a rundown of perturbation prediction benchmark papers, describing each study's datasets, baselines, and major findings.
050
Eric Kernfeld @ekernf01.bsky.social · 25/05/2025
AI has been superhuman in genomics for a long time. How many of you can look at a vcf (or a FASTQ, lol) and tell me what part of England someone is from? peopleofthebritishisles.web.ox.ac.uk/population-g... Keep in mind, for text and images, our brains process the RAW data.
peopleofthebritishisles.web.ox.ac.uk
Population genetics
000
Eric Kernfeld @ekernf01.bsky.social · 25/05/2025
... I kinda did this with PEREGGRN, though txpert now claims some real progress. This study threads the needle nicely by finding a data split that will probably favor the more causally correct models, but also isn't so hard that the signal evaporates.
000