Sign in

Kaitlin Samocha

@ksamocha.bsky.social
1.5K followers 244 following 70 posts

Assistant Investigator @ MGH / Broad / HMS. Focus on human genomics and modeling rare variation. She/her

PostsRepliesMedia
Reposted by Kaitlin Samocha
Aaron Quinlan (he/him) @aaronquinlan.bsky.social · 28/08/2026
I am delighted to announce that The Department of Human Genetics at the University of Utah is continuing a multi-year recruiting initiative for multiple tenure-track faculty positions at the rank of Assistant Professor. Please share and consider applying! utah.peopleadmin.com/postings/208...
184102
Reposted by Kaitlin Samocha
Danny Miller, MD, PhD @danrdanny.bsky.social · 06/07/2026
Hi everyone! I'm excited to announce that our lab at UW will soon be recruiting a postdoc to work on DNA methylation signatures in rare disease. This will be a mostly computational position. Please share and reach out if interested! millerlaboratory.com
millerlaboratory.com
Miller Lab | University of Washington
Led by Danny E. Miller, MD, PhD, the Miller Lab uses long-read Nanopore sequencing to investigate the significance of structural genomic variation, methylation, and RNA in human disease, and to impro...
02317
Reposted by Kaitlin Samocha
Zornitza Stark @zornitza.bsky.social · 25/06/2026
📣 Out now @naturemedicine.bsky.social 🧬🤖 👉 rdcu.be/fqasl Our automated reanalysis tool Talos enables broad adoption and delivers timely + equitable #raredisease #diagnosis at scale! ♻️ Runs monthly on >10K datasets 🎯High specificity 💰Low running costs 🤗Open source @dgmacarthur.bsky.social
02311
Reposted by Kaitlin Samocha
Jack Kosmicki @jakphd.bsky.social · 25/06/2026
After 4 years, it's rather nice to finally present our work on genetic's model trait, height, in >1.4M WES/WGS samples (826k discovery; led by Adam Locke & Goncalo Abecasis where we found (amongst many other things) 207 genes (P<1.75e-9). A thread of findings below⬇️ www.medrxiv.org/content/10.6...
1187
Reposted by Kaitlin Samocha
prathitha.bsky.social @prathitha.bsky.social · 23/06/2026
Excited to share our new preprint, "Inference of elevated mutation rates and variant effects using 700k exomes"! - www.biorxiv.org/content/10.6... Using gnomAD v4, we estimate per-variant missense selection coefficients, and find loss-of-function (LoF) mutations with enhanced mutation rates.
biorxiv.org
Inference of elevated mutation rates and variant effects using 700k exomes
Genomic sequencing is now widely accessible for genetic diagnostics and is emerging as a component of newborn screening. This technological development generates the need to characterize incoming muta...
1127
Reposted by Kaitlin Samocha
European Society of Human Genetics @eshg.bsky.social · 13/06/2026
Attend the Leena Peltonen Award lecture now at #eshg2026 with Nicola Whiffin from Oxford University and learn About her recent discovieries in rare diseases and the broader inclusion of non-coding variants in clinical genetic testing that led to the Leena Peltonen Award. Congratulations!
0299
Reposted by Kaitlin Samocha
Zornitza Stark @zornitza.bsky.social · 13/06/2026
👏👏👏 fantastic to see @nickywhiffin.bsky.social win the Leena Peltonen award at #eshg2026 @eshg.bsky.social 👏👏👏
0176
Kaitlin Samocha @ksamocha.bsky.social · 12/06/2026
Excited to be at #ESHG2026 in the beautiful Gothenburg! For those of you attending and up for an extra game, check out the bingo cards 👇
010
Kaitlin Samocha @ksamocha.bsky.social · 30/03/2026
We are excited to share our gnomAD v4.1.1 release gnomad.broadinstitute.org/news/2026-03... Major changes: * Constraint scores on X and Y * Improved coverage correction * LOFTEE fix * Guidance on constraint cut-offs * New quality flag for low coverage/mappability genes @gnomad-project.bsky.social
gnomad.broadinstitute.org
gnomAD v4.1.1 | gnomAD browser
The Genome Aggregation Database (gnomAD) is a resource developed by an international coalition of investigators, with the goal of aggregating and harmonizing both exome and genome sequencing data from...
1112
Reposted by Kaitlin Samocha
Konrad @konradjk.bsky.social · 26/03/2026
To nominate disease genes, we introduce two discovery scores: ΔPEPPER flags genes where biological features predicts clinical impact beyond whats published, and DisPo highlights genes under strong constraint with limited literature. Together they prioritize hundreds of candidate genes for follow-up
101
Reposted by Kaitlin Samocha
Konrad @konradjk.bsky.social · 26/03/2026
We train a new model trained on biomedical literature (PEPPER_XGB). Mix LOEUF and PEPPER to make an OMELET, which outperforms either each individual model in identifying disease genes
222
Reposted by Kaitlin Samocha
Konrad @konradjk.bsky.social · 26/03/2026
We introduce LOEUF-MIS, combining pLoF and top 1% predicted deleterious missense constraint. This captures not just LoF but also gain-of-function and dominant-negative signals.
Precision-recall curves showing LOEUF-MIS outperforming other metrics
111
Reposted by Kaitlin Samocha
Konrad @konradjk.bsky.social · 26/03/2026
gnomAD v4's 5x sample increase benefits both common and rare disease: more common variants observed across ancestries improve diagnostic filtering, while more rare variants strengthen constraint metrics for disease gene detection
Figure 1a-b, growth of variation with sample size
111
Reposted by Kaitlin Samocha
Konrad @konradjk.bsky.social · 26/03/2026
Excited to share our new preprint on gnomAD v4! We present the full analysis of 730,947 exomes — new constraint metrics, improved LoF annotation (LOFTEE-2), LLM-based literature curation, and a unified framework for gene discovery and rare disease diagnosis. www.medrxiv.org/content/10.6...
medrxiv.org
Integrating 730,947 exome sequences with clinical literature improves gene discovery
Accurate estimates of allele frequencies aid in genetic discovery, including rare disease diagnosis, common disease investigations, and population genetics. Here, we present the Genome Aggregation Dat...
23717
Reposted by Kaitlin Samocha
Abraham Palmer @abepalmer.bsky.social · 06/01/2026
If you are submitting an NIH grant in February, you will be required to use SciENcv to prepare you biosketch. IT IS MUCH WORSE THAN YOU CAN POSSIBLY IMAGINE. Set aside *at least* 4 hours just to transfer an existing an biosketch into SciENcv.
2717792
Reposted by Kaitlin Samocha
Nik Baya @nbaya.bsky.social · 06/01/2026
Why do some individuals defy their polygenic score? In the largest study of its kind (402k UKB individuals; 7 continuous traits + 3 diseases), we asked: If your phenotype deviates from common-variant polygenic score prediction, what's driving that difference? www.medrxiv.org/content/10.6...
25027
Reposted by Kaitlin Samocha
Sasha Gusev @sashagusevposts.bsky.social · 30/11/2025
Massive single-cell study by Kanai et al (www.medrxiv.org/content/10.1...): - Once statistical power is high, constrained genes have more (though weaker) eQTLs. - Chromatin-QTLs near constrained genes have "normal" effect sizes, colocalize more with disease, but exhibit attenuated peak-gene effects.
1349
Reposted by Kaitlin Samocha
Kartik Chundru @chundru.bsky.social · 08/11/2025
New paper on everyone’s favourite topic, QC! We show why you should do genotype-level QC on your WGS data www.biorxiv.org/content/10.1... Very real quotes about this paper - “The most exciting, mind-blowing paper of the year!” “On a par with Fisher 1918” “I read it every night. Just so beautiful”
biorxiv.org
Genotype-level quality control substantially reduces error rates in population-scale whole-genome sequencing
Population-scale whole-genome sequencing data will contain many individual-level genotype errors, even after allele-level quality control (QC). We establish the need for genotype-level QC using UK Bio...
24619
Reposted by Kaitlin Samocha
European Society of Human Genetics @eshg.bsky.social · 04/11/2025
New study of 800K+ genomes from gnomAD reveals most “pathogenic” variants in healthy people aren’t truly disease-tolerant. They are explained by annotation errors, mosaicism, or compensatory variants. 🧬 A big step for precision medicine! www.nature.com/articles/s41...
nature.com
Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database - Nature Communications
Here the authors provide an explanation for 95% of examined predicted loss of function variants found in disease-associated haploinsufficient genes in the Genome Aggregation Database (gnomAD),…
02315
Reposted by Kaitlin Samocha
Pragati Kore @pragskore.bsky.social · 06/10/2025
📃 We’re excited to share our latest work, now published in Nature Communications — a major update to the Genome Aggregation Database (gnomAD) that improves allele frequency resolution for two gnomAD-defined genetic ancestry groups using local ancestry inference (LAI).
nature.com
Improved allele frequencies in gnomAD through local ancestry inference - Nature Communications
This study incorporates local ancestry into the Genome Aggregation Database (gnomAD) to improve allele frequency estimates for admixed populations, enhancing variant interpretation and enabling more accurate and equitable genomic research and clinical care.
1298
Reposted by Kaitlin Samocha
Matthew Neville @mattneville.bsky.social · 08/10/2025
Now published! Our paper on: (1) Accurate sequencing of sperm at scale (2) Positive selection of spermatogenesis driver mutations across the exome (3) Offspring disease risks from male reproductive aging [1/n] www.nature.com/articles/s41...
nature.com
Sperm sequencing reveals extensive positive selection in the male germline - Nature
A combination of whole-genome NanoSeq with deep whole-exome and targeted NanoSeq is used to accurately characterize mutation rates and genes under positive selection in sperm cells.
38650
Reposted by Kaitlin Samocha
Nicky Whiffin @nickywhiffin.bsky.social · 31/07/2025
📣 We are recruiting! Please share!! Are you a bioinformatician / computational scientist who wants to apply your skills to understanding regulatory biology and improving rare disease diagnosis and treatment? 🧠 💻 🧬 🩺 We have two roles available 👇 🧵 1/4
Image of an old building in Oxford with the heading 'postdoc opportunities' and the text 'computational approaches to improve rare disease diagnosis and treatment' and 'Big Data Institute, University of Oxford'
14443
Reposted by Kaitlin Samocha
Nicky Whiffin @nickywhiffin.bsky.social · 18/08/2025
Isn't genetics cool??? Within only 145 nucleotides(!) of a non-coding RNA (RNU4-2) - different variants in distinct regions / structures cause three distinct disorders!!! (all discovered within the last 18 months) 🤯🤓🧬❤️
Schematic of the U4 and U6 snRNAs with coloured annotations to note nucleotides linked to different disorders:
- Teal in the T-loop and Stem III for ReNU syndrome (Chen et al. Nature 2024 and Greene et al. Nature Medicine 2024)
- Red for variants causing a recessive NDD in Stem II, the k-turn and Sm protein binding sites (De Jonghe et al. medRxiv 2025 and Rius & Blakes medRxiv 2025)
- Yellow for the central loop and Retinitis pigmentosa (Quinodoz et al. medRxiv 2025)
16213
Reposted by Kaitlin Samocha
James Fasham @jamesfasham.bsky.social · 24/05/2025
🗣️ Quote of #ESHG2025 (so far) "Who licks bone !?!" 🦴 - Johannes Krause Anyone have that on your bingo card? Well apparently archeologists do, to distinguish bone from stones and it causes problems in DNA sequencing. 🤔
12810
Kaitlin Samocha @ksamocha.bsky.social · 24/05/2025
We are just wrapping up day 1 at #ESGH2025 in beautiful Milan. For those who want some extra fun while listening to the great science, you can play bingo.👇 I know multiple of these have already occurred.
150
Reposted by Kaitlin Samocha
Alex Hoischen @ahoischen.bsky.social · 24/05/2025
Buongiorno Milano! Ready for a great day 1 of #eshg2025? Packed program of excellent science 8.30am-8.00pm - plus networking event till 9.30pm to meet many friends, colleagues and collaborators! …andiamo @eshg.bsky.social @eshgyoung.bsky.social
0156
Reposted by Kaitlin Samocha
Zornitza Stark @zornitza.bsky.social · 23/05/2025
🤗 Hugely excited to share our work on automating iterative reanalysis in #raredisease, preprint out: www.medrxiv.org/content/10.1... 🤖🧬 github.com/populationge... A superb collaboration with @dgmacarthur.bsky.social @cassimons.bsky.social @heidirehm.bsky.social @ksamocha.bsky.social and many more!
medrxiv.org
Scalable automated reanalysis of genomic data in research and clinical rare disease cohorts
Reanalysis of genomic data in rare disease is highly effective in increasing diagnostic yields but remains limited by manual approaches. Automation and optimization for high specificity will be necess...
12413
Reposted by Kaitlin Samocha
deciphergenomics.bsky.social @deciphergenomics.bsky.social · 07/05/2025
Human Developmental Cell Atlas (HDCA) expression data is now displayed. Expression is displayed in 12 sections of a 6-7 post-conception week human embryo, alongside a sagittal view which displays the region of the embryo represented by each section @mhaniffa.bsky.social
0177
Reposted by Kaitlin Samocha
Nicky Whiffin @nickywhiffin.bsky.social · 02/03/2025
A few weeks ago, I had an incredibly emotional call with James Coney, a writer for the Sunday Times whose son Charlie was in the @genomicsengland.bsky.social 100k project and was recently diagnosed with ReNU syndrome. This beautiful article tells their story ❤️ www.thetimes.com/article/0bcc...
thetimes.com
My son Charlie — and the breakthrough that changed our lives
James Coney and his wife, Sarah, struggled not knowing why their 12-year-old was born with a severe learning disability. In their darkest moments, they blamed themselves. Then, out of the blue, came a...
511040
Kaitlin Samocha @ksamocha.bsky.social · 20/04/2024
Recently out on #bioRxiv: our updated approach to identify regional variability in missense mutation intolerance (“constraint”) in protein-coding genes using the gnomAD database. www.biorxiv.org/content/10.1... 1/10
162
Kaitlin Samocha @ksamocha.bsky.social · 08/03/2024
Some updated guidance on our gnomAD v4 constraint scores: gnomad.broadinstitute.org/news/2024-03... The @gnomad-project.bsky.social team is hard at work on v4.1 and improvements across the board, so expect more updates. Thanks to Katherine Chao for spearheading this blogpost.
021
Kaitlin Samocha @ksamocha.bsky.social · 08/12/2023
Our paper describing a way to infer the phase of rare variant pairs using gnomAD v2 is out now in Nature Genetics. We hope that the resource we generated will be useful when interpreting rare co-occurring variants in the context of recessive disease. www.nature.com/articles/s41...
1168
Kaitlin Samocha @ksamocha.bsky.social · 17/11/2023
It’s the final day of giving thanks for the teams that make gnomAD possible. Today is focused on the participants in studies, data contributors (>300!), the Scientific Advisory Board, and our steering committee. 1/6
100
Kaitlin Samocha @ksamocha.bsky.social · 16/11/2023
Day four of giving thanks to the teams that make gnomAD happen is focused on the CNV and SV teams! v4 is the first time we released structural variants at the same time as SNVs/indels, specifically: - CNVs from 464,297 exomes - SVs from 63,046 genomes 1/4
110
Kaitlin Samocha @ksamocha.bsky.social · 15/11/2023
Continuing the week of thanking the teams that make gnomAD possible, today I’m thanking the data generation and operations teams! As a reminder, I'm only highlighting a few individuals of the many that contribute. You can see more on our team page: gnomad.broadinstitute.org/team 1/5
120
Kaitlin Samocha @ksamocha.bsky.social · 14/11/2023
Next up in my week of thanks to the teams that make gnomAD: the gnomAD browser team! With 150k+ views a week, the browser is a crucial part of making gnomAD accessible. Quickly loading data from >800k samples + presenting it in a user-friendly format is no small feat. 1/6
110
Kaitlin Samocha @ksamocha.bsky.social · 13/11/2023
Now that the dust has settled on the gnomAD v4 release, which hopefully many of you have already checked out, I wanted to take this week to thank many of the members of the team who made this possible. First up this week is the amazing production team. 1/7
120
Reposted by Kaitlin Samocha
Molly Przeworski @mollyprz.bsky.social · 09/11/2023
Helpful description of the choices of population designators in gnomAD: gnomad.broadinstitute.org/news/2023-11...
184
Kaitlin Samocha @ksamocha.bsky.social · 03/11/2023
Constraint scores are now up for v4!
031
Reposted by Kaitlin Samocha
Genome Aggregation Database (gnomAD) @gnomad-project.bsky.social · 03/11/2023
To learn more about the impact of diversity on variant discovery and gene constraint please attend Katherine Chao’s #ASHG23 talk tomorrow (11/4) at 11am in rm 202A
042
Reposted by Kaitlin Samocha
Genome Aggregation Database (gnomAD) @gnomad-project.bsky.social · 02/11/2023
As part of #gnomAD v4, in collaboration with the Talkowski Lab, we have released 1,199,117 genome SVs and 66,903 rare exome CNVs. These data represent the first gnomAD SV dataset released native to the GRCh38 reference genome. (1/2)
173
Kaitlin Samocha @ksamocha.bsky.social · 02/11/2023
First of our gnomAD talks at #ASHG2023 is Jack Fu on copy number variants (CNVs) in >630k individuals. Starts by discussing major challenges of calling CNVs from exome data. False positives are an issue -- using the newly described GATK-gCNV (nature.com/articles/s41...)
163
Kaitlin Samocha @ksamocha.bsky.social · 01/11/2023
So, so hyped to have gnomAD v4 out for #ASHG23. We'll be discussing it at a few talks this meeting. Please see: - Jack Fu, Nov 2, 1:45pm - Julia Goodrich, Nov 4, 10:30am - Katherine Chao, Nov 4, 11am cc: @gnomad-project.bsky.social
130
Reposted by Kaitlin Samocha
Genome Aggregation Database (gnomAD) @gnomad-project.bsky.social · 01/11/2023
The #gnomAD team is proud to announce the release of gnomAD v4! The v4 dataset includes 730,947 exomes & 76,215 genomes, which is ~5x larger than the v2 & v3 releases combined, & includes nearly 120K indivs of non-European genetic ancestry broad.io/gnomad #ASHG23 (1/11)
23221
Reposted by Kaitlin Samocha
Kristin @klewis.bsky.social · 27/10/2023
🧬 🖥️ For our genetics/genomics friends visiting #ASHG23, you can follow the conference's conversation on this feed: bsky.app/profile/did:... Just use #ASHG23 (or #ASHG2023) to post to the feed.
21314
Kaitlin Samocha @ksamocha.bsky.social · 26/10/2023
You know exciting things are afoot when Slack threads are >150 messages long within a few hours and no one can open them.
020
Reposted by Kaitlin Samocha
Nicky Whiffin @nickywhiffin.bsky.social · 12/10/2023
The resolution to this (from that other site) and a warning to anyone using these data: The bgen and plink files for the 200k genome release in UK Biobank are missing all multi-allelic sites, meaning there are a large proportion of variants missing from these files! 😮 The pVCF files look OK 🧬🖥️
1147
Kaitlin Samocha @ksamocha.bsky.social · 10/10/2023
Big welcome to @gnomad-project.bsky.social! Keep an eye on this space in the coming weeks for updates. #AcademicSky #Genomics
1142
Reposted by Kaitlin Samocha
Jonathan Pritchard @jkpritch.bsky.social · 01/10/2023
I'm delighted to release the first half of my new textbook in human genetics: web.stanford.edu/group/pritch... "An Owner's Guide to the Human Genome: an introduction to human population genetics, variation and disease"
web.stanford.edu
An Owner's Guide to the Human Genome
An Owner's Guide to the Human Genome
8296175