Sign in

Andrew Carroll

@acarroll.bsky.social
2.8K followers 220 following 27 posts

Product lead Genomics Google Research

PostsRepliesMedia
Reposted by Andrew Carroll
Guillaume Holley @guillaumeolesan.bsky.social · 20/08/2026
I am delighted to present with @hannespetur.bsky.social our new study on the Icelandic pangenome reference HPRC-ICE. This work goes all the way from pangenome construction to disease association! (1/7) www.nature.com/articles/s41...
nature.com
An Icelandic pangenome reference - Nature
Newly developed methods enable construction of Icelandic haplotypes and mapping of population-scale short reads to a pangenome by reducing reference bias and improving discovery in low-mappability reg...
13217
Andrew Carroll @acarroll.bsky.social · 02/07/2026
OK, that was indeed interesting. BWA MEM3 is also very fast. MiniBWA is a median of 40% faster than BWA MEM3 and faster in 85% of genomes, with the smallest of genomes favoring BWA MEM3. Thanks for the suggestion. BWA MEM3 is great work as well.
152
Andrew Carroll @acarroll.bsky.social · 02/07/2026
That is interesting. Let me see if I can quickly do that with the faster tests. I can't test the methylation, but the other components I could.
120
Andrew Carroll @acarroll.bsky.social · 01/07/2026
How good is MiniBWA, the successor to BWA? To test it, I ran MiniBWA on sequencing from 76 different species, comparing mapping speed, rate and accuracy with BWA MEM. In short, it's really good. If you map short reads, it's well worth your time. andrewcarroll.github.io/2026/06/30/t...
andrewcarroll.github.io
The Best of Both Worlds - Assessing MiniBWA
Recently, Heng Li released MiniBWA (GitHub) alongside a paper by Heng Li and Nils Homer describing the method (paper). MiniBWA builds on the approaches in Minimap2 (also by Heng Li), but falls back on...
19862
Andrew Carroll @acarroll.bsky.social · 28/05/2026
This blog shares some thoughts on protein and genome foundation models. The first part explains some of the concepts by training models for example tasks. The second part is opinion on the state of the field. andrewcarroll.github.io/2026/05/26/g...
054
Andrew Carroll @acarroll.bsky.social · 07/03/2026
Release led by Kishwar Shafin. Special thanks to Ehud Amitai and Ultima Genomics for a contribution that improves accuracy (~4% error reduction) for all technologies. Core team: Alexey Kolesnikov, Daniel Cook, Lucas Brambrink & manager Pi-Chuan Chang. 20%er work in release notes
020
Andrew Carroll @acarroll.bsky.social · 07/03/2026
Release of DeepVariant v1.10 Phased VCF output for long-reads Accuracy improvements for multi-allelic variants Pangenome accuracy improvements (18% fewer errors) Most technologies ~10% faster RNA-seq is a full supported mode DeepSomatic is 12-40% faster github.com/google/deepv...
github.com
Release DeepVariant 1.10.0 · google/deepvariant
DeepVariant: Continuous phasing: Long-read variant calls (PacBio and ONT) are now natively phased and phased output is generated for both vcf and gvcf formats. Fuzzy channels: Added “fuzzy channel...
1144
Reposted by Andrew Carroll
Chris Saunders @ctsa.bsky.social · 20/02/2026
What if you could improve small variant accuracy, CNV inference, and interpretability of your HiFi WGS data by taking a different approach to read mapping? Our new preprint describes portello, a method which demonstrates the potential for such improvements. (1/5)
Comparison of read mappings at HG002 chr4:40,294,825-40,295,700, showing conventional (pbmm2) read mappings (above) and portello mappings (below). The same set of unaligned input reads were input into each mapping process.
12511
Reposted by Andrew Carroll
Maria Nattestad 🧬💻 @omgenomics.com · 12/02/2026
Lab tour and takeaways: www.youtube.com/watch?v=nS2o...
youtube.com
The "Why Not?" Era of Sequencing Has Begun
YouTube video by OMGenomics
131
Andrew Carroll @acarroll.bsky.social · 10/02/2026
I wrote up some thoughts on the automation of lab work, in particular how it relates to how people will work in the lab. In short, it will deliver a lot of value for assays run at scale, but there is a long tail of experiments where humans are essential. andrewcarroll.github.io/2026/02/09/f...
andrewcarroll.github.io
For Automation The Wet Lab Has An Incredibly Long Tail
Disclaimer: These are solely my views.
041
Reposted by Andrew Carroll
Vertebrate Genomes Project @vertebrategenomes.bsky.social · 03/02/2026
Thanks to the support of @wcs.org and Google Research, we have sequenced and assembled the genomes of nine endangered species, with more on the way! To learn more: blog.google/innovation-a...
African penguin
Source: Wildlife Conservation SocietyCotton top tamarin
Source: Wildlife Conservation SocietyEld's deer
Source: Wildlife Conservation SocietyElongated tortoise
Source: Wich’yanan (Jay) Limparungpatthanakij, via inaturalist.org and Wikimedia Commons
062
Andrew Carroll @acarroll.bsky.social · 03/02/2026
This blog talks about the great work of the @ebpgenome.bsky.social. To support it Google.org has funded sequencing and open release of 13 genomes, with a $3M commit to sequence 150 more and develop methods to improve assembly finishing and other bottlenecks. blog.google/innovation-a...
blog.google
How we’re helping preserve the genetic information of endangered species with AI
Scientists are working to sequence the genome of every known species on Earth.
0112
Reposted by Andrew Carroll
Barack Obama @barackobama.bsky.social · 25/01/2026
The killing of Alex Pretti is a heartbreaking tragedy. It should also be a wake-up call to every American, regardless of party, that many of our core values as a nation are increasingly under assault.
30555981119370
Andrew Carroll @acarroll.bsky.social · 24/12/2025
I've been thinking about the "virtual cell" concept and wanted to write up a few thoughts. Specifically on how I think the prior experience in GWAS informs the most likely way these models will be useful. andrewcarroll.github.io/2025/12/23/t...
andrewcarroll.github.io
The Virtual Cell Will Be More Like Gwas Than Alphafold
There has been significant discussion recently on the concept of the “virtual cell.” I want to summarize the key concepts regarding what the field wants from a virtual cell and the challenges we face....
03717
Reposted by Andrew Carroll
Joseph Guhlin @josephguhlin.bsky.social · 28/10/2025
🐧We researched one of the world’s rarest #penguins. The yellow‑eyed penguin (aka hoiho/takaraka) isn’t one homogeneous species after all! www.biorxiv.org/content/10.1... #hoiho #conservation #genomics #birds #nzwildlife #endangered #wildlife #nature
Hoiho - the world’s rarest penguin, fewer than 150 mainland pairs left
36422
Reposted by Andrew Carroll
Stephen Turner @stephenturner.us · 16/10/2025
Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic www.nature.com/articles/s41... (read free: rdcu.be/eLny0) github.com/google/deeps...
0115
Reposted by Andrew Carroll
Adam Phillippy @aphillippy.bsky.social · 22/09/2025
Delighted to finally announce a preprint describing the Q100 project! “A complete diploid human genome benchmark for personalized genomics” For which we finished HG002 to near-perfect accuracy: www.biorxiv.org/content/10.1... 🧵[1/14]
biorxiv.org
A complete diploid human genome benchmark for personalized genomics
Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and ...
49757
Andrew Carroll @acarroll.bsky.social · 21/08/2025
I'll be speaking in this webinar (go.roche.com/sbx-d) on September 10, where I'll share our benchmarks and observations for Roche's SBX sequencing instrument, as well as models developed by our team for SBX data.
go.roche.com
Germline Small Variant Calling Workflow for SBX Duplex Data
Wednesday, September 10, 2025 at 12:00 PM Eastern Daylight Time.
1104
Andrew Carroll @acarroll.bsky.social · 13/05/2025
Also thanks to 20% contributors: Ben Soudry, Mike Kruskal, Sowmiya Nagarajan, Suchismita Tripathy, Francisco Unda, Vasiliy Strelnikov And community contributions from Sam Yadav and Seraj Ahmad at Roche improving the code for custom model training
010
Andrew Carroll @acarroll.bsky.social · 13/05/2025
Release led by Kishwar Shafin, contributions by Daniel Cook, Alexey Kolesnikov, Lucas Brambrink, and Pi-Chuan Chang as engineering manager. Thanks to student researcher contributions from Farica Zhuang and Mobin Asri. DeepSomatic release page: github.com/google/deeps...
github.com
Release DeepSomatic 1.9.0 · google/deepsomatic
DeepSomatic: In this release, we are introducing FFPE_WGS_TUMOR_ONLY and FFPE_WES_TUMOR_ONLY models. The WGS and WGS_TUMOR_ONLY models have been retrained with all datasets described in the manusc...
120
Andrew Carroll @acarroll.bsky.social · 13/05/2025
Release of DeepVariant and DeepSomatic v1.9 DV: Now train on HG002 T2T-Q100. Error reduction of 12% for Illumina and 30% for PacBio on this truth set. 25% faster. DeepTrio is 5x faster (20h -> 4h). DS: New models FFPE_TUMOR_ONLY for {WGS, WES}. Much improved WGS models. github.com/google/deepv...
github.com
Release DeepVariant 1.9.0 · google/deepvariant
DeepVariant: In this version we have updated our training scheme for the HG002 sample with the newly released HG002-T2T truth set which improves accuracy against that truth set. Our labeling metho...
1209
Reposted by Andrew Carroll
Adam Keiper @adamkeiper.com · 02/02/2025
Incredibly moving Justin Trudeau remarks: "We have fought and died alongside you....During your darkest hours...we were always there. Standing with you, grieving with you, the American people....Canadians are a little perplexed as to why our closest friends and neighbors are choosing to target us."
680225396182
Andrew Carroll @acarroll.bsky.social · 01/02/2025
You have some additional control on memory use by the number of threads you run with. For running on GPU, I am not sure if you've seen this - github.com/google/deepv... Which requires a little more configuration, but can let you better manage CPU-GPU tradeoffs. Definitely expert use.
github.com
000
Andrew Carroll @acarroll.bsky.social · 01/02/2025
Hi Eric, sorry to not notice till now. From the DV FAQ, we see the Keras model takes 16GB of memory (github.com/google/deepv...). It's possible that pangenome-aware models will take more memory, and we do observe more memory per thread used for that. Definitely not lower than 16GB.
github.com
200
Andrew Carroll @acarroll.bsky.social · 06/12/2024
Great question. We were talking recently about L40S benchmarks. We don't have that data immediately on hand, but are planning to generate runtime stats for it.
110
Andrew Carroll @acarroll.bsky.social · 06/12/2024
They're very close - to the point that small changes of coverage or the inclusion of PCR in preparation would tip between one and the other.
030
Andrew Carroll @acarroll.bsky.social · 05/12/2024
Release led by DeepVariant tech lead Kishwar Shafin. Team Engineering manager Pi-Chuan Chang. Small model work led by Lucas Brambrink. Pangenome-aware led by Mobin Asri and Juan Carlos Mier. Fast pipeline by Alexey Kolesnikov. Kinnex/MAS-Seq model by Daniel Cook and Shiyi Yin from Verily. 3/3
020
Andrew Carroll @acarroll.bsky.social · 05/12/2024
Added SPRQ to PacBio training, reducing Indel error on SPRQ by 26%. Added Platinum Pedigree training data for PacBio model, reducing errors by 34% on more extensive Platinum truth. New model and case study for Kinnex/Mas-Seq/Iso-Seq. Additional speed options for GPU pipelines 2/3
Plots of SNP and Indel error numbers for DeepVariant models. Shows a Indel error reduction of 26% for PacBio and a ~50% SNP error reduction for ONT.
384
Andrew Carroll @acarroll.bsky.social · 05/12/2024
Release of DeepVariant 1.8. Large speed improvement (~67% faster) via small model for easy sites. New Pangenome-aware option. Reduces error by ~30% for vg-mapped WGS, ~10% for BWA WGS, ~5% BWA exome. New config for custom model users, see release notes. (github.com/google/deepv...)
Runtime figure for new version of DeepVariant with and without small model. Showing reduction in runtime of 155 minutes to 101 minutes with Illumina, 174 minutes to 71 minutes with PacBio, and 295 minutes to 114 minutes with Oxford Nanopore.
13814
Reposted by Andrew Carroll
Benedict Paten @benedictpaten.bsky.social · 15/12/2023
How do we make a pangenome maximally relevant for the study of a new sample? www.biorxiv.org/content/10.1...
biorxiv.org
Personalized Pangenome References
bioRxiv - the preprint server for biology, operated by Cold Spring Harbor Laboratory, a research and educational institution
0103
Andrew Carroll @acarroll.bsky.social · 26/10/2023
Release by Kishwar Shafin Major contributions from Pi-Chuan Chang, Daniel Cook, Alexey Kolesnikov Google 20%ers: Will Kwan, Pauline Sho, Lucas Brambrink, Mo Samman, Atilla Kiraly UCSC for vg: @benedictpaten.bsky.social, Shloka Negi, Jimin Park, Mobin Asri Pacbio: Billy Rowell, Nathaniel Echols
010
Andrew Carroll @acarroll.bsky.social · 26/10/2023
There are now custom models and case studies for CompleteGenomics instruments. T7: github.com/google/deepv... G400: github.com/google/deepv... For now, these are stand alone models. We'll likely consider whether we can jointly include these in the broad WGS model later.
120
Andrew Carroll @acarroll.bsky.social · 26/10/2023
The changes to DeepTrio for de novo detection are substantial. We now in two steps - first for overall accuracy and then a weighted fine tuning for de novos. Our benchmarks show large improvements in de novo calling relative to the prior DeepTrio. github.com/google/deepv...
github.com
110
Andrew Carroll @acarroll.bsky.social · 26/10/2023
Want to benefit from pangenomes and want a recipe? github.com/google/deepv... Shows a step by step process, with Docker images for how to map to a Pangenome reference w/ vg and calls w/ DeepVariant. Final calls are more accurate and in GRCh38 coordinates. Thanks to the UCSC team for co-development
120
Andrew Carroll @acarroll.bsky.social · 26/10/2023
Release of DeepVariant v1.6. Support for haploid regions, chrX/Y. Workflow for Pangenome FASTQ-to-VCF. Major DeepTrio improvements for de novo variants. Models for CompleteGenomics T7, G400 Add NovaSeqX to training data Release by Kishwar Shafin github.com/google/deepv...
github.com
Release DeepVariant 1.6.0 · google/deepvariant
Improved support for haploid regions, chrX and chY. Users can specify haploid regions with a flag. Updated case studies show usage and metrics. Added pangenome workflow (FASTQ-to-VCF mapping with V...
174
Reposted by Andrew Carroll
Mikhail Kolmogorov @mishakolmogorov.bsky.social · 24/10/2023
Proud of Ayse Keskus and Asher Bryant in my group for making this happen! This work is a collaboration with Children's Mercy, UCSC and Google Health - who are also releasing the first version of DeepSomatic today: github.com/google/deeps...
142
Andrew Carroll @acarroll.bsky.social · 24/10/2023
TThanks to Jimin Park and Benedict Paten from UCSC, Mikhail Kolmogorov from NCI for analysis and testing. This group has also developed Severus for somatic SV calling, and we've worked closely with them (github.com/KolmogorovLa...) Thanks to Khi Pin Chua and Billy Rowell from PacBio
github.com
GitHub - KolmogorovLab/Severus: A tool for somatic structural variant calling using long reads
A tool for somatic structural variant calling using long reads - GitHub - KolmogorovLab/Severus: A tool for somatic structural variant calling using long reads
020
Andrew Carroll @acarroll.bsky.social · 24/10/2023
Initial release of DeepSomatic, which identifies subclonal variants when given tumor and normal BAM files. Pre-trained models and case studies available for Illumina and PacBio. Development led by Kishwar Shafin which built off a framework by Pi-Chuan Chang. (github.com/google/deeps...)
github.com
GitHub - google/deepsomatic
Contribute to google/deepsomatic development by creating an account on GitHub.
23617
Reposted by Andrew Carroll
Alex Rubinsteyn @alexr.bsky.social · 14/09/2023
Best resource for getting extracellular domain localization from Ensembl gene/protein IDs? I tried the subcellular locations in HPA but the membrane annotation is mostly fully intracellular proteins. 🧪🧬🖥️🔬
384