Sign in

Jim Shaw

@jimshaw.bsky.social
1.2K followers 576 following 125 posts

Postdoc at Dana-Farber and Harvard Med with Heng Li (@lh3lh3.bsky.social). Prev: UBC / UofT. I like thinking about biological sequence analysis and its applications to metagenomics / microbial genomics. jim-shaw-bluenote.github.io

PostsRepliesMedia
Reposted by Jim Shaw
Matt Olm @mattolm.bsky.social · 07/10/2026
My final postdoc project is out today in Nature! We show microbes co-migrated with humans out of Africa; three independent methods agree, and timing matches archaeological / genetic evidence (Fig. 4c). Seeing Fig. 4c for 1st time was one of my most memorable "eureka" moments! doi.org/10.1038/s415...
doi.org
Prehistoric global migration of vanishing gut microbes with humans - Nature
Microbial species disappearing from industrialized populations have evolved with humans over millennia and migrated worldwide with them—this may have consequences for human health.
310745
Reposted by Jim Shaw
Saria McKeithen-Mead, PhD @sciria.bsky.social · 05/10/2026
Pleased to share my first preprint from my postdoc, where I ask what drives mobile DNA movement between bacterial cells in complex communities like the gut microbiome. To answer this, we built a way to capture mobile DNA in the act of moving between cells, using short-read metagenomics. 🧵
doi.org
23517
Reposted by Jim Shaw
jakobheinz.bsky.social @jakobheinz.bsky.social · 01/10/2026
I’m excited to share our new preprint introducing HipHap, a tool for assigning long reads to haplotype-resolved diploid reference assemblies! With Max Marin, Matthew Meyerson, and @lh3lh3.bsky.social Preprint: www.biorxiv.org/content/10.6... 🧵 1/6
biorxiv.org
HipHap: Haplotype Assignment and Confidence Scoring for Diploid Reference Genomes
Diploid genome assemblies are now routinely available, but most read aligners were designed for haploid references, which have long been the gold standard. When reads are aligned to a diploid assembly...
22818
Reposted by Jim Shaw
Milot Mirdita @milot.bsky.social · 16/09/2026
ColabFold 1.6.3 is out! 2.5x faster, pip-installable, ipSAE+pDockQ2 scores. Thanks Choonghwan Lee, Marielle Russo, Gyuri Kim 🐍pip install colabfold[alphafold] CF2 Sneak Peak with AF3/Boltz/Protenix/ESMFold2… 🐍pip install "colabfold[alphafold3]@git+https://github.com/sokrypton/ColabFold@af3-preview"
29735
Jim Shaw @jimshaw.bsky.social · 15/09/2026
The sylph metagenome profiler is v1.0.0! sylph-docs.github.io A new DB format + approach --> huge performance gains: GTDB-R232 (200k species) now takes < 5 GB of RAM and ~30s (2GB fq.gz). Huge thanks to @benjwoodcroft.bsky.social and his ongoing performance efforts (github.com/wwood/weebill)
sylph-docs.github.io
Documentation for sylph - ultrafast, precise metagenomic profiling
26330
Reposted by Jim Shaw
Tami Lieberman @contaminatedsci.bsky.social · 08/09/2026
Two announcements regarding AccuSNV, which calls high precision SNVs across microbial genomes: Version 1.1 is now incredibly easy to install, run, and perform downstream analyses with (see image of typical output!). The manuscript was also published this summer in Genome Research! Links next...
15625
Reposted by Jim Shaw
Oxford Nanopore @nanoporetech.com · 04/09/2026
Still relying on legacy approaches for microbial analysis? Join this webinar to discover the latest tools for resolving strain-level diversity, recovering complete genomes, and unlocking deeper insights from complex microbial communities. Register here: bit.ly/3UnBy0r
142
Jim Shaw @jimshaw.bsky.social · 03/09/2026
Sylph's cANI can have slight biases at very low coverages (< 1x). In principle you're right; it strongly depends on the magnitude of the coverage. Here's a simple test in the supp. figs. for very small coverage values (x-axis; note the scale). More testing would be good, though.
110
Reposted by Jim Shaw
Gerry Tonkin-Hill @gerrythill.bsky.social · 02/09/2026
Glad to share StrainSpy, a new method developed by Sudaraka and co which identifies strain-level associations in metagenomics using ANI estimates from tools like Sylph. www.biorxiv.org/content/10.6...
biorxiv.org
Accurate detection of metagenomic strain-level associations using average nucleotide identity with StrainSpy
Genetic variation among microbial strains of the same species can profoundly influence their phenotypes, ecological functions, and impacts on human health. Traditionally, the relative abundance of a s...
12112
Jim Shaw @jimshaw.bsky.social · 02/09/2026
Cool preprint about strain-level metagenome association studies using k-mer containment. Mallawaarachchi et al. from @gerrythill.bsky.social's group. I tried tackling this in our original paper for sylph (sylph-docs.github.io), but this seems to be much more sophisticated. Excited to read!
2238
Reposted by Jim Shaw
bioRxiv Microbiology @biorxiv-microbiol.bsky.social · 01/09/2026
Accurate detection of metagenomic strain-level associations using average nucleotide identity with StrainSpy www.biorxiv.org/content/10.64898/20…
0102
Reposted by Jim Shaw
Kristoffer Sahlin @ksahlin.bsky.social · 01/09/2026
strobealign v0.18.0 is out, and it's now written in Rust. Handwritten by Marcel, started 2y ago. No LLM involved (yet*). Contributions by Nicolas and Ivan. 1/3
github.com
Release Version 0.18.0 · ksahlin/strobealign
Changes Strobealign has been ported to Rust. The Rust version is at least as fast and accurate as the C++ version. There are currently some regressions due to the port: The Rust version cannot sat...
2195
Reposted by Jim Shaw
Falk Hildebrand @bioinf.bsky.social · 07/08/2026
After a very long time, I'm very proud that we have now Joachim "magnus opum" on www.biorxiv.org/content/10.6... Protal is a ultra fast metagenomic species profiler. More precise than other profiled software, faster than most (not sylph though), but generates strain resolution instead.
02812
Reposted by Jim Shaw
Niranjan Nagarajan @niranjantw.bsky.social · 07/08/2026
Thank you for featuring our work National University Health System! Its great to have this article describe in simple terms what we have been doing for nearly a decade now in terms of microbial genomic surveillance and metagenomics 🥲 nuhsplus.edu.sg/article/dna-...
nuhsplus.edu.sg
DNA detectives: Hunting down hidden hospital bacteria | NUHS+
NUHS researchers use a genetic fingerprint to expose invisible microbial reservoirs, aiding outbreak prevention across local hospitals.
031
Reposted by Jim Shaw
George Bouras @gbouras13.bsky.social · 07/08/2026
If you are looking for web-based phage annotation, phage-annotation.org is live. Upload a genome to run Pharokka, Phold and Phynteny sequentially (see our recent protocols paper doi.org/10.1002/cpz1...). If you have feedback, please reach out Thanks to @ardc.edu.au for making this possible!
phage-annotation.org
Phage Annotation Server -- phage genome annotation
Free automated phage genome annotation -- pharokka, phold and phynteny in one pipeline.
46130
Reposted by Jim Shaw
Ragnar {Groot Koerkamp} @curiouscoding.nl · 05/08/2026
Travis has been cooking something up since IGGSy ;) A fast algorithm for SMEM finding in haplotype panels that uses space proportional to the number of 1s in the (sparse) matrix, inspired by the jump index [13].
[13] Ragnar Groot Koerkamp. Personal communication, 2026.
031
Reposted by Jim Shaw
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 01/08/2026
Riboseek is a fast RNA/DNA search. More sensitive than nhmmer at 250x speed. Structure-aware realignment produces MSAs approaching rMSA quality. Plus 1.7M precomputed RNA MSAs, and an API to search your own 📄 www.biorxiv.org/content/10.6... 💾 github.com/steineggerla... 🌐 search.foldseek.com/riboseek
217975
Reposted by Jim Shaw
Tami Lieberman @contaminatedsci.bsky.social · 31/07/2026
The manuscript describing SimPhyNI is now published! www.microbiologyresearch.org/content/jour... A few updates since last post, including some method tweaks that improve performance and speed. Also even easier to run. Try it for your microbial GWAS or epistasis needs!
microbiologyresearch.org
High-precision binary trait association on phylogenetic trees
Traditional methods for identifying associations between genomic features and traits, or between pairs of genomic traits, struggle when applied to bacterial genomes. While several microbial genome-wid...
14413
Reposted by Jim Shaw
Ryan Wick @rrwick.bsky.social · 24/07/2026
New blog post! I reran the Autocycler paper benchmarks on some new tools/versions/pipelines: rrwick.github.io/2026/07/24/b... (1/3)
rrwick.github.io
Benchmark update: Ilesta, Autocycler-fast and new versions
a blog for miscellaneous bioinformatics stuff
23416
Reposted by Jim Shaw
Steven Robbins @stevenjrobbins.bsky.social · 22/07/2026
It's out! Excited to present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), a comprehensive DB of 1000s of high-quality prokaryote, virus, plasmid, and chromosome-level eukaryote MAGs using Nanopore long reads. Subthreads incoming. Please share widely. 🙂 www.nature.com/articles/s41...
nature.com
The planktonic microbiome of the Great Barrier Reef - Nature
The Great Barrier Reef Microbial Genomes Database compiles prokaryotic, viral and eukaryotic genomes from seawater collected from the Great Barrier Reef, providing a rich resource for the study of mar...
88443
Reposted by Jim Shaw
Zamin Iqbal @zaminiqbal.bsky.social · 20/07/2026
Significant update to the AllTheBacteria paper, including discovering new antimicrobial peptides and testing in vitro and vivo. This has grown into a fantastic collaboration!
17231
Reposted by Jim Shaw
Vikram Shivakumar @vikramshivakumar.bsky.social · 15/07/2026
1/ Excited to share the newest tool in the pangenome MUMiverse: Shredtools! Shredtools enables a user to navigate the pangenome coordinate system with multi-MUMs. More in the thread🧵 Code: github.com/vikshiv/shredtools Interactive tool for querying HPRC assemblies: vikshiv.github.io/shredtools
biorxiv.org
Navigating the pangenome coordinate system with Shredtools
Existing notions of pangenome coordinates rely on hard-to-compute multiple sequence alignments. On the other hand, pangenome-wide exact unique matches (multi-MUMs) can be computed efficiently, and rep...
13822
Reposted by Jim Shaw
Heng Li @lh3lh3.bsky.social · 13/07/2026
minibwa-0.4 released with minor improvement and a few fixes to typos. Also a new blog post on "minibwa is the new bwa-mem": lh3.github.io/2026/07/04/m...
lh3.github.io
Minibwa is the new bwa-mem
0199
Reposted by Jim Shaw
Sina Majidian @sinamajidian.bsky.social · 09/07/2026
JOB ALERT! I'm hiring two postdocs in Computational Genomics to join our lab in beautiful Gothenburg, Sweden. Please share and repost! CGRLab.github.io/research/
01417
Reposted by Jim Shaw
Sergey Nurk @sergeynurk.bsky.social · 07/07/2026
Senior Bioinformatics Software Developer opening in my team in ONT's Applications department! Hybrid with a few days a week in our Oxford headquarters. We can sponsor visas! www.linkedin.com/jobs/view/44...
linkedin.com
Oxford Nanopore Technologies hiring Senior Bioinformatics Software Engineer, Applications in Oxfordshire, England, United Kingdom | LinkedIn
Posted 3:04:10 PM. Job DescriptionOur goal is to bring the widest benefits to society through enabling the analysis of…See this and similar jobs on LinkedIn.
12218
Reposted by Jim Shaw
A. Murat Eren (Meren) @merenbey.bsky.social · 04/07/2026
New study by Alexander Henoch (@ahenoch.bsky.social), a PhD student in our group @hifmb.de and @awi.de, shows what it takes to bring gene synteny into microbial pangenomes, and what we learn about the variability landscape of genomes when we do that. See the pre-print here: doi.org/10.64898/202...
16624
Reposted by Jim Shaw
Andrew Carroll @acarroll.bsky.social · 01/07/2026
How good is MiniBWA, the successor to BWA? To test it, I ran MiniBWA on sequencing from 76 different species, comparing mapping speed, rate and accuracy with BWA MEM. In short, it's really good. If you map short reads, it's well worth your time. andrewcarroll.github.io/2026/06/30/t...
andrewcarroll.github.io
The Best of Both Worlds - Assessing MiniBWA
Recently, Heng Li released MiniBWA (GitHub) alongside a paper by Heng Li and Nils Homer describing the method (paper). MiniBWA builds on the approaches in Minimap2 (also by Heng Li), but falls back on...
19862
Reposted by Jim Shaw
Daan Speth @daanspeth.bsky.social · 26/06/2026
I'm happy to announce the release of GlobDB r232! This version contains 346,233 bacterial and archaeal genomes, based on 26 datasets. More info globdb.org 🦠🖥️🧬
globdb.org
home | GlobDB
24623
Reposted by Jim Shaw
Amy Zamora @amycrobes.bsky.social · 25/06/2026
We're excited to share our work on how prophages influence the evolution of resistance against DNA-damaging antibiotics! This has been a fun project with @sianowen.bsky.social, @baym.lol, @theshreyaspai.bsky.social, @kepatitis-c.bsky.social, & @fernpizza.bsky.social biorxiv.org/content/10.6... 1/
biorxiv.org
SOS-mediated prophage induction constrains resistance evolution to DNA-damaging antibiotics
Most naturally occurring bacteria are lysogens, encoding one or more temperate phages (prophages) integrated into their genome. As prophages are induced by the bacterial SOS response, DNA-damaging ant...
15923
Reposted by Jim Shaw
Morten Kam Dahl Dueholm @mkddueholm.bsky.social · 23/06/2026
Congratulations to my PhD student, Stefania Andrea Rosso Villanelo, on publishing her first first-author paper! 🎉In this study, she used antibiotics to facilitate the isolation of previously uncultured bacterial species from activated sludge. Check it out! 🦠🧫💊 journals.asm.org/doi/10.1128/...
journals.asm.org
Application of antibiotics for the selective isolation of previously uncultured species from activated sludge | Microbiology Spectrum
Biological wastewater treatment relies on diverse microbial communities to degrade pollutants and drive nutrient transformations. Understanding the physiology and metabolism of these microorganisms is essential for improving the efficiency and cost-effectiveness of treatment processes. Much of our current knowledge is derived from 16S rRNA gene amplicon sequencing and metagenomic analyses. However, validating these sequencing- and genome-based insights requires bacterial species as pure cultures, and only a limited number of taxa common in wastewater treatment plants are currently available in culture. Here, we present an isolation strategy that uses antibiotics as a selective pressure to reduce microbial complexity and alleviate competitive exclusion during cultivation, while full-length 16S rRNA gene amplicon sequencing is used to monitor enrichment and guide targeted isolation, thereby facilitating the recovery of process-relevant activated sludge bacteria, including potentially uncultured taxa. These isolates can serve as model organisms for experimental validation of genome-based predictions.
073
Reposted by Jim Shaw
Heng Li @lh3lh3.bsky.social · 16/06/2026
Minibwa is a hybrid of bwa-mem and minimap2 and the successor of bwa-mem for short-read mapping. ~4X/2.5X as fast as bwa-mem/bwa-mem2 for WGS reads at comparable accuracy. Native support of directional bisulfite-seq. Applicable to long reads. Preprint at arxiv.org/abs/2606.15357
1193109
Reposted by Jim Shaw
Johanna von Wachsmann @johannavw.bsky.social · 16/06/2026
🧬 New preprint! We clustered 5.6 million bacterial genomes into genomically cohesive units (GCUs) 500× faster than existing tools. (In just 14 hours, 16.5 GB RAM using 48 CPUs). 🦠🐙Meet gemsparcl 💎✨! www.biorxiv.org/content/10.6...
06224
Reposted by Jim Shaw
Johannes Köster @johanneskoester.bsky.social · 12/06/2026
#rustbio 4.0 has been released. It harmonizes the error handling, improves the API, makes gap-open/extend behavior in pairwise alignment more intuitive and in-line with the literature, improves GFF parsing, and allows incremental building of the rank-select datastructure. github.com/rust-bio/rus...
github.com
Release v4.0.0 · rust-bio/rust-bio
4.0.0 (2026-06-12) ⚠ BREAKING CHANGES Replace anyhow with typed thiserror errors (#674) Change Phase conversion methods to use TryFrom for better error handling (#625) for pairwise alignment, only...
0197
Reposted by Jim Shaw
Ryan Wick @rrwick.bsky.social · 11/06/2026
New blog post! I analyse the new hac@v6.0.0 basecalling model from @nanoporetech.com and discuss the conspicuous lack of a new sup model: rrwick.github.io/2026/06/11/d...
rrwick.github.io
Dorado v2.0.0: no more sup?
a blog for miscellaneous bioinformatics stuff
15526
Reposted by Jim Shaw
Tami Lieberman @contaminatedsci.bsky.social · 10/06/2026
Excited to have Alex Crits Christoph @acritschristoph.bsky.social join the lab today as a staff scientist! I had a hunch it was good time to find an experienced computational microbiologist to join the lab, but still feel very lucky to have him joining us. Looking forward to cool science!
5474
Reposted by Jim Shaw
Rasmus Kirkegaard @kirk3gaard.bsky.social · 08/06/2026
Is @nanoporetech.com hac v6 better than 5.2.0 sup? The short answer is no. But are a few errors in a genome worth 5 times more basecalling compute? github.com/Kirk3gaard/M...
0127
Reposted by Jim Shaw
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 06/06/2026
Folddisco is now published @natbiotech.nature.com. It’s a fast motif search for similar 3D DISCOntinuous residues like catalytic sites or zinc fingers across the entire protein universe. 📄 www.nature.com/articles/s41... 💾 folddisco.foldseek.com​​​​​​​​​​​​​​​​ 🌐 search.foldseek.com/folddisco
nature.com
Structural motif search across the protein universe with Folddisco - Nature Biotechnology
Folddisco enables protein structural motif search in million scale databases.
213455
Reposted by Jim Shaw
Yuya Kiguchi @ykiguchi.bsky.social · 01/06/2026
Excited to share our review relating to the giant extrachromosomal elements (ECEs) We cover: why long-read metagenomics enables their discovery, genomic comparisons across recently characterized examples, and limitations of current classification tools. www.cell.com/trends/genet...
cell.com
Giants within: a new class of microbial mobile elements
Prokaryotes harbor a diverse spectrum of extrachromosomal elements (ECEs), which are intracellular replicons maintained independently of the primary chromosome. Historically, the ECE research field ha...
0115
Jim Shaw @jimshaw.bsky.social · 01/06/2026
The method works with any set of primers. The datasets we benchmarked included a V1-V8 dataset (see www.biorxiv.org/content/10.6...), or newly-released 16S primers from ONT (see epi2me.nanoporetech.com/zymo_16s_202...). Maybe @mkddueholm.bsky.social has more general thoughts on FL 16S primers.
biorxiv.org
110
Jim Shaw @jimshaw.bsky.social · 30/05/2026
Savont is available on github and conda. Written in Rust as usual. We've integrated databases (Greengenes2, SILVA, EMU's db), multi-sample merging, QIIME2 compatible outputs, etc. We're very open to feedback, so please let us know of feature requests / bugs! github.com/bluenote-157...
github.com
GitHub - bluenote-1577/savont: Amplicon sequencing variants from 16s ONT R10.4 / HiFi long reads
Amplicon sequencing variants from 16s ONT R10.4 / HiFi long reads - bluenote-1577/savont
040
Jim Shaw @jimshaw.bsky.social · 30/05/2026
An advantage of ASVs ---> more confident species-level profiling. Read mapping can lead to overconfident species classifications: the "best hit" database species isn't necessarily correct. ASVs helps avoid this, since you remove seq. error from the equation. 6/6
The best read alignment (minimum divergence) against a species' 16S genes. Each dot = one species. Fecal 16S ONT sample.
130
Jim Shaw @jimshaw.bsky.social · 30/05/2026
On real 16S ONT amplicon data, savont gets a lot more diversity and ASVs compared to existing methods. But these aren't just false positives. The dataset (from @mkddueholm.bsky.social and team) had paired PacBio HiFi data as a orthogonal reference: ~98% of the ONT ASVs mapped perfectly. 5/6
141
Jim Shaw @jimshaw.bsky.social · 30/05/2026
For long 16S nanopore amplicons (R10.4, sup-basecalled), savont requires 5-16x less depth for capturing ASVs. This improvement is more stark for longer amplicons (e.g. rRNA operon). For HiFi, savont + existing ASV methods are comparable, although savont misses a few intragenomic 16S copies. 4/6
162
Jim Shaw @jimshaw.bsky.social · 30/05/2026
Savont instead uses a high-resolution clustering approach: cluster similar reads --> create an error-free consensus. The problem is resolving clusters at the single nucleotide level to get true ASVs. To do this, we adopted the "SNPmer" technique: cluster reads by their polymorphic k-mers. 3/6
111
Jim Shaw @jimshaw.bsky.social · 30/05/2026
Why savont? Short-read ASV-generation techniques rely on finding error-free reads and "denosing". For noisy long reads, this doesn't work as well. If a 16S amplicon has length 1.4 kbp & a read is 99.5% accurate => it has (0.995)^1400 ~ exp(-7) = 0.09% chance of being perfect. Not so great. 2/6
110
Jim Shaw @jimshaw.bsky.social · 30/05/2026
Our method, savont, for generating amplicon sequence variants (ASVs) for long-read amplicons is now on bioRxiv. Work with @lh3lh3.bsky.social and help from @mkddueholm.bsky.social and team (Marie Riisgaard-Jensen, @kirk3gaard.bsky.social, Kasper Skytte Andersen) github.com/bluenote-157... 1/6
24219
Reposted by Jim Shaw
Heng Li @lh3lh3.bsky.social · 30/05/2026
Jeremy Wang developed rammap, a minimap2 rewrite in Rust. It achieves comparable or better performance than minimap2 and produces identical output to minimap2. During rewrite, Jeremy found two long-existing bugs in minimap2 which are fixed in v2.31. www.biorxiv.org/content/10.6...
biorxiv.org
310944
Reposted by Jim Shaw
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 30/05/2026
Sensitive long-read amplicon sequence variant recovery with savont www.biorxiv.org/content/10.64898/20…
0135
Jim Shaw @jimshaw.bsky.social · 29/05/2026
Thanks steven :)
010
Reposted by Jim Shaw
Christine He @christinehe.bsky.social · 29/05/2026
Come to ONT's industry event at ASM to hear about two exciting new tools from @jimshaw.bsky.social! SNPmers (polymorphic kmers) are leveraged for impressive performance in both myloasm (metagenomic assembly) and savont (16S ASVs). Registration link 👇
0116