Sign in

Jim Shaw

@jimshaw.bsky.social
1.2K followers 575 following 125 posts

Postdoc at Dana-Farber and Harvard Med with Heng Li (@lh3lh3.bsky.social). Prev: UBC / UofT. I like thinking about biological sequence analysis and its applications to metagenomics / microbial genomics. jim-shaw-bluenote.github.io

PostsRepliesMedia
Reposted by Jim Shaw
Milot Mirdita @milot.bsky.social · 16/09/2026
ColabFold 1.6.3 is out! 2.5x faster, pip-installable, ipSAE+pDockQ2 scores. Thanks Choonghwan Lee, Marielle Russo, Gyuri Kim 🐍pip install colabfold[alphafold] CF2 Sneak Peak with AF3/Boltz/Protenix/ESMFold2… 🐍pip install "colabfold[alphafold3]@git+https://github.com/sokrypton/ColabFold@af3-preview"
29735
Jim Shaw @jimshaw.bsky.social · 15/09/2026
The sylph metagenome profiler is v1.0.0! sylph-docs.github.io A new DB format + approach --> huge performance gains: GTDB-R232 (200k species) now takes < 5 GB of RAM and ~30s (2GB fq.gz). Huge thanks to @benjwoodcroft.bsky.social and his ongoing performance efforts (github.com/wwood/weebill)
sylph-docs.github.io
Documentation for sylph - ultrafast, precise metagenomic profiling
26330
Reposted by Jim Shaw
Tami Lieberman @contaminatedsci.bsky.social · 08/09/2026
Two announcements regarding AccuSNV, which calls high precision SNVs across microbial genomes: Version 1.1 is now incredibly easy to install, run, and perform downstream analyses with (see image of typical output!). The manuscript was also published this summer in Genome Research! Links next...
15320
Reposted by Jim Shaw
Oxford Nanopore @nanoporetech.com · 04/09/2026
Still relying on legacy approaches for microbial analysis? Join this webinar to discover the latest tools for resolving strain-level diversity, recovering complete genomes, and unlocking deeper insights from complex microbial communities. Register here: bit.ly/3UnBy0r
142
Reposted by Jim Shaw
Gerry Tonkin-Hill @gerrythill.bsky.social · 02/09/2026
Glad to share StrainSpy, a new method developed by Sudaraka and co which identifies strain-level associations in metagenomics using ANI estimates from tools like Sylph. www.biorxiv.org/content/10.6...
biorxiv.org
Accurate detection of metagenomic strain-level associations using average nucleotide identity with StrainSpy
Genetic variation among microbial strains of the same species can profoundly influence their phenotypes, ecological functions, and impacts on human health. Traditionally, the relative abundance of a s...
12112
Jim Shaw @jimshaw.bsky.social · 02/09/2026
Cool preprint about strain-level metagenome association studies using k-mer containment. Mallawaarachchi et al. from @gerrythill.bsky.social's group. I tried tackling this in our original paper for sylph (sylph-docs.github.io), but this seems to be much more sophisticated. Excited to read!
2238
Reposted by Jim Shaw
bioRxiv Microbiology @biorxiv-microbiol.bsky.social · 01/09/2026
Accurate detection of metagenomic strain-level associations using average nucleotide identity with StrainSpy www.biorxiv.org/content/10.64898/20…
0102
Reposted by Jim Shaw
Kristoffer Sahlin @ksahlin.bsky.social · 01/09/2026
strobealign v0.18.0 is out, and it's now written in Rust. Handwritten by Marcel, started 2y ago. No LLM involved (yet*). Contributions by Nicolas and Ivan. 1/3
github.com
Release Version 0.18.0 · ksahlin/strobealign
Changes Strobealign has been ported to Rust. The Rust version is at least as fast and accurate as the C++ version. There are currently some regressions due to the port: The Rust version cannot sat...
2195
Reposted by Jim Shaw
Falk Hildebrand @bioinf.bsky.social · 07/08/2026
After a very long time, I'm very proud that we have now Joachim "magnus opum" on www.biorxiv.org/content/10.6... Protal is a ultra fast metagenomic species profiler. More precise than other profiled software, faster than most (not sylph though), but generates strain resolution instead.
02812
Reposted by Jim Shaw
Niranjan Nagarajan @niranjantw.bsky.social · 07/08/2026
Thank you for featuring our work National University Health System! Its great to have this article describe in simple terms what we have been doing for nearly a decade now in terms of microbial genomic surveillance and metagenomics 🥲 nuhsplus.edu.sg/article/dna-...
nuhsplus.edu.sg
DNA detectives: Hunting down hidden hospital bacteria | NUHS+
NUHS researchers use a genetic fingerprint to expose invisible microbial reservoirs, aiding outbreak prevention across local hospitals.
031
Reposted by Jim Shaw
George Bouras @gbouras13.bsky.social · 07/08/2026
If you are looking for web-based phage annotation, phage-annotation.org is live. Upload a genome to run Pharokka, Phold and Phynteny sequentially (see our recent protocols paper doi.org/10.1002/cpz1...). If you have feedback, please reach out Thanks to @ardc.edu.au for making this possible!
phage-annotation.org
Phage Annotation Server -- phage genome annotation
Free automated phage genome annotation -- pharokka, phold and phynteny in one pipeline.
46129
Reposted by Jim Shaw
Ragnar {Groot Koerkamp} @curiouscoding.nl · 05/08/2026
Travis has been cooking something up since IGGSy ;) A fast algorithm for SMEM finding in haplotype panels that uses space proportional to the number of 1s in the (sparse) matrix, inspired by the jump index [13].
[13] Ragnar Groot Koerkamp. Personal communication, 2026.
031
Reposted by Jim Shaw
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 01/08/2026
Riboseek is a fast RNA/DNA search. More sensitive than nhmmer at 250x speed. Structure-aware realignment produces MSAs approaching rMSA quality. Plus 1.7M precomputed RNA MSAs, and an API to search your own 📄 www.biorxiv.org/content/10.6... 💾 github.com/steineggerla... 🌐 search.foldseek.com/riboseek
217874
Reposted by Jim Shaw
Tami Lieberman @contaminatedsci.bsky.social · 31/07/2026
The manuscript describing SimPhyNI is now published! www.microbiologyresearch.org/content/jour... A few updates since last post, including some method tweaks that improve performance and speed. Also even easier to run. Try it for your microbial GWAS or epistasis needs!
microbiologyresearch.org
High-precision binary trait association on phylogenetic trees
Traditional methods for identifying associations between genomic features and traits, or between pairs of genomic traits, struggle when applied to bacterial genomes. While several microbial genome-wid...
14413
Reposted by Jim Shaw
Ryan Wick @rrwick.bsky.social · 24/07/2026
New blog post! I reran the Autocycler paper benchmarks on some new tools/versions/pipelines: rrwick.github.io/2026/07/24/b... (1/3)
rrwick.github.io
Benchmark update: Ilesta, Autocycler-fast and new versions
a blog for miscellaneous bioinformatics stuff
23416
Reposted by Jim Shaw
Steven Robbins @stevenjrobbins.bsky.social · 22/07/2026
It's out! Excited to present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), a comprehensive DB of 1000s of high-quality prokaryote, virus, plasmid, and chromosome-level eukaryote MAGs using Nanopore long reads. Subthreads incoming. Please share widely. 🙂 www.nature.com/articles/s41...
nature.com
The planktonic microbiome of the Great Barrier Reef - Nature
The Great Barrier Reef Microbial Genomes Database compiles prokaryotic, viral and eukaryotic genomes from seawater collected from the Great Barrier Reef, providing a rich resource for the study of mar...
88342
Reposted by Jim Shaw
Zamin Iqbal @zaminiqbal.bsky.social · 20/07/2026
Significant update to the AllTheBacteria paper, including discovering new antimicrobial peptides and testing in vitro and vivo. This has grown into a fantastic collaboration!
17231
Reposted by Jim Shaw
Vikram Shivakumar @vikramshivakumar.bsky.social · 15/07/2026
1/ Excited to share the newest tool in the pangenome MUMiverse: Shredtools! Shredtools enables a user to navigate the pangenome coordinate system with multi-MUMs. More in the thread🧵 Code: github.com/vikshiv/shredtools Interactive tool for querying HPRC assemblies: vikshiv.github.io/shredtools
biorxiv.org
Navigating the pangenome coordinate system with Shredtools
Existing notions of pangenome coordinates rely on hard-to-compute multiple sequence alignments. On the other hand, pangenome-wide exact unique matches (multi-MUMs) can be computed efficiently, and rep...
13822
Reposted by Jim Shaw
Heng Li @lh3lh3.bsky.social · 13/07/2026
minibwa-0.4 released with minor improvement and a few fixes to typos. Also a new blog post on "minibwa is the new bwa-mem": lh3.github.io/2026/07/04/m...
lh3.github.io
Minibwa is the new bwa-mem
0199
Reposted by Jim Shaw
Sina Majidian @sinamajidian.bsky.social · 09/07/2026
JOB ALERT! I'm hiring two postdocs in Computational Genomics to join our lab in beautiful Gothenburg, Sweden. Please share and repost! CGRLab.github.io/research/
01417
Reposted by Jim Shaw
Sergey Nurk @sergeynurk.bsky.social · 07/07/2026
Senior Bioinformatics Software Developer opening in my team in ONT's Applications department! Hybrid with a few days a week in our Oxford headquarters. We can sponsor visas! www.linkedin.com/jobs/view/44...
linkedin.com
Oxford Nanopore Technologies hiring Senior Bioinformatics Software Engineer, Applications in Oxfordshire, England, United Kingdom | LinkedIn
Posted 3:04:10 PM. Job DescriptionOur goal is to bring the widest benefits to society through enabling the analysis of…See this and similar jobs on LinkedIn.
12218
Reposted by Jim Shaw
A. Murat Eren (Meren) @merenbey.bsky.social · 04/07/2026
New study by Alexander Henoch (@ahenoch.bsky.social), a PhD student in our group @hifmb.de and @awi.de, shows what it takes to bring gene synteny into microbial pangenomes, and what we learn about the variability landscape of genomes when we do that. See the pre-print here: doi.org/10.64898/202...
16623
Reposted by Jim Shaw
Andrew Carroll @acarroll.bsky.social · 01/07/2026
How good is MiniBWA, the successor to BWA? To test it, I ran MiniBWA on sequencing from 76 different species, comparing mapping speed, rate and accuracy with BWA MEM. In short, it's really good. If you map short reads, it's well worth your time. andrewcarroll.github.io/2026/06/30/t...
andrewcarroll.github.io
The Best of Both Worlds - Assessing MiniBWA
Recently, Heng Li released MiniBWA (GitHub) alongside a paper by Heng Li and Nils Homer describing the method (paper). MiniBWA builds on the approaches in Minimap2 (also by Heng Li), but falls back on...
19862
Reposted by Jim Shaw
Daan Speth @daanspeth.bsky.social · 26/06/2026
I'm happy to announce the release of GlobDB r232! This version contains 346,233 bacterial and archaeal genomes, based on 26 datasets. More info globdb.org 🦠🖥️🧬
globdb.org
home | GlobDB
24623
Reposted by Jim Shaw
Amy Zamora @amycrobes.bsky.social · 25/06/2026
We're excited to share our work on how prophages influence the evolution of resistance against DNA-damaging antibiotics! This has been a fun project with @sianowen.bsky.social, @baym.lol, @theshreyaspai.bsky.social, @kepatitis-c.bsky.social, & @fernpizza.bsky.social biorxiv.org/content/10.6... 1/
biorxiv.org
SOS-mediated prophage induction constrains resistance evolution to DNA-damaging antibiotics
Most naturally occurring bacteria are lysogens, encoding one or more temperate phages (prophages) integrated into their genome. As prophages are induced by the bacterial SOS response, DNA-damaging ant...
15923
Reposted by Jim Shaw
Morten Kam Dahl Dueholm @mkddueholm.bsky.social · 23/06/2026
Congratulations to my PhD student, Stefania Andrea Rosso Villanelo, on publishing her first first-author paper! 🎉In this study, she used antibiotics to facilitate the isolation of previously uncultured bacterial species from activated sludge. Check it out! 🦠🧫💊 journals.asm.org/doi/10.1128/...
journals.asm.org
Application of antibiotics for the selective isolation of previously uncultured species from activated sludge | Microbiology Spectrum
Biological wastewater treatment relies on diverse microbial communities to degrade pollutants and drive nutrient transformations. Understanding the physiology and metabolism of these microorganisms is essential for improving the efficiency and cost-effectiveness of treatment processes. Much of our current knowledge is derived from 16S rRNA gene amplicon sequencing and metagenomic analyses. However, validating these sequencing- and genome-based insights requires bacterial species as pure cultures, and only a limited number of taxa common in wastewater treatment plants are currently available in culture. Here, we present an isolation strategy that uses antibiotics as a selective pressure to reduce microbial complexity and alleviate competitive exclusion during cultivation, while full-length 16S rRNA gene amplicon sequencing is used to monitor enrichment and guide targeted isolation, thereby facilitating the recovery of process-relevant activated sludge bacteria, including potentially uncultured taxa. These isolates can serve as model organisms for experimental validation of genome-based predictions.
073
Reposted by Jim Shaw
Heng Li @lh3lh3.bsky.social · 16/06/2026
Minibwa is a hybrid of bwa-mem and minimap2 and the successor of bwa-mem for short-read mapping. ~4X/2.5X as fast as bwa-mem/bwa-mem2 for WGS reads at comparable accuracy. Native support of directional bisulfite-seq. Applicable to long reads. Preprint at arxiv.org/abs/2606.15357
1193109
Reposted by Jim Shaw
Johanna von Wachsmann @johannavw.bsky.social · 16/06/2026
🧬 New preprint! We clustered 5.6 million bacterial genomes into genomically cohesive units (GCUs) 500× faster than existing tools. (In just 14 hours, 16.5 GB RAM using 48 CPUs). 🦠🐙Meet gemsparcl 💎✨! www.biorxiv.org/content/10.6...
06124
Reposted by Jim Shaw
Johannes Köster @johanneskoester.bsky.social · 12/06/2026
#rustbio 4.0 has been released. It harmonizes the error handling, improves the API, makes gap-open/extend behavior in pairwise alignment more intuitive and in-line with the literature, improves GFF parsing, and allows incremental building of the rank-select datastructure. github.com/rust-bio/rus...
github.com
Release v4.0.0 · rust-bio/rust-bio
4.0.0 (2026-06-12) ⚠ BREAKING CHANGES Replace anyhow with typed thiserror errors (#674) Change Phase conversion methods to use TryFrom for better error handling (#625) for pairwise alignment, only...
0197
Reposted by Jim Shaw
Ryan Wick @rrwick.bsky.social · 11/06/2026
New blog post! I analyse the new hac@v6.0.0 basecalling model from @nanoporetech.com and discuss the conspicuous lack of a new sup model: rrwick.github.io/2026/06/11/d...
rrwick.github.io
Dorado v2.0.0: no more sup?
a blog for miscellaneous bioinformatics stuff
15526
Reposted by Jim Shaw
Tami Lieberman @contaminatedsci.bsky.social · 10/06/2026
Excited to have Alex Crits Christoph @acritschristoph.bsky.social join the lab today as a staff scientist! I had a hunch it was good time to find an experienced computational microbiologist to join the lab, but still feel very lucky to have him joining us. Looking forward to cool science!
5474
Reposted by Jim Shaw
Rasmus Kirkegaard @kirk3gaard.bsky.social · 08/06/2026
Is @nanoporetech.com hac v6 better than 5.2.0 sup? The short answer is no. But are a few errors in a genome worth 5 times more basecalling compute? github.com/Kirk3gaard/M...
0127
Reposted by Jim Shaw
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 06/06/2026
Folddisco is now published @natbiotech.nature.com. It’s a fast motif search for similar 3D DISCOntinuous residues like catalytic sites or zinc fingers across the entire protein universe. 📄 www.nature.com/articles/s41... 💾 folddisco.foldseek.com​​​​​​​​​​​​​​​​ 🌐 search.foldseek.com/folddisco
nature.com
Structural motif search across the protein universe with Folddisco - Nature Biotechnology
Folddisco enables protein structural motif search in million scale databases.
213455
Reposted by Jim Shaw
Yuya Kiguchi @ykiguchi.bsky.social · 01/06/2026
Excited to share our review relating to the giant extrachromosomal elements (ECEs) We cover: why long-read metagenomics enables their discovery, genomic comparisons across recently characterized examples, and limitations of current classification tools. www.cell.com/trends/genet...
cell.com
Giants within: a new class of microbial mobile elements
Prokaryotes harbor a diverse spectrum of extrachromosomal elements (ECEs), which are intracellular replicons maintained independently of the primary chromosome. Historically, the ECE research field ha...
0115
Jim Shaw @jimshaw.bsky.social · 30/05/2026
Our method, savont, for generating amplicon sequence variants (ASVs) for long-read amplicons is now on bioRxiv. Work with @lh3lh3.bsky.social and help from @mkddueholm.bsky.social and team (Marie Riisgaard-Jensen, @kirk3gaard.bsky.social, Kasper Skytte Andersen) github.com/bluenote-157... 1/6
24219
Reposted by Jim Shaw
Heng Li @lh3lh3.bsky.social · 30/05/2026
Jeremy Wang developed rammap, a minimap2 rewrite in Rust. It achieves comparable or better performance than minimap2 and produces identical output to minimap2. During rewrite, Jeremy found two long-existing bugs in minimap2 which are fixed in v2.31. www.biorxiv.org/content/10.6...
biorxiv.org
310944
Reposted by Jim Shaw
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 30/05/2026
Sensitive long-read amplicon sequence variant recovery with savont www.biorxiv.org/content/10.64898/20…
0135
Reposted by Jim Shaw
Christine He @christinehe.bsky.social · 29/05/2026
Come to ONT's industry event at ASM to hear about two exciting new tools from @jimshaw.bsky.social! SNPmers (polymorphic kmers) are leveraged for impressive performance in both myloasm (metagenomic assembly) and savont (16S ASVs). Registration link 👇
0116
Jim Shaw @jimshaw.bsky.social · 29/05/2026
Excited to speak at ASM Microbe 2026 in the Oxford Nanopore session about new tools for long-read metagenomics + 16S sequencing. If you're attending ASM Microbe June 4-7 in D.C. and want to chat, let me know!
1347
Reposted by Jim Shaw
Samuel Lampa @smllmp.bsky.social · 23/05/2026
TaxProfiler, the @nf-co.re #metagenomics pipeline had a new major release, adding multiple cool new classifiers (metacache, sylph and melon)! Great work by Sofia, @lilianderssonli.bsky.social , @jfy133.genomic.social.ap.brid.gy and others! 🙌 github.com/nf-core/taxp...
github.com
Release v2.0.0 - Crazy Corgi · nf-core/taxprofiler
Added #682 Added metacache classifier and improved nf-tests (added by @sofstam) #559 Profiling of long reads with motus (added by @LilyAnderssonLee and @sofstam ) #591 Add options to enable the ab...
193
Reposted by Jim Shaw
Rauf Salamzade @raufs.bsky.social · 23/05/2026
New versions of skDER (github.com/raufs/skDER) and zol (github.com/Kalan-Lab/zol) are now on Bioconda. Both have automated downloading of genomes based on species / genus name according to classifications in GTDB R232. The latest zol also incorporates FAMSA2 which provides considerable speed boosts!
0123
Reposted by Jim Shaw
Piotr Rozwalak @prozwalak.bsky.social · 11/05/2026
Mushuvirus is the most widespread phage genus in the human gut. 🌍 Together with other family members, these viruses occur in 89% of humans worldwide, including the Iceman Ötzi! How is it possible that they were hidden in previous metagenomic analyses? [1/7] Read: doi.org/10.64898/202...
23114
Reposted by Jim Shaw
Paul Medvedev @pashadag.bsky.social · 05/05/2026
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
doi.org
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
13019
Reposted by Jim Shaw
Christine He @christinehe.bsky.social · 02/05/2026
Released today: ~1.5 Tbp of ONT metagenomic sequencing from compost. Assembly, MAGs, and downstream analysis are available to play with epi2me.nanoporetech.com/compost_mgx_...
epi2me.nanoporetech.com
Metagenomic Assembly Sheds Light on Microbial Diversity in Compost
Overview We are pleased to release a metagenomic dataset from deep sequencing of a mature compost…
01310
Reposted by Jim Shaw
Yan Shao @yanshao.bsky.social · 28/04/2026
Great to see TRACS out in press - a new metagenomic strain transmission tool, led by @gerrythill.bsky.social. With improved precision and sensitivity, and unique support for custom reference and minor-strain detection, TRACS enables more rigorous and comprehensive strain-level metagenomics analysis.
094
Reposted by Jim Shaw
Mile Sikic @msikic.bsky.social · 28/04/2026
HERRO has been published in @nature.com nature.com/articles/s41.... This achievement is a result of the great work by Dominik Stanojevic, with contributions from Dehui Lin, @sergeynurk.bsky.social, and Paola Florez de Sessions Welcome to the era of high-quality genome assemblies supported by AI.
nature.com
Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads - Nature
Nature - Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads
1279
Reposted by Jim Shaw
Niranjan Nagarajan @niranjantw.bsky.social · 25/04/2026
Finally out in @natmethods.nature.com 🎉 Our work with Wan Yue's lab shows that signal alignments from direct RNA @nanopore sequencing can reveal structural heterogeneity in RNA. Huge implications for RNA therapeutics! Read all about it here: www.nature.com/articles/s41...
nature.com
Direct RNA sequencing and signal alignment reveal RNA structure ensembles in a eukaryotic cell - Nature Methods
sm-PORE-cupine combines SHAPE-based chemical probing with nanopore-based direct RNA sequencing to identify RNA structural ensembles in the SARS-CoV-2 genome and the Candida albicans transcriptome.
082
Reposted by Jim Shaw
Katharina Hoff @katharinahoff.bsky.social · 24/04/2026
Tiberius 2.0.0 is out 🎉 Now supports 7 eukaryotic clades, covering ~92% of NCBI assemblies. Modular rewrite + ~30% faster runtime. Benchmarks included, more soon. Thanks to Lars Gabriel, Richard Krieg & Felix Becker 🙌 github.com/Gaius-August... #bioinformatics #genomics #genomeannotation
0135
Reposted by Jim Shaw
Scott V. Edwards @scottvedwards.bsky.social · 22/04/2026
Ver excited to share my just-published Darwin Review with @lh3lh3.bsky.social on population-scale long-read sequencing! royalsocietypublishing.org/rspb/article...
royalsocietypublishing.org
Population-scale long-read DNA sequencing: peering under the hood of the new evolutionary genomics
Abstract. Population-scale long-read DNA sequencing (PLRS) is rapidly reshaping our understanding of genomic variation in humans and non-model species. In
02919
Reposted by Jim Shaw
Tami Lieberman @contaminatedsci.bsky.social · 22/04/2026
Interested in a *staff computational scientist* position? We are looking for an experienced computational biologist (ideally with microbiology experience) to support published packages while also driving new research. Pay range $73k-$111k w excellent benefits. careers.peopleclick.com/careerscp/cl...
careers.peopleclick.com
Computational Biology Program Scientist
MIT - Computational Biology Program Scientist - Cambridge MA 02139
14574