Sign in

Gu Zhenhao

@guzhenhao.bsky.social
35 followers 71 following 0 posts

PhD student at NUS Computing / Genome Institute of Singapore. Alto clef enjoyer. My work playlist: www.youtube.com/playlist?list=PLPgS…

PostsRepliesMedia
Reposted by Gu Zhenhao
Steven Salzberg @stevensalzberg.bsky.social · 24/09/2026
Check our new @biorxiv-genomic.bsky.social preprint on a method for masking low-complexity regions in genome databases, led by Yuchin (Peter) Ge in my lab. Also includes a new release of a very large, pre-masked microbial database: www.biorxiv.org/content/10.6...
biorxiv.org
Improving Metagenomics Classification with Kmask: Entropy-Based Masking of Low-Complexity Regions
Abstract Accurate taxonomic classification in metagenomics is often compromised by low-complexity sequences, which lead to chance matches that in turn cause sequences to be misclassified. Here we pres...
063
Reposted by Gu Zhenhao
Jim Shaw @jimshaw.bsky.social · 15/09/2026
The sylph metagenome profiler is v1.0.0! sylph-docs.github.io A new DB format + approach --> huge performance gains: GTDB-R232 (200k species) now takes < 5 GB of RAM and ~30s (2GB fq.gz). Huge thanks to @benjwoodcroft.bsky.social and his ongoing performance efforts (github.com/wwood/weebill)
sylph-docs.github.io
Documentation for sylph - ultrafast, precise metagenomic profiling
26330
Reposted by Gu Zhenhao
Gerry Tonkin-Hill @gerrythill.bsky.social · 02/09/2026
Glad to share StrainSpy, a new method developed by Sudaraka and co which identifies strain-level associations in metagenomics using ANI estimates from tools like Sylph. www.biorxiv.org/content/10.6...
biorxiv.org
Accurate detection of metagenomic strain-level associations using average nucleotide identity with StrainSpy
Genetic variation among microbial strains of the same species can profoundly influence their phenotypes, ecological functions, and impacts on human health. Traditionally, the relative abundance of a s...
12112
Reposted by Gu Zhenhao
Rob Patro @robp.bsky.social · 23/08/2026
The first public release of seqproc is out! 🎉 v0.1.1, powered by ANTISEQUENCE v0.1.0. seqproc let's you describe a sequencing protocol in EFGDL, compile it into a fast Rust pipeline for matching, filtering, correction, and FASTQ transformation. A powerful read pre-processing engine! 1/6
1197
Reposted by Gu Zhenhao
Niranjan Nagarajan @niranjantw.bsky.social · 24/08/2026
Are there robust associations between the gut microbiome and obesity? We explored this extensively with population-scale metagenomic data for >800 South-East Asians in the HELIOS cohort www.medrxiv.org/content/10.6...
medrxiv.org
Population-scale analysis reveals limited and non-generalizable associations between the gut microbiome and obesity in Asian adults
Background The gut microbiome has been widely studied in the context of obesity, and yet the reported associations vary widely across populations and analytical approaches. In Asian populations where ...
132
Reposted by Gu Zhenhao
Rob Edwards @linsalrob.bsky.social · 17/08/2026
We're having fun developing post-cloud bioinformatics tools: reimagining some of the bioinformatics canon as local, browser-native applications so your data never has to leave your laptop. Here's the first set: edwards.flinders.edu.au/viz/ Let us know if you want other tools local to your laptop
edwards.flinders.edu.au
-viz
Welcome to the -viz suite of tools: classic bioinformatics tools, reimagined for the browser. agviz · bamviz · genbankviz · phispyviz Each of our tools is a classic bioinformatics tool, but reimagi…
14814
Reposted by Gu Zhenhao
Rob Patro @robp.bsky.social · 15/08/2026
Cuttlefish 3 is on bioconda! 🦑 A parallel, external-memory algorithm for building colored compacted de Bruijn graphs at collection scale — a RECOMB 2026 paper, and as of today a production release: v3.0.0. A thread on the algorithm, the numbers, and why the released tool is a Rust rewrite. 🧵 1/10
34412
Reposted by Gu Zhenhao
Rob Patro @robp.bsky.social · 10/08/2026
2.3 billion read pairs of 10x Flex v2 (281 GB of gzipped FASTQ) mapped in under 2 minutes on one machine (-t 64) with piscem-rs. That's ~20M read pairs/sec at ~87% mapped. One interesting part is the mapper, but what I want to talk about here is who gets the threads. 1/8
1247
Reposted by Gu Zhenhao
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 10/08/2026
PanGBank: a large-scale resource of precomputed microbial pangenomes built with PPanGGOLiN www.biorxiv.org/content/10.64898/20…
036
Reposted by Gu Zhenhao
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 01/08/2026
Riboseek is a fast RNA/DNA search. More sensitive than nhmmer at 250x speed. Structure-aware realignment produces MSAs approaching rMSA quality. Plus 1.7M precomputed RNA MSAs, and an API to search your own 📄 www.biorxiv.org/content/10.6... 💾 github.com/steineggerla... 🌐 search.foldseek.com/riboseek
217975
Reposted by Gu Zhenhao
Rob Patro @robp.bsky.social · 29/07/2026
New preprint🚨(*long* time coming)! We introduce seqproc: an efficient, flexible, and concise tool for describing and transforming sequencing-read geometry. The aim: make complex protocols easy to specify & transform without giving up speed or accuracy. www.biorxiv.org/content/10.6... 🧵
Crop of the preprint's first page. The title is "seqproc: An
efficient, flexible, and concise tool for sequence geometry description and
transformation." Authors are Noah Cape, Elan Fisher, Daniel Liu, and Rob Patro.
The abstract introduces seqproc, EFGDL, and the antisequence execution engine.
1369
Reposted by Gu Zhenhao
Zamin Iqbal @zaminiqbal.bsky.social · 20/07/2026
Significant update to the AllTheBacteria paper, including discovering new antimicrobial peptides and testing in vitro and vivo. This has grown into a fantastic collaboration!
17231
Reposted by Gu Zhenhao
danydoerr.bsky.social @danydoerr.bsky.social · 16/07/2026
New Nature Reviews Genetics paper out! How are graph-based pangenomes removing reference bias and opening up new possibilities in GWAS and rare- and common disease genetics? 🔗 nature.com/articles/s41576-026-00987-7 #pangenome #geneticdiversity #referencebias
nature.com
Building and applying pangenome references to capture genetic diversity - Nature Reviews Genetics
Pangenomes are genome references that integrate sequences from multiple individuals into graph-based or multi-haplotype representations, capturing genetic variation beyond a single linear reference. H...
0104
Reposted by Gu Zhenhao
Vikram Shivakumar @vikramshivakumar.bsky.social · 15/07/2026
1/ Excited to share the newest tool in the pangenome MUMiverse: Shredtools! Shredtools enables a user to navigate the pangenome coordinate system with multi-MUMs. More in the thread🧵 Code: github.com/vikshiv/shredtools Interactive tool for querying HPRC assemblies: vikshiv.github.io/shredtools
biorxiv.org
Navigating the pangenome coordinate system with Shredtools
Existing notions of pangenome coordinates rely on hard-to-compute multiple sequence alignments. On the other hand, pangenome-wide exact unique matches (multi-MUMs) can be computed efficiently, and rep...
13822
Reposted by Gu Zhenhao
haonanwu.bsky.social @haonanwu.bsky.social · 11/07/2026
Our #ISMB2026 paper is now online! I’m excited to present it at HiTSeq on July 13. Many thanks to my advisor @pashadag.bsky.social. I’m also grateful to @iscb.bsky.social for awarding me the Conference Fellowship. If you’ll be at ISMB, feel free to stop by my talk and say hi. I’d love to connect!
academic.oup.com
The gift of novelty: repeat-robust k-mer-based estimators of mutation rates
AbstractMotivation. Estimating mutation rates between evolutionarily related sequences is a central problem in molecular evolution. Due to the rapid expans
1134
Reposted by Gu Zhenhao
Bede Constantinides @bede.im · 13/07/2026
Ever wanted to quickly check host content of DNA sequences? bede.im/sapiometer
12010
Reposted by Gu Zhenhao
JS Gounot @jsgounot.bsky.social · 13/07/2026
It's my pleasure to introduce for the first time r-news.net, a project I had in my head for quite a long time (since my PhD!). With AI improving literature review quite substantially, we still miss a critical step: What does the community as a whole like the most?
r-news.net
Top Stories — RNews
211
Reposted by Gu Zhenhao
Shuai Wang @wshuai.bsky.social · 10/07/2026
So proud to see this study out — a real honor to be involved in work from my PhD lab!
0114
Reposted by Gu Zhenhao
Ben Langmead @benlangmead.bsky.social · 22/06/2026
Movi 2 has appeared (as an advance article) in Bioinformatics 🧬 Faster, leaner pangenome queries — half the memory of Movi 1, ~30% faster. Paper: academic.oup.com/bioinformati... Code: github.com/mohsenzakeri/Movi (1/6)
academic.oup.com
Validate User
34317
Reposted by Gu Zhenhao
Heng Li @lh3lh3.bsky.social · 16/06/2026
Minibwa is a hybrid of bwa-mem and minimap2 and the successor of bwa-mem for short-read mapping. ~4X/2.5X as fast as bwa-mem/bwa-mem2 for WGS reads at comparable accuracy. Native support of directional bisulfite-seq. Applicable to long reads. Preprint at arxiv.org/abs/2606.15357
1193109
Reposted by Gu Zhenhao
Igor Martayan @imartayan.bsky.social · 15/06/2026
Just submitted my PhD thesis on algorithms for fast, large-scale k-mer-based sequence analysis. It's now available to read at phd.martayan.org Take a look and feel free to share! #Bioinformatics #PhDone
phd.martayan.org
Algorithm design and implementation for the scale of sequencing data
26822
Reposted by Gu Zhenhao
Jonathan Göke @jonathangoeke.bsky.social · 09/06/2026
Nanopore sequencing provides not just long reads, but the the raw signal data can also be used to identify RNA and DNA modifications. This repository (and the associated review) lists some of the great tools that have been developed www.cell.com/trends/genet...
cell.com
Beyond sequencing: machine learning algorithms extract biology hidden in Nanopore signal data
Nanopore sequencing provides signal data corresponding to the nucleotide motifs sequenced. Through machine learning-based methods, these signals are translated into long-read sequences that overcome t...
096
Reposted by Gu Zhenhao
hbkgenomics.bsky.social @hbkgenomics.bsky.social · 06/06/2026
Does your designed active site already exist in nature? Is an uncharacterized protein hiding a catalytic site or a pocket? Folddisco answers both, searching millions of structures for a 3D motif in seconds. @natbiotech.nature.com 🧬 📄 www.nature.com/articles/s41... 🧵1/7👇
nature.com
Structural motif search across the protein universe with Folddisco - Nature Biotechnology
Folddisco enables protein structural motif search in million scale databases.
14715
Reposted by Gu Zhenhao
Roland Faure @rfaure.bsky.social · 04/06/2026
Glad to announce that our new long-read metagenomic SNP caller, SNooPy, is published in NAR! Existing long-read SNP callers are not designed for metagenomic data, check out our new solution 👉https://academic.oup.com/nar/article/54/10/gkag556/8700491 @narjournal.bsky.social
1177
Reposted by Gu Zhenhao
Jim Shaw @jimshaw.bsky.social · 30/05/2026
Our method, savont, for generating amplicon sequence variants (ASVs) for long-read amplicons is now on bioRxiv. Work with @lh3lh3.bsky.social and help from @mkddueholm.bsky.social and team (Marie Riisgaard-Jensen, @kirk3gaard.bsky.social, Kasper Skytte Andersen) github.com/bluenote-157... 1/6
24219
Reposted by Gu Zhenhao
Travis Wheeler @wheelerlab.org · 31/05/2026
Introducing nail - a Rust implementation of profile HMM sequence alignment for proteins. Near-HMMER sensitivity, but a lot faster: www.biorxiv.org/content/10.1... github.com/TravisWheele...
biorxiv.org
14418
Reposted by Gu Zhenhao
Ben Langmead @benlangmead.bsky.social · 19/05/2026
Revamped the teaching materials site. Light/dark mode switch at the top. More consistent, compact presentation of all the materials. + Icons and such! www.langmead-lab.org/teaching.html
View of https://www.langmead-lab.org/teaching.htmlView of https://www.langmead-lab.org/teaching.html with the Suffix Indexing section expanded
22611
Reposted by Gu Zhenhao
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 17/05/2026
KaryoScope: rapid, alignment-free sequence annotation for the pangenome era www.biorxiv.org/content/10.64898/20…
042
Reposted by Gu Zhenhao
Sina Majidian @sinamajidian.bsky.social · 17/05/2026
Population-level structural variant characterization using pangenome graphs www.nature.com/articles/s41... Swave introduces 'projection waves` to summarize the dotplot images, capturing mapping patterns in pangenomes. A recurrent neural networks distinguishes true SV from background noise/repeats.
Fig. 1 | Schematic overview of Swave. a, Pangenome graph construction and allele extraction. Swave extracts allele paths for both reference and individual assemblies from a pangenome graph. b, Sequence-to-image transformation. Dotplot images depicting structural differences between the REF sequence and ALT allele sequences are projected into waves to extract both background and SV-indicating signals. c, Deep learning-based SV classification. A RNN takes in the wave signals and assigns SV types based on learned sequence-context patterns.
Fig. 2 | Benchmarking the performance of Swave and comparative methods. a, Two key metrics, including Mendelian consistency and genotyping missing rates were evaluated among Swave, assembly- and LRS-based callers. Mendelian consistency was averaged across three parent–child trios (CHS, PUR and YRI). Genotyping missing rates were calculated using the HGSVC samples.
b, Comparison of the detected inversion numbers among HGSVC 65 samples. c, Comparison of the detected complex inversion numbers among HGSVC 65 samples.Extended Data Fig. 3 | Recurrent Neural Network for SV classification in Swave. a, Projected wave signals are encoded as four-element tuples per genomic segment, comprising span length, background average wave value, and the differences between SV-implying and background waves for both forward and reverse matches. These tuples serve as the input for the RNN classification model. b, A one-layer Bi-LSTM with 64 hidden units forms the core of the RNN, enabling context-aware classification of SV components across the sequence.
1114
Reposted by Gu Zhenhao
Bioinformatics Advances @bioinfoadv.bsky.social · 07/05/2026
🔍 New paper in Bioinformatics Advances: "Kaminari: A frugal colored index for approximate k-mer queries"  Read it here: doi.org/10.1093/bioadv/vbag120 Authors include: @yhhshb.bsky.social, @yoann.bsky.social, @robp.bsky.social, @pierrepeterlongo.bsky.social, @jermp.bsky.social
21510
Reposted by Gu Zhenhao
Paul Medvedev @pashadag.bsky.social · 05/05/2026
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
doi.org
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
13019
Reposted by Gu Zhenhao
Mile Sikic @msikic.bsky.social · 28/04/2026
HERRO has been published in @nature.com nature.com/articles/s41.... This achievement is a result of the great work by Dominik Stanojevic, with contributions from Dehui Lin, @sergeynurk.bsky.social, and Paola Florez de Sessions Welcome to the era of high-quality genome assemblies supported by AI.
nature.com
Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads - Nature
Nature - Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads
1279
Reposted by Gu Zhenhao
Erkison Odih @erkison.bsky.social · 27/04/2026
Excited to finally share our new preprint on bioRxiv describing Verticall (github.com/rrwick/Verti...), a robust & efficient tool for building recombination-free bacterial phylogenies. Huge thanks to @rrwick.bsky.social & @katholt.bsky.social for this incredible work! www.biorxiv.org/content/10.6...
biorxiv.org
0259
Reposted by Gu Zhenhao
Rob Patro @robp.bsky.social · 17/04/2026
QCatch is now published in Bioinformatics (academic.oup.com/bioinformati...)! Great work from Yuan and Dongze for quality control and analysis downstream of simpleaf/alevin-fry (taking advantage of its structured AnnData output). Give it a try: github.com/COMBINE-lab/...
academic.oup.com
QCatch: A framework for quality control assessment and analysis of single-cell sequencing data
AbstractMotivation. Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstrea
0184
Reposted by Gu Zhenhao
Sebastian Schmidt @tsbschm.bsky.social · 14/04/2026
This simulation-cum-benchmark study on MAG making by @tkorem.bsky.social & team looks really interesting. Loads of plots and results to work through! www.biorxiv.org/content/10.6...
biorxiv.org
1194
Reposted by Gu Zhenhao
Giulio Ermanno Pibiri @jermp.bsky.social · 13/04/2026
Accepted to ISMB'26. Revised paper is here: jermp.github.io/assets/pdf/p.... I'd like to thank @robp.bsky.social once again and all the received feedback from the reviewers. To me, ISMB has had the highest quality review process over the past few years!
jermp.github.io
1126
Reposted by Gu Zhenhao
Josipa Lipovac @jlipovac.bsky.social · 06/04/2026
Following up on this - MADRe is now officially published 🎉 Very grateful for the guidance of @msikic.bsky.social @rvicedomini.bsky.social and Kresimir Krizanovic 🔗 academic.oup.com/gigascience/...
1106
Reposted by Gu Zhenhao
Jim Shaw @jimshaw.bsky.social · 27/03/2026
Myloasm, our long-read metagenome assembler, is now published! w/ @mgmarin.bsky.social and @lh3lh3.bsky.social Very rewarding after > a year of development and countless hours thinking about assembly. Thanks to beta testers, Li lab, and reviewers who gave very helpful feedback. rdcu.be/famFj
rdcu.be
High-resolution metagenome assembly for modern long reads with myloasm
Nature Biotechnology - A long-read metagenome assembly method recovers circular and complete genomes better than existing tools.
410056
Reposted by Gu Zhenhao
Antoine Limasset @npmalfoy.bsky.social · 27/03/2026
Preprint alert! TLDR: Super Bloom is a Bloom-filter variant for streaming k-mer queries. It uses minimizers to group adjacent k-mers into super-k-mers and map them to the same memory block. Result: much better locality, faster queries, and with the findere trick, dramatically fewer false positives.
doi.org
2208
Reposted by Gu Zhenhao
Heng Li @lh3lh3.bsky.social · 23/03/2026
Long reads carry multiple small vars and SVs and their phasing. LongcallD is the only caller that tightly integrates germline/mosaic small/structural vars/MEIs and their phasing in a single C program. One command line to get competitive small variant calls and better SVs. Led by Yan Gao.
04522
Reposted by Gu Zhenhao
Nicolae Sapoval @nsapoval.bsky.social · 13/03/2026
Our new preprint on quantifying microbial sample diversity/complexity in a way that accounts for both metagenome architecture and taxonomic composition is now live on bioRxiv: www.biorxiv.org/content/10.6... #metagenomics #bioinformatics #dataanalysis #graphdata
biorxiv.org
2148
Reposted by Gu Zhenhao
Ragnar {Groot Koerkamp} @curiouscoding.nl · 13/03/2026
It's a good day when the first item in your feed is your own work :) @rickbitloo.bsky.social was annoyed that scanning reads for all 96 rapid kit barcodes is bottleneck in Barbell, so he made Sassy2: 13x (150bp) to 4.6x (8kbp) faster than v1 by batch-searching patterns, and >100Gbp/s on 16 threads!
22411
Reposted by Gu Zhenhao
Matthew Nguyen @mnguyen667.bsky.social · 09/03/2026
1/ Excited to share my first first-author preprint from my PhD! We introduce Perseus, a lineage-aware confidence estimation framework for taxonomic classification in long-read metagenomics. Preprint: www.biorxiv.org/content/10.6... Code: github.com/matnguyen/Pe...
biorxiv.org
1148
Reposted by Gu Zhenhao
Mile Sikic @msikic.bsky.social · 10/03/2026
Transformer-based AI has boosted @nanoporetech.com sequencing accuracy, but at a cost to portability due to GPU demands. Our new work, spearheaded by Sara Bakic, introduces Campolina link.springer.com/article/10.1... to improve nanopore signal segmentation for event-based mappers.
link.springer.com
Campolina: a deep neural framework for accurate segmentation of nanopore signals - Genome Biology
Nanopore sequencing enables real-time, long-read analysis by processing raw signals as they are produced. A key step, segmentation of signals into events, is typically handled algorithmically, struggl...
22314
Reposted by Gu Zhenhao
Kristoffer Sahlin @ksahlin.bsky.social · 09/03/2026
1/ Our paper on Multi-Context Seeds is now out, with @tolyan.bsky.social spearheading the work and contributions from Nicolas and @marcelm.net. We introduce a new seeding concept that improves read alignment accuracy while maintaining speed. link.springer.com/article/10.1...
link.springer.com
Multi-context seeds enable fast and high-accuracy read mapping - Genome Biology
A key step in sequence similarity search is to identify shared seeds between a query and a reference sequence. A well-known tradeoff is that longer seeds offer fast searches but reduce sensitivity in ...
11912
Reposted by Gu Zhenhao
Páll Melsted @pmelsted.bsky.social · 06/03/2026
Excited to share this preprint that describes my latest work on using GPUs to accelerate processing of RNA-seq data. The title says it all: "RNA-seq analysis in seconds using GPUs" now on biorxiv www.biorxiv.org/content/10.6... and github github.com/pachterlab/k... Figure 1 shows they key result
618887
Reposted by Gu Zhenhao
Niranjan Nagarajan @niranjantw.bsky.social · 06/03/2026
🌟 We finally have a brand new lab webpage, with blog-style articles describing our work, and exciting *open positions* in culturomics and metagenomics! Please spread the word 🙏 mtms-lab.github.io csb5.github.io/open-positions
mtms-lab.github.io
Redirecting to https://csb5.github.io/ ...
071
Reposted by Gu Zhenhao
James Ferguson @psy-fer.bsky.social · 01/03/2026
Introducing kuva: A scientific plotting library in rust, along with cli binary with the option to plot directly into the terminal. Feel free to drop me some feedback as an issue on the repo github.com/Psy-Fer/kuva crates.io/crates/kuva/...
github.com
GitHub - Psy-Fer/kuva: A scientific plotting library in Rust
A scientific plotting library in Rust. Contribute to Psy-Fer/kuva development by creating an account on GitHub.
24311
Reposted by Gu Zhenhao
Niranjan Nagarajan @niranjantw.bsky.social · 20/02/2026
Are you still relying on quality values to do QC for genome sequencing data? What if there was an ultra-fast method that does not need quality values or alignments to genomes? If this sounds interesting, check out our latest preprint w/ @guzhenhao.bsky.social: www.biorxiv.org/content/10.6...
biorxiv.org
052