Sign in

Wei Shen 沈 伟

@shenwei356.bsky.social
2.4K followers 991 following 46 posts

Associate professor of Bioinformatics at Chongqing Medical University, China. Lab: mbio.info

PostsRepliesMedia
Wei Shen 沈 伟 @shenwei356.bsky.social · 24/09/2026
SeqKit is now faster and uses 40% less memory when reading large FASTA records, such as human genomes, though it still uses more RAM than tools written in C/C++/Rust. Worth upgrading! github.com/shenwei356/s...
github.com
GitHub - shenwei356/seqkit: A cross-platform and ultrafast toolkit for FASTA/Q file manipulation
A cross-platform and ultrafast toolkit for FASTA/Q file manipulation - shenwei356/seqkit
0146
Wei Shen 沈 伟 @shenwei356.bsky.social · 24/09/2026
csvtk brings several new commands and lots of new features! It's getting harder to find the time to meet all users' expectations (new features) without the help of AI agents. I hope these toolkits will remain useful and continue to benefit people in the AI era. github.com/shenwei356/c...
github.com
GitHub - shenwei356/csvtk: A cross-platform, efficient and practical CSV/TSV toolkit in Golang
A cross-platform, efficient and practical CSV/TSV toolkit in Golang - shenwei356/csvtk
171
Wei Shen 沈 伟 @shenwei356.bsky.social · 24/09/2026
Lots of improvements to Rush! (vibe-coded) Closing all those GitHub issues feels so good 😀 github.com/shenwei356/r...
github.com
GitHub - shenwei356/rush: A cross-platform command-line tool for executing jobs in parallel
A cross-platform command-line tool for executing jobs in parallel - shenwei356/rush
020
Reposted by Wei Shen 沈 伟
Heng Li @lh3lh3.bsky.social · 17/09/2026
Single-page, interactive overlap assembler for teaching purposes: lh3.github.io/teach-demo/o... (vibe coded)
lh3.github.io
Overlap Graph Assembly Demo
14218
Reposted by Wei Shen 沈 伟
Jim Shaw @jimshaw.bsky.social · 15/09/2026
The sylph metagenome profiler is v1.0.0! sylph-docs.github.io A new DB format + approach --> huge performance gains: GTDB-R232 (200k species) now takes < 5 GB of RAM and ~30s (2GB fq.gz). Huge thanks to @benjwoodcroft.bsky.social and his ongoing performance efforts (github.com/wwood/weebill)
sylph-docs.github.io
Documentation for sylph - ultrafast, precise metagenomic profiling
26330
Reposted by Wei Shen 沈 伟
Zamin Iqbal @zaminiqbal.bsky.social · 09/09/2026
More than a decade since Mykrobe first came out - v happy @martibartfast.bsky.social has developed a Go version, which means we could revive the old app, which always made it much more accessible. Same results as Mykrobe but faster.
0123
Reposted by Wei Shen 沈 伟
Igor Martayan @imartayan.bsky.social · 04/09/2026
This is in 2 hours! In the meantime, you can look at the slides here: igor.martayan.org/slides-phd.pdf
igor.martayan.org
352
Wei Shen 沈 伟 @shenwei356.bsky.social · 03/09/2026
It's really a great opportunity!
010
Wei Shen 沈 伟 @shenwei356.bsky.social · 23/08/2026
Vibe-coded a tiny tool for cold-cache benchmarks: evict cached data for specific files or directories—no root required! hyperfine --prepare 'drop_file_cache data/' 'tool data/' github.com/shenwei356/d...
github.com
GitHub - shenwei356/drop_file_cache: Recursively evicts cached data for specified files or directories. It is designed for cold-cache benchmarks and does not require root privileges on Linux, macOS, o...
Recursively evicts cached data for specified files or directories. It is designed for cold-cache benchmarks and does not require root privileges on Linux, macOS, or BSD. - shenwei356/drop_file_cache
1151
Reposted by Wei Shen 沈 伟
Rob Patro @robp.bsky.social · 15/08/2026
Cuttlefish 3 is on bioconda! 🦑 A parallel, external-memory algorithm for building colored compacted de Bruijn graphs at collection scale — a RECOMB 2026 paper, and as of today a production release: v3.0.0. A thread on the algorithm, the numbers, and why the released tool is a Rust rewrite. 🧵 1/10
34412
Reposted by Wei Shen 沈 伟
Ryan Wick @rrwick.bsky.social · 24/07/2026
New blog post! I reran the Autocycler paper benchmarks on some new tools/versions/pipelines: rrwick.github.io/2026/07/24/b... (1/3)
rrwick.github.io
Benchmark update: Ilesta, Autocycler-fast and new versions
a blog for miscellaneous bioinformatics stuff
23416
Reposted by Wei Shen 沈 伟
Zamin Iqbal @zaminiqbal.bsky.social · 25/06/2026
Honoured to be awarded the Mary Lyon medal by the Genetics Society today, here in Edinburgh. V grateful to mentors and all my previous group members and colleagues. Medal is an X chromosome!!
Gold/bronze coloured medal, with depicted X chromosome Second side of medal says "Mary Lyon Medal" awarded to Prof Zam Iqbal. By the Genetics Society
710510
Reposted by Wei Shen 沈 伟
Daan Speth @daanspeth.bsky.social · 26/06/2026
I'm happy to announce the release of GlobDB r232! This version contains 346,233 bacterial and archaeal genomes, based on 26 datasets. More info globdb.org 🦠🖥️🧬
globdb.org
home | GlobDB
24623
Reposted by Wei Shen 沈 伟
Ragnar {Groot Koerkamp} @curiouscoding.nl · 26/06/2026
2/2 ESA-E accepted: - SimdQuickHeap [w/ Johannes Breitling, Marvin Williams] - non-minimal k-PHF [Stefan Hermann, Peter Sanders, Stefan Walzer] and 2/2 WABI accepted: - O(n lg lg n) chaining [Nicola Rizzo] - anti-lex SUS-anchors arxiv: arxiv.org/search/?quer...
1111
Reposted by Wei Shen 沈 伟
Johanna von Wachsmann @johannavw.bsky.social · 16/06/2026
🧬 New preprint! We clustered 5.6 million bacterial genomes into genomically cohesive units (GCUs) 500× faster than existing tools. (In just 14 hours, 16.5 GB RAM using 48 CPUs). 🦠🐙Meet gemsparcl 💎✨! www.biorxiv.org/content/10.6...
06124
Reposted by Wei Shen 沈 伟
Heng Li @lh3lh3.bsky.social · 16/06/2026
Minibwa is a hybrid of bwa-mem and minimap2 and the successor of bwa-mem for short-read mapping. ~4X/2.5X as fast as bwa-mem/bwa-mem2 for WGS reads at comparable accuracy. Native support of directional bisulfite-seq. Applicable to long reads. Preprint at arxiv.org/abs/2606.15357
1193109
Reposted by Wei Shen 沈 伟
Kristian G. Andersen @kgandersen.bsky.social · 21/05/2026
Just a shout-out to the folks over at @pathoplexus.org and, importantly, all the scientists sharing genomic data with such amazing speed. It's great to see the field having moved to a new platform, with non-transparent platforms being a thing of the past. Fantastic.
16225
Reposted by Wei Shen 沈 伟
Igor Martayan @imartayan.bsky.social · 21/04/2026
New blog post! I use ntHash all the time to hash k-mers, yet it turns out it has some unexpected flaws (collision propagation, bias on leading zeros...). The good news: each of them can be fixed! igor.martayan.org/posts/breaki...
igor.martayan.org
Breaking ntHash (to better fix it)
NtHash is a popular method for hashing k-mers in bioinformatics, yet it has some surprising flaws. In this post, I walk through a few of them, and show that they can arise naturally, without an advers...
13016
Reposted by Wei Shen 沈 伟
ace-gtdb.bsky.social @ace-gtdb.bsky.social · 15/04/2026
GTDB release 11 based on RefSeq 232 (R11-RS232) is live at gtdb.ecogenomic.org. This release covers 901,341 genomes (23% increase) and has 199,923 species clusters (39% increase). Release notes at: forum.gtdb.ecogenomic.org/t/announcing.... Release statistics at: gtdb.ecogenomic.org/stats/r232.
gtdb.ecogenomic.org
GTDB - Genome Taxonomy Database
The Genome Taxonomy Database (GTDB) is an initiative to establish a standardised microbial taxonomy based on genome phylogeny.
15331
Reposted by Wei Shen 沈 伟
LaurieWired @lauriewired.bsky.social · 07/04/2026
Modern DRAM is based on a brilliant design from IBM. But, we're still paying for a latency penalty that's existed since the 60s! In this video, I'm introducing my research project (Tailslayer) that immensely reduces p99.99 latency on traditional RAM!
319041
Wei Shen 沈 伟 @shenwei356.bsky.social · 13/03/2026
LexicMap v0.9.0 has been released with - a few bug fixes: CIGAR, bitscore and evalue calculation. - new features: better support for big genomes like human. - a new command to convert output to SAM format github.com/shenwei356/L...
github.com
Release LexicMap v0.9.0 · shenwei356/LexicMap
v0.9.0 - 2026-03-13 New commands: lexicmap utils 2sam: Convert the default search output to SAM format (#26). Attention: This command requires search results generated by the current LexicMap ver...
081
Wei Shen 沈 伟 @shenwei356.bsky.social · 27/02/2026
Can't wait to release a 10-year-old birthday version for SeqKit! - 10 years - 2 papers, 3500 citations - 20 contributors - 40 subcommands - 880 commits - 500 issues - 685.5K Bioconda total downloads Thank you all, dear contributors and users! I'll keep maintaining it. github.com/shenwei356/s...
github.com
Release SeqKit v2.13.0 (10-year-old birthday version) · shenwei356/seqkit
Changelog SeqKit is 10 years old! SeqKit v2.13.0 - 2026-02-28 seqkit: add support for reading and writing LZ4 compression format. new command: seqkit sample2: improved seqkit sample by @stahiga....
612735
Reposted by Wei Shen 沈 伟
Zamin Iqbal @zaminiqbal.bsky.social · 26/02/2026
I should have said, this takes us to about 2.8 million genomes in total. We don't have annotations, etc for the latest data yet, this will be an ongoing process
1147
Reposted by Wei Shen 沈 伟
Zamin Iqbal @zaminiqbal.bsky.social · 09/02/2026
A long time ago in a galaxy far away, there was a SARS-CoV-2 pandemic. Our paper, led by @martibartfast.bsky.social a) correcting errors in 4.5 million genomes & their phylogeny b) improving representation of the Global South in public data www.nature.com/articles/s41... (thread 1/n)
nature.com
Addressing pandemic-wide systematic errors in the SARS-CoV-2 phylogeny - Nature Methods
This Resource paper presents a global SARS-CoV-2 phylogenetic tree of 4,471,579 high-quality genomes consistently constructed by Viridian, an efficient amplicon-aware assembler.
313766
Reposted by Wei Shen 沈 伟
George Bouras @gbouras13.bsky.social · 14/01/2026
Phold's manuscript is now available @narjournal.bsky.social thanks to @susiegriggo.bsky.social @npbhavya.bsky.social @vijinim.bsky.social @linsalrob.bsky.social @martinsteinegger.bsky.social @milot.bsky.social @eunbelivable.bsky.social & others not on bsky #phagesky academic.oup.com/nar/article/...
academic.oup.com
Protein structure-informed bacteriophage genome annotation with Phold
Abstract. Bacteriophage (phage) genome annotation is essential for understanding their functional potential and suitability for use as therapeutic agents.
18644
Reposted by Wei Shen 沈 伟
Zamin Iqbal @zaminiqbal.bsky.social · 19/01/2026
For those of us interested in software development, data structure design etc in science, this is a must-read. A taste of what is happening in communities letting AI agents go wild writing code, creating PRs, writing documentation : spoiler - humans get addicted, lose perspective, slop everywhere.
0206
Wei Shen 沈 伟 @shenwei356.bsky.social · 04/01/2026
Phage therapy of perinephric abscess in kidney transplantation recipients caused by drug‐resistant Pseudomonas aeruginosa - Liu - 2025 - mLife - Wiley Online Library onlinelibrary.wiley.com/doi/10.1002/...
onlinelibrary.wiley.com
Phage therapy of perinephric abscess in kidney transplantation recipients caused by drug‐resistant Pseudomonas aeruginosa
Click on the article title to read more.
031
Reposted by Wei Shen 沈 伟
Qiyun Zhu @zhuqiyun.bsky.social · 11/12/2025
The scikit-bio paper in online in Nature Methods! Many thanks to our collaborators, community contributors and reviewers! We couldn’t have done it without you. www.nature.com/articles/s41... #Bioinformatics #OpenSource
nature.com
Scikit-bio: a fundamental Python library for biological omic data analysis - Nature Methods
Nature Methods - Scikit-bio: a fundamental Python library for biological omic data analysis
39852
Wei Shen 沈 伟 @shenwei356.bsky.social · 11/12/2025
My first Rust toy tool (just for learning purposes)!!! The performance (both time and memory) is good! Rust is hard to learn, and I still need more practice! github.com/shenwei356/f...
Performance on reading and writing plain FASTA/Q files
2272
Reposted by Wei Shen 沈 伟
Zamin Iqbal @zaminiqbal.bsky.social · 20/11/2025
Most exciting study have seen for ages, and Fernando the most excited speaker. Much anticipated. Highly recommended, a lot of food for thought (and quite a dense paper - lots to think about)
1185
Reposted by Wei Shen 沈 伟
Microbiome Virtual International Forum @microbiomevif.bsky.social · 17/11/2025
It's Monday! ...and a new #MVIF program is out! 🤩 Free registration: cassyni.com/s/mvif-44 ⭐️ Highlights: 🇺🇸 Vanessa Hale 🇰🇷 Jun Hyung Cha ⭐️ Keynote: 🇺🇸 Katherine Lemon @kathlemon.bsky.social ⭐️ Talks: 🇺🇸 Meenakshi Chakraborty 🇨🇳 Wei Shen @shenwei356.bsky.social 🇺🇸 Johanna Gutleben
MVIF 44
044
Reposted by Wei Shen 沈 伟
Vivek Mutalik @vivekmutalik.bsky.social · 16/11/2025
📣 New preprint from us at phagefoundry.org 📣 A solid machine learning framework & to predict strain-level phage-host interactions across diverse bacterial genera from genome sequences alone. Avery Noonan from the Arkin Lab led this massive effort www.biorxiv.org/content/10.1...
phagefoundry.org
Phage Foundry
12817
Reposted by Wei Shen 沈 伟
Zamin Iqbal @zaminiqbal.bsky.social · 14/11/2025
Honoured and quite blown-over to receive this award. I have been, and continue to be, very lucky - first with great mentors, and then really prodigious students, postdocs and collaborators. Working with them has been a joy.
4220119
Reposted by Wei Shen 沈 伟
Michael Hall @mbhall88.bsky.social · 07/11/2025
Our method for genome size estimation from long-read overlaps is now published 🥳 academic.oup.com/bioinformati...
academic.oup.com
Genome size estimation from long read overlaps
AbstractMotivation. Accurate genome size estimation is an important component of genomic analyses such as assembly and coverage calculation, though existin
13716
Reposted by Wei Shen 沈 伟
Camille Marchet ⚡ @camillemrcht.bsky.social · 06/11/2025
Thread on #GI2025 's second day! 👇🏻
0115
Reposted by Wei Shen 沈 伟
Sina Majidian @sinamajidian.bsky.social · 06/11/2025
Ben Langmead @benlangmead.bsky.social delivers the official opening for this year's Genome Informatics Conference #GI2025 at Cold Spring Harbor Laboratory. List of talks and posters: meetings.cshl.edu/abstracts.as...
1347
Reposted by Wei Shen 沈 伟
Ragnar {Groot Koerkamp} @curiouscoding.nl · 04/11/2025
Cool paper new paper from Lorién López-Villellas, @santiagomarco.bsky.social and others! Super cute and simple idea: In Gotoh's affine-cost alignment, only the M matrix is needed during tracing: we can just search for a gap-length x such that M[i][j] = M[i-x][j]+o+x*e or M[i][j] = M[i][j-x]+o+x*e.
182
Reposted by Wei Shen 沈 伟
Titus Brown @titus.idyll.org · 29/10/2025
I also have serious concerns about the consolidation of roles (one person is now publisher, chief editor, and also a frequent author) as exemplified in a recent paper that was fast-tracked for publication.
161
Reposted by Wei Shen 沈 伟
Ragnar {Groot Koerkamp} @curiouscoding.nl · 23/10/2025
Really exciting that the preprint on Barbell, a new demultiplexer, is finally out! It's the first tool that builds on Sassy, the approximate-DNA-searching tool that @rickbitloo.bsky.social and myself developed earlier this year, specifically with this application in mind.
22015
Reposted by Wei Shen 沈 伟
Rick Beeloo @rickbitloo.bsky.social · 23/10/2025
Around 10% of your Nanopore reads (SQK-RBK114) are incorrectly trimmed. Here is why, and how our new tool Barbell solves it: www.biorxiv.org/content/10.1... Want to get started? github.com/rickbeeloo/b...
35231
Reposted by Wei Shen 沈 伟
Mohsen Zakeri @mohsenzakeri.bsky.social · 21/10/2025
1/6 Movi 2 is here: faster and more space-efficient for pangenome queries. Its fastest mode uses half the memory of Movi 1 while running ~30% faster. github.com/mohsenzakeri...
github.com
GitHub - mohsenzakeri/Movi: Fast, Cache-Efficient, and Scalable Queries on Pangenomes
Fast, Cache-Efficient, and Scalable Queries on Pangenomes - mohsenzakeri/Movi
14424
Reposted by Wei Shen 沈 伟
Zamin Iqbal @zaminiqbal.bsky.social · 17/10/2025
Podcast with me and @turiking.bsky.social for the @milnerevolution.bsky.social series, on plasmid evolution over the last 100 years, talking about our ( @cazares-adr.bsky.social , Nick Thomson, @sarah1alexander.bsky.social & co) recent paper www.science.org/doi/10.1126/... youtu.be/Mzr3TD4ijs0?...
youtu.be
How the Vectors of Antibiotic Resistance Have Evolved - Professor Zamin Iqbal
YouTube video by Milner Centre for Evolution
14416
Reposted by Wei Shen 沈 伟
Jim Shaw @jimshaw.bsky.social · 08/09/2025
Preprint out for myloasm, our new nanopore / HiFi metagenome assembler! Nanopore's getting accurate, but 1. Can this lead to better metagenome assemblies? 2. How, algorithmically, to leverage them? with co-author Max Marin @mgmarin.bsky.social, supervised by Heng Li @lh3lh3.bsky.social 1 / N
511380
Reposted by Wei Shen 沈 伟
Roland Faure @rfaure.bsky.social · 03/10/2025
Our preprint on our new metagenomic HiFi assembler Alice is out 🥳 Based on a *new sketching method* (🧵1/6) 👉 Preprint www.biorxiv.org/content/10.1... 👉 Github github.com/rolandfaure/...
biorxiv.org
Alice: fast and haplotype-aware assembly of high-fidelity reads based on MSR sketching
We introduce Mapping-friendly Sequence Reduction (MSR) sketches, a sketching method for high-fidelity (HiFi) long reads, and Alice, an assembler that operates directly on these sketches. MSR produces ...
22521
Reposted by Wei Shen 沈 伟
Ben Langmead @benlangmead.bsky.social · 14/10/2025
New tool "bwt-svg" for making illustrations of the BWT and the many auxiliary arrays and other structures related to it. Pyodide-based no-installation-necessary interface here: benlangmead.github.io/bwt-svg/. (H/t to @robert.bio for pointing me to pyodide!) Full repo: github.com/benlangmead/....
Illustration of Burrows-Wheeler Transform and many auxiliary structures from the input string how$now$brown$cow$#
34021
Reposted by Wei Shen 沈 伟
Andre Kahles @akkah21.bsky.social · 08/10/2025
After years of research and continuous refinement, we’re thrilled to share that our paper on the MetaGraph framework — enabling Petabase-scale search across sequencing data — has been published today in Nature (www.nature.com/articles/s41...)
nature.com
Efficient and accurate search in petabase-scale sequence repositories - Nature
MetaGraph enables scalable indexing of large sets of DNA, RNA or protein sequences using annotated de Bruijn graphs.
33016
Reposted by Wei Shen 沈 伟
Stephen Turner @stephenturner.us · 09/10/2025
Efficient and accurate search in petabase-scale sequence repositories www.nature.com/articles/s41... 🧬🖥️🧪 MetaGraph: metagraph.ethz.ch Code: github.com/ratschlab/me...
0197
Reposted by Wei Shen 沈 伟
Robert Aboukhalil @robert.bio · 09/10/2025
Just published an interactive article about a magical algorithm known as the Burrows-Wheeler Transform, which powers sequence alignment tools like bowtie and bwa: sandbox.bio/concepts/bwt It's also notoriously unintuitive so I'm hoping this article helps you build that intuition.
39829
Reposted by Wei Shen 沈 伟
EMBL-EBI @ebi.embl.org · 30/09/2025
There are millions of openly available microbial genomes, but searching them can be slow. Until now 🥁 Introducing LexicMap, a new alignment tool that lets scientists search these data in minutes, helping track antibiotic resistance, trace outbreaks, and more. www.ebi.ac.uk/about/news/r... 🦠
ebi.ac.uk
How to rapidly search the world’s microbial DNA
By making the world’s microbial DNA easier to explore, LexicMap helps researchers track outbreaks, study antibiotic resistance, and understand microbial diversity.
14116
Reposted by Wei Shen 沈 伟
Paul Medvedev @pashadag.bsky.social · 25/09/2025
Thank you folks for your feedback on our survey about Hash functions in genomic sequence analysis. We've updated the paper and you can see the new version here: tinyurl.com/4kk9ccmt.
tinyurl.com
Dropbox
0116