Sign in

Gaëtan Benoit

@gaetanbenoit.bsky.social
131 followers 61 following 15 posts

Postdoc researcher in bioinformatics at Pasteur institute. Scalable methods and software for metagenomics. github.com/GaetanBenoitDev

PostsRepliesMedia
Reposted by Gaëtan Benoit
Jim Shaw @jimshaw.bsky.social · 15/09/2026
The sylph metagenome profiler is v1.0.0! sylph-docs.github.io A new DB format + approach --> huge performance gains: GTDB-R232 (200k species) now takes < 5 GB of RAM and ~30s (2GB fq.gz). Huge thanks to @benjwoodcroft.bsky.social and his ongoing performance efforts (github.com/wwood/weebill)
sylph-docs.github.io
Documentation for sylph - ultrafast, precise metagenomic profiling
26330
Reposted by Gaëtan Benoit
Ryan Wick @rrwick.bsky.social · 24/07/2026
Notable finding #1: the current version of metaMDBG is quite good! Fast, very memory efficient and now one of the most accurate assemblers. Nice work, @gaetanbenoit.bsky.social! github.com/GaetanBenoit... (2/3)
261
Reposted by Gaëtan Benoit
Ryan Wick @rrwick.bsky.social · 24/07/2026
New blog post! I reran the Autocycler paper benchmarks on some new tools/versions/pipelines: rrwick.github.io/2026/07/24/b... (1/3)
rrwick.github.io
Benchmark update: Ilesta, Autocycler-fast and new versions
a blog for miscellaneous bioinformatics stuff
23416
Reposted by Gaëtan Benoit
Steven Robbins @stevenjrobbins.bsky.social · 22/07/2026
It's out! Excited to present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), a comprehensive DB of 1000s of high-quality prokaryote, virus, plasmid, and chromosome-level eukaryote MAGs using Nanopore long reads. Subthreads incoming. Please share widely. 🙂 www.nature.com/articles/s41...
nature.com
The planktonic microbiome of the Great Barrier Reef - Nature
The Great Barrier Reef Microbial Genomes Database compiles prokaryotic, viral and eukaryotic genomes from seawater collected from the Great Barrier Reef, providing a rich resource for the study of mar...
88342
Reposted by Gaëtan Benoit
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 09/07/2026
Lossless compression of k-mer matrices enabling random row access www.biorxiv.org/content/10.64898/20…
075
Reposted by Gaëtan Benoit
Ryan Wick @rrwick.bsky.social · 19/06/2026
Follow-up to my last blog post: I now look at assembly polishing with Dorado v2 and the new hac@v6.0.0 model from @nanoporetech.com. rrwick.github.io/2026/06/19/d...
rrwick.github.io
Dorado v2.0.0 part 2: assembly polishing
a blog for miscellaneous bioinformatics stuff
1196
Reposted by Gaëtan Benoit
Johanna von Wachsmann @johannavw.bsky.social · 16/06/2026
🧬 New preprint! We clustered 5.6 million bacterial genomes into genomically cohesive units (GCUs) 500× faster than existing tools. (In just 14 hours, 16.5 GB RAM using 48 CPUs). 🦠🐙Meet gemsparcl 💎✨! www.biorxiv.org/content/10.6...
06124
Reposted by Gaëtan Benoit
Heng Li @lh3lh3.bsky.social · 16/06/2026
Minibwa is a hybrid of bwa-mem and minimap2 and the successor of bwa-mem for short-read mapping. ~4X/2.5X as fast as bwa-mem/bwa-mem2 for WGS reads at comparable accuracy. Native support of directional bisulfite-seq. Applicable to long reads. Preprint at arxiv.org/abs/2606.15357
1193109
Reposted by Gaëtan Benoit
Igor Martayan @imartayan.bsky.social · 15/06/2026
Just submitted my PhD thesis on algorithms for fast, large-scale k-mer-based sequence analysis. It's now available to read at phd.martayan.org Take a look and feel free to share! #Bioinformatics #PhDone
phd.martayan.org
Algorithm design and implementation for the scale of sequencing data
26822
Reposted by Gaëtan Benoit
Ryan Wick @rrwick.bsky.social · 11/06/2026
New blog post! I analyse the new hac@v6.0.0 basecalling model from @nanoporetech.com and discuss the conspicuous lack of a new sup model: rrwick.github.io/2026/06/11/d...
rrwick.github.io
Dorado v2.0.0: no more sup?
a blog for miscellaneous bioinformatics stuff
15526
Reposted by Gaëtan Benoit
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
🌎 🧬 🖥️ logan-search.org the tool to query all SRA sequences (Dec 2023 snapshot) — just got major updates. Here's what's new. 🧵 1/12
13626
Reposted by Gaëtan Benoit
Steven Robbins @stevenjrobbins.bsky.social · 04/06/2026
Thought i'd highlight that the ONT London Calling tech talk is now up. Points of interest for the microbiome community: 1) Direct RNA multiplexing now available. Can now run 24 samples per flow cell, recover full-length transcripts with 8 base pair modifications... www.youtube.com/watch?v=CE69...
youtube.com
London Calling 2026 Technology update
YouTube video by Oxford Nanopore Technologies
1116
Reposted by Gaëtan Benoit
Roland Faure @rfaure.bsky.social · 04/06/2026
Glad to announce that our new long-read metagenomic SNP caller, SNooPy, is published in NAR! Existing long-read SNP callers are not designed for metagenomic data, check out our new solution 👉https://academic.oup.com/nar/article/54/10/gkag556/8700491 @narjournal.bsky.social
1177
Reposted by Gaëtan Benoit
Jim Shaw @jimshaw.bsky.social · 30/05/2026
Our method, savont, for generating amplicon sequence variants (ASVs) for long-read amplicons is now on bioRxiv. Work with @lh3lh3.bsky.social and help from @mkddueholm.bsky.social and team (Marie Riisgaard-Jensen, @kirk3gaard.bsky.social, Kasper Skytte Andersen) github.com/bluenote-157... 1/6
24219
Reposted by Gaëtan Benoit
Heng Li @lh3lh3.bsky.social · 30/05/2026
Jeremy Wang developed rammap, a minimap2 rewrite in Rust. It achieves comparable or better performance than minimap2 and produces identical output to minimap2. During rewrite, Jeremy found two long-existing bugs in minimap2 which are fixed in v2.31. www.biorxiv.org/content/10.6...
biorxiv.org
310944
Reposted by Gaëtan Benoit
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 28/05/2026
Fast Set Operations for Compact k-mer Sets www.biorxiv.org/content/10.64898/20…
01510
Reposted by Gaëtan Benoit
Camille Marchet ⚡ @camillemrcht.bsky.social · 17/05/2026
More and better human assemblies. Now annotate them to the minute. Special kudos @trhyker.bsky.social @jnalanko.bsky.social and @florisbarthel.bsky.social
0104
Gaëtan Benoit @gaetanbenoit.bsky.social · 11/05/2026
Note that I released a new version of metaMDBG (v1.4) last week focusing on scalability. You can now process such dataset in 6 days and 130 GB of memory
084
Reposted by Gaëtan Benoit
Bioinformatics Advances @bioinfoadv.bsky.social · 07/05/2026
🔍 New paper in Bioinformatics Advances: "Kaminari: A frugal colored index for approximate k-mer queries"  Read it here: doi.org/10.1093/bioadv/vbag120 Authors include: @yhhshb.bsky.social, @yoann.bsky.social, @robp.bsky.social, @pierrepeterlongo.bsky.social, @jermp.bsky.social
21510
Reposted by Gaëtan Benoit
Paul Medvedev @pashadag.bsky.social · 05/05/2026
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
doi.org
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
13019
Reposted by Gaëtan Benoit
recombseq.bsky.social @recombseq.bsky.social · 03/05/2026
The RECOMB-Seq 2026 program is now available! Join us May 24–25 in Thessaloniki, Greece, for two days of cutting-edge biological sequence analysis, with keynotes by Camille Marchet (CNRS) and Manolis Kellis (MIT). Full schedule: recomb-seq.github.io/seq2026/prog... #RECOMBseq
recomb-seq.github.io
Program
RECOMB-Seq 2026 Web Page
11610
Reposted by Gaëtan Benoit
Christine He @christinehe.bsky.social · 02/05/2026
Released today: ~1.5 Tbp of ONT metagenomic sequencing from compost. Assembly, MAGs, and downstream analysis are available to play with epi2me.nanoporetech.com/compost_mgx_...
epi2me.nanoporetech.com
Metagenomic Assembly Sheds Light on Microbial Diversity in Compost
Overview We are pleased to release a metagenomic dataset from deep sequencing of a mature compost…
01310
Reposted by Gaëtan Benoit
Ragnar {Groot Koerkamp} @curiouscoding.nl · 29/04/2026
New preprint: The SimdQuickHeap is the fastest priority queue by far! 2x faster than a radix heap and up to 10x faster than binary heaps. arxiv.org/abs/2604.25681 with Marvin Williams and Johannes Breitling:
arxiv.org
SimdQuickHeap: The QuickHeap Reconsidered
Priority queues are data structures that maintain a dynamic collection of elements and allow inserting new elements and removing the smallest element. The most widely known and used priority queue is ...
1195
Reposted by Gaëtan Benoit
Floris Barthel @florisbarthel.bsky.social · 29/04/2026
Very excited for our lab's first paper to be printed!
1155
Reposted by Gaëtan Benoit
Mile Sikic @msikic.bsky.social · 28/04/2026
HERRO has been published in @nature.com nature.com/articles/s41.... This achievement is a result of the great work by Dominik Stanojevic, with contributions from Dehui Lin, @sergeynurk.bsky.social, and Paola Florez de Sessions Welcome to the era of high-quality genome assemblies supported by AI.
nature.com
Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads - Nature
Nature - Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads
1279
Reposted by Gaëtan Benoit
Erkison Odih @erkison.bsky.social · 27/04/2026
Excited to finally share our new preprint on bioRxiv describing Verticall (github.com/rrwick/Verti...), a robust & efficient tool for building recombination-free bacterial phylogenies. Huge thanks to @rrwick.bsky.social & @katholt.bsky.social for this incredible work! www.biorxiv.org/content/10.6...
biorxiv.org
0259
Reposted by Gaëtan Benoit
Scott V. Edwards @scottvedwards.bsky.social · 22/04/2026
Ver excited to share my just-published Darwin Review with @lh3lh3.bsky.social on population-scale long-read sequencing! royalsocietypublishing.org/rspb/article...
royalsocietypublishing.org
Population-scale long-read DNA sequencing: peering under the hood of the new evolutionary genomics
Abstract. Population-scale long-read DNA sequencing (PLRS) is rapidly reshaping our understanding of genomic variation in humans and non-model species. In
02919
Reposted by Gaëtan Benoit
Ragnar {Groot Koerkamp} @curiouscoding.nl · 21/04/2026
Turns out that the usual NtHash is not as random as one might think?!?! At least not for minimizers. Seq-hash (and simd-minimizers) already has this fixed by default ;) github.com/rust-seq/seq...
084
Reposted by Gaëtan Benoit
Igor Martayan @imartayan.bsky.social · 21/04/2026
New blog post! I use ntHash all the time to hash k-mers, yet it turns out it has some unexpected flaws (collision propagation, bias on leading zeros...). The good news: each of them can be fixed! igor.martayan.org/posts/breaki...
igor.martayan.org
Breaking ntHash (to better fix it)
NtHash is a popular method for hashing k-mers in bioinformatics, yet it has some surprising flaws. In this post, I walk through a few of them, and show that they can arise naturally, without an advers...
13016
Gaëtan Benoit @gaetanbenoit.bsky.social · 20/04/2026
I was investigating the genomes that I didn't manage to convert to near-complete MAGs in my assembly graph (the components in gray). The circle on top left is actually a complete genome but with 40% completeness (both in metaMDBG and myloasm)
293
Reposted by Gaëtan Benoit
recombseq.bsky.social @recombseq.bsky.social · 18/04/2026
Two more days to submit your abstract for a short talk or poster at RECOMB-Seq 2026. See instructions at recomb-seq.github.io/seq2026/call...
0107
Reposted by Gaëtan Benoit
A. Murat Eren (Meren) @merenbey.bsky.social · 17/04/2026
This is from 7 years ago (merenlab.org/2019/02/24/f...). We are talking about the same things today. We will be talking about the same things 7 years from now. There is no one to blame for this apart from ourselves. I find it very depressing.
12711
Reposted by Gaëtan Benoit
Sebastian Deorowicz @sdeorowicz.bsky.social · 14/04/2026
10 years after the first FAMSA paper, its successor is now published in Nat Biotech! We believe that FAMSA2 can enable analyses of large protein collections that were previously unattainable. Thank you, Andrzej and Cedric, for great collaboration www.nature.com/articles/s41...
nature.com
Fast and accurate multiple-protein-sequence alignment at scale with FAMSA2 - Nature Biotechnology
FAMSA2 accurately aligns millions of protein sequences at high speed.
35923
Reposted by Gaëtan Benoit
Giulio Ermanno Pibiri @jermp.bsky.social · 13/04/2026
Accepted to ISMB'26. Revised paper is here: jermp.github.io/assets/pdf/p.... I'd like to thank @robp.bsky.social once again and all the received feedback from the reviewers. To me, ISMB has had the highest quality review process over the past few years!
jermp.github.io
1126
Reposted by Gaëtan Benoit
George Bouras @gbouras13.bsky.social · 07/04/2026
Whenever I presented Phold, I was frequently asked "can you do the same beyond phages?" We ( @oschwengers.bsky.social @linsalrob.bsky.social @binomicalabs.org et al) finally did it with Baktfold github.com/gbouras13/ba... www.biorxiv.org/content/10.6...
github.com
GitHub - gbouras13/baktfold: Rapid & standardized genome annotation using protein structural information
Rapid & standardized genome annotation using protein structural information - gbouras13/baktfold
15623
Reposted by Gaëtan Benoit
Josipa Lipovac @jlipovac.bsky.social · 06/04/2026
Following up on this - MADRe is now officially published 🎉 Very grateful for the guidance of @msikic.bsky.social @rvicedomini.bsky.social and Kresimir Krizanovic 🔗 academic.oup.com/gigascience/...
1106
Reposted by Gaëtan Benoit
Sebastian Schmidt @tsbschm.bsky.social · 03/04/2026
Our work on 'hidden diversity' in unbinned contigs is now published in @natmicrobiol.nature.com : www.nature.com/articles/s41... See the linked threads for more details!
nature.com
Unbinned contigs expand known diversity in the global microbiome - Nature Microbiology
Re-analysis of over 92,000 metagenomes reveals hundreds of thousands of previously undescribed Bacterial and Archaeal clades hidden in plain sight.
36941
Reposted by Gaëtan Benoit
Sina Majidian @sinamajidian.bsky.social · 30/03/2026
A run-length-compressed skiplist data structure for dynamic GBWTs supports time and space efficient pangenome operations over syncmers doi.org/10.64898/202...
0175
Reposted by Gaëtan Benoit
Heng Li @lh3lh3.bsky.social · 30/03/2026
LongcallR for competitive SNP calling and haplotype phasing, and simplified allele-specific analysis with long RNA-seq reads. Found ~100 junctions affected by SNPs per sample with most junctions novel. Developed by Neng Huang. Published in @natmethods.nature.com. Read at rdcu.be/faKhL
rdcu.be
SNP calling, haplotype phasing and allele-specific analysis with long RNA-seq reads
Nature Methods - In this study, long-read RNA sequencing achieves accurate single-nucleotide polymorphism calling, haplotype phasing and allele-specific expression analysis.
04418
Reposted by Gaëtan Benoit
Igor Martayan @imartayan.bsky.social · 30/03/2026
A quick rant on people vibe-translating our Rust libraries to other languages That's the second time in a week that I see new bioinformatics tools with a vibe-coded translation of our Rust libraries to C/C++. I have two major issues with that:
23310
Reposted by Gaëtan Benoit
Jim Shaw @jimshaw.bsky.social · 27/03/2026
Myloasm, our long-read metagenome assembler, is now published! w/ @mgmarin.bsky.social and @lh3lh3.bsky.social Very rewarding after > a year of development and countless hours thinking about assembly. Thanks to beta testers, Li lab, and reviewers who gave very helpful feedback. rdcu.be/famFj
rdcu.be
High-resolution metagenome assembly for modern long reads with myloasm
Nature Biotechnology - A long-read metagenome assembly method recovers circular and complete genomes better than existing tools.
410056
Reposted by Gaëtan Benoit
Antoine Limasset @npmalfoy.bsky.social · 27/03/2026
Preprint alert! TLDR: Super Bloom is a Bloom-filter variant for streaming k-mer queries. It uses minimizers to group adjacent k-mers into super-k-mers and map them to the same memory block. Result: much better locality, faster queries, and with the findere trick, dramatically fewer false positives.
doi.org
2208
Reposted by Gaëtan Benoit
Hajk-Georg Drost @hajkdrost.bsky.social · 24/03/2026
How much protein diversity can Life on Earth actually generate? With DIAMOND DeepClust, we show how billions of proteins across the tree of life can be clustered at low-identity for downstream analytics tasks. 📚Paper: www.nature.com/articles/s41... 💻Code: github.com/bbuchfink/di...
16529
Reposted by Gaëtan Benoit
Jim Shaw @jimshaw.bsky.social · 24/03/2026
_720 Gbp_ marine nanopore metagenome -> 328 circular prokaryotic contigs: using myloasm! Insane work by Lui and Nielsen. Also shows how modern long read assemblies can disentangle coexisting strains and reveal ecological insights.
24914
Reposted by Gaëtan Benoit
Adam Phillippy @aphillippy.bsky.social · 23/03/2026
Recently amplified gene arrays are a super interesting phenomenon, but many still resist our attempts to assemble them. @dantipov.bsky.social has developed a new method (Trivial Tangle Traverser) that resolves assembly graph tangles caused by such sequences (1/4) www.biorxiv.org/content/10.6...
biorxiv.org
12812
Reposted by Gaëtan Benoit
Heng Li @lh3lh3.bsky.social · 23/03/2026
Long reads carry multiple small vars and SVs and their phasing. LongcallD is the only caller that tightly integrates germline/mosaic small/structural vars/MEIs and their phasing in a single C program. One command line to get competitive small variant calls and better SVs. Led by Yan Gao.
04522
Reposted by Gaëtan Benoit
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 19/03/2026
Super Bloom: Fast and precise filter for streaming k-mer queries www.biorxiv.org/content/10.64898/20…
02213
Reposted by Gaëtan Benoit
Ragnar {Groot Koerkamp} @curiouscoding.nl · 13/03/2026
It's a good day when the first item in your feed is your own work :) @rickbitloo.bsky.social was annoyed that scanning reads for all 96 rapid kit barcodes is bottleneck in Barbell, so he made Sassy2: 13x (150bp) to 4.6x (8kbp) faster than v1 by batch-searching patterns, and >100Gbp/s on 16 threads!
22411
Reposted by Gaëtan Benoit
Loïs Maignien, PhD. @loimai.bsky.social · 10/03/2026
Very happy to share the latest paper of our group (et al.)! rdcu.be/e7zyX . This one has a special place… 1/n
rdcu.be
Water mass specific genes dominate the Southern Ocean microbiome
Nature Communications - Southern Ocean microbial communities are less well studied. Here, the authors generate a circumpolar-scale gene catalog from 218 metagenomics samples revealing broadscale...
1217
Reposted by Gaëtan Benoit
Kristoffer Sahlin @ksahlin.bsky.social · 09/03/2026
1/ Our paper on Multi-Context Seeds is now out, with @tolyan.bsky.social spearheading the work and contributions from Nicolas and @marcelm.net. We introduce a new seeding concept that improves read alignment accuracy while maintaining speed. link.springer.com/article/10.1...
link.springer.com
Multi-context seeds enable fast and high-accuracy read mapping - Genome Biology
A key step in sequence similarity search is to identify shared seeds between a query and a reference sequence. A well-known tradeoff is that longer seeds offer fast searches but reduce sensitivity in ...
11912