Sign in

Pierre Peterlongo

@pierrepeterlongo.bsky.social
391 followers 89 following 68 posts

Inria Senior researcher. Head of the team.inria.fr/genscale at Inria and Irisa. Algorithmics for sequencing data analyses, genomics and metagenomics.

PostsRepliesMedia
Pierre Peterlongo @pierrepeterlongo.bsky.social · 09/07/2026
7/ Work by Sébastien Bellenous, Kamil S. Jaron @kamilsjaron.bsky.social (Genscale, Univ. Rennes / Inria / CNRS / IRISA + Wellcome Sanger Institute). Paper (JOSS submission draft): github.com/sebllns/kmhe...
000
Pierre Peterlongo @pierrepeterlongo.bsky.social · 09/07/2026
6/ kmhelpers wraps kmindex (Lemane et al.) and kmtricks, is Python ≥3.8, Conda-installable, with a CLI (Click) and matching Python API. GPLv3, open source. Code: github.com/sebllns/kmhe... Docs: sebllns.github.io/kmhelpers/
github.com
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 09/07/2026
5/ Real-world use: • Tree of Life @ Wellcome Sanger Institute: 7,394 samples indexed, now powering sample-swap detection across the Darwin Tree of Life project. • Logan-Search (~50 petabases from SRA): next update will use kmhelpers to re-optimize the index.
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 09/07/2026
4/ It also handles indexes grow. As new samples arrive constantly kmhelpers supports incremental updates — and `manage` federates several independently built indexes into one logical, queryable resource.
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 09/07/2026
3/ kmhelpers automates the whole chain: list → profile → compose → plan → apply → query It picks BF sizes and sample-to-sub-index assignments that minimize total index size under a false-positive-rate constraint, estimates disk/RAM upfront, and generates ready-to-run pipelines.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 09/07/2026
2/ The problem: k-mer indexes (à la BIGSI, COBS, kmindex...) let you search for a sequence across thousands of samples using Bloom filters. But building an *efficient* one means hand-tuning BF sizes, memory, data distribution & sub-index grouping — for every sample, every time.
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 09/07/2026
1/ 🧬 & 🖥️ New tool: kmhelpers — a Python toolkit dedicated to kmindex, automating the building, updating & querying of large-scale k-mer genomic indexes. Turning raw sequencing samples into a fast, size-optimized search engine — without needing to be a Bloom-filter expert. 🧵
141
Pierre Peterlongo @pierrepeterlongo.bsky.social · 08/07/2026
Might be useful. Here is a short blog entry describing a tool that annotates a PDF using a text file. I couldn't find anything that simply does this single task, so I created this. Blog post: pierrepeterlongo.github.io Git repo: github.com/pierrepeterl...
062
Reposted by Pierre Peterlongo
Heng Li @lh3lh3.bsky.social · 28/06/2026
minibwa v0.3 released with a few minor bug fixes and two missing bwa-mem features (XA tag for secondary hits and option -H to inject header lines). Also added the "mem" subcommand to mimic "bwa mem" CLI to some extent. github.com/lh3/minibwa/...
github.com
Release Minibwa-0.3 (r391) · lh3/minibwa
Notable changes: New feature: added the mem subcommand to mimic the bwa-mem command-line interface (CLI). Most input/output options and commonly used options are retained; unsupported or incompat...
03613
Reposted by Pierre Peterlongo
Ben Langmead @benlangmead.bsky.social · 22/06/2026
Movi 2 has appeared (as an advance article) in Bioinformatics 🧬 Faster, leaner pangenome queries — half the memory of Movi 1, ~30% faster. Paper: academic.oup.com/bioinformati... Code: github.com/mohsenzakeri/Movi (1/6)
academic.oup.com
Validate User
34317
Pierre Peterlongo @pierrepeterlongo.bsky.social · 11/06/2026
Happy to see that K2Rmini was recommended today by PCI Mathematical & Computational Biology. "quickly evaluate whether an arbitrary sequence has a number of k-mer [of interest] matches above or below a threshold." by @imartayan.bsky.social and colleagues: www.biorxiv.org/content/10.1...
074
Reposted by Pierre Peterlongo
Federico Lopez @flopezos.bsky.social · 09/06/2026
📄 OrthoFinder: improved phylogenetic orthology inference with enhanced accuracy and scalability 🧬🖥️ ⬇️ www.nature.com/articles/s41...
nature.com
OrthoFinder: improved phylogenetic orthology inference with enhanced accuracy and scalability - Nature Methods
The updated OrthoFinder v3 software boosts accuracy and scalability in phylogenetic orthology inference with massive and diverse datasets.
0124
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
🙏 Thanks: rayanchikhi.bsky.social tlemane.bsky.social rnalab.bsky.socialand the whole logan team! Paper: www.biorxiv.org/content/10.1...
rayanchikhi.bsky.social
Rayan Chikhi (@rayanchikhi.bsky.social)
000
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
🛠️ logan_blaster: a command-line tool that automates the alignment step locally, with no timeout constraints. github.com/pierrepeterl... conda install -c bioconda logan_blaster
130
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
🔍 Raising your attention to the BLAST-like alignment tab: you can perform proper alignment of the query against assembled contigs or unitigs directly in the browser. It is limited to one sample at a time and may timeout on large inputs. Which is why we built the next tool.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
9/12 🎬 A 5-min tutorial covers both basic usage and deeper features for new and returning users. See docs.logan-search.org
120
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
8/12 🤖 The AI summary of query results has been improved and now provides much more information.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
7/12 📄 When publicly available, the scientific paper(s) associated with matched samples are provided.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
6/12 📊 Results now come with an e-value, a p-value, and an ANI estimation.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
5/12 🧬 Matches to RefSeq reference genomes are now included in results.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
4/12 🎨 The interface got a refresh: new query history, email input is optional.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
3/12 ⚡ New fast mode: queries 99.5% of non-viral SRA samples, around 2x faster.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
2/12 📏 Query length limit raised: 1000 → 2500 bp. One of the most-requested changes.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
🌎 🧬 🖥️ logan-search.org the tool to query all SRA sequences (Dec 2023 snapshot) — just got major updates. Here's what's new. 🧵 1/12
13626
Pierre Peterlongo @pierrepeterlongo.bsky.social · 29/05/2026
—8/8 🙏 Thanks to tlemane.bsky.social for key contributions, and the whole logan team!
tlemane.bsky.social
Téo Lemane (@tlemane.bsky.social)
000
Pierre Peterlongo @pierrepeterlongo.bsky.social · 29/05/2026
📊 Synthesis files show per-position query coverage using letter codes (a–z, then Z for 26+ contigs) — a compact view of how contigs pile up along your query sequence.
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 29/05/2026
📁 Output structure: alignments/ — BLAST results per sample logan_data/ — downloaded contigs or unitigs and recruited sequences failed_accessions.txt — samples that couldn’t be downloaded
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 29/05/2026
—5/8 🎛️ Key options: Using a logan-search session ID (received by email) -s Logan-search session ID Or for any sequence with any accessions list: -q query FASTA -a accessions file General options: -u align against unitigs instead of contigs -l limit number of accessions processed
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 29/05/2026
—4/8 📦 Install: conda install -c conda-forge -c bioconda logan_blaster
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 29/05/2026
—3/8 ⚙️ Under the hood, three steps per accession: Download Logan contigs or unitigs Recruit contigs or unitigs sharing k-mers with your query (uses github.com/pierrepeterlongo/back_to_sequences) Run local BLAST between query and recruited contigs or unitigs
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 29/05/2026
🔗 It picks up where the Logan-search blast-like browser tab leaves off. Two modes: provide a Logan-search session ID directly or bring your own query FASTA + accessions list
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 29/05/2026
1/8 🛠️ Introducing logan_blaster — a command-line tool to locally align your sequences against Logan contigs or unitigs. github.com/pierrepeterlongo/logan_blaster
github.com
GitHub - pierrepeterlongo/logan_blaster: Given either a kmviz id (obtained from a Logan-Search query) or a query (fasta sequence), and a file containing a list of SRA accessions (provided or not by Lo...
Given either a kmviz id (obtained from a Logan-Search query) or a query (fasta sequence), and a file containing a list of SRA accessions (provided or not by Logan-Search results) run a local blast ...
184
Reposted by Pierre Peterlongo
Bioinformatics Advances @bioinfoadv.bsky.social · 07/05/2026
🔍 New paper in Bioinformatics Advances: "Kaminari: A frugal colored index for approximate k-mer queries"  Read it here: doi.org/10.1093/bioadv/vbag120 Authors include: @yhhshb.bsky.social, @yoann.bsky.social, @robp.bsky.social, @pierrepeterlongo.bsky.social, @jermp.bsky.social
21510
Pierre Peterlongo @pierrepeterlongo.bsky.social · 21/11/2025
Good Friday Evening news: we updated back_to_sequences (find the origin of kmers) - Faster - Can consider multiline fasta files - Much easier installation: see github.com/pierrepeterl...
041
Reposted by Pierre Peterlongo
Rob Patro @robp.bsky.social · 09/10/2025
The Metagraph paper is out in Nature; it showed up in my feeds today! Congratulations to Mikhail Karasikov, @gxxxr.bsky.social, @akkah21.bsky.social and all of the other authors (whom I'd love to follow on Bluesky if I can find you ;P) www.nature.com/articles/s41...
nature.com
Efficient and accurate search in petabase-scale sequence repositories - Nature
MetaGraph enables scalable indexing of large sets of DNA, RNA or protein sequences using annotated de Bruijn graphs.
13615
Reposted by Pierre Peterlongo
Jim Shaw @jimshaw.bsky.social · 08/09/2025
Preprint out for myloasm, our new nanopore / HiFi metagenome assembler! Nanopore's getting accurate, but 1. Can this lead to better metagenome assemblies? 2. How, algorithmically, to leverage them? with co-author Max Marin @mgmarin.bsky.social, supervised by Heng Li @lh3lh3.bsky.social 1 / N
511380
Pierre Peterlongo @pierrepeterlongo.bsky.social · 04/09/2025
❗ I clearly consider this result as THE most important result achieved over this last decade for exploiting and democratizing genomic data. I think there will be a "before" and an "after" logan and logan-search github.com/IndexThePlan... logan-search.org Have a look at this thread
0103
Pierre Peterlongo @pierrepeterlongo.bsky.social · 27/05/2025
🤝 Amazing collaboration with @jermp.bsky.social, @yhhshb.bsky.social, @robp.bsky.social, Victor Levallois, and Bertrand Le Gal, and the help of ‪@yoann.bsky.social‬. 8/8
030
Pierre Peterlongo @pierrepeterlongo.bsky.social · 27/05/2025
🌊 On metagenomic data, other tools such as kmindex are good alternatives. At the same time, Kaminari consistently ranks as one of the fastest tools across all data types, generating the smallest indexes (or the lower FPR). 7/8
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 27/05/2025
💾 For fixed False Positive rates, it uses up to 37x less space than COBS while being an order of magnitude faster to build and query. 6/8
120
Pierre Peterlongo @pierrepeterlongo.bsky.social · 27/05/2025
📊 Experimental results show Kaminari's superiority in index size and query performance across various genomic datasets. 5/8
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 27/05/2025
🧬 Kaminari's design leverages properties of k-mer minimizers for compact space and fast query time, as inspired by the techniques proposed in Fulgor. 4/8
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 27/05/2025
💻 We implemented Kaminari in C++17, available under the MIT license at github.com/yhhshb/kaminari. Additional results and reproducibility info at github.com/vicLeva/benchmarks_kaminari. 3/8
github.com
GitHub - yhhshb/kaminari: 雷 - kaminari (thunder/lightning)
雷 - kaminari (thunder/lightning). Contribute to yhhshb/kaminari development by creating an account on GitHub.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 27/05/2025
🔍 Key findings include: - Use of minimizers and integer compression for indexing. - Lower memory footprint and faster query times. - Minimal impact of false positives on result ranking, using the Rank-Biased Overlap (RBO) metric. 2/8
120
Pierre Peterlongo @pierrepeterlongo.bsky.social · 27/05/2025
📜 Excited to share insights from our recent paper: "Kaminari: a resource-frugal index for approximate colored k-mer queries". The study aims to efficiently identify documents containing a query string, focusing on DNA strings. www.biorxiv.org/content/10.1... 🧬 🖥️ 1/8
12516
Pierre Peterlongo @pierrepeterlongo.bsky.social · 25/03/2025
Thanks guys for your precious feedback. I modified the code accordingly.
010
Pierre Peterlongo @pierrepeterlongo.bsky.social · 25/03/2025
Hi @imartayan.bsky.social I wanted to run distinct-kmers, but I faced limitations as my input data contains non-ACGTacgt characters. Thus I created this github.com/pierrepeterl... (again extremely simple)
github.com
GitHub - pierrepeterlongo/hyperloglog_kmer_counter
Contribute to pierrepeterlongo/hyperloglog_kmer_counter development by creating an account on GitHub.
120
Pierre Peterlongo @pierrepeterlongo.bsky.social · 25/03/2025
That's correct. I just created this github.com/pierrepeterl... This is yet a new hll kmer counter, but hyper simple. And I did not find a way to accumulate the kmer counts for several input datasets.
github.com
GitHub - pierrepeterlongo/hyperloglog_kmer_counter
Contribute to pierrepeterlongo/hyperloglog_kmer_counter development by creating an account on GitHub.
100
Pierre Peterlongo @pierrepeterlongo.bsky.social · 24/03/2025
@imartayan.bsky.social I needed a version of distinct_kmers for multiple fasta/fastq. I created this fork github.com/pierrepeterl... I'm almost ashamed that this code modification is public, but maybe it can be useful.
github.com
GitHub - pierrepeterlongo/distinct-kmers: How many distinct k-mers are there in a sequence?
How many distinct k-mers are there in a sequence? Contribute to pierrepeterlongo/distinct-kmers development by creating an account on GitHub.
110
Pierre Peterlongo @pierrepeterlongo.bsky.social · 20/03/2025
I added the notion of insertion order (mentioning your name). However, I don't get the point of the mergeability issue.
100