Sign in

Roland Faure

@rfaure.bsky.social
73 followers 133 following 22 posts

Sequence bioinfomatician, algorithms, methods. Postdoc in Institut Pasteur in Rayan Chikhi's lab

PostsRepliesMedia
Reposted by Roland Faure
Mel Andrews @bayesianboy.bsky.social · 17/06/2026
When utilized in literature review, LLMs consistently 1. fail to mention female authors in female-led literatures, 2. insist that men are more influential or more heavily cited when this is contradicted by objective citation counts, and 3. attribute women’s work to hallucinated male scholars.
8153382302
Roland Faure @rfaure.bsky.social · 04/06/2026
Most modern SNP callers use neural networks, but this software takes a different path—no NN at all! That leaves room for big improvements by blending both approaches.
010
Roland Faure @rfaure.bsky.social · 04/06/2026
We use a new statistical test, which I find beautiful in its simplicity. The idea is to exploit the length of the reads to look at multiple loci at the same time. The p-value of a *group of* variants being due only to sequencing errors is 0.1^(ab)*C(m,b)*C(n,a), C denoting binomial coefficients.
130
Roland Faure @rfaure.bsky.social · 04/06/2026
The key is distinguishing SNPs from sequencing errors. In metagenomic data, you cannot rely on "50% of the reads should carry the variant", as you could with human/diploid reads.
100
Roland Faure @rfaure.bsky.social · 04/06/2026
SNooPy calls much more SNPs than other tools, including DeepVariant. This is especially true for datasets with high coverage and several strains of the same species, typically human gut metagenomes sequencing 🦠
100
Roland Faure @rfaure.bsky.social · 04/06/2026
Glad to announce that our new long-read metagenomic SNP caller, SNooPy, is published in NAR! Existing long-read SNP callers are not designed for metagenomic data, check out our new solution 👉https://academic.oup.com/nar/article/54/10/gkag556/8700491 @narjournal.bsky.social
1177
Reposted by Roland Faure
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 22/05/2026
Min-frame transformation enables more sensitive viral genome alignment www.biorxiv.org/content/10.64898/20…
284
Reposted by Roland Faure
Paul Medvedev @pashadag.bsky.social · 05/05/2026
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
doi.org
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
13019
Reposted by Roland Faure
Igor Martayan @imartayan.bsky.social · 21/04/2026
New blog post! I use ntHash all the time to hash k-mers, yet it turns out it has some unexpected flaws (collision propagation, bias on leading zeros...). The good news: each of them can be fixed! igor.martayan.org/posts/breaki...
igor.martayan.org
Breaking ntHash (to better fix it)
NtHash is a popular method for hashing k-mers in bioinformatics, yet it has some surprising flaws. In this post, I walk through a few of them, and show that they can arise naturally, without an advers...
13016
Reposted by Roland Faure
Zamin Iqbal @zaminiqbal.bsky.social · 27/03/2026
what's the current status of HERRO and (bacterial) nanopore? Last I saw was Ryan Wick's blog evaluating it. How many people now use it (or did it actually get folded into a basecaller, or did it get dropped in favour of something else)?
111
Reposted by Roland Faure
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 19/03/2026
Super Bloom: Fast and precise filter for streaming k-mer queries www.biorxiv.org/content/10.64898/20…
02213
Reposted by Roland Faure
Société Française de BioInformatique @handle.invalid · 02/02/2026
📌 Les #soumissions pour @jobim2026.bsky.social sont ouvertes jusqu'au 15/03/26. 📝 Soumission de travaux originaux, articles longs (+ PCI), activités de plateformes et de service, posters et démonstrations 📍 Plus d’infos sur : jobim2026.sfbi.fr #JOBIM2026 #bioinfo #Strasbourg
032
Reposted by Roland Faure
Torsten Seemann @torstenseemann.bsky.social · 20/01/2026
🗜️⚡ If you use gzip/gunzip a lot in your pipelines, switch to the faster"libdeflate" versions instead! They use modern CPU capabilities to achieve a 2-3x speedup. libdeflate is in conda, and "libdeflate-gzip" and "libdeflate-gunzip" are drop-in replacements. #unix github.com/ebiggers/lib...
github.com
GitHub - ebiggers/libdeflate: Heavily optimized library for DEFLATE/zlib/gzip compression and decompression
Heavily optimized library for DEFLATE/zlib/gzip compression and decompression - ebiggers/libdeflate
17123
Reposted by Roland Faure
Zamin Iqbal @zaminiqbal.bsky.social · 20/12/2025
"..based on a common wavefront design that can be adapted to support a variety of dynamic programming algorithms: local, global, and semi-global alignment of genomic and protein sequences with a variety of commonly used scoring schemes" from @martinsteinegger.bsky.social andco
0146
Reposted by Roland Faure
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 11/12/2025
Inverted colored de Bruijn Graph for practical kmer sets storage www.biorxiv.org/content/10.64898/20…
082
Reposted by Roland Faure
Giulio Ermanno Pibiri @jermp.bsky.social · 10/12/2025
The 12th edition of the 2-days workshop “Data Structures in Bioinformatics” (DSB) will take place in Venice (Italy) on February 18-19th, 2026: dsb-meeting.github.io/DSB2026/
dsb-meeting.github.io
DSB 2026 Venice - February 18-19
Workshop Data Structures in Bioinformatics
1109
Reposted by Roland Faure
Karel Břinda @brinda.eu · 05/12/2025
1/9 Just out: k-mer indexes are the backbone of fast search in genomic data, but many degrade under small k, subsampling, or high diversity. With Ondřej Sladký and @pavelvesely.bsky.social we asked: can we build one that works efficiently for any k-mer set?
12713
Roland Faure @rfaure.bsky.social · 04/12/2025
Sorry for the first figure, got a problem of background, here it is:
000
Roland Faure @rfaure.bsky.social · 04/12/2025
Coming up with a name was pretty hard, we had a lot of good candidates. We settled on SNooPy, which is a reference to SNPs and to the fact that the tool is 100% python. We thought about metaSNooPy but this went too far 😅. The github: github.com/RolandFaure/SNooPy
github.com
GitHub - RolandFaure/SNooPy: metagenomic SNP caller
metagenomic SNP caller. Contribute to RolandFaure/SNooPy development by creating an account on GitHub.
210
Roland Faure @rfaure.bsky.social · 04/12/2025
Tests show that: 1/ SNooPy has the best recall in our tests 2/ Using genomic long-read SNP callers does not (always) work well: most tools have very low recall, but DeepVariant perform much better than other tested methods 3/ The recall of all tools is still far from 100%
100
Roland Faure @rfaure.bsky.social · 04/12/2025
We propose a new statistical framework. The idea to distinguish artefacts from SNPs is to look at several loci simultaneously: artefacts will occur on random reads, while SNPs will occur systematically on the reads that come from the same strain.
120
Roland Faure @rfaure.bsky.social · 04/12/2025
Existing long-read SNP callers (DeepVariant, longshot...) have been developed for diploid genomes. Deep-learning methods are trained on [human] genomic data. Statistical methods contain assumption that do not hold for metagenomics.
120
Roland Faure @rfaure.bsky.social · 04/12/2025
Preprint out! Check out our new long-read metagenomic SNP-caller, SNooPy 😀. Work with Chris Quince. Thread 🧵 👉 www.biorxiv.org/content/10.6...
1138
Reposted by Roland Faure
Antoine Limasset @npmalfoy.bsky.social · 01/12/2025
Preprint Alert! We present new strategies to accelerate large-scale document comparison using MinHash-like sketches. A thread:
1127
Roland Faure @rfaure.bsky.social · 03/10/2025
🧵6/ 6 Since MSRs sketches are sequence, they are super easy to use. I think they could be useful for many other problems, e.g. SNP calling, pangenome graphs, indexing, etc.
010
Roland Faure @rfaure.bsky.social · 03/10/2025
🧵5/6 The sketching makes assembly extremely fast: a gut metagenome sample of 138Gbp of sequencing data was assembled in less that 2h and 10G RAM on 8 threads ⚡. And thanks to MSRs, *highly similar strains are not collapsed*
111
Roland Faure @rfaure.bsky.social · 03/10/2025
🧵4/6 Two key properties that make MSRs sketches really cool: 👉 They are alignable sequences: you can just feed them in existing assembler 👉 MSR sketches can *keep all the SNPs*, i.e. two highly similar sequences are (almost) always reduced to different sketches -> useful to separate similar strains
120
Roland Faure @rfaure.bsky.social · 03/10/2025
🧵3/ 6 MSRs have been defined by @lblassel.bsky.social @rayanchikhi.bsky.social and @pashadag.bsky.social in pmc.ncbi.nlm.nih.gov/articles/PMC.... Take a sequence, a value of k, and stream all k-mers through a function that output either a base or the empty character, and you got your sketch
110
Roland Faure @rfaure.bsky.social · 03/10/2025
🧵2/6 Conceptually, the assembler is on the same lines as metaMDBG: 1. sketching reads 2. assembly procedure on the sketches 3. reversing to base-space to obtain the final assembly The main difference is the sketching scheme: we introduce *Mapping-friendly Sequence Reductions (MSR) sketching*
110
Roland Faure @rfaure.bsky.social · 03/10/2025
Our preprint on our new metagenomic HiFi assembler Alice is out 🥳 Based on a *new sketching method* (🧵1/6) 👉 Preprint www.biorxiv.org/content/10.1... 👉 Github github.com/rolandfaure/...
biorxiv.org
Alice: fast and haplotype-aware assembly of high-fidelity reads based on MSR sketching
We introduce Mapping-friendly Sequence Reduction (MSR) sketches, a sketching method for high-fidelity (HiFi) long reads, and Alice, an assembler that operates directly on these sketches. MSR produces ...
22521
Reposted by Roland Faure
Rayan Chikhi @rayanchikhi.bsky.social · 03/09/2025
🌎👩‍🔬 For 15+ years biology has accumulated petabytes (million gigabytes) of🧬DNA sequencing data🧬 from the far reaches of our planet.🦠🍄🌵 Logan now democratizes efficient access to the world’s most comprehensive genetics dataset. Free and open. doi.org/10.1101/2024...
3218118
Reposted by Roland Faure
Jim Shaw @jimshaw.bsky.social · 08/09/2025
Preprint out for myloasm, our new nanopore / HiFi metagenome assembler! Nanopore's getting accurate, but 1. Can this lead to better metagenome assemblies? 2. How, algorithmically, to leverage them? with co-author Max Marin @mgmarin.bsky.social, supervised by Heng Li @lh3lh3.bsky.social 1 / N
511380
Roland Faure @rfaure.bsky.social · 16/05/2025
Congrats! Nice results 🎉
100
Reposted by Roland Faure
Josipa Lipovac @jlipovac.bsky.social · 16/05/2025
I am happy to share our new preprint introducing MADRe - a pipeline for Metagenomic Assembly-Driven Database Reduction, enabling accurate and computationally efficient strain-level metagenomic classification. 🔗https://www.biorxiv.org/content/10.1101/2025.05.12.653324v1 1/9
3158
Reposted by Roland Faure
Camille Marchet ⚡ @camillemrcht.bsky.social · 24/04/2025
Starting #RECOMBseq with @rayanchikhi.bsky.social 's keynote. Here stressing our responsibility as scientists to enable access to a common good: genomic data
13010
Reposted by Roland Faure
Michael Baym @baym.lol · 09/04/2025
Side note: you could, speaking purely theoretically, also fit every microbe onto an SD card, which is within the weight limit for a carrier pigeon. For some distances, it would be faster than the internet for transmitting sequence libraries 7/
44511
Reposted by Roland Faure
Zamin Iqbal @zaminiqbal.bsky.social · 09/04/2025
So glad this is finally out. The method has been instrumental in allowing us to compress the AllTheBacteria data - ~2 million bacterial genomes shrink from 3Terabytes (gzipped) to 100Gb using phylogenetic compression. Great work by @brinda.eu
412652
Reposted by Roland Faure
Ryan Wick @rrwick.bsky.social · 27/03/2025
Do you (like me) create a bunch of conda environments, then later forget what they're for, when they were last updated, or which tools are in them? If so, you might this little project: github.com/rrwick/conda...
github.com
GitHub - rrwick/condaenvlist: a simple tool for listing conda environments with descriptions
a simple tool for listing conda environments with descriptions - rrwick/condaenvlist
17840
Roland Faure @rfaure.bsky.social · 07/03/2025
So glad to have participated in #DSB2025, what a great workshop! For some mysterious reason it was the first time I attended after 3 years of sequence research. Thanks to all participants & organizers 😃
120
Reposted by Roland Faure
Igor Martayan @imartayan.bsky.social · 13/12/2024
Ragnar's made some incredible optimizations on the computation of minimizers, can't wait to see how these improvements will benefit bioinfo tools!
042
Roland Faure @rfaure.bsky.social · 13/12/2024
Really cool work!
010
Reposted by Roland Faure
Pierre Peterlongo @pierrepeterlongo.bsky.social · 04/12/2024
Amazing ideas here www.biorxiv.org/content/bior... from @yoann.bsky.social and collaborators. Reorganize minimizers to allow kmers dichotomic search. That's brilliant. #bioinformatics 🧬🖥️
Toy example of the AAAAAAA bucket associated to four super-k-mer turned into their interleaved representation.
2235
Roland Faure @rfaure.bsky.social · 04/12/2024
So glad to have successfully defended my Ph.D. last week 😀 Work on producing haplotype-resolved metagenomic assemblies using noisy long reads (HairSplitter) and high-fidelity long reads (Alice assembler, unpublished yet). Thanks to my advisors Dominique Lavenier and Jean-François Flot ❤️
193
Roland Faure @rfaure.bsky.social · 02/12/2024
Congrats @firtinac.bsky.social ! I enjoyed thouroughly reading the BLEND paper 😄
110