Sign in

Paul Medvedev

@pashadag.bsky.social
1.9K followers 160 following 92 posts

Algorithmic Bioinformatics Researcher and Teacher. Posts about research results and educational/mentorship topics (for details, see bit.ly/380vX22).

PostsRepliesMedia
Reposted by Paul Medvedev
haonanwu.bsky.social @haonanwu.bsky.social · 11/07/2026
Our #ISMB2026 paper is now online! I’m excited to present it at HiTSeq on July 13. Many thanks to my advisor @pashadag.bsky.social. I’m also grateful to @iscb.bsky.social for awarding me the Conference Fellowship. If you’ll be at ISMB, feel free to stop by my talk and say hi. I’d love to connect!
academic.oup.com
The gift of novelty: repeat-robust k-mer-based estimators of mutation rates
AbstractMotivation. Estimating mutation rates between evolutionarily related sequences is a central problem in molecular evolution. Due to the rapid expans
1134
Paul Medvedev @pashadag.bsky.social · 05/05/2026
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
doi.org
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
13019
Reposted by Paul Medvedev
Kristoffer Sahlin @ksahlin.bsky.social · 16/03/2026
𝗣𝗼𝘀𝘁𝗱𝗼𝗰 𝗮𝗻𝗱 𝗣𝗵𝗗 𝗽𝗼𝘀𝗶𝘁𝗶𝗼𝗻𝘀 𝗶𝗻 𝗖𝗼𝗺𝗽𝘂𝘁𝗮𝘁𝗶𝗼𝗻𝗮𝗹 𝗚𝗲𝗻𝗼𝗺𝗶𝗰𝘀 / 𝗔𝗹𝗴𝗼𝗿𝗶𝘁𝗵𝗺𝗶𝗰 𝗕𝗶𝗼𝗶𝗻𝗳𝗼𝗿𝗺𝗮𝘁𝗶𝗰𝘀 I am currently recruiting for both: 🔹 Postdoc position su.varbi.com/what:job/job... 🔹 PhD position su.varbi.com/en/what:job/... Please share with anyone who might be interested!
su.varbi.com
Postdoktor i Beräkningsbiologi
Matematiska institutionen består av cirka 120 forskare, lärare och administrativ personal och är organiserad i tre huvudsakliga avdelningar: Matematik, Matematisk statistik och Beräkningsmatematik
21919
Reposted by Paul Medvedev
Rob Patro @robp.bsky.social · 17/02/2026
How would you design a *multithreaded*, *concurrent* & *dynamic* hash table if you are focused specifically on common k-mer workloads, where streaming query & insertion are common? Jamshed, Prashant and I explore this in kache-hash, a cache-friendly k-mer hash table! www.biorxiv.org/content/10.6...
biorxiv.org
02013
Reposted by Paul Medvedev
Rob Patro @robp.bsky.social · 16/02/2026
I feel it quite possible that those relying on AI from the very start may form a different skill set that will accelerate some software, but result in key regressions (without the expertise to address them) in other types. combine-lab.github.io/blog/2026/02... see the caveats section here…
combine-lab.github.io
COMBINE-lab - The skeptic’s guide to generative AI assisted coding
An easy-to-use, flexible website template for labs, with automatic citations, GitHub tag imports, pre-built components, and more.
1124
Reposted by Paul Medvedev
Sam Horsfield @samuelhorsfield.bsky.social · 07/02/2026
At long last, my final PhD chapter is out: we developed a novel evolutionary simulator of bacterial pangenomes, Pansim, fitting it to data from >600K genomes using a likelihood-free framework, PopPUNK-mod, to explore neutral and adaptive pangenome dynamics www.biorxiv.org/content/10.6...
biorxiv.org
24518
Reposted by Paul Medvedev
RECOMB Conference Series @recombconf.bsky.social · 05/02/2026
🚨UPCOMING DEADLINES🚨 RECOMB-CG: 13 February RECOMB-RSG: 15 February RECOMB-Privacy: 9 March RECOMB-Seq: 12 March (abstract registration) RECOMB-Arch: 12 March (abstract registration) RECOMB-Genetics: 13 March #RECOMB2026 #deadlines
065
Reposted by Paul Medvedev
Adam Phillippy @aphillippy.bsky.social · 02/02/2026
Time for a thread on our Christmas preprint “Origin and evolution of acrocentric chromosomes in human and great apes”. I had so much fun with this project and paper. It will be hard to summarize in a thread, but I’ll try www.biorxiv.org/content/10.6... [1/21]
14229
Reposted by Paul Medvedev
Antoine Limasset @npmalfoy.bsky.social · 04/02/2026
Preprint alert! arxiv.org/abs/2602.03525 TLDR: ZOR filters are STATIC filters with false positives. -Almost memory optimal: <1% overhead over the theoretical lower bound (!!!) -Fast queries: ~100 ns -Construction cannot fail A thread:
arxiv.org
ZOR filters: fast and smaller than fuse filters
Probabilistic membership filters support fast approximate membership queries with a controlled false-positive probability $\varepsilon$ and are widely used across storage, analytics, networking, and b...
13413
Reposted by Paul Medvedev
Michael Baym @baym.lol · 15/01/2026
If you are an Israeli PhD student and are interested in a postdoc at Harvard Medical (my lab included!), I strongly recommend looking into the Kalaniyot fellowship program, providing 2-3 years of full support: globalprograms.hms.harvard.edu/kalaniyot-hm...
globalprograms.hms.harvard.edu
Programs
154
Reposted by Paul Medvedev
Giulio Ermanno Pibiri @jermp.bsky.social · 10/12/2025
The 12th edition of the 2-days workshop “Data Structures in Bioinformatics” (DSB) will take place in Venice (Italy) on February 18-19th, 2026: dsb-meeting.github.io/DSB2026/
dsb-meeting.github.io
DSB 2026 Venice - February 18-19
Workshop Data Structures in Bioinformatics
1109
Paul Medvedev @pashadag.bsky.social · 03/12/2025
This thread gives really interesting and relevant history!
030
Reposted by Paul Medvedev
Ben Langmead @benlangmead.bsky.social · 03/12/2025
Kraken 2 (K2) community: we are giving more attention to our new `k2` wrapper, and a NEW functionality since 2.17.0 is: you can build several component K2 indexes, e.g. each covering a different Refseq database, and then query them all at once... github.com/DerrickWood/... 1/6
github.com
144
Reposted by Paul Medvedev
Antoine Limasset @npmalfoy.bsky.social · 02/12/2025
Preprint alert! We introduce new ideas to revisit the notion of sampling with window guarantees, also known as minimizers. A thread:
1156
Reposted by Paul Medvedev
Yaron Orenstein @yaronorenstein.bsky.social · 10/11/2025
Interested in a post-doc in Israel? The deadline for the Azrieli International Postdoctoral Fellowship is November 19. The fellowship offers generous funding for postdocs to conduct research in any academic discipline at eligible Israeli institutions: azrielifoundation.org/fellows/inte...
azrielifoundation.org
International Postdoctoral Fellowship - The Azrieli Foundation
The Azrieli Fellows Program is an elite group of academics who cultivate a network of leading professionals in Israel and around the world.
011
Reposted by Paul Medvedev
Sina Majidian @sinamajidian.bsky.social · 06/11/2025
Haonan Wu gives a talk on "A k-mer-based estimator of the substitution rate between repetitive sequences" www.biorxiv.org/content/10.1... This work tackles the issue of Mash which ignores repeats in the genome, providing better distance estimation #GI2025
194
Reposted by Paul Medvedev
Andre Kahles @akkah21.bsky.social · 08/10/2025
After years of research and continuous refinement, we’re thrilled to share that our paper on the MetaGraph framework — enabling Petabase-scale search across sequencing data — has been published today in Nature (www.nature.com/articles/s41...)
nature.com
Efficient and accurate search in petabase-scale sequence repositories - Nature
MetaGraph enables scalable indexing of large sets of DNA, RNA or protein sequences using annotated de Bruijn graphs.
33016
Reposted by Paul Medvedev
Rob Patro @robp.bsky.social · 08/10/2025
And it's posted! If you're interested and eligible, please consider applying through the UMD portal: umd.wd1.myworkdayjobs.com/en-US/UMCP/j.... If you're a PI working in algorithmic genomics (& you can recommend my lab to your top graduating students ;P), please let them know!
umd.wd1.myworkdayjobs.com
Postdoctoral Associate
Job Description Summary Organization's Summary Statement: The postdoctoral research associate is responsible for developing novel computational methodology for high-throughput sequence genomics tasks,...
02221
Reposted by Paul Medvedev
Ben Langmead @benlangmead.bsky.social · 07/10/2025
I've added 7 videos to my Burrows-Wheeler indexing playlist (www.youtube.com/playlist?lis...), rounding out the r-index series and adding a 5-part series on the move structure. Now 27 videos in that playlist. I aim to add videos on prefix-free parsing, PBWT, Wheeler languages/automata in the future.
youtube.com
Burrows-Wheeler Indexing - YouTube
Videos on : (a) the Burrows-Wheeler Transform (BWT), (b) the FM Index, which uses the BWT to construct a full-text index, (c) Wheeler graphs, (d) r-index, an...
26215
Reposted by Paul Medvedev
Roland Faure @rfaure.bsky.social · 03/10/2025
Our preprint on our new metagenomic HiFi assembler Alice is out 🥳 Based on a *new sketching method* (🧵1/6) 👉 Preprint www.biorxiv.org/content/10.1... 👉 Github github.com/rolandfaure/...
biorxiv.org
Alice: fast and haplotype-aware assembly of high-fidelity reads based on MSR sketching
We introduce Mapping-friendly Sequence Reduction (MSR) sketches, a sketching method for high-fidelity (HiFi) long reads, and Alice, an assembler that operates directly on these sketches. MSR produces ...
22521
Reposted by Paul Medvedev
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 01/10/2025
Alice: fast and haplotype-aware assembly of high-fidelity reads based on MSR sketching www.biorxiv.org/content/10.1101/202…
076
Reposted by Paul Medvedev
RECOMB Conference Series @recombconf.bsky.social · 26/09/2025
#RECOMB2026 will be in Thessaloniki, Greece on May 26-29, 2026. Satellites on May 24-25. Save the date! Το συνέδριο #RECOMB2026 θα πραγματοποιηθεί στη Θεσσαλονίκη, στις 26-29 Μαΐου 2026. Οι δορυφορικές εκδηλώσεις θα διεξαχθούν στις 24-25 Μαΐου 2026. Σημειώστε την ημερομηνία!
02313
Paul Medvedev @pashadag.bsky.social · 25/09/2025
If you're wondering why we're hosting the pre-print via dropbox, its because arXiv (and bioRxiv) did not accept it (because it is a review). Its a bit disconcerting, because a review is precisely the type of paper that would benefit a lot from pre-publication dissemination and feedback.
9133
Paul Medvedev @pashadag.bsky.social · 25/09/2025
Thank you folks for your feedback on our survey about Hash functions in genomic sequence analysis. We've updated the paper and you can see the new version here: tinyurl.com/4kk9ccmt.
tinyurl.com
Dropbox
0116
Reposted by Paul Medvedev
Sina Majidian @sinamajidian.bsky.social · 21/09/2025
Excited to share our EvANI benchmarking workflow, published in Briefings in Bioinformatics doi.org/10.1093/bib/... Computing average nucleotide identity (ANI) is neither conceptually nor computationally trivial. Its definition has evolved over years, with different meanings and assumptions (1/5)
Figure 1(A) ANI quantifies the similarity between two genomes. ANI can be defined as the number of aligned positions where the two aligned bases are identical, divided by the total number of aligned bases. Historically, ANI was calculated using a single gene family for multiple sequence alignment. Another approach finds orthologous genes between two genomes and reports the average similarity between their CDSs. This method was later extended to whole-genome alignment by identifying local alignments and excluding supplementary alignments with lower similarity. (B) Different ANI tools employ various approaches in calculating ANI values. ANIm, OrthoANI, and FastANI use aligners to identify homologous regions, whereas Mash uses k-mer hashing to estimate similarities. Only alignments with higher similarity represented by green arrows are included in ANI calculations, while red arrows, corresponding to paralogs, are excluded. (C) The proposed benchmarking method evaluates the performance of different tools using both real and simulated data. It assumes that more distantly related species on the phylogenetic tree should have lower ANI similarities. This is measured by calculating the statistics of Spearman rank correlation. We expect a negative correlation between ANI and the tree distance (scatter plot on the right).
https://academic.oup.com/bib/article/doi/10.1093/bib/bbaf267/8160681
13012
Reposted by Paul Medvedev
Jim Shaw @jimshaw.bsky.social · 08/09/2025
Preprint out for myloasm, our new nanopore / HiFi metagenome assembler! Nanopore's getting accurate, but 1. Can this lead to better metagenome assemblies? 2. How, algorithmically, to leverage them? with co-author Max Marin @mgmarin.bsky.social, supervised by Heng Li @lh3lh3.bsky.social 1 / N
511380
Reposted by Paul Medvedev
Rayan Chikhi @rayanchikhi.bsky.social · 03/09/2025
🌎👩‍🔬 For 15+ years biology has accumulated petabytes (million gigabytes) of🧬DNA sequencing data🧬 from the far reaches of our planet.🦠🍄🌵 Logan now democratizes efficient access to the world’s most comprehensive genetics dataset. Free and open. doi.org/10.1101/2024...
3218118
Reposted by Paul Medvedev
Tobias Marschall @tobiasmar.bsky.social · 23/07/2025
Two papers in today's issue of @nature.com ‬: 1) we assemble 65 genomes to near completion, including centromeres and the MHC. tinyurl.com/3huhax6w. 2) we sequence 1,019 genomes from the 1kGP with long reads, revealing SVs down to low allele frequencies tinyurl.com/wbx3we9x.
tinyurl.com
Complex genetic variation in nearly complete human genomes - Nature
Using sequencing and haplotype-resolved assembly of 65&nbsp;diverse human genomes, complex regions including the major histocompatibility complex and centromeres are analysed.
15424
Reposted by Paul Medvedev
Sebastian Deorowicz @sdeorowicz.bsky.social · 19/07/2025
Interested in a tool that aligns millions of proteins in minutes with quality similar to or better than the state-of-the-art utilities? Please take a look at our FAMSA2 paper: www.biorxiv.org/content/10.1... and GH repo: github.com/refresh-bio/...
biorxiv.org
FAMSA2 enables accurate multiple sequence alignment at protein-universe scale
We introduce FAMSA2, an algorithm that produces high-accuracy multiple protein sequence alignments with unprecedented speed. Across structural, phylogenetic, and functional benchmarks, FAMSA2 matches ...
34928
Reposted by Paul Medvedev
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 26/07/2025
Sassy: Searching Short DNA Strings in the 2020s www.biorxiv.org/content/10.1101/202…
073
Reposted by Paul Medvedev
Institut Pasteur | 130 years of biomedical research @pasteur.fr · 24/07/2025
Congratulations to Rayan Chiki, (Institut Pasteur) head of the “Sequence Bioinformatics” unit, for securing the ERC Proof of Concept 2025 for his project ENZYMINER! 👏 ‪@rayan.chiki.bsky.social #Bioinformatics
46013
Reposted by Paul Medvedev
Richard Sever @richardsever.bsky.social · 11/07/2025
After Tim Hunt won the Nobel, he said, "We do science because we like discovering things about the world...and then boasting about what we found". Any one individual can argue about their own motivation, but it would naive to dispute that's an accurate description of many people.
2154
Reposted by Paul Medvedev
Antoine Limasset @npmalfoy.bsky.social · 05/07/2025
Paper Alert! Our preprint on the K2R index, being able to efficiently associate kmers to the reads containing them is finally out there! A thread! academic.oup.com/bioinformati...
academic.oup.com
K2R: Tinted de Bruijn graphs implementation for efficient read extraction from sequencing datasets
AbstractSummary. Biological sequence analysis often relies on reference genomes, but producing accurate assemblies remains a challenge. As a result, de nov
1179
Reposted by Paul Medvedev
Antoine Limasset @npmalfoy.bsky.social · 03/07/2025
Paper alert! We present Oreo a tools that reorder long reads datasets in a way to compress them efficiently with ANY universal compressor like gz, zstd, xz ... TLDR: You can get state of the art compression WITHOUT a dedicated compressor/decompressor! academic.oup.com/bioinformati... A thread!
academic.oup.com
OReO: optimizing read order for practical compression
AbstractMotivation. Recent advances in high-throughput and third-generation sequencing technologies have created significant challenges in storing and mana
12318
Reposted by Paul Medvedev
Kristoffer Sahlin @ksahlin.bsky.social · 02/07/2025
I worked with Thomas during a three months research visit during his PhD, and it resulted in a paper in NAR. I highly recommend him. doi.org/10.1093/nar/...
doi.org
Improved sub-genomic RNA prediction with the ARTIC protocol
Abstract. Viral subgenomic RNA (sgRNA) plays a major role in SARS-COV2’s replication, pathogenicity, and evolution. Recent sequencing protocols, such as th
088
Reposted by Paul Medvedev
Antoine Limasset @npmalfoy.bsky.social · 02/07/2025
Preprint alert! We present K2Rmini, an ultra-fast, grep-like tool that extracts sequences of interest from FASTA/FASTQ files based on their k-mer content. www.biorxiv.org/content/10.1... A thread
biorxiv.org
Accelerating k-mer-based sequence filtering
The exponential growth of global sequencing data repositories presents both analytical challenges and opportunities. While k - mer-based indexing has improved scalability over traditional alignment fo...
13819
Reposted by Paul Medvedev
Rob Patro @robp.bsky.social · 25/06/2025
🖥️🧬 WABI '25 will not only have excellent keynotes, but an exciting program of papers. The titles and abstracts of all accepted WABI '25 papers are now available on the conference website (wabiconf.github.io/2025/talks/). I'm looking forward to seeing these talks!
wabiconf.github.io
Talks
WABI Conference on Algorithms in Bioinformatics
193
Paul Medvedev @pashadag.bsky.social · 25/06/2025
🧵1/n Estimating mutation rates using k-mers is fast—but what happens when repeats dominate the genome? In a new preprint, Haonan Wu, Antonio Blanca, and myself propose a *repeat-aware* estimator that's accurate even in centromeres.
13014
Reposted by Paul Medvedev
Rob Patro @robp.bsky.social · 23/06/2025
🚀 We are thrilled to introduce QCatch — a fast, command-line QC reporting tool built for alevin-fry & simpleaf single-cell data! Led by @ygao61.bsky.social & in collaboration with Dongze He 🧬🖥️ . The Preprint 📖 is available at bit.ly/4neSznl. Read more below: 1/3
github.com
GitHub - COMBINE-lab/QCatch: Quality Control downstream of alevin-fry / simpleaf
Quality Control downstream of alevin-fry / simpleaf - COMBINE-lab/QCatch
1216
Reposted by Paul Medvedev
Antoine Limasset @npmalfoy.bsky.social · 19/06/2025
Preprint alert! 🦌 Our new abundance index, REINDEER2, is out! It's cheap to build and update, offers tunable abundance precision at kmer level, and delivers very high query throughput. Short thread! www.biorxiv.org/content/10.1... github.com/Yohan-Hernan...
biorxiv.org
12313
Reposted by Paul Medvedev
Ragnar {Groot Koerkamp} @curiouscoding.nl · 20/06/2025
Also: what are the bottlenecks in your data processing? Specifically, I'm looking for reasonably well defined & understood and widely used methods that could use a fresh high-throughput implementation. Stuff like sketching, maybe assembly, ... Surely, many pipelines could be sped up 10x ;)
243
Paul Medvedev @pashadag.bsky.social · 12/06/2025
1/4 Hash functions in genomic sequence analysis (tinyurl.com/4kk9ccmt) : a new survey written together with Ke Chen, Xiang Li, Qian Shi, and Mingfu Shao. Before submitting it, we are posting it online to get feedback from the community.
tinyurl.com
Dropbox
12816
Reposted by Paul Medvedev
Rayan Chikhi @rayanchikhi.bsky.social · 03/06/2025
Slides from my talk (with @kamilsjaron.bsky.social) on an history of k-mers in bioinformatics: rayan.chikhi.name/pdf/2025-kme...
14424
Reposted by Paul Medvedev
Jouni Sirén @jltsiren.bsky.social · 15/05/2025
A new preprint on indexing pangenome graphs using an FM-index of the haplotypes and a tag array. Joint work with Parsa Eskandar and @benedictpaten.bsky.social.
biorxiv.org
Lossless Pangenome Indexing Using Tag Arrays
Pangenome graphs represent the genomic variation by encoding multiple haplotypes within a unified graph structure. However, efficient and lossless indexing of such structures remains challenging due t...
13615
Reposted by Paul Medvedev
Kristoffer Sahlin @ksahlin.bsky.social · 08/05/2025
@alexanderjpetri.bsky.social's isONclust3 algorithm is now published doi.org/10.1093/bioi.... isONclust3 performs de novo clustering of long-read cDNA sequencing data. A key step in reference-free transcriptome analysis.
doi.org
De novo clustering of large long-read transcriptome datasets with isONclust3
AbstractMotivation. Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription proc
1116
Reposted by Paul Medvedev
Rob Patro @robp.bsky.social · 07/05/2025
The deadline for WABI 2025 has been extended (but is still rapidly approaching) wabiconf.github.io/2025/ * abstract deadline: May 12 (AoE) * paper deadline: May 15 (AoE) Consider submitting your exciting algorithmic bioinformatics work to the WABI conference!
wabiconf.github.io
WABI 2025
WABI Conference on Algorithms in Bioinformatics
01011