Sign in

Igor Martayan

@imartayan.bsky.social
830 followers 370 following 96 posts

PhD in algorithmic bioinformatics at @bonsaiseqbioinfo.bsky.social. Interested in space-efficient data structures, sketching algorithms & high-performance computing igor.martayan.org

PostsRepliesMedia
Reposted by Igor Martayan
Clément Canonne @ccanonne.github.io · 20h
A very important message from Omer Reingold to our TCS community, especially us (arg, already) senior researchers: "snap out of it." Please, digest, and share. theorydish.blog/2026/09/29/s...
As for depression: the next time you have the urge to lament, or even celebrate, being “the last generation of human mathematicians,” perhaps keep it to yourself. Contemplating the end of your profession from the relative comfort of an established, tenured career is a privilege, and it comes with responsibilities. Senior academics are not merely individual researchers; we are stewards of our field, and we owe our junior colleagues active leadership rather than abandonment.
47020
Reposted by Igor Martayan
Rob Patro @robp.bsky.social · 27/09/2026
Tangentially, @jermp.bsky.social also points out @recombconf.bsky.social is listed as B tier in CORE portal.core.edu.au/conf-ranks/?.... This is something that we should figure out how to address. RECOMB is a premiere conference in comp bio, & it is highly selective. It is `A` tier without question.
portal.core.edu.au
043
Reposted by Igor Martayan
Rob Patro @robp.bsky.social · 25/09/2026
Ok comp. bio / bioinformatics / genomics friends. Let's talk about @recombconf.bsky.social and, specifically, the way in which the "proceedings" have evolved. This seems like something that we, as a community, should address. RECOMB now has no "published" proceedings in the classic sense.
195
Reposted by Igor Martayan
Jim Shaw @jimshaw.bsky.social · 15/09/2026
The sylph metagenome profiler is v1.0.0! sylph-docs.github.io A new DB format + approach --> huge performance gains: GTDB-R232 (200k species) now takes < 5 GB of RAM and ~30s (2GB fq.gz). Huge thanks to @benjwoodcroft.bsky.social and his ongoing performance efforts (github.com/wwood/weebill)
sylph-docs.github.io
Documentation for sylph - ultrafast, precise metagenomic profiling
26330
Reposted by Igor Martayan
Tommi Mäklin @themaklin.bsky.social · 14/09/2026
I've been working on processing pseudoalignments from different tools to make downstream methods agnostic to the pseudoaligner choice, give it a try if you use themisto / fulgor / etc: docs.rs/ahda Also provides compression, conversion, and set operations.
docs.rs
ahda - Rust
ahda is a library and a command-line client for:
152
Reposted by Igor Martayan
Igor Martayan @imartayan.bsky.social · 28/08/2026
Happy to announce that I'll defend my PhD next Friday (Sept 4) at 2pm CEST! There will be a live stream on Zoom, just send me a message if you'd like to join!
2224
Igor Martayan @imartayan.bsky.social · 28/08/2026
Happy to announce that I'll defend my PhD next Friday (Sept 4) at 2pm CEST! There will be a live stream on Zoom, just send me a message if you'd like to join!
2224
Reposted by Igor Martayan
Gautam Kamath @gautamkamath.com · 18/08/2026
Obviously this doesn't scale, but I chose to send this junior researcher a (hopefully polite) email with some of my thoughts (written hastily in 30 minutes, not polished). More practicing mathematicians & researchers should advocate for the culture they want, so here's my bit.
1234
Reposted by Igor Martayan
arXiv cs.DS Data Structures and Algorithms @csds-bot.bsky.social · 03/08/2026
Xilin Tang (Cornell University), Yuqi Mai (Cornell University), William Kuszmaul (Carnegie Mellon University), Alex Conway (Cornell Tech): Succinct and Fast Tiny Pointer Hash Tables arxiv.org/abs/2607.28892 arxiv.org/pdf/2607.28892 arxiv.org/html/2607.28892
061
Reposted by Igor Martayan
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 01/08/2026
Riboseek is a fast RNA/DNA search. More sensitive than nhmmer at 250x speed. Structure-aware realignment produces MSAs approaching rMSA quality. Plus 1.7M precomputed RNA MSAs, and an API to search your own 📄 www.biorxiv.org/content/10.6... 💾 github.com/steineggerla... 🌐 search.foldseek.com/riboseek
217874
Reposted by Igor Martayan
Vikram Shivakumar @vikramshivakumar.bsky.social · 15/07/2026
1/ Excited to share the newest tool in the pangenome MUMiverse: Shredtools! Shredtools enables a user to navigate the pangenome coordinate system with multi-MUMs. More in the thread🧵 Code: github.com/vikshiv/shredtools Interactive tool for querying HPRC assemblies: vikshiv.github.io/shredtools
biorxiv.org
Navigating the pangenome coordinate system with Shredtools
Existing notions of pangenome coordinates rely on hard-to-compute multiple sequence alignments. On the other hand, pangenome-wide exact unique matches (multi-MUMs) can be computed efficiently, and rep...
13822
Reposted by Igor Martayan
Bede Constantinides @bede.im · 13/07/2026
Ever wanted to quickly check host content of DNA sequences? bede.im/sapiometer
12010
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 09/07/2026
My secret 4th ALGO paper is out! We show a tight space lower on non-minimal k-perfect hash functions, generalize PtrHash into a non-minimal k-PHF, and then use it to develop a hash set implementation that is up to 1.6x faster than other hash sets! With Stefan {Hermann, Walzer} and Peter Sanders
2168
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 09/07/2026
Lossless compression of k-mer matrices enabling random row access www.biorxiv.org/content/10.64898/20…
075
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 04/07/2026
Binary search and and set operations on compacted k-mer lists www.biorxiv.org/content/10.64898/20…
093
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 03/07/2026
synpact: accurate, memory-light PacBio HiFi read mapping via a hierarchy of locally-consistent syncmer blocks www.biorxiv.org/content/10.64898/20…
051
Reposted by Igor Martayan
arXiv cs.DS Data Structures and Algorithms @csds-bot.bsky.social · 01/07/2026
Francisco Olivares, Gonzalo Navarro: Practical Linear-Time Computation of Smallest Suffixient Sets arxiv.org/abs/2606.31034 arxiv.org/pdf/2606.31034 arxiv.org/html/2606.31034
044
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 24/06/2026
In other news: Great SEA talk by Nathaniel Brown on Orbit, an efficient implementation of the move structure for run-length encoded permutations. Also, congrats on winning a best paper award with this work!
Nathaniel standing in front of his title slide at the start of the presentation.
2104
Reposted by Igor Martayan
Antoine Limasset @npmalfoy.bsky.social · 24/06/2026
Oh my, a GPU implementation of Super Bloom filters!
071
Reposted by Igor Martayan
Igor Martayan @imartayan.bsky.social · 15/06/2026
Just submitted my PhD thesis on algorithms for fast, large-scale k-mer-based sequence analysis. It's now available to read at phd.martayan.org Take a look and feel free to share! #Bioinformatics #PhDone
phd.martayan.org
Algorithm design and implementation for the scale of sequencing data
26822
Igor Martayan @imartayan.bsky.social · 15/06/2026
Just submitted my PhD thesis on algorithms for fast, large-scale k-mer-based sequence analysis. It's now available to read at phd.martayan.org Take a look and feel free to share! #Bioinformatics #PhDone
phd.martayan.org
Algorithm design and implementation for the scale of sequencing data
26822
Reposted by Igor Martayan
Pierre Peterlongo @pierrepeterlongo.bsky.social · 11/06/2026
Happy to see that K2Rmini was recommended today by PCI Mathematical & Computational Biology. "quickly evaluate whether an arbitrary sequence has a number of k-mer [of interest] matches above or below a threshold." by @imartayan.bsky.social and colleagues: www.biorxiv.org/content/10.1...
074
Reposted by Igor Martayan
hbkgenomics.bsky.social @hbkgenomics.bsky.social · 06/06/2026
Does your designed active site already exist in nature? Is an uncharacterized protein hiding a catalytic site or a pocket? Folddisco answers both, searching millions of structures for a 3D motif in seconds. @natbiotech.nature.com 🧬 📄 www.nature.com/articles/s41... 🧵1/7👇
nature.com
Structural motif search across the protein universe with Folddisco - Nature Biotechnology
Folddisco enables protein structural motif search in million scale databases.
14715
Reposted by Igor Martayan
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
🌎 🧬 🖥️ logan-search.org the tool to query all SRA sequences (Dec 2023 snapshot) — just got major updates. Here's what's new. 🧵 1/12
13626
Reposted by Igor Martayan
Roland Faure @rfaure.bsky.social · 04/06/2026
Glad to announce that our new long-read metagenomic SNP caller, SNooPy, is published in NAR! Existing long-read SNP callers are not designed for metagenomic data, check out our new solution 👉https://academic.oup.com/nar/article/54/10/gkag556/8700491 @narjournal.bsky.social
1177
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 02/06/2026
Now also on arxiv arxiv.org/abs/2606.01190
arxiv.org
The anti-lexicographic SUS-anchor: a near-optimal k=1 sampling scheme
In recent years, there has been a renewed interest in the search for low density minimizer schemes. These schemes take a window of $w$ consecutive $k$-mers, and sample one of them: the smallest under ...
172
Igor Martayan @imartayan.bsky.social · 31/05/2026
Congrats to @leoackermann.bsky.social on winning RECOMB's best poster award for his work on compressing pairwise distance matrices! Check it out here: lacker.gitlab.io/pdf/research...
lacker.gitlab.io
0165
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 30/05/2026
Memory-safe high-performance sequence mapping with rammap www.biorxiv.org/content/10.64898/20…
12613
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 28/05/2026
Fast Set Operations for Compact k-mer Sets www.biorxiv.org/content/10.64898/20…
01510
Reposted by Igor Martayan
Rob Patro @robp.bsky.social · 21/05/2026
Is anyone attending #RECOMB2026 with a talk in the Sequencing 2 session who would be willing to switch slots with a talk in Sequencing 1 a day earlier? We’d be very grateful. Please reach out if you might be willing to do this!
045
Reposted by Igor Martayan
Camille Marchet ⚡ @camillemrcht.bsky.social · 17/05/2026
More and better human assemblies. Now annotate them to the minute. Special kudos @trhyker.bsky.social @jnalanko.bsky.social and @florisbarthel.bsky.social
0104
Reposted by Igor Martayan
terence @tterence.bsky.social · 15/05/2026
Rivers of France. #rayshader adventures, an #rstats tale
A visualisation of France's rivers
0235
Reposted by Igor Martayan
anil oza @aniloza.bsky.social · 14/05/2026
wow — the preprint host, arxiv, is banning authors for a year if they submit papers with hallucinated citations 🤖
Twitter thread from Thomas Dietterich, reading "Attention 
@arxiv
 authors: Our Code of Conduct states that by signing your name as an author of a paper, each author takes full responsibility for all its contents, irrespective of how the contents were generated. 1/
3:03 PM · May 14, 2026
·
66.7K
 Views
Relevant
View quotes

Thomas G. Dietterich
@tdietterich
·
1h
If generative AI tools generate inappropriate language, plagiarized content, biased content, errors, mistakes, incorrect references, or misleading content, and that output is included in scientific works, it is the responsibility of the author(s). 2/
Thomas G. Dietterich
@tdietterich
·
1h
We have recently clarified our penalties for this. If a submission contains incontrovertible evidence that the authors did not check the results of LLM generation, this means we can't trust anything in the paper. 3/
Thomas G. Dietterich
@tdietterich
·
1h
The penalty is a 1-year ban from arXiv followed by the requirement that subsequent arXiv submissions must first be accepted at a reputable peer-reviewed venue. 4/
Thomas G. Dietterich
@tdietterich
·
1h
Examples of incontrovertible evidence: hallucinated references, meta-comments from the LLM ("here is a 200 word summary; would you like me to make any changes?"; "the data in this table is illustrative, fill it in with the real numbers from your experiments") end/"
9458131915
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 12/05/2026
Me in group meeting update: Sassy was accepted at Bioinformatics, and Barbell should be accepted soon. Me at end of group meeting update: Both Sassy and Barbell are now accepted at Bioinformatics 🎉 Who will be the first to cite? All thanks to the wonderful work of @rickbitloo.bsky.social!!!
1183
Reposted by Igor Martayan
Bioinformatics Advances @bioinfoadv.bsky.social · 07/05/2026
🔍 New paper in Bioinformatics Advances: "Kaminari: A frugal colored index for approximate k-mer queries"  Read it here: doi.org/10.1093/bioadv/vbag120 Authors include: @yhhshb.bsky.social, @yoann.bsky.social, @robp.bsky.social, @pierrepeterlongo.bsky.social, @jermp.bsky.social
21510
Reposted by Igor Martayan
Paul Medvedev @pashadag.bsky.social · 05/05/2026
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
doi.org
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
13019
Reposted by Igor Martayan
recombseq.bsky.social @recombseq.bsky.social · 03/05/2026
The RECOMB-Seq 2026 program is now available! Join us May 24–25 in Thessaloniki, Greece, for two days of cutting-edge biological sequence analysis, with keynotes by Camille Marchet (CNRS) and Manolis Kellis (MIT). Full schedule: recomb-seq.github.io/seq2026/prog... #RECOMBseq
recomb-seq.github.io
Program
RECOMB-Seq 2026 Web Page
11610
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 29/04/2026
New preprint: The SimdQuickHeap is the fastest priority queue by far! 2x faster than a radix heap and up to 10x faster than binary heaps. arxiv.org/abs/2604.25681 with Marvin Williams and Johannes Breitling:
arxiv.org
SimdQuickHeap: The QuickHeap Reconsidered
Priority queues are data structures that maintain a dynamic collection of elements and allow inserting new elements and removing the smallest element. The most widely known and used priority queue is ...
1195
Reposted by Igor Martayan
Igor Martayan @imartayan.bsky.social · 21/04/2026
New blog post! I use ntHash all the time to hash k-mers, yet it turns out it has some unexpected flaws (collision propagation, bias on leading zeros...). The good news: each of them can be fixed! igor.martayan.org/posts/breaki...
igor.martayan.org
Breaking ntHash (to better fix it)
NtHash is a popular method for hashing k-mers in bioinformatics, yet it has some surprising flaws. In this post, I walk through a few of them, and show that they can arise naturally, without an advers...
13016
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 21/04/2026
Turns out that the usual NtHash is not as random as one might think?!?! At least not for minimizers. Seq-hash (and simd-minimizers) already has this fixed by default ;) github.com/rust-seq/seq...
084
Igor Martayan @imartayan.bsky.social · 21/04/2026
New blog post! I use ntHash all the time to hash k-mers, yet it turns out it has some unexpected flaws (collision propagation, bias on leading zeros...). The good news: each of them can be fixed! igor.martayan.org/posts/breaki...
igor.martayan.org
Breaking ntHash (to better fix it)
NtHash is a popular method for hashing k-mers in bioinformatics, yet it has some surprising flaws. In this post, I walk through a few of them, and show that they can arise naturally, without an advers...
13016
Reposted by Igor Martayan
Daniel Lemire @lemire.bsky.social · 19/04/2026
The fastest way to match characters on ARM processors? lemire.me/blog/2026/04/19/the-faste…
081
Reposted by Igor Martayan
Dion Dokter @diondokter.nl · 17/04/2026
I think this blog post comes closest to my current thinking on AI than any other I've read: fransskarman.com/im_not_using...
fransskarman.com
I am starting this post off right in the middle, with a paragraph that comes later:
1143
Reposted by Igor Martayan
Kai Blin @kblin.bsky.social · 09/04/2026
I'm not looking forward to a future where all the tools are being vibe-rewritten into languages people don't want to learn. Who will maintain all this? The original maintainers won't. It's not the language they were comfortable with. Does the prompter understand the tool well enough?
3206
Igor Martayan @imartayan.bsky.social · 10/04/2026
I kept being rate limited by doi2bib, so I made a small CLI to replace it locally: github.com/imartayan/bi... It can output bibtex from DOIs and arxiv/biorxiv links, redirect to the published version when it's available and copy the result to your clipboard
github.com
GitHub - imartayan/bibelot: A command-line tool adapted from doi2bib to fetch BibTeX entries from DOIs and more
A command-line tool adapted from doi2bib to fetch BibTeX entries from DOIs and more - imartayan/bibelot
1166
Reposted by Igor Martayan
LaurieWired @lauriewired.bsky.social · 07/04/2026
Modern DRAM is based on a brilliant design from IBM. But, we're still paying for a latency penalty that's existed since the 60s! In this video, I'm introducing my research project (Tailslayer) that immensely reduces p99.99 latency on traditional RAM!
319041
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 30/03/2026
Yeah; I have a bunch of incoherent thoughts about this... 1) maintaining bindings is annoying and purely a service. I already have the rust code and don't need bindings myself. 2) somebody needs to own them. Typically the main dev (me), even though someone else is the user and expert. (1/10)
172
Igor Martayan @imartayan.bsky.social · 30/03/2026
A quick rant on people vibe-translating our Rust libraries to other languages That's the second time in a week that I see new bioinformatics tools with a vibe-coded translation of our Rust libraries to C/C++. I have two major issues with that:
23310
Reposted by Igor Martayan
Jim Shaw @jimshaw.bsky.social · 27/03/2026
Myloasm, our long-read metagenome assembler, is now published! w/ @mgmarin.bsky.social and @lh3lh3.bsky.social Very rewarding after > a year of development and countless hours thinking about assembly. Thanks to beta testers, Li lab, and reviewers who gave very helpful feedback. rdcu.be/famFj
rdcu.be
High-resolution metagenome assembly for modern long reads with myloasm
Nature Biotechnology - A long-read metagenome assembly method recovers circular and complete genomes better than existing tools.
410056
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 30/03/2026
A run-length-compressed skiplist data structure for dynamic GBWTs supports time and space efficient pangenome operations over syncmers www.biorxiv.org/content/10.64898/20…
063