Sign in

Igor Martayan

@imartayan.bsky.social
835 followers 372 following 96 posts

PhD in algorithmic bioinformatics at @bonsaiseqbioinfo.bsky.social. Interested in space-efficient data structures, sketching algorithms & high-performance computing igor.martayan.org

PostsRepliesMedia
Reposted by Igor Martayan
Clément Canonne @ccanonne.github.io · 29/09/2026
A very important message from Omer Reingold to our TCS community, especially us (arg, already) senior researchers: "snap out of it." Please, digest, and share. theorydish.blog/2026/09/29/s...
As for depression: the next time you have the urge to lament, or even celebrate, being “the last generation of human mathematicians,” perhaps keep it to yourself. Contemplating the end of your profession from the relative comfort of an established, tenured career is a privilege, and it comes with responsibilities. Senior academics are not merely individual researchers; we are stewards of our field, and we owe our junior colleagues active leadership rather than abandonment.
47321
Reposted by Igor Martayan
Rob Patro @robp.bsky.social · 27/09/2026
Tangentially, @jermp.bsky.social also points out @recombconf.bsky.social is listed as B tier in CORE portal.core.edu.au/conf-ranks/?.... This is something that we should figure out how to address. RECOMB is a premiere conference in comp bio, & it is highly selective. It is `A` tier without question.
portal.core.edu.au
043
Reposted by Igor Martayan
Rob Patro @robp.bsky.social · 25/09/2026
Ok comp. bio / bioinformatics / genomics friends. Let's talk about @recombconf.bsky.social and, specifically, the way in which the "proceedings" have evolved. This seems like something that we, as a community, should address. RECOMB now has no "published" proceedings in the classic sense.
195
Igor Martayan @imartayan.bsky.social · 16/09/2026
On the negative side, I find the placement of the fingerprint sensor a bit annoying as it tends to unlock whenever I grab the phone, so I prefer facial unlock instead. Also the phone is probably a bit bulkier than other flagship models but that's the cost of repairability I guess.
010
Igor Martayan @imartayan.bsky.social · 16/09/2026
Bought a 6+ a couple weeks ago bc the battery of my old phone was dying and I couldn't replace it. So far I'm very happy with it, the experience is pretty smooth and it's very easy to take it apart. I remember ppl criticizing the photos a while back, but I find them perfectly fine with this one.
110
Reposted by Igor Martayan
Jim Shaw @jimshaw.bsky.social · 15/09/2026
The sylph metagenome profiler is v1.0.0! sylph-docs.github.io A new DB format + approach --> huge performance gains: GTDB-R232 (200k species) now takes < 5 GB of RAM and ~30s (2GB fq.gz). Huge thanks to @benjwoodcroft.bsky.social and his ongoing performance efforts (github.com/wwood/weebill)
sylph-docs.github.io
Documentation for sylph - ultrafast, precise metagenomic profiling
26330
Reposted by Igor Martayan
Tommi Mäklin @themaklin.bsky.social · 14/09/2026
I've been working on processing pseudoalignments from different tools to make downstream methods agnostic to the pseudoaligner choice, give it a try if you use themisto / fulgor / etc: docs.rs/ahda Also provides compression, conversion, and set operations.
docs.rs
ahda - Rust
ahda is a library and a command-line client for:
152
Igor Martayan @imartayan.bsky.social · 05/09/2026
Thank you so much! Of course, I'm more than happy if you reuse them!
011
Igor Martayan @imartayan.bsky.social · 05/09/2026
Thanks Ben!
000
Igor Martayan @imartayan.bsky.social · 05/09/2026
Thank you Rob, I'm also looking forward to this!
010
Igor Martayan @imartayan.bsky.social · 04/09/2026
This is in 2 hours! In the meantime, you can look at the slides here: igor.martayan.org/slides-phd.pdf
igor.martayan.org
352
Reposted by Igor Martayan
Igor Martayan @imartayan.bsky.social · 28/08/2026
Happy to announce that I'll defend my PhD next Friday (Sept 4) at 2pm CEST! There will be a live stream on Zoom, just send me a message if you'd like to join!
2224
Igor Martayan @imartayan.bsky.social · 30/08/2026
It's nice to see the continuous improvement you're doing on SSHash, such a great work!
230
Igor Martayan @imartayan.bsky.social · 30/08/2026
Looks like the same technique I discussed in my thesis (breaking ties with the innermost candidate): github.com/jermp/sshash...
github.com
centre-closest tie-break (Proposition 24): the scheme is now forward · jermp/sshash@b1d0706
Replace the tie-break of the canonical minimizer with the one of [Cologni and Pibiri, &quot;Canonical Schemes and Minimizers&quot;, Proposition 24]: among the loci tied at the minimum, take the one...
230
Igor Martayan @imartayan.bsky.social · 28/08/2026
Happy to announce that I'll defend my PhD next Friday (Sept 4) at 2pm CEST! There will be a live stream on Zoom, just send me a message if you'd like to join!
2224
Reposted by Igor Martayan
Gautam Kamath @gautamkamath.com · 18/08/2026
Obviously this doesn't scale, but I chose to send this junior researcher a (hopefully polite) email with some of my thoughts (written hastily in 30 minutes, not polished). More practicing mathematicians & researchers should advocate for the culture they want, so here's my bit.
1234
Reposted by Igor Martayan
arXiv cs.DS Data Structures and Algorithms @csds-bot.bsky.social · 03/08/2026
Xilin Tang (Cornell University), Yuqi Mai (Cornell University), William Kuszmaul (Carnegie Mellon University), Alex Conway (Cornell Tech): Succinct and Fast Tiny Pointer Hash Tables arxiv.org/abs/2607.28892 arxiv.org/pdf/2607.28892 arxiv.org/html/2607.28892
061
Reposted by Igor Martayan
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 01/08/2026
Riboseek is a fast RNA/DNA search. More sensitive than nhmmer at 250x speed. Structure-aware realignment produces MSAs approaching rMSA quality. Plus 1.7M precomputed RNA MSAs, and an API to search your own 📄 www.biorxiv.org/content/10.6... 💾 github.com/steineggerla... 🌐 search.foldseek.com/riboseek
217975
Igor Martayan @imartayan.bsky.social · 31/07/2026
can relate, it's even worse when you're not a native speaker
010
Reposted by Igor Martayan
Vikram Shivakumar @vikramshivakumar.bsky.social · 15/07/2026
1/ Excited to share the newest tool in the pangenome MUMiverse: Shredtools! Shredtools enables a user to navigate the pangenome coordinate system with multi-MUMs. More in the thread🧵 Code: github.com/vikshiv/shredtools Interactive tool for querying HPRC assemblies: vikshiv.github.io/shredtools
biorxiv.org
Navigating the pangenome coordinate system with Shredtools
Existing notions of pangenome coordinates rely on hard-to-compute multiple sequence alignments. On the other hand, pangenome-wide exact unique matches (multi-MUMs) can be computed efficiently, and rep...
13822
Reposted by Igor Martayan
Bede Constantinides @bede.im · 13/07/2026
Ever wanted to quickly check host content of DNA sequences? bede.im/sapiometer
12010
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 09/07/2026
My secret 4th ALGO paper is out! We show a tight space lower on non-minimal k-perfect hash functions, generalize PtrHash into a non-minimal k-PHF, and then use it to develop a hash set implementation that is up to 1.6x faster than other hash sets! With Stefan {Hermann, Walzer} and Peter Sanders
2168
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 09/07/2026
Lossless compression of k-mer matrices enabling random row access www.biorxiv.org/content/10.64898/20…
075
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 04/07/2026
Binary search and and set operations on compacted k-mer lists www.biorxiv.org/content/10.64898/20…
093
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 03/07/2026
synpact: accurate, memory-light PacBio HiFi read mapping via a hierarchy of locally-consistent syncmer blocks www.biorxiv.org/content/10.64898/20…
051
Reposted by Igor Martayan
arXiv cs.DS Data Structures and Algorithms @csds-bot.bsky.social · 01/07/2026
Francisco Olivares, Gonzalo Navarro: Practical Linear-Time Computation of Smallest Suffixient Sets arxiv.org/abs/2606.31034 arxiv.org/pdf/2606.31034 arxiv.org/html/2606.31034
044
Igor Martayan @imartayan.bsky.social · 26/06/2026
It's almost midnight AoE
220
Igor Martayan @imartayan.bsky.social · 26/06/2026
Reminds me of the famous Xbox numbering system: 1 -> 360 -> 1 -> X
051
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 24/06/2026
In other news: Great SEA talk by Nathaniel Brown on Orbit, an efficient implementation of the move structure for run-length encoded permutations. Also, congrats on winning a best paper award with this work!
Nathaniel standing in front of his title slide at the start of the presentation.
2104
Reposted by Igor Martayan
Antoine Limasset @npmalfoy.bsky.social · 24/06/2026
Oh my, a GPU implementation of Super Bloom filters!
071
Igor Martayan @imartayan.bsky.social · 21/06/2026
Thank you Zam!
000
Reposted by Igor Martayan
Igor Martayan @imartayan.bsky.social · 15/06/2026
Just submitted my PhD thesis on algorithms for fast, large-scale k-mer-based sequence analysis. It's now available to read at phd.martayan.org Take a look and feel free to share! #Bioinformatics #PhDone
phd.martayan.org
Algorithm design and implementation for the scale of sequencing data
26822
Igor Martayan @imartayan.bsky.social · 15/06/2026
bsky.app/profile/imar...
000
Igor Martayan @imartayan.bsky.social · 15/06/2026
Just submitted my PhD thesis on algorithms for fast, large-scale k-mer-based sequence analysis. It's now available to read at phd.martayan.org Take a look and feel free to share! #Bioinformatics #PhDone
phd.martayan.org
Algorithm design and implementation for the scale of sequencing data
26822
Reposted by Igor Martayan
Pierre Peterlongo @pierrepeterlongo.bsky.social · 11/06/2026
Happy to see that K2Rmini was recommended today by PCI Mathematical & Computational Biology. "quickly evaluate whether an arbitrary sequence has a number of k-mer [of interest] matches above or below a threshold." by @imartayan.bsky.social and colleagues: www.biorxiv.org/content/10.1...
074
Reposted by Igor Martayan
hbkgenomics.bsky.social @hbkgenomics.bsky.social · 06/06/2026
Does your designed active site already exist in nature? Is an uncharacterized protein hiding a catalytic site or a pocket? Folddisco answers both, searching millions of structures for a 3D motif in seconds. @natbiotech.nature.com 🧬 📄 www.nature.com/articles/s41... 🧵1/7👇
nature.com
Structural motif search across the protein universe with Folddisco - Nature Biotechnology
Folddisco enables protein structural motif search in million scale databases.
14715
Reposted by Igor Martayan
Pierre Peterlongo @pierrepeterlongo.bsky.social · 03/06/2026
🌎 🧬 🖥️ logan-search.org the tool to query all SRA sequences (Dec 2023 snapshot) — just got major updates. Here's what's new. 🧵 1/12
13626
Reposted by Igor Martayan
Roland Faure @rfaure.bsky.social · 04/06/2026
Glad to announce that our new long-read metagenomic SNP caller, SNooPy, is published in NAR! Existing long-read SNP callers are not designed for metagenomic data, check out our new solution 👉https://academic.oup.com/nar/article/54/10/gkag556/8700491 @narjournal.bsky.social
1177
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 02/06/2026
Now also on arxiv arxiv.org/abs/2606.01190
arxiv.org
The anti-lexicographic SUS-anchor: a near-optimal k=1 sampling scheme
In recent years, there has been a renewed interest in the search for low density minimizer schemes. These schemes take a window of $w$ consecutive $k$-mers, and sample one of them: the smallest under ...
172
Igor Martayan @imartayan.bsky.social · 31/05/2026
Congrats to @leoackermann.bsky.social on winning RECOMB's best poster award for his work on compressing pairwise distance matrices! Check it out here: lacker.gitlab.io/pdf/research...
lacker.gitlab.io
0165
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 30/05/2026
Memory-safe high-performance sequence mapping with rammap www.biorxiv.org/content/10.64898/20…
12613
Reposted by Igor Martayan
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 28/05/2026
Fast Set Operations for Compact k-mer Sets www.biorxiv.org/content/10.64898/20…
01510
Igor Martayan @imartayan.bsky.social · 26/05/2026
I usually do it with GitHub runners
120
Reposted by Igor Martayan
Rob Patro @robp.bsky.social · 21/05/2026
Is anyone attending #RECOMB2026 with a talk in the Sequencing 2 session who would be willing to switch slots with a talk in Sequencing 1 a day earlier? We’d be very grateful. Please reach out if you might be willing to do this!
045
Reposted by Igor Martayan
Camille Marchet ⚡ @camillemrcht.bsky.social · 17/05/2026
More and better human assemblies. Now annotate them to the minute. Special kudos @trhyker.bsky.social @jnalanko.bsky.social and @florisbarthel.bsky.social
0104
Reposted by Igor Martayan
terence @tterence.bsky.social · 15/05/2026
Rivers of France. #rayshader adventures, an #rstats tale
A visualisation of France's rivers
0235
Reposted by Igor Martayan
anil oza @aniloza.bsky.social · 14/05/2026
wow — the preprint host, arxiv, is banning authors for a year if they submit papers with hallucinated citations 🤖
Twitter thread from Thomas Dietterich, reading "Attention 
@arxiv
 authors: Our Code of Conduct states that by signing your name as an author of a paper, each author takes full responsibility for all its contents, irrespective of how the contents were generated. 1/
3:03 PM · May 14, 2026
·
66.7K
 Views
Relevant
View quotes

Thomas G. Dietterich
@tdietterich
·
1h
If generative AI tools generate inappropriate language, plagiarized content, biased content, errors, mistakes, incorrect references, or misleading content, and that output is included in scientific works, it is the responsibility of the author(s). 2/
Thomas G. Dietterich
@tdietterich
·
1h
We have recently clarified our penalties for this. If a submission contains incontrovertible evidence that the authors did not check the results of LLM generation, this means we can't trust anything in the paper. 3/
Thomas G. Dietterich
@tdietterich
·
1h
The penalty is a 1-year ban from arXiv followed by the requirement that subsequent arXiv submissions must first be accepted at a reputable peer-reviewed venue. 4/
Thomas G. Dietterich
@tdietterich
·
1h
Examples of incontrovertible evidence: hallucinated references, meta-comments from the LLM ("here is a 200 word summary; would you like me to make any changes?"; "the data in this table is illustrative, fill it in with the real numbers from your experiments") end/"
9458111914
Reposted by Igor Martayan
Ragnar {Groot Koerkamp} @curiouscoding.nl · 12/05/2026
Me in group meeting update: Sassy was accepted at Bioinformatics, and Barbell should be accepted soon. Me at end of group meeting update: Both Sassy and Barbell are now accepted at Bioinformatics 🎉 Who will be the first to cite? All thanks to the wonderful work of @rickbitloo.bsky.social!!!
1183
Reposted by Igor Martayan
Bioinformatics Advances @bioinfoadv.bsky.social · 07/05/2026
🔍 New paper in Bioinformatics Advances: "Kaminari: A frugal colored index for approximate k-mer queries"  Read it here: doi.org/10.1093/bioadv/vbag120 Authors include: @yhhshb.bsky.social, @yoann.bsky.social, @robp.bsky.social, @pierrepeterlongo.bsky.social, @jermp.bsky.social
21510
Reposted by Igor Martayan
Paul Medvedev @pashadag.bsky.social · 05/05/2026
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
doi.org
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
13019