Rob Patro @robp.bsky.social · 21/09/2026It's one binary, aie: ingest, replay under any annotation, region/junction/APA queries, federated collections that search a cohort by junction shape, cohort-wide discovery, and GQ, a composable query language with three-valued logic that runs unchanged over one archive or a whole federation. 120
Rob Patro @robp.bsky.social · 21/09/2026Case study 4: pooling evidence across cells inside an EM recovers 75–98% of withheld multi-gene molecule labels. That's mass conventional unique counting simply discards. 110
Rob Patro @robp.bsky.social · 21/09/2026Case study 1: the archive answers what the matrix threw away. Across four PBMC archives it recovers the RT-PCR-validated FYB1 splicing switch: cassette inclusion of 6–9% in T cells vs 98–99% in monocytes, in every dataset. 110
Rob Patro @robp.bsky.social · 21/09/2026How faithful is gravlax? A matrix replayed from the archive deviates from a fresh STARsolo run by 0.24–0.75% of UMI mass. Switching GENCODE v32→v49 moves 5–10× more. And Gene replay is 34–82× faster than realigning, because alignment is paid once at ingest. 110
Rob Patro @robp.bsky.social · 21/09/2026How small? 11–18 bits per read. That's 9–13× smaller than tag-preserving CRAM and 32–49× smaller than FASTQ. Small enough to keep an entire cohort online instead of banishing reads to cold storage. But to be useful, the representation should be accurate/faithful. 120
Rob Patro @robp.bsky.social · 21/09/2026It's here! new preprint: Gravlax, an annotation-independent molecular evidence archive for scRNA-seq. A count matrix freezes one annotation. Molecules never change. Instead, gravlax keeps the evidence. Requantify, query, and discover under any annotation. 🧵 www.biorxiv.org/content/10.6... 33413
Rob Patro @robp.bsky.social · 16/09/2026The part that surprised us: no general tool for this existed. STAR has an internal projection module, but only for STAR's own alignments. Salmon has decoys (but doesn't do spliced alignment). For long reads, nothing at all. So people pick a side of the fork, or quietly roll their own. 112
Rob Patro @robp.bsky.social · 16/09/2026New preprint led by Zoe Rudnick: bramble 🌿 RNA-seq quantification makes you pick a side. Align to the transcriptome and your quantifier is happy, but reads from unannotated transcripts get misassigned to annotated ones. Align to the genome and you keep discovery, but limit quantification choices. 12913
Rob Patro @robp.bsky.social · 13/09/2026Derrick Henry will come for you, and you will do nothing, because you can do nothing. 🐦⬛ 020
Rob Patro @robp.bsky.social · 12/09/2026Because the molecules are kept, you can ask what the matrix discarded. Gravlax recovers the FYB1 T-cell/monocyte splicing switch and an 8-donor NTRK2 3′-isoform shift from astrocytes to neurons — straight from the archive, no reprocessing. 110
Rob Patro @robp.bsky.social · 12/09/2026Is deferring the annotation lossy? Barely. A matrix replayed from the archive deviates from a fresh STARsolo run by 0.24–0.75% of UMI mass — while switching GENCODE v32→v49 moves 5–10× more. And replay is 34–82× faster than realigning. 110
Rob Patro @robp.bsky.social · 12/09/2026How small? A gravlax .aie (annotation independent evidence) archive is 11–18 bits per read: 9–13× smaller than tag-preserving CRAM and 32–49× smaller than FASTQ. Small enough to keep a whole cohort online instead of banishing reads to cold storage. 130
Rob Patro @robp.bsky.social · 12/09/2026Single-cell RNA-seq count matrices bake in one gene annotation the day you build them. Annotations change several times a year; your molecules never do. Gravlax captures the molecular evidence once, so you can requantify under any future annotation. Align once, query forever. 🧵 1143
Rob Patro @robp.bsky.social · 07/09/2026For the Monday morning crowd: We've now build a domain specific query language for gravlax! This gives a powerful and concise syntax to ask questions about the molecular evidence in your scRNA-seq experiments! 031
Rob Patro @robp.bsky.social · 06/09/2026The same query text runs over one archive or a federation of donors. Results are typed tables with provenance: query digest, plan digest, archive roots. validate, explain, run, from the CLI or Python, and explain tells you what it will read before it reads it. 100
Rob Patro @robp.bsky.social · 06/09/2026The archive keeps two representatives per read chain, not every read. Some questions can't be decided from that. GQ returns unknown instead of guessing, and a tally keeps every true/false/unknown combination, so your denominators survive. 100
Rob Patro @robp.bsky.social · 06/09/2026Why a language? Because the good questions are relational. "Does one read show exon inclusion AND this junction?" and "does one cell show both?" are different questions with different answers. GQ makes you say which one you mean. Leaving out the quantifier is a type error. 110
Rob Patro @robp.bsky.social · 06/09/2026Gravlax 0.2.0 ships GQ, a query language for single-cell molecular evidence! Ask questions of the molecules in an experiment, per cell, per UMI class, or per aligned observation, straight from a compact .aie archive. No reads, no re-alignment, no annotation required. 🧵 171
Rob Patro @robp.bsky.social · 01/09/2026Big update to the seqproc paper (www.biorxiv.org/content/10.6...)! We've made seqproc even faster & more capable. We've also extended the analysis & tweaked & improved the configurations we used for other tools as well. Seqproc now has the beginnings of a graph optimization backend for its DSL! 1132
Rob Patro @robp.bsky.social · 25/08/2026Well, not with that attitude it's not 😂. arxiv.org/pdf/2608.23502 030
Rob Patro @robp.bsky.social · 23/08/2026SPLiT-seq PE shows why a DSL helps. R2 is UMI(10)+BC3(8)+L1+BC2(8)+L2+BC1(8). Linkers allow edit distance 3; barcodes allow Hamming distance 1 to a named whitelist. We then emit R1 plus the corrected 34-nt UMI+BC3+BC2+BC1 product. Config attached ↓ 5/6 100
Rob Patro @robp.bsky.social · 23/08/2026The first public release of seqproc is out! 🎉 v0.1.1, powered by ANTISEQUENCE v0.1.0. seqproc let's you describe a sequencing protocol in EFGDL, compile it into a fast Rust pipeline for matching, filtering, correction, and FASTQ transformation. A powerful read pre-processing engine! 1/6 1197
Rob Patro @robp.bsky.social · 22/08/2026Oh wow, I knew people like Fulgor, but this journal is more interested in our esoteric work “Span Class U-Sans-Serif Fulgor Span” cc @jermp.bsky.social 🤣🤣😂 120
Rob Patro @robp.bsky.social · 15/08/2026…but a real one: …on ~150K Salmonella genomes (k=31, 64 threads), the released Rust build vs the validated C++: 11% faster uncolored, 22% faster colored, at 60% less peak RAM (9.2 vs 23.3 GB) uncolored — identical graphs, verified unitig-for-unitig. 8/10 100
Rob Patro @robp.bsky.social · 15/08/2026Here's the part we're most pleased with: the released v3.0.0 isn't the paper's C++ implementation. It's a ground-up Rust rewrite. We made that call for ease of future development and maintenance — the performance improvement turned out to be an additional benefit… 7/10 110
Rob Patro @robp.bsky.social · 15/08/2026Performance, from the paper: 3.29–4.09× faster than GGCAT (v2.0.0) — the prior state of the art — at comparable memory, across large-scale genomic datasets. 6/10 110
Rob Patro @robp.bsky.social · 15/08/2026Cuttlefish 3 is on bioconda! 🦑 A parallel, external-memory algorithm for building colored compacted de Bruijn graphs at collection scale — a RECOMB 2026 paper, and as of today a production release: v3.0.0. A thread on the algorithm, the numbers, and why the released tool is a Rust rewrite. 🧵 1/10 34412
Rob Patro @robp.bsky.social · 15/08/2026👀 github.com/bioconda/bio... combine-lab.github.io/cuttlefish/ 150
Rob Patro @robp.bsky.social · 10/08/2026salmon 2.5.0 is out: 28 PRs. Headline: -p is now one measured budget — gzip decode + mapping share it, and a live controller (thread-broker) solves the split from measurement instead of guessing. Quantification results unchanged; no index rebuild. github.com/COMBINE-lab/salmon 1/4 1112
Rob Patro @robp.bsky.social · 10/08/2026All open source: piscem-rs, rapidgzip-core & thread-broker (standalone crate; bring your own producer/consumer). Now also in salmon 2.5.0 & coming to the alevin-fry ecosystem to turbocharge workflows. Maybe rapidgzip & thread-broker can help you too! Props to paraseq for making consumers easy! 8/8 032
Rob Patro @robp.bsky.social · 10/08/2026Our fix, thread-broker: don't search for the split, solve for it (à la DS2, Kalavri et al., OSDI '18). With N = thread budget, P = busy time decoding, C = busy time mapping, the decode share is d* = N·P/(P+C). Converged in 2.1 s & beat every fixed split we swept. crates.io/crates/threa... 6/8 141
Rob Patro @robp.bsky.social · 10/08/2026And guessing wrong is expensive: we swept fixed splits on the full 2.3B-pair run. The surface is sharply peaked. A few slots off the optimum costs 5–46%, and a bad guess costs +116%. Worse, the optimum moves with the thread budget, so you can't fix it statically. 5/8 110
Rob Patro @robp.bsky.social · 10/08/2026Problem 1: gzip. A gzip stream is inherently serial; one decoder thread, no matter how many cores you have. Hand a fast 64-thread mapper two gzipped FASTQs and you've built a sports car with a garden-hose fuel line. The mapping threads starve while waiting for inflation. 2/8 120
Rob Patro @robp.bsky.social · 10/08/20262.3 billion read pairs of 10x Flex v2 (281 GB of gzipped FASTQ) mapped in under 2 minutes on one machine (-t 64) with piscem-rs. That's ~20M read pairs/sec at ~87% mapped. One interesting part is the mapper, but what I want to talk about here is who gets the threads. 1/8 1247
Rob Patro @robp.bsky.social · 01/08/2026The most popularly repeated “stochastic parrot” interpretation (that LLMs merely remix surface text & can't reason or create genuinely new abstractions) is no longer tenable. Novel open-problem solutions with machine-checkable Lean proofs are evidence of systematic generalization, not quotation.1/3 1150
Rob Patro @robp.bsky.social · 31/07/2026I got frustrated that using piscem-rs to map 10x Flex reads was entirely dominated by decompressing the single pair of input files. So, GPT5.6-Sol & I built a thing! Rapidgzip-rust (rapid-gzip algo in rust) github.com/COMBINE-lab/... 2.3 billion flex reads in <2 min? Parsing w paraseq. Yes plz! 1182
Rob Patro @robp.bsky.social · 29/07/2026Did preprocessing choices change the biology? On full SPLiT-seq data, per-gene Pearson correlations were 0.994–0.999 and per-barcode correlations 0.974–0.992. Of 220 cells called by all tools, 89.1% received the same cell-type label across all three. All of these methods are rather good, IMO. 120
Rob Patro @robp.bsky.social · 29/07/2026Each tool used 32 threads. Every tool–dataset combination was run 3 times on the full data; runtime and memory are means. Read recovery was checked against a protocol-specific, tool-neutral reference set built from the raw reads. 110
Rob Patro @robp.bsky.social · 29/07/2026Under the hood, antisequence represents the work as a directed acyclic graph of primitive operations. Reads stream through matching, filtering, cutting, trimming, reordering, and output nodes. Independent branches can run in parallel, and common work can be shared. 120
Rob Patro @robp.bsky.social · 29/07/2026An EFGDL specification has up to three parts: 1. definitions that name intervals 2. a read structure that describes valid geometry 3. an optional transformation that removes, normalizes, or reorders matched intervals Only the read structure is required. 110
Rob Patro @robp.bsky.social · 29/07/2026New preprint🚨(*long* time coming)! We introduce seqproc: an efficient, flexible, and concise tool for describing and transforming sequencing-read geometry. The aim: make complex protocols easy to specify & transform without giving up speed or accuracy. www.biorxiv.org/content/10.6... 🧵 1369
Rob Patro @robp.bsky.social · 29/07/2026Yea; being a Ravens fan is tough. But whenever you want to feel better about the Ravens, it always helps to also have the Orioles as your home team... 010
Rob Patro @robp.bsky.social · 14/07/2026Want to further improve your transcript quantification (esp. long read)? Come see our poster on Bramble! Presented by Zoe, who has done a phenomenal job on this method & tool. #ISMB2026 1171