Sign in

Nils Homer

@nilshomer.com
2.4K followers 215 following 171 posts

I write software for Biologists. Founder @fulcrumgenomics, Bioinformatician, Computer Scientist, Coder, Father of 2xGirls. Opinions are my own.

PostsRepliesMedia
Nils Homer @nilshomer.com · 14/09/2026
bwa-mem3 v0.12.0 is out 🧬 Since v0.10.0: faster on both Arm & x86 (--fast is now ~2× minibwa and stock is within ~15% on x86) while staying a byte-identical (--compat). It also brings ~30–40% faster methylation and a big memory-safety pass. github.com/fg-labs/bwa-... #bioinformatics #genomics
github.com
Release v0.12.0 · fg-labs/bwa-mem3
0.12.0 — faster on both architectures, a faster --meth, and a large safety pass A broad optimization release that speeds up both Arm and x86 (unlike 0.11.0, which concentrated on Arm), plus a methy...
0132
Nils Homer @nilshomer.com · 25/08/2026
Thanks @jpelbers.bsky.social for pointing out the link was broken: reposted!
000
Nils Homer @nilshomer.com · 25/08/2026
fgumi v0.7.0 is released: github.com/fulcrumgenom... 1. sort is 28% faster; 2-3x faster than samtools 2. improved CODEC consensus calling 3. dedup now outputs metrics closer to picard MarkDuplicates and dupblaster 4. retag is a new tool that can copy/move/delete SAM tags
github.com
Release v0.7.0 · fulcrumgenomics/fgumi
For those running the fgumi command line tools v0.7.0 makes sort faster, hardens CODEC consensus, and brings dedup metrics to Picard/dupblaster parity. Sort is up to 28% faster — and 2–3× faster th...
173
Nils Homer @nilshomer.com · 11/08/2026
Want to see a mistake I made?
050
Reposted by Nils Homer
Rob Patro @robp.bsky.social · 10/08/2026
2.3 billion read pairs of 10x Flex v2 (281 GB of gzipped FASTQ) mapped in under 2 minutes on one machine (-t 64) with piscem-rs. That's ~20M read pairs/sec at ~87% mapped. One interesting part is the mapper, but what I want to talk about here is who gets the threads. 1/8
1247
Nils Homer @nilshomer.com · 08/08/2026
> then around 850-950 ad, their ceremonial centers were suddenly abandoned, and researchers still have no idea why
static.klipy.com
Ragnar Onurcem: Charismatic Leader
ALT: Ragnar Onurcem: Charismatic Leader
000
Nils Homer @nilshomer.com · 08/08/2026
Along with tricord, I am getting much more accurate profiling of my rewrites.  Give them a try! github.com/fg-labs/tric...
github.com
GitHub - fg-labs/tricord: Run a command and report its process tree's CPU, memory, and I/O usage
Run a command and report its process tree's CPU, memory, and I/O usage - fg-labs/tricord
030
Nils Homer @nilshomer.com · 08/08/2026
I worked with Claude built tachyon to quickly measure how contended the system is so I can normalize cpu and run time statistics over time.  github.com/nh13/tachyon
github.com
GitHub - nh13/tachyon: Detect the cloaked: measure a host's memory access rate under current contention
Detect the cloaked: measure a host's memory access rate under current contention - nh13/tachyon
130
Nils Homer @nilshomer.com · 08/08/2026
When running benchmarks on shared tenancy systems (e.g. AWS), it's hard to compare cpu/wall times over time, even with many replicates.  have seen performance differences up to 20% just based on fighting others for access to the memory sub-system.
121
Reposted by Nils Homer
Alison Meynert @ameynert.bsky.social · 05/08/2026
DivRef is a resource for including common human variants and haplotypes in CRISPR off-target searches. I rebuilt the generation workflow to make its inputs, assumptions and population choices easier to inspect and change. blog.fulcrumgenomics.com/p/a-crispr-o...
blog.fulcrumgenomics.com
A CRISPR off-target search is only as good as the sequences it searches
Rebuilding DivRef so CRISPR off-target searches can account for human variation
144
Reposted by Nils Homer
Tim Dunn @timd.one · 22/07/2026
CRISPR off-target search can find the right locus while hiding plausible alignments within it. I wrote about how Sassy v0.2.5 uses recursive backtracking to report every reasonable alignment—more than 9× as many in under 30 seconds. blog.fulcrumgenomics.com/p/why-crispr...
blog.fulcrumgenomics.com
Why CRISPR Off-Target Search Should Report Multiple Alignments Per Locus
How Sassy enumerates every reasonable alignment without sacrificing runtime
054
Nils Homer @nilshomer.com · 04/08/2026
The rest of the release is speed — about 13% off the default path, 18% with --fast, and nothing got slower. And for methylation: NM and MD were counting bisulfite conversions as mismatches, inflating edit distance on nearly every read. They aren't anymore.
010
Nils Homer @nilshomer.com · 04/08/2026
Note: bwa-mem2: doesn't reproduce itself across thread counts. Insert-size statistics are estimated per batch, and batch size scales with -t, so the same input at -t 16 and -t 32 gives you slightly different BAMs. Under --compat we pin the batch, so thread count stops affecting your results.
100
Nils Homer @nilshomer.com · 04/08/2026
We measured it rather than claiming it: 1.57 billion alignment records, across 7 datasets, on 6 CPU types, 5 replicates each. Every record matched. Zero differences.
110
Nils Homer @nilshomer.com · 04/08/2026
Add --compat=bwa-mem2 and you get exactly what bwa-mem2 v2.2.1 would have produced. Same positions, same CIGARs, same mapping qualities, same tags, same header.
100
Nils Homer @nilshomer.com · 04/08/2026
bwa-mem3 v0.8.0 is out. New --compat=bwa-mem2: byte-identical to bwa-mem2 v2.2.1, verified across 1.57B records on 6 CPU types — including ARM. Also ~13% faster, and methylation NM/MD no longer count bisulfite conversions as mismatches. github.com/fg-labs/bwa-mem3/releases/tag/v0.8.0
github.com
Release v0.8.0 · fg-labs/bwa-mem3
0.8.0 — a verified drop-in for bwa-mem2, ~13% faster, and corrected methylation tags 1. --compat=bwa-mem2 — swap in bwa-mem3 and get the same BAM The big one. Run bwa-mem3 mem --compat=bwa-mem2 ......
190
Reposted by Nils Homer
Tim Fennell @tfenne.bsky.social · 01/08/2026
Maybe three weeks ago, I* refactored some code out of dupblaster and methylsieve for using dedicated threads with large buffers for read-ahead and write-behind. Why? Because it turns out, even in Rust, it's still hard to write threading code that "just works". github.com/tfenne/rawb-io
github.com
GitHub - tfenne/rawb-io: Read-ahead / write-behind byte IO: threaded Reader/Writer that decouple a pipeline stage from the kernel pipe via a background thread and a byte ring buffer
Read-ahead / write-behind byte IO: threaded Reader/Writer that decouple a pipeline stage from the kernel pipe via a background thread and a byte ring buffer - tfenne/rawb-io
242
Nils Homer @nilshomer.com · 01/08/2026
I've been doing a ton of SIMD optimization in bwa-mem3, in case you want to see if there any new ideas?
120
Nils Homer @nilshomer.com · 01/08/2026
I want to see them!
010
Nils Homer @nilshomer.com · 01/08/2026
I mean some software "just works", so not many releases that often, especially in the rust ecosystem, so I think both at least. But the gradient goes from grad/post-doc who wrote something for their publication to folks who are still active day to day.
110
Nils Homer @nilshomer.com · 31/07/2026
So if you plan to do a rewrite of an actively maintained software, please do reach out to those folks first.
280
Nils Homer @nilshomer.com · 31/07/2026
Meritocracy and competition is great, and the moat due to implementation effort is gone, but we also have to remember the relational side of bioinformatics.
152
Nils Homer @nilshomer.com · 31/07/2026
It's not only about speed, it's about doing it right the second time (see @tfenne.bsky.social's Riker: the successor to Picard). It's about being able to finally explore all those ideas that were left on the cutting board due to time and money constraints.
130
Nils Homer @nilshomer.com · 31/07/2026
It has involved years of late-night insights and hard fought wins. Even recently with AI, I've spent so long pouring my own ideas and discernment into it. Years diligently and genuinely supporting, maintaining and developing this software. Building a livelihood on the reputation of our work.
140
Nils Homer @nilshomer.com · 31/07/2026
Bioinformatics rewrites miss the relational impact tofolks that are still actively maintaining and developing the software. I've been guilty of this myself, and I'll be sharing my story soon so others can learn from it.
1216
Nils Homer @nilshomer.com · 23/07/2026
It's start with M versus start with H in the recurrence, there's a PR to close the gap on main.
000
Nils Homer @nilshomer.com · 23/07/2026
2. CPU-specific optimized code (AVX2, AVX-512, Arm) but the per-CPU (AVX2/AVX-512/Arm) kernels can score the runner-up alignment slightly differently. And since MAPQ comes from the best-vs-second-best score gap, it moves on <0.1% of reads. Plus a few extra split alignments.
100
Nils Homer @nilshomer.com · 23/07/2026
Why not byte-identical to bwa-mem2? By design. 1. bwa-mem3 breaks score ties deterministically, bwa-mem2 didn't. Reproducibility is restored. That's a feature, not a bug.
110
Nils Homer @nilshomer.com · 23/07/2026
bwa-mem3 v0.7.0 is out. 5–24% faster than v0.6.0, output byte-identical; upgrade and it's free. If you do methylation: reworked bisulfite path with TAPS support too. github.com/fg-labs/bwa-mem3/releases/tag/v0.7.0
github.com
Release v0.7.0 · fg-labs/bwa-mem3
0.7.0 — methylation accuracy, a sharper --fast, and broad speedups This is primarily a methylation release. The bisulfite path was reworked to a chemistry-aware contract — TAPS support, a NEUTRAL -...
151
Reposted by Nils Homer
Fulcrum Genomics @fulcrumgenomics.com · 22/07/2026
Reporting one alignment per locus can commit a CRISPR off-target analysis to one scoring model too early. @timd.one explains how Sassy v0.2.5 enumerates every reasonable alignment while keeping runtime fast: >9× as many alignments in under 30 seconds. blog.fulcrumgenomics.com/p/why-crispr...
blog.fulcrumgenomics.com
Why CRISPR Off-Target Search Should Report Multiple Alignments Per Locus
How Sassy enumerates every reasonable alignment without sacrificing runtime
1162
Reposted by Nils Homer
Robert Aboukhalil @robert.bio · 22/07/2026
grep is a fantastic tool, but it doesn't really work on sequencing data: it breaks when the pattern is in the read name, and doesn't support paired-end FASTQs. Thanks to @nilshomer.com, we have a new interactive guide on using fqgrep to find patterns in FASTQ files: ➡️ sandbox.bio/tutorials/fq...
sandbox.bio
Interactive bioinformatics tutorials
Learn bioinformatics from your browser, no setup required. Everything runs in a sandbox, so you can experiment all you want.
0197
Nils Homer @nilshomer.com · 16/07/2026
Start from the patient, name the real gap, then build the analysis that closes it. @fulcrumgenomics.com VP of Translational Research Juliann Chmielecki, PhD nails it. Worth your time if you work anywhere near translational oncology. www.linkedin.com/pulse/from-e...
linkedin.com
From equations to patients: where computational biology meets translational science
Building translational strategies marries new methodologies with an unanswered clinical question, then works with the data and analyses needed to answer it When I think about what translational scienc...
030
Nils Homer @nilshomer.com · 14/07/2026
Part of rewriting a tool is being able to make opinionated choices about it's development, providing maintenance, and that's what I am doing here. E Pluribus Unum. #reright
010
Nils Homer @nilshomer.com · 14/07/2026
What started out as "make OptiType" faster turned into a multi-month saga that discovered that "yes, a Heng Li lab tool is faster and better on the data we care about".
110
Nils Homer @nilshomer.com · 14/07/2026
unum: a pure-Rust HLA/KIR genotyper It's a port of github.com/mourisl/T1K, offering significant speedups, ergonomics, and opinionated improvements. I am looking for folks who want to give it a try and give constructive feedback. github.com/fg-labs/unum
github.com
GitHub - fg-labs/unum: Fulcrum-owned Rust HLA/KIR genotyper (strangler port of T1K)
Fulcrum-owned Rust HLA/KIR genotyper (strangler port of T1K) - fg-labs/unum
192
Nils Homer @nilshomer.com · 12/07/2026
I’ll rebase later today and see what I get. And my apologies for the LLM-based PRs, figured it was better to get you _something_ rather than nothing given your initial call to action 🤝
010
Nils Homer @nilshomer.com · 12/07/2026
I also made a PR just now to speed up GPU across all # of patterns: github.com/RagnarGrootK... $1 of GPU spent :p
github.com
perf: text-direction tiling + parallel PEQ — GPU wins from p=100 (2.7x on the default workload) by nh13 · Pull Request #75 · RagnarGrootKoerkamp/sassy
Two GPU-side performance changes that make the WGPU backend faster than the CPU SIMD path across the whole pattern-count range — including the low-p regime where it previously lost. Measured on an ...
110
Nils Homer @nilshomer.com · 09/07/2026
✅ like the Continuum Transfunctioner, my mystery is only exceeded by my power (for failure)
010
Nils Homer @nilshomer.com · 09/07/2026
You miss 100% of the shots you don't take.
100
Nils Homer @nilshomer.com · 09/07/2026
🚀 ferro-hgvs 0.7.0 is out: our biggest release yet. Major upgrades to HGVS normalization, parsing & projection: spec-compliant 3′ shifting, mosaic/compound alleles, multi-axis projection (g/c/n/p/r), Ensembl support, plus ~1.7× faster parsing. github.com/fulcrumgenom...
github.com
Release v0.7.0 · fulcrumgenomics/ferro-hgvs
Added (reference) validate manifest schema/version at load, fail loud on an incompatible reference (#1003) (mosaic) parse predicted-wrapper and whole-entity-LHS =/ forms (#992) (protein) parse ins...
021
Nils Homer @nilshomer.com · 07/07/2026
Save the compute, save the world!
030
Reposted by Nils Homer
Steven Salzberg @stevensalzberg.bsky.social · 07/07/2026
Our new genome annotation method relies almost entirely on transcriptome and alignment evidence, and as a result outperforms pretty much all other de novo pipelines. Check out the just-published paper led by Aleksey Zimin: rdcu.be/frSOg
rdcu.be
Efficient evidence-based genome annotation with EviAnn
Nature Methods - EviAnn surpasses existing genome annotation methods by leveraging gene expression and protein sequence homology evidence to achieve higher accuracy and efficiency.
04420
Nils Homer @nilshomer.com · 06/07/2026
#bioinformatics re-writes are more often about doing it "right" this time, rather than faster this time. #genomics #rust #rewrites
141
Reposted by Nils Homer
Tim Fennell @tfenne.bsky.social · 06/07/2026
New riker release this morning - version 0.4.0 is out! Major updates are: - New "rna" tool that ports picard CollectRnaSeqMetrics, fgbio EstimateRnaInsertSize, and much more - New global --threads option to multithread input BAM/CRAM decoding - Big performance improvements in "wgs" and "hybcap"
181
Nils Homer @nilshomer.com · 05/07/2026
When can I have T2T?
010
Nils Homer @nilshomer.com · 04/07/2026
🔍 It's not byte-identical to default; --fast changes the seeding/extension heuristics. Against simulated truth, placement accuracy and recall hold (WGS within 0.06pp; high-MAPQ unchanged). Different alignments, same answers. Docs + cross-arch numbers:  github.com/fg-labs/bwa-...
github.com
docs(allowances): authorize v0.5.0 (9dd30dd0) golden bless by nh13 · Pull Request #26 · fg-labs/bwa-mem3-bench
Records the Gate #2 sign-off for promoting v0.5.0 (9dd30dd0) to the golden, per the full bless_release run (6 arches × reps=3 + holodeck accuracy + minibwa; golden gate vs v0.4.0). Golden already p...
010
Nils Homer @nilshomer.com · 04/07/2026
🚄 How fast? 📐 We use minibwa as our speed yardstick. On WGS, --fast lands within ~3% of it on Arm/NEON (a head-to-head SIMD comparison) and in the same ballpark on x86.
100
Nils Homer @nilshomer.com · 04/07/2026
🚀 bwa-mem3 v0.5.0 is out 🎉 One flag to rule them all. The new --fast preset makes whole-genome alignment ~2× faster while preserving accuracy and recall. 🧵 github.com/fg-labs/bwa-...
github.com
Release v0.5.0 · fg-labs/bwa-mem3
0.5.0 (2026-07-04) Features add opt-in --seed-order seed reordering (default off, byte-identical) (#186) (04749a1) add opt-in --smem-dedup (dedup identical SMEMs before chaining) (#187) (1384972) ...
1205
Nils Homer @nilshomer.com · 03/07/2026
I’ll have a release soon where bwa-meme3 will be as fast as minibwa. Probably over the weekend!
000
Reposted by Nils Homer
Fulcrum Genomics @fulcrumgenomics.com · 30/06/2026
New on the Fulcrum blog: minibwa, a faster mapper from @lh3lh3.bsky.social and our @nilshomer.com Its speed is great, yes, but more interesting is the decision to revisit BWA-MEM as infrastructure – keep what still works, change what limits performance, then test downstream impact. shorturl.at/xxqeI
blog.fulcrumgenomics.com
Minibwa: alignment is never solved
Heng Li and Nils Homer revisit BWA-MEM with a faster mapper for short reads, accurate long reads, and bisulfite sequencing data.
0195