Reposted by Rick BeelooRagnar {Groot Koerkamp} @curiouscoding.nl · 26/05/2026Turns out I presented Sassy yesterday without realising it has been published! Finally officially a coauthor with @rickbitloo.bsky.social 😆 doi.org/10.1093/bioi... 4177
Reposted by Rick BeeloobioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 12/03/2026Sassy2: Batch Searching of Short DNA Patterns www.biorxiv.org/content/10.64898/20… 0116
Reposted by Rick BeelooRagnar {Groot Koerkamp} @curiouscoding.nl · 13/03/2026It's a good day when the first item in your feed is your own work :) @rickbitloo.bsky.social was annoyed that scanning reads for all 96 rapid kit barcodes is bottleneck in Barbell, so he made Sassy2: 13x (150bp) to 4.6x (8kbp) faster than v1 by batch-searching patterns, and >100Gbp/s on 16 threads! 22411
Reposted by Rick BeelooRagnar {Groot Koerkamp} @curiouscoding.nl · 10/12/2025If you ever need to fuzzy search some DNA, sassy is your tool. Please spread the word; I think many people just outside my own circle could benefit from this :) cc @rickbitloo.bsky.social github.com/RagnarGrootK... 44024
Reposted by Rick BeelooRagnar {Groot Koerkamp} @curiouscoding.nl · 30/10/2025Following ish's `filter` and bqtools' `grep`, Sassy now also has initial support for grep and filter! Grep mode shows all matches, grouped per record, and is meant for human consumption. Filter mode prints full matching (or non-matching) records to stdout or output files. 1155
Rick Beeloo @rickbitloo.bsky.social · 25/10/2025Was just checking if we could add the adapter search directly, then "slice" out all sub-reads so splitting does still look at the barcodes. And thanks for the discussions, nice to get ideas from what else to improve :) 020
Rick Beeloo @rickbitloo.bsky.social · 25/10/2025Indeed does happen, but more often they are duplex reads that are not detected as such by Dorado. Then the sequences on either side of the mid strand adapter are reverse complement (example pic) 102
Rick Beeloo @rickbitloo.bsky.social · 25/10/2025With Flye? Aah in the preprint we actually used Flye with Dorado trimmed reads 😅. You can see a list of "contaminated" assemblies here: zenodo.org/records/1739..., still quite a few, and this is very strict filtering. In reality there will be more 000
Rick Beeloo @rickbitloo.bsky.social · 24/10/2025Yeah Porechop is probably "necessary" when using Dorado, though it shouldn't be. "Appreciable" heavily depends on the application. For assembly a few reads wrong is not always an issue, but for diagnostics a few reads wrong can be a big issue. Of course one should also check scores then, etc. 140
Rick Beeloo @rickbitloo.bsky.social · 24/10/2025Hey Misha, on a 10K read sample around 60% is the expected dual-end (top), but often a single-end is enough to assign already. That puts it at around 90% (top 3 in image). Double barcode ligations do happen (bottom) but not sure about the exact stats of bleeding because of that. 020
Rick Beeloo @rickbitloo.bsky.social · 24/10/2025To add, you can then add the desired pattern to barbell filter where you put “cut” marks (<<, >>) such that only the long reads is kept and the short contam is cut off 010
Rick Beeloo @rickbitloo.bsky.social · 24/10/2025Agree, Barbell will just give a "complex" pattern for those reads which are not added to the output by default. Though, if the reads are very long, and contain little contamination (i.e. long read + short concat read at the end) it might still be worth it including it for assembly 120
Rick Beeloo @rickbitloo.bsky.social · 24/10/2025Yeah most of NCBI is down, or outdated at the moment. We will upload them to Zenodo 031
Rick Beeloo @rickbitloo.bsky.social · 24/10/2025Hey Kevin, that comes down to the difference between demultiplexing correct and removing contamination. If you had a read with NB01-NB02--read--. Dorado will demux to NB01, and you can remove NB02 contamination using Porechop, but that does mean your read was still demultiplexed incorrect initially 130
Reposted by Rick BeelooRagnar {Groot Koerkamp} @curiouscoding.nl · 23/10/2025Really exciting that the preprint on Barbell, a new demultiplexer, is finally out! It's the first tool that builds on Sassy, the approximate-DNA-searching tool that @rickbitloo.bsky.social and myself developed earlier this year, specifically with this application in mind. 22015
Rick Beeloo @rickbitloo.bsky.social · 23/10/2025See the thread for a quick summary! bsky.app/profile/rick... 000
Reposted by Rick BeeloobioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 23/10/2025Barbell Resolves Demultiplexing and Trimming Issues in Nanopore Data www.biorxiv.org/content/10.1101/202… 163
Rick Beeloo @rickbitloo.bsky.social · 23/10/2025In the pre-print we also discuss a specific tagmentation pattern causing partial loss of the second barcode, a custom barcode scoring scheme, and more: www.biorxiv.org/content/10.1... 030
Rick Beeloo @rickbitloo.bsky.social · 23/10/2025In Barbell we solve this by first annotating all the reads, and then detecting all patterns, which looks like this: 130
Rick Beeloo @rickbitloo.bsky.social · 23/10/2025Why is this a problem? Remaining adapters/barcodes not only contaminated assemblies but also created "artificial" links in taxonomic annotation, to Enterobacteriaceae, and to contaminated assemblies in NCBI. 150
Rick Beeloo @rickbitloo.bsky.social · 23/10/2025Many tools, including the widely used Dorado, almost always only detected the *first* occurrence, leaving the rest untrimmed. In the figure black lines indicate matches to adapters + barcodes in trimmed reads. 120
Rick Beeloo @rickbitloo.bsky.social · 23/10/2025For rapid barcoding only ~83% of the reads contain the expected single barcode on the left. The rest? Barcodes on both sides (6.1%), two barcodes on the left side (3.5%), and so on. 120
Rick Beeloo @rickbitloo.bsky.social · 23/10/2025Around 10% of your Nanopore reads (SQK-RBK114) are incorrectly trimmed. Here is why, and how our new tool Barbell solves it: www.biorxiv.org/content/10.1... Want to get started? github.com/rickbeeloo/b... 35231
Reposted by Rick BeelooRagnar {Groot Koerkamp} @curiouscoding.nl · 26/07/2025Now also on biorxiv :) 152
Reposted by Rick BeelooRagnar {Groot Koerkamp} @curiouscoding.nl · 18/07/2025Sassy is out now! Ever need to search for approximate matches of short DNA strings? Sassy is the tool to use! Available now wherever you get your code With @rickbitloo.bsky.social curiouscoding.nl/papers/sassy... github.com/ragnarGrootK... 23922