Sign in

Eddie Ruiz

@ed2uiz.bsky.social
58 followers 674 following 17 posts

ed2.dev

PostsRepliesMedia
Reposted by Eddie Ruiz
SMT @pp0196.bsky.social · 15/08/2026
This is great work, happy to see DuckDB usage in this space. Related self plug, i made a HTS file formats reader extension rgenomicsetl.r-universe.dev/Rduckhts that allow fast imports directly into duckb and avoid RSamtools and al. There are community readers for single cell formats
rgenomicsetl.r-universe.dev
Rduckhts: 'DuckDB' High Throughput Sequencing File Formats Reader Extension
Bundles the 'duckhts' 'DuckDB' extension for reading High Throughput Sequencing file formats with 'DuckDB'. The 'DuckDB' C extension API <https://duckdb.org/docs/stable/clients/c/api> and its 'htslib'...
111
Eddie Ruiz @ed2uiz.bsky.social · 15/08/2026
Thanks and very cool package, I'll have to check it out. Do you know how it compares versus wheretrue/exon or abdenlab/oxbow, both of which I think build on zaeleus/noodles and datafusion?
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
I’m deeply grateful to my coauthors and collaborators for making this work possible, and to our funding sources for their support. Questions, feedback, and contributions are welcome. Thanks for reading! 🪐 16/16
000
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
dbverse builds on many open-source projects, including @duckdb.org, #DataFusion, #ApacheArrow, and the #RStats ecosystem (dbplyr, dplyr, DBI, and more). If you're interested in contributing to dbverse for Python or Julia, I’d love to connect (narwhals and @tidierjl.bsky.social teams👀) 15/🧵
210
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
The dbverse framework is free and open-source. Feedback, contributions, and PRs are welcome. Preprint: www.biorxiv.org/content/10.6... Source: github.com/dbverse-org/... Docs: dbverse-org.github.io/dbverse/ GiottoDB: github.com/giotto-suite... GiottoSeq: github.com/giotto-suite... 14/🧵
github.com
GitHub - dbverse-org/dbverse
Contribute to dbverse-org/dbverse development by creating an account on GitHub.
120
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
Larger scientific datasets often push us toward specialized hardware or unfamiliar software stacks. These can be difficult to adopt. OLAP engines offer a practical alternative. dbverse shows how we can use them to power scientific APIs and enable complex analyses on ordinary computers. 13/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
These operations enabled us to explore poly(A)-site usage in the Visium HD 3' OCCC sample. We observed patterns of 3'UTR shortening in tumor-associated epithelial cells (TAEPi), as well as shifts in poly(A)-site usage for cancer associated fibroblasts (CAFs) closer to or farther from TAEPi. 12/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
Compared to movAPA (Ye et al 2021), GiottoSeq was up to 3× faster >=100K cells at calculating the RUD score in the Visium HD 3' data. At 2.1M cells, GiottoSeq completed and movAPA failed (OOM; out of memory). 11/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
We next turned to alternative polyadenylation (APA) analysis in the Visium HD 3' ovarian sample. We evaluated computation of the relative usage of the distal poly(A) site (RUD). 3'UTRs harbor regulatory motifs that can impact mRNA and protein fate through co/post-transcriptional reg. 10/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
We first tested end-to-end spatial omics analysis, spanning many operations. GiottoDB completed 10M cells. The in-memory Giotto failed (OOM; out of memory). The HDF5-backed TENxMatrix path was ~39× slower than GiottoDB at 100K cells and exceeded the 1,000s time limit for 1M+ cells. 9/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
We next integrated dbverse into Giotto Suite to demonstrate its ability to scale an existing #spatialomics analysis framework. We achieved this by developing GiottoDB + GiottoSeq, which enabled joint analysis of gene expression and alt. polyadenylation (APA) on a Visium HD 3' ovarian sample. 8/🧵
110
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
Genomic operations showed similar scaling improvements. dbSequence was able to compute the overlap between a fixed chromosomal region and up to one billion genomic intervals within our 600s time window. It outperformed GenomicRanges and plyranges after one million genomic intervals. 7/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
The speedups were even larger for spatial operations. We evaluated the time to intersect up to one billion points (e.g. molecules in a tissue) with 100K polygons (e.g. cells). The sf R package timed out within our 600s window at the two largest tested sizes, whereas dbSpatial completed both. 6/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
But do OLAP databases actually improve performance of large-scale scientific operations? To answer this, we performed several benchmarks. First, for column means calculation in sparse matrices, dbMatrix was the fastest method at 2 billion non-zero values, up to 3x faster than BPCells. 5/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
Standalone libraries represent matrix (dbMatrix), spatial (dbSpatial), and genomic data (dbSequence) as compressed tables within an OLAP database. Because of the ORM design, dbverse functions can serve as drop-in replacements for over 100 operations from commonly used #RStats packages! 4/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
To solve this problem, the dbverse framework represents scientific data as object-relational mappings (ORMs) within embedded OLAP engines like @duckdb.org and #apachedatafusion The key idea 💡 Scientific APIs ➡️ optimized SQL queries ➡️ local OLAP databases scale to larger-than-memory workloads. 3/🧵
100
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
Scientific datasets now frequently exceed RAM on ordinary computers. Spatial omics datasets illustrate this problem clearly, with datasets often exceeding RAM at 10–100× the memory footprint of earlier bulk and single-cell omics datasets. 2/🧵
110
Eddie Ruiz @ed2uiz.bsky.social · 14/08/2026
Can databases scale scientific analyses beyond RAM? I’m excited to share dbverse, a composable database framework built to do just that. Across benchmarks, dbverse delivers ~1000x speedups and scales #spatialomics analyses to millions of cells. Preprint: www.biorxiv.org/content/10.6... 1/🧵
biorxiv.org
dbverse scales spatial omics analysis with embedded analytical databases
Spatial omics datasets are increasing in size and complexity, exceeding the memory of standard computers and thereby limiting data analysis. Here we present dbverse, a framework for larger-than memory...
242