Sign in

Gregor Sturm

@grst.bsky.social
985 followers 358 following 44 posts

Single Cell/Spatial. Cancer Immunology. Outdoor activities. Core developer @scverse.bsky.social. Working in Clinical Bioinformatics at Boehringer Ingelheim. Formerly PhD student at Medical University of Innsbruck. My private account. github.com/grst

PostsRepliesMedia
Gregor Sturm @grst.bsky.social · 05/08/2026
It turns out @nextflow.io is useful outside bioinformatics. Also kudos to the karttapullautin devs for their awesome tool, which the pipeline relies on for the tile rendering step. github.com/karttapullau...
github.com
GitHub - karttapullautin/karttapullautin: Source code of the rust implementation of Karttapullautin
Source code of the rust implementation of Karttapullautin - karttapullautin/karttapullautin
010
Gregor Sturm @grst.bsky.social · 05/08/2026
Source code and data are freely available here: github.com/grst/mapant-...
github.com
GitHub - grst/mapant-bayern: Automatically generated orienteering map of Bavaria
Automatically generated orienteering map of Bavaria - grst/mapant-bayern
100
Gregor Sturm @grst.bsky.social · 05/08/2026
Using a Nextflow pipeline, I processed 15 TB of LIDAR data + OpenStreetMap shapes into a 180GB tile pyramid.
100
Gregor Sturm @grst.bsky.social · 05/08/2026
This time not bioinformatics, but geoinformatics: I created MapAnt Bayern, an automatically generated orienteering map of Bavaria: mapant.orienteering-allgaeu.de
131
Gregor Sturm @grst.bsky.social · 27/07/2026
@daweonline.bsky.social does that mean that schist becomes usable on larger datasets now?
220
Gregor Sturm @grst.bsky.social · 29/05/2026
Also thanks to @francescafinotello.bsky.social for supporting this and @thehumanborch.bsky.social for providing early feedback!
020
Gregor Sturm @grst.bsky.social · 29/05/2026
This started as a collaboration with biocypher (@slobentanzer.bsky.social) at a @scverse.bsky.social hackathon, was then advanced by @valeriiadragan.bsky.social as a master thesis project and finally pushed over the finish line by Raphael De Gottardi
222
Gregor Sturm @grst.bsky.social · 29/05/2026
Happy to announce the release of IggyTop, a metadatabase for immune receptor–epitope interactions. It's integrated with scirpy v0.24 so you can readily use it to annotate your single-cell TCR datasets. iggytop.readthedocs.io/en/latest/
198
Gregor Sturm @grst.bsky.social · 17/04/2026
sorry, I meant min, not max
000
Gregor Sturm @grst.bsky.social · 17/04/2026
Thanks for the response! The thing is that the original TCRdist uses max(4, 4-score) so it caps scores at 4. With TCRblosum that would be max(2, 2-score) which would destroy all the strong signal that you have in e.g. the C residue. So you'd suggest to use 2-score without the max instead?
200
Gregor Sturm @grst.bsky.social · 14/04/2026
See github.com/scverse/scir... for more detais.
github.com
Add tcrblosum support to TCRdist by felixpetschko · Pull Request #685 · scverse/scirpy
So far, the TCRdist metric used a distance matrix derived from the blosum62 substitution matrix. This PR extends TCRdistDistanceCalculator with a new base_matrix="tcrblosum" option alongs...
000
Gregor Sturm @grst.bsky.social · 14/04/2026
Hi @pmeysman.bsky.social‬, we are trying to integrate tcrBLOSUM into scirpy, our library for scTCRseq analysis. Specifically, we want to adapt the TCRdist algorithm to use tcrBLOSUM substitution values. How did you turn the substitution matrix into a distance matrix in your paper?
200
Gregor Sturm @grst.bsky.social · 22/08/2025
There's another scverse conference this year and it will be amazing! Register now: www.eventbrite.com/e/scverse-co...
021
Gregor Sturm @grst.bsky.social · 19/08/2025
AFAIK, these differences are minor, numeric differences. I would consider them equivalent.
130
Gregor Sturm @grst.bsky.social · 13/08/2025
Our benchmark + guidelines for atlas-level differential gene expression of single cells is online: academic.oup.com/bib/article/... Bottom line: Use pseudobulk + DESeq2 in simple and pseudobulk + DREAM in more complex settings. Collab w/ @leonhafner.bsky.social @itisalist.bsky.social
1156
Gregor Sturm @grst.bsky.social · 05/08/2025
Register now for the best conference of the year!
000
Reposted by Gregor Sturm
scverse @scverse.bsky.social · 12/05/2025
📣 Mark your calendars! The 2025 edition of the scverse conference will take place on 17-19 November at Stanford University (US) scverse.org/conference20... Call for abstracts and registrations coming soon!
scverse.org
scverse conference 2025
Follow us on our channels to learn more details in the coming weeks
1129
Gregor Sturm @grst.bsky.social · 02/04/2025
Just released a new version of the @scverse.bsky.social cookiecutter template: github.com/scverse/cook... Some highlights: 🔃 improved template sync (merge conflicts now show up as such) 🚀 use hatch as project manager 🔧 lots of fixes and documentation updates
github.com
Release v0.5.0 · scverse/cookiecutter-scverse
New template sync We re-implemented template sync from scratch instead on relying on cruft. This allows us to create real merge conflicts that show up as such on GitHub instead of .rej files. Gene...
042
Reposted by Gregor Sturm
Stephen Turner @stephenturner.us · 14/03/2025
rogue-scholar.org
rogue-scholar.org
Rogue Scholar
011
Gregor Sturm @grst.bsky.social · 14/03/2025
Nice post! How did you generate the doi-link for a blog post?
100
Reposted by Gregor Sturm
Wolfgang Huber @wkhuber.bsky.social · 04/03/2025
Blog post by @const-ae.bsky.social with a simple explanation of the manifold regression algorithm & code that underlies our paper “Analysis of multi-condition single-cell data with latent embedding multivariate regression” (doi.org/10.1002/eji....). const-ae.name/post/2025-01...
const-ae.name
LEMUR simplified | const-ae
A simplified implementation of the LEMUR algorithm.
1255
Gregor Sturm @grst.bsky.social · 25/02/2025
Just released scirpy v0.21 -- Now with GPU Support for Hamming sequence distance and a brand new tutorial for working with scTCR datasets >1M cells: scirpy.scverse.org/en/latest/tu... @scverse.bsky.social
scirpy.scverse.org
Working with >1M cells
Scirpy scales to millions of cells on a single workstation. This page is a work-in-progess collection with advice how to work with large datasets. Distance metrics: Computing pairwise sequence dist...
020
Reposted by Gregor Sturm
scverse @scverse.bsky.social · 14/02/2025
🎉 Scanpy 1.11.0 is out! 🎉 just after reaching 2000 stars on GitHub! - sc.pp.sample replaces subsample with many new features - Sparse Dask support pca - session-info2 package for more reproducible notebooks See the release notes:
buff.ly
Release notes
Version 1.11: 1.11.0 2025-02-14: Release candidates: rc2 2025-01-24, rc1 2024-12-20. Features: rc1 sample() supports both upsampling and downsampling of observations and variables. subsample() is n...
14919
Reposted by Gregor Sturm
Edmund Miller @edmundmiller.dev · 09/02/2025
Been looking forward to this talk since @alexpeltzer.bsky.social told me about DSO in October!
041
Reposted by Gregor Sturm
Gregor Sturm @grst.bsky.social · 05/02/2025
I'd like to share DSO, a command line helper to build reproducible data science projects with ease. It is an opinionated way to organize data science projects, built around data version control (DVC). github.com/Boehringer-I...
github.com
GitHub - Boehringer-Ingelheim/dso: Data Science Operations (dso) command line tool
Data Science Operations (dso) command line tool. Contribute to Boehringer-Ingelheim/dso development by creating an account on GitHub.
1124
Gregor Sturm @grst.bsky.social · 05/02/2025
We try to avoid that by using this with preprocessed data only. All the heavy lifting is done with nextflow pipelines before. Datasets up to tens of GBs have worked well so far.
020
Gregor Sturm @grst.bsky.social · 05/02/2025
Finally, many thanks to my colleagues @alexpeltzer.bsky.social, Daniel Schreyer and Tom Schwarzl for testing, adopting, and contributing to DSO.
110
Gregor Sturm @grst.bsky.social · 05/02/2025
If you want to learn more, I'll be presenting this at a @nf-co.re bytesize talk: nf-co.re/events/2025/...
nf-co.re
Bytesize: data science operations (DSO)
Gregor Sturm, Boehringer Ingelheim
152
Gregor Sturm @grst.bsky.social · 05/02/2025
We built this at @boehringerglobal.bsky.social to meet the quality standards required for biomarker analysis in clinical trials. But I think this is useful for any kind of data analysis project.
110
Gregor Sturm @grst.bsky.social · 05/02/2025
One of my favorite features: automated watermarking of all plots in a quarto report. Nobody gonna publish my plots anymore before I think they are ready.
An exemplary PCA plot with a "preliminary" watermark.
130
Gregor Sturm @grst.bsky.social · 05/02/2025
It brings together the best tools: - git, for code versioning - dvc, for data versioning and tracking inputs and outputs - jinja2, for templates - uv, for Python dep mgmt - quarto, for authoring reports - hiyapyco, for hierarchical YAML config - pre-commit, for linting
110
Gregor Sturm @grst.bsky.social · 05/02/2025
I'd like to share DSO, a command line helper to build reproducible data science projects with ease. It is an opinionated way to organize data science projects, built around data version control (DVC). github.com/Boehringer-I...
github.com
GitHub - Boehringer-Ingelheim/dso: Data Science Operations (dso) command line tool
Data Science Operations (dso) command line tool. Contribute to Boehringer-Ingelheim/dso development by creating an account on GitHub.
1124
Reposted by Gregor Sturm
Stefano Mangiola @stemang.bsky.social · 22/01/2025
We (Chen Zhan!) just launched #sccomp for #Python! Testing for differences in cell-type proportion in #singlecell #spatial data? #sccomp is a mixed-effect Bayesian model - Use sum-constrained BetaBinomial distribution - Outliers detect. - Remove unwanted effects github.com/MangiolaLabo...
1123
Gregor Sturm @grst.bsky.social · 19/01/2025
(2) Finding the mistake, tracing it back to its origin, and fixing it was only possible because the data and scripts for building the atlas are publicly available and fully reproducible. github.com/icbi-lab/luca
github.com
GitHub - icbi-lab/luca: Single-cell Lung Cancer Atlas with 1.2M cells
Single-cell Lung Cancer Atlas with 1.2M cells. Contribute to icbi-lab/luca development by creating an account on GitHub.
020
Gregor Sturm @grst.bsky.social · 19/01/2025
(1) Maintaining a data resource is very much like maintaining software. It is never "done" but constantly improving.
120
Gregor Sturm @grst.bsky.social · 19/01/2025
Two years after publication of our single-cell lung cancer atlas, a user found a mistake in the annotation of the EGFR-status of some patients. We fixed the issue and the atlas is now updated on cell-x-gene: cellxgene.cziscience.com/collections/... What are the takeaways from that? (1/3)
cellxgene.cziscience.com
Cellxgene Data Portal
Find, download, and visually explore curated and standardized single cell datasets.
273
Reposted by Gregor Sturm
Lukas Heumos @lukasheumos.bsky.social · 17/01/2025
I am Stoked about our upcoming @scverse.bsky.social and @owkin.bsky.social hackathon, focused on spatial omics data analysis. 📅 March 17-19, 2025 📍 Owkin office, Paris Apply now: docs.google.com/forms/d/e/1F...
docs.google.com
Scverse x Owkin Hackathon in Paris
We're pleased to announce the next Scverse Hackathon will take place in the Owkin offices in Paris from 17/03/2025 9am to 19/03/2025 1:30pm. This hackathon is a joint initiative between the scverse c...
0108
Gregor Sturm @grst.bsky.social · 14/01/2025
protein sequencing 👀
020
Reposted by Gregor Sturm
Constantin Ahlmann-Eltze @const-ae.bsky.social · 03/01/2025
After 4y in the making, I am super excited that my main PhD project is published 🎉🥳🎉🎉🥳 www.nature.com/articles/s41... LEMUR is a tool to analyze multi-condition single-cell data and model differential expression as a continuous function of the cell-state space. Some highlights⬇️
Overview of the LEMUR steps: (1) subspace alignment, (2) differential expression, (3) DE neighborhoods, (4) pseudobulking.
815435
Gregor Sturm @grst.bsky.social · 02/01/2025
The big issue here in Germany is that we pay ~20 ct/kWh in fixed network fees and tax. Really limits how much you can save.
100
Gregor Sturm @grst.bsky.social · 02/01/2025
While definitely interesting, dynamic tariffs have a much higher cost-saving potential for devices that consume a lot of energy and are easy to regulate automatically. Such as a heat pump or electric car - of which we have neither, for now.
100
Gregor Sturm @grst.bsky.social · 02/01/2025
There's a certain price risk, though. In the energy crisis 2021-2022 prices increased significantly. However prices in 2023 were down to normal, while many fixed price tariffs increased their rates.
100
Gregor Sturm @grst.bsky.social · 02/01/2025
Dynamic electricity tariffs are an incentive to use energy when it's abundant and emits little CO2. But are they also cheaper? ✅ For 2024, we would have saved 10-13% compared to our current fixed price tariff. Without any optimization. Full post (in German): grst.github.io/dynamischer-...
Monthly comparison of fixed price tariff (AÜW) with dynamic tariff (tado).
120
Gregor Sturm @grst.bsky.social · 30/12/2024
Modern tar detects the compression automatically when reading from a file. So `tar xvf` covers most of the cases already.
040
Reposted by Gregor Sturm
khrovatin.bsky.social @khrovatin.bsky.social · 16/12/2024
To bring to light data science topics that usually don’t make it into publications I started a blog on this topic: hrovatin.github.io By interviewing different researchers, I plan to find out what is going on in the community.
hrovatin.github.io
Karin Hrovatin
Data science blog on topics that don’t get published.
1106
Reposted by Gregor Sturm
Gregor Sturm @grst.bsky.social · 07/12/2024
Formulaic is the go-to way to specify design formulas in Python, e.g. ~treatment + timepoint. To compare sth, one needs to specify a contrast, e.g "on treatment vs baseline". To make this easier, we developed "formulaic-contrasts": formulaic-contrasts.readthedocs.io/en/latest/
# Define model with interaction term
model = MyModel(data, "~ treatment * timepoint")

# compare timepoints
contrast = model.cond(timepoint="on_treatment") - model.cond(timepoint="baseline")

# compare timepoints within drugA only
contrast = (
  model.cond(treatment="drugA", timepoint="on_treatment") - 
  model.cond(treatment="drugA", timepoint="baseline")
)

# compare interaction of timepoint with treatment 
# (= difference of changes between both treatments)
contrast = (
    mod.cond(treatment="drugB", timepoint="on_treatment")
    - mod.cond(treatment="drugB", timepoint="baseline")
) - (
    mod.cond(treatment="drugA", timepoint="on_treatment")
    - mod.cond(treatment="drugA", timepoint="baseline")
)
1186
Gregor Sturm @grst.bsky.social · 07/12/2024
Formulaic-contrasts was originally conceived at the @scverse.bsky.social hackathon in Cambridge based on an idea by @const-ae.bsky.social. It is currently (being) integrated into #pyDESeq2, #pertpy and #pyLemur.
020
Gregor Sturm @grst.bsky.social · 07/12/2024
If you develop a model that understands a design matrix and contrasts vectors, it's easy to integrate with formulaic and fomulaic-contrasts: formulaic-contrasts.readthedocs.io/en/latest/mo...
formulaic-contrasts.readthedocs.io
Usage in custom model
Via Inheritance: The most straightforward way to use formulaic-contrasts with a custom model is to use FormulaicContrasts as a base class or mixin class. As an example, let’s wrap an Ordinary Least...
100
Gregor Sturm @grst.bsky.social · 07/12/2024
Formulaic is the go-to way to specify design formulas in Python, e.g. ~treatment + timepoint. To compare sth, one needs to specify a contrast, e.g "on treatment vs baseline". To make this easier, we developed "formulaic-contrasts": formulaic-contrasts.readthedocs.io/en/latest/
# Define model with interaction term
model = MyModel(data, "~ treatment * timepoint")

# compare timepoints
contrast = model.cond(timepoint="on_treatment") - model.cond(timepoint="baseline")

# compare timepoints within drugA only
contrast = (
  model.cond(treatment="drugA", timepoint="on_treatment") - 
  model.cond(treatment="drugA", timepoint="baseline")
)

# compare interaction of timepoint with treatment 
# (= difference of changes between both treatments)
contrast = (
    mod.cond(treatment="drugB", timepoint="on_treatment")
    - mod.cond(treatment="drugB", timepoint="baseline")
) - (
    mod.cond(treatment="drugA", timepoint="on_treatment")
    - mod.cond(treatment="drugA", timepoint="baseline")
)
1186