Sign in

Luke Zappia

@lazappi.bsky.social
1.1K followers 335 following 55 posts

Bioinformatician, data scientist, software developer Also @_lazappi_ and @lazappi@mastodon.au

PostsRepliesMedia
Reposted by Luke Zappia
Yihui Xie @yihui.org · 21/09/2026
After a massive cleanup of knitr issues/PRs in the past few days, I can finally stop being jealousy of Will Landau for his zero open Github issues in {targets}: yihui.org/en/2026/09/k... For the record, he has one open issue at the moment and so do I (debating with myself whether I should close it).
yihui.org
A Four-Day knitr Issue/PR Backlog Sprint - Yihui Xie | 谢益辉
After knitr v1.52 went to CRAN on 2026-09-06, I did something I don’t do often but have been hoping to do: I sat down and worked through the GitHub issue tracker more or less top to bottom. Over...
2252
Reposted by Luke Zappia
Bioconductor @bioconductor.bsky.social · 04/08/2026
📢 Join BioC2026 online! Although in-person registration for BioC2026 is now closed, you can still join the conference virtually from anywhere in the world. 💻 Virtual registration closes: August 12 Secure your virtual ticket today: 🎟️ ti.to/nf-projects/...
074
Luke Zappia @lazappi.bsky.social · 10/06/2026
Thanks to my co-maintainers @louisedck.bsky.social and @rcannood.bsky.social, @scverse.bsky.social and Data Intuitive for their support and everyone in the community for their contributions!
010
Luke Zappia @lazappi.bsky.social · 10/06/2026
Excited that the {anndataR} 📦 publication is out academic.oup.com/bioinformati... 🎉! I hope this makes it easier to move single-cell data between R and Python and encourages people to take advantage of both languages. Check it out on @bioconductor.bsky.social bioconductor.org/packages/ann....
13011
Reposted by Luke Zappia
Luke Zappia @lazappi.bsky.social · 12/05/2026
My quick wrap up of the recent @bioconductor.bsky.social 3.23 release focusing on single-cell and other things I'm interested in lazappi.id.au/posts/2026-0...
lazappi.id.au
Bioconductor 3.23 wrap-up – lazappi
My wrap-up of the Bioconductor 3.23 release
051
Luke Zappia @lazappi.bsky.social · 12/05/2026
My quick wrap up of the recent @bioconductor.bsky.social 3.23 release focusing on single-cell and other things I'm interested in lazappi.id.au/posts/2026-0...
lazappi.id.au
Bioconductor 3.23 wrap-up – lazappi
My wrap-up of the Bioconductor 3.23 release
051
Reposted by Luke Zappia
Dr Di Cook @visnut.bsky.social · 17/04/2026
#rstats SSA SCV section is excited to bring a panel discussion on “where do R packages live?” May 20, 5pm AEST Sign up to join at statsoc.org.au/event-6653060 @fontikar.bsky.social
1106
Luke Zappia @lazappi.bsky.social · 23/04/2026
I mainly worked on trying to finalise Zarr support in anndataR so it can be used by R SpatialData. Check out the PR here, hopefully merged soon github.com/scverse/annd....
github.com
Add Zarr support by keller-mark · Pull Request #190 · scverse/anndataR
Fixes #91 These changes are from both me and @Artur-man The main public-facing changes here are: The ZarrAnnData class read_zarr and write_zarr top-level functions Support for from_Seurat(output_c...
000
Luke Zappia @lazappi.bsky.social · 23/04/2026
Last week I attended a SpatialData hackathon with members of the @scverse.bsky.social, @bioconductor.bsky.social and @openmicroscopy.org communities in Padua, Italy. Here is a short blog post about the experience www.data-intuitive.com/insights/blo....
data-intuitive.com
2nd SpatialData Hackathon – Data Intuitive
Frameworks, formats and Interoperability
163
Reposted by Luke Zappia
Maëlle Salmon @masalmon.eu · 21/04/2026
New in Git, ✨ git history ✨ (I don't mean the Git history 🫠 ) 😰 Oh shit, that old commit's message had a typo! 😌 git history reword <commit-id> 😰 Oh shit, that old commit is too big! 😌 git history split <commit-id> (to split it into two commits) github.blog/open-source/... (h/t Hugo Gruson)
github.blog
Highlights from Git 2.54
The open source Git project just released Git 2.54. Here is GitHub’s look at some of the most interesting features and changes introduced since last time.
2369
Reposted by Luke Zappia
Surag Nair @suragnair.bsky.social · 13/04/2026
Agents for comp bio are advancing rapidly, but evals are lagging. Current benchmarks can be overly prescriptive. Full analysis vignettes are hard to verify. We introduce CompBioBench: 100 diverse, challenging, verifiable tasks. We benchmark Codex and Claude Code. biorxiv.org/content/10.6... 1/9
196
Luke Zappia @lazappi.bsky.social · 07/04/2026
Adapting to all the constant new features in the Python package led us to come up with a flexible wrapping approach via reticulate. This lets us expose new functionality automatically while adapting or replacing when needed to work nicely in R.
010
Luke Zappia @lazappi.bsky.social · 07/04/2026
LaminR 📦 has been officially announced! This is a project @rcannood.bsky.social and I have worked on to create an R client for @laminlabs.bsky.social. 📦 Repo github.com/laminlabs/la... Announcement blog blog.lamin.ai/intro-to-lam... DI blog www.data-intuitive.com/insights/blo...
github.com
GitHub - laminlabs/laminr: An R package for working with LaminDB instances.
An R package for working with LaminDB instances. Contribute to laminlabs/laminr development by creating an account on GitHub.
142
Reposted by Luke Zappia
Dr. Jean Fan @jef.works · 25/02/2026
We previously showed how changing training data alone can improve deep learning prediction of spatial transcriptomics gene expression from histology images by +38% (without any changes to model architecture). We've now updated our preprint w/ expanded results: www.biorxiv.org/content/10.1... 🧵1/n
132
Luke Zappia @lazappi.bsky.social · 17/02/2026
Nice blog post from @thomas-sandmann.genomic.social.ap.brid.gy testing out {anndataR} 📦 tomsing1.github.io/blog/posts/a... 🎉! Cool to see people trying it out, find it on @bioconductor.bsky.social if you want to give it a go bioconductor.org/packages//re...
tomsing1.github.io
anndataR: converting single-cell data, from R to python and back – Thomas Sandmann’s blog
041
Luke Zappia @lazappi.bsky.social · 05/02/2026
If you work with Docker a lot, here is your friendly reminder to occasionally run `docker system prune`. I just freed up 417 GB so I have space to work again 🗑️.
010
Luke Zappia @lazappi.bsky.social · 23/01/2026
I'll been at #FOG2026 in London next week 🧬. If you are attending and want to have a chat feel free to get in touch 🗣️.
Festival of genomics social media banner showing I am attending
021
Reposted by Luke Zappia
Andrew Heiss @andrew.heiss.phd · 13/01/2026
This was fun! I just posted all the links and resources we talked about at www.andrewheiss.com/blog/2026/01... - check it out to learn more about Raycast, Espanso, and neat Positron/VS Code extensions like Peacock and Pastum and Project Manager #rstats #databs
andrewheiss.com
How to make your data analysis life easier using Positron, Raycast, and Espanso | Andrew Heiss
Links and resources from my time on Posit’s Data Science Lab in January 2026
56012
Reposted by Luke Zappia
James Bonfield @jbonfield.bsky.social · 15/09/2025
Heads up: ignore samtools dot org, similarly minimap2 dot com and likely others. It's owned by a known phishing site and while the binaries they offer look valid currently (but note they may be serving us different binaries to others), that could change. Ie: it's not us (Samtools team)! Be warned
2144126
Reposted by Luke Zappia
Stephen Turner @stephenturner.us · 29/08/2025
anndataR enables seamless R-Python interoperability for single-cell RNA-seq by reading, writing, and converting H5AD files, supporting Seurat and SingleCellExperiment formats www.biorxiv.org/content/10.1... 🧬🖥️🧪 #Rstats github.com/scverse/annd...
0197
Reposted by Luke Zappia
Louise Deconinck @louisedck.bsky.social · 25/08/2025
A very big thank you to all co-authors and collaborators! @lazappi.bsky.social @rcannood.bsky.social Martin Morgan @scverse.bsky.social @ivirshup.bsky.social Chananchida Sang-aram Danila Bredikhin Brian Schilder Ruth Seurinck @yvansaeys.bsky.social @saeyslab.bsky.social
031
Luke Zappia @lazappi.bsky.social · 26/08/2025
Hopefully coming soon to @bioconductor.bsky.social!
031
Luke Zappia @lazappi.bsky.social · 26/08/2025
It's taken more than 2 years but we can officially announce anndataR! There's still a lot of features to add but we hope a robust, "official", R-native H5AD reader/writer will unlock the defacto single-cell storage format for R users by avoiding issues with current solutions like zellkonverter.
1205
Reposted by Luke Zappia
Louise Deconinck @louisedck.bsky.social · 25/08/2025
We're excited to share that our preprint on anndataR, a new package bringing Python's AnnData to R, is now available on bioRxiv 🎉 🔗 Read the paper: www.biorxiv.org/content/10.1... 💻 Check the package in action: anndatar.data-intuitive.com
2257
Reposted by Luke Zappia
David Shiffman, Ph.D. 🦈 @whysharksmatter.bsky.social · 17/06/2025
This week in my graduate-level Science Communication course, we are discussing communicating science to the public through public speaking (i.e., museum "evening with a scientist" nights, Nerd Nite lectures, speaking at community events, etc). What advice do you have to share with my students? 🧪
20563
Reposted by Luke Zappia
Michael Love @mikelove.bsky.social · 30/05/2025
Bioconductor moving to Zulip for community chat stat.ethz.ch/pipermail/bi...
stat.ethz.ch
[Bioc-devel] Bioconductor is moving from Slack to Zulip – 2 June
2169
Reposted by Luke Zappia
Marnie Blewitt @marnieblewitt.bsky.social · 14/05/2025
www.wehi.edu.au/careers/make... WEHI is looking for new lab heads! New and experienced lab heads welcome. Check out the ad, consider applying and share. We’re a collaborative bunch, in Australia, with fabulous technology, great colleagues at a values-driven organisation. @wehi-research.bsky.social
wehi.edu.au
Make your future Melbourne | WEHI
Lead the next wave of biomedical discovery
02528
Reposted by Luke Zappia
Lisa Sikkema @lisasikkema.bsky.social · 03/06/2025
Analyzing your single-cell data by mapping to a reference atlas? Then how do you know the mapping actually worked, and you’re not analyzing mapping-induced artifacts? We developed mapQC, a mapping evaluation tool www.biorxiv.org/content/10.1... from the ‪@fabiantheis lab. Let’s dive in🧵
22410
Reposted by Luke Zappia
Matthias Stahl @higsch.com · 27/05/2025
We digitized the AfD report of the Federal Office for the Protection of the Constitution. This (secret) document was created to deliver proofs for the party‘s extreme right nature. Now you can explore it interactively. ➡️ 🎁 🇩🇪 www.spiegel.de/politik/deut...
26722
Reposted by Luke Zappia
Yvan Saeys @yvansaeys.bsky.social · 01/05/2025
🔥 Just published! Thrilled to share that our work on funkyheatmap has just been published in the Journal of Open Source Software 🎉
174
Reposted by Luke Zappia
khrovatin.bsky.social @khrovatin.bsky.social · 28/04/2025
There is no reason to stay bound to one programming language. I discussed ways to ease R-Python interoperability with Luke Zappia, Philipp Angerer, Tomasz Kalinowski. Their tips and tricks are collected in this blog: hrovatin.github.io/posts/r_pyth... @lazappi.bsky.social @t-kalinowski.bsky.social
hrovatin.github.io
From R to Python with minimal baggage
Getting the best of both worlds.
5193
Luke Zappia @lazappi.bsky.social · 16/04/2025
If they don't agree then it depends which you find more convincing I guess. I think there are lots of reasons two benchmarks might not after without either if them being "wrong" though.
120
Luke Zappia @lazappi.bsky.social · 20/03/2025
Very on brand 😹. I guess I missed a space 🤦🏻. At least the mystery is solved.
000
Luke Zappia @lazappi.bsky.social · 19/03/2025
Yeah, it surprised me a bit as well. This wasn't the main aim of the paper though so it's a fairly limited comparison. There is a lot more you could do if you wanted to properly compare them.
000
Luke Zappia @lazappi.bsky.social · 19/03/2025
Thanks!
010
Reposted by Luke Zappia
Luke Zappia @lazappi.bsky.social · 18/03/2025
Our paper benchmarking feature selection for scRNA-seq integration and reference usage is out now www.nature.com/articles/s41...! Keep reading for more about how we did the study and what we found out 🧵 👇 1/16
nature.com
64017
Luke Zappia @lazappi.bsky.social · 19/03/2025
Thanks!
000
Luke Zappia @lazappi.bsky.social · 19/03/2025
Thanks!
010
Luke Zappia @lazappi.bsky.social · 19/03/2025
Hmmm...It's broken for me too 😿. I'm not sure what happened there but this is working www.nature.com/articles/s41.... Thanks for letting me know!
nature.com
Feature selection methods affect the performance of scRNA-seq data integration and querying - Nature Methods
This Registered Report presents a benchmarking study evaluating the impact of feature selection on scRNA-seq integration.
120
Reposted by Luke Zappia
Philipp Bayer @philippbayer.bsky.social · 17/03/2025
We're hiring for a bioinformatics lead at the OceanOmics Centre! L7, come work with all of the cool genomics data! HiFi, HiC, Illumina, a PromethION, it's all there. Sequence *all* of the marine vertebrates! #bioinformatics external.jobs.uwa.edu.au/cw/en/job/51...
external.jobs.uwa.edu.au
Prospective staff : Jobs at UWA : IN DEVELOPMENT
22319
Luke Zappia @lazappi.bsky.social · 18/03/2025
Big thank you to everyone who contributed to the study 🎉! And thank you for reading 📖! 16/16
160
Luke Zappia @lazappi.bsky.social · 18/03/2025
The project took 2.5 years from initial commit to publication. Some things could have been quicker but that’s not a complaint. I think that’s how long science takes and we should set realistic expectations, especially for junior researchers. It also emphasises the need for continuous benchmarking.
120
Luke Zappia @lazappi.bsky.social · 18/03/2025
You can also find the code on GitHub github.com/theislab/atl... and the data on figshare figshare.com/projects/Ben.... 14/16
github.com
GitHub - theislab/atlas-feature-selection-benchmark: Code for "Feature selection methods affect the performance of scRNA-seq data integration and querying"
Code for "Feature selection methods affect the performance of scRNA-seq data integration and querying" - theislab/atlas-feature-selection-benchmark
110
Luke Zappia @lazappi.bsky.social · 18/03/2025
I have expanded on these points in this blog post lazappi.id.au/posts/2025-0... and of course all the detail is in the paper doi.org/10.1038/s415.... 13/16
lazappi.id.au
Benchmarking feature selection for integration – lazappi
A summary of our recently published paper “Feature selection methods affect the performance of scRNA-seq data integration and querying”
100
Luke Zappia @lazappi.bsky.social · 18/03/2025
We focused on feature selection methods, but we compared scANVI and Harmony/Symphony to our baseline of scVI. Feature selection methods performed similarly but scANVI scored higher overall and Symphony worse, particularly at unseen population detection. More work is needed to understand why. 12/16
Figure showing the comparison between scVI, scANVI and Harmony/Symphony integration methods. a) metric category scores for each feature selection and integration method. b) difference in metric scores for scANVI and Symphony compared to scVI. c) metric category ranks for each feature selection and integration method. d) difference in ranks for scANVI and Symphony compared to scVI.
110
Luke Zappia @lazappi.bsky.social · 18/03/2025
What about lineage-specific integration? Using subsets of the Human Lung Cell Atlas we saw poorer performance overall on lineages compared to the full dataset, particularly for unseen population detection, but a full study is needed to properly answer this. 11/16
Figure showing performance on subsets of the Human Lung Cell Atlas. a) shows scores for metric categories on the full HLCA, the immune lineage and the epithelial lineage. b) is a heatmap of the Jaccard index of the features selected on different subsets. c) shows the proportion of marker genes selected by each method on each subset. d) shows Milo scores for identifying unseen populations on the full HLCA and lineage subsets, as well as the difference between lineages compared to the full dataset.
110
Luke Zappia @lazappi.bsky.social · 18/03/2025
Do you need to select features in a batch-aware way? We didn’t see a clear effect of doing that so it depends on your dataset and computational resources. 10/16
100
Luke Zappia @lazappi.bsky.social · 18/03/2025
Highly variable features performed consistently well, especially the Seurat VST method. Supervised marker genes also score highly but are more variable and require cell labels. Check out triku for an alternative approach that performs similarly. 9/16
Figure showing the comparison of feature selection methods. a) shows the overall scores and ranks for each metric category. b) is a heatmap of Jaccard index between feature sets selected by each method. c) shows the number of common features selected by different numbers of methods on each dataset. d) shows the number of features selected by methods that automatically choose the number of features. e) is a heatmap of the difference in metric category shows for batch-aware variants of methods compared to the standard version.
110
Luke Zappia @lazappi.bsky.social · 18/03/2025
Most methods require setting a number of features. We tried different numbers for some common methods and used 2000 for the rest of the benchmark. Slightly more features improves query metrics while slightly less improves the integration, but this should be tuned to your dataset and use case. 8/16
Figure showing the selection of the number of features to use. a) shows the overall trend in metric category scores as the number of selected features increases. b) and c) shows heatmaps of the trends for each dataset and selection method.
130
Luke Zappia @lazappi.bsky.social · 18/03/2025
Even well-designed metrics have different effective ranges. We used a set of positive and negative baseline methods to scale each metric to a range that was meaningful for this task, providing extra context. Scaled scores were combined to summarise each metric category. 7/16
Figure showing metric baselines and scaling. a) shows the effective range for each metric calculated using the baseline methods. b) shows the scaling procedure including measuring metrics, scaling using baselines, averaging by metric type and calculating an overall score.
100