Sign in

Paul Harrison

@paulfharrison.bsky.social
459 followers 128 following 101 posts

Bioinformatician at Monash University, Melbourne, Australia. I also use mastodon: @pfh@mastondon.online mastodon.online/@pfh My homepage is: logarithmic.net/pfh On Twitter I was: @paulfharrison

PostsRepliesMedia
Paul Harrison @paulfharrison.bsky.social · 20/09/2026
Saw an interesting video of a Go player analyzing an AlphaGo Zero game. He didn't much care for it, as the play style was incomprehensible. He was however more satisfied by a game by a more recent and superior Go engine. So human incomprehensibility seems to have at most marginal benefit to Go play.
010
Paul Harrison @paulfharrison.bsky.social · 24/08/2026
x <- c(rep(0,19),1) cor.test(x,x, method="...something rank-based and exact...") Surely the p value is no smaller than 1/20, but I have tried a variety of packages and they are all hopelessly optimistic. :-(
000
Paul Harrison @paulfharrison.bsky.social · 23/08/2026
A transistor allows the rewiring of an electrical pathway. (The analogy being the insufficiency of linear ODEs to describe some electrical circuits and some metabolic pathways.)
001
Reposted by Paul Harrison
Frank Harrell @f2harrell.bsky.social · 23/06/2026
#Statistics thought of the day: At the heart of much nonsense research lies data analysis (including machine learning) of a complexity that the effective sample size cannot support: hbiostat.org/hdata #StatsSky
hbiostat.org
Challenges of High-Dimensional Data Analysis
1236
Paul Harrison @paulfharrison.bsky.social · 10/06/2026
He's comparing the Michaelis-Menten and Duckwork-Lewis models now.
000
Paul Harrison @paulfharrison.bsky.social · 10/06/2026
I'm greatly appreciating this series of videos about biological modelling with Karthik Raman. I started with the flux-balance models, but now I'm going back to the beginning. www.youtube.com/playlist?lis...
youtube.com
Computational Systems Biology | IIT Madras - YouTube
"This lecture series, part of the ""Computational Systems Biology"" course, focuses on the mathematical modeling of biological systems. It introduces various...
110
Reposted by Paul Harrison
Wolfgang Huber @wkhuber.bsky.social · 29/05/2026
Thanks to input from @helucro.bsky.social, Daria Lazic, Hugo Gruson, Artür Manukyan and many others, there is now a whole new level of maturity and ease-of-use for spatial omics data, see helenalc.github.io/SpatialData....
helenalc.github.io
csama – SpatialData
0226
Paul Harrison @paulfharrison.bsky.social · 24/05/2026
Where is the model of the yeast cell in the yeast cell, and how does it solve the linear program of its planned economy?
010
Paul Harrison @paulfharrison.bsky.social · 04/05/2026
"The cow is homeostatic for blood-temperature ... If, however, a sensitive temperature-recorder be inserted in the brain and then a stream of ice-cold air driven past the animal the temperature rises without any preliminary fall." The cow must have a model. pespmc1.vub.ac.be/books/Conant...
pespmc1.vub.ac.be
100
Paul Harrison @paulfharrison.bsky.social · 09/04/2026
It's slightly maddening that cybernetic negative-feedback control loops *in particular* are not observable from correlation.
010
Paul Harrison @paulfharrison.bsky.social · 05/04/2026
I've recently been exploring an Ornstein-Uhlenbeck SDE model of scRNA-Seq and Perturb-Seq. It's a simplistic model, but is enough to encounter many concrete problems that come up in stochastic models, such as parameter identifiability. Some rough slides here. logarithmic.net/2026/sde/sde...
logarithmic.net
A Simple SDE Model from Yeast Perturb-Seq
121
Paul Harrison @paulfharrison.bsky.social · 19/03/2026
This is a case where you could do a TREAT test (T-test RElative to A Threshold) to test for a biologically meaningful difference. limma, edgeR, and DESeq2 all support this, for example. Or the FDR has an under-used confidence interval form, the FCR.
000
Paul Harrison @paulfharrison.bsky.social · 03/02/2026
A question: How do biologists learn experimental design? Any suggestions for good textbooks or notable people responsible for current practices? (I recently failed to convince some collaborators doing an experiment about data-visualization of the importance of a positive control.)
110
Paul Harrison @paulfharrison.bsky.social · 29/12/2025
For example parquet is a high performance format, but writing was a bottleneck until I wrote to multiple files at once. The arrow library then makes it easy to read across multiple files at once. Zarr is another format that seems to lean-in to this idea, but I don't have much experience with it.
000
Paul Harrison @paulfharrison.bsky.social · 29/12/2025
Minor theme this year: There's a rough right number of files to split datasets into. For my data, on the order of 100s of files. Too few, can't process in parallel. Want to be writing in parallel now! Too many, filesystems go slow or break, especially network/cloud storage. HPC team becomes sad.
100
Reposted by Paul Harrison
Jayani Lakshika @jayanigamage.bsky.social · 07/12/2025
Shared my talk “Visualise your fitted non-linear reduction model in high-dimensional space” at ASC2025, Curtin University, Perth. Great conversations, great people, and so much inspiration. ✨ Slides: jayani-asc2025.netlify.app #ASC2025 #Statistics #DataScience #Visualisation
031
Paul Harrison @paulfharrison.bsky.social · 07/12/2025
This is my simplified version which I use to illustrate different types of error bars: logarithmic.net/2017/dance/
logarithmic.net
Dance of the CIs
000
Paul Harrison @paulfharrison.bsky.social · 07/12/2025
Love this type of thing. I particularly like repeatedly sampling Confidence Intervals as a teaching tool. There's a web app that goes with a book caled "The New Statistics" where they call this the "Dance of the CIs". esci.thenewstatistics.com
esci.thenewstatistics.com
esci-web Main Menu
110
Paul Harrison @paulfharrison.bsky.social · 29/11/2025
I should also note Shiny's own version of this, ExtendedTask, similarly allows an app to remain responsive. If inputs to an ExtendedTask change while it is running, the new computation is delayed until the current one finishes. This isn't ideal for my application. shiny.posit.co/r/articles/i...
shiny.posit.co
Shiny - Non-blocking operations
Shiny is a package that makes it easy to create interactive web apps using R and Python.
000
Paul Harrison @paulfharrison.bsky.social · 29/11/2025
I'm trying using background workers in Shiny. Pattern: Cache results on disk. On a cache miss, launch-and-forget a background worker, tell Shiny to invalidate later, and throw an error for Shiny to display. Workers use file locks to avoid doubling up work. App remains responsive! #R #Shiny
100
Reposted by Paul Harrison
Jovana Maksimovic @jovmaksimovic.bsky.social · 26/11/2025
It's not just because @torstenseemann.bsky.social is a bioinformatics legend but this may be my favourite #abacbs2025 poster because I'm a #trek tragic 🖖
03810
Reposted by Paul Harrison
Jovana Maksimovic @jovmaksimovic.bsky.social · 25/11/2025
Congrats @lonsbio.bsky.social on being awarded the inaugural @abacbs.bsky.social Nick Wong Community Building Award at #abacbs2025! @hdashnow.bsky.social was absolutely right - you really are the Ted Lasso of Australian bioinformatics and computational biology 😃
media.tenor.com
a man pointing to a sign that says believe
ALT: a man pointing to a sign that says believe
2233
Reposted by Paul Harrison
Julia Evans @b0rk.jvns.ca · 20/10/2025
just added the MEGA TERMINAL CHEAT SHEET from "The Secret Rules of the Terminal" to our list of posters at wizardzines.com#posters
321145
Paul Harrison @paulfharrison.bsky.social · 11/10/2025
The splats can also be used as weights for local model fitting.
000
Paul Harrison @paulfharrison.bsky.social · 11/10/2025
Pondering k-Nearest Neighbor density estimation. There's some subtlety making the density smooth and integrate to 1. Here is a simple scheme: - Assign each point a radius from the distance to its kth nearest neighbor. - The density is the sum of a set of Gaussian splats with those radii.
110
Paul Harrison @paulfharrison.bsky.social · 26/08/2025
"... a fundamental question that often remains overlooked is whether or not model parameters can be confidently estimated from the available data."
000
Reposted by Paul Harrison
Ming Tommy Tang @tommytang.bsky.social · 06/08/2025
Heatmap in ggplot2 yunuuuu.github.io/ggalign/ind... I always use complexheatmap, but this seems to be a good alternative if you want to stay within the ggplot
0192
Paul Harrison @paulfharrison.bsky.social · 03/08/2025
*Grimes
000
Paul Harrison @paulfharrison.bsky.social · 03/08/2025
If you've ever wondered why medical research has so many Chesterton's Fences, or what it would actually take to "do your own research", this would be a good starting point. (The target audience of the book is doctors seeking to use published medical research.)
000
Paul Harrison @paulfharrison.bsky.social · 03/08/2025
Happened to pick up "The Lancet Handbook of Essential Concepts in Clinical Research" (Schulz and Graves). A short read and surprisingly excellent. You can really tell that the authors have seen it all. www.elsevierhealth.com.au/essential-co...
elsevierhealth.com.au
Essential Concepts in Clinical Research: 2nd edition | Kenneth Schulz | ISBN: 9780702073946 | Elsevier Australia Bookstore
This practical guide speaks to two audiences: those who read and those who conduct research. Clinicians are medical detectives by training. For each patient, they assemble clinical clues to establish ...
220
Paul Harrison @paulfharrison.bsky.social · 31/07/2025
I did get the AMS. Not quite sure what I'm doing with it. Maybe some interesting possibilities mixing soft and hard materials. I've heard good things about marble PLA.
110
Paul Harrison @paulfharrison.bsky.social · 27/07/2025
uvx demakein I've finally updated my wind instrument design program to Python 3. It only took me 10 years to get around to. I was pleased to find there is now a fairly solid python library for 3D boolean operations (manifold3d). github.com/pfh/demakein
A picture of a 3D printed whistle in front of a 3D printer.
230
Paul Harrison @paulfharrison.bsky.social · 11/07/2025
with apologies to grugbrain.dev
grugbrain.dev
The Grug Brained Developer
000
Paul Harrison @paulfharrison.bsky.social · 11/07/2025
grug brain bioinformatician not trust maximum a posteriori estimate. big brained bioinformatics shaman develop map estimate. danger! noise demon hide deeper in data! grug prefer count matrix. grug know what to do when have count matrix.
110
Reposted by Paul Harrison
Robert Aboukhalil @robert.bio · 10/06/2025
Excited to announce our first interactive article on sandbox.bio, about genomic ranges: sandbox.bio/concepts/gen... Move & resize the ranges to see how that affects bedtools operations like merge and intersect in real time!
15018
Reposted by Paul Harrison
MACSYS @macsys.bsky.social · 23/06/2025
🚨 Exciting PhD Opportunities with MACSYS! The MACSYS team at Monash University is offering multiple fully funded #PhD scholarships for students eager to explore the cutting edge of computational biology, microbiology, and systems modelling. 👉More info/apply: macsys.org/monash-phd-s...
012
Reposted by Paul Harrison
Kim-Anh Lê Cao @mixomics.org · 23/06/2025
👩‍💻 We’re hiring! Lê Cao Lab at Uni Melbourne @mig-unimelb.bsky.social needs an R dev to power the next-gen of mixOmics 🚀 Love #RStats, #Bioconductor & multi-omics? Help expand mixOmics, run workshops & publish cutting-edge methods. Apply: unimelb.wd105.myworkdayjobs.com/en-US/UoM_Ex...
01112
Reposted by Paul Harrison
Davis McCarthy @davisjmcc.bsky.social · 16/06/2025
📢 PostDoc opportunity in our Bioinformatics & Cellular Genomics lab at SVI! 🧬 You’d join a welcoming, supportive, and brilliant team. Why not spend a few years in Melbourne and be part of something exciting? Apply here: www.seek.com.au/job/84737876 #ScienceCareers #PostDoc #Bioinformatics
seek.com.au
Research Officer - Bioinformatics Job in Fitzroy, Melbourne VIC - SEEK
Seeking a Postdoc to develop computational toolkits to enable large-scale studies of single-cell and spatial 'omics and statistical genetics
23425
Paul Harrison @paulfharrison.bsky.social · 17/06/2025
Parquet format and the arrow library have been life changing. (I suspect I should be getting on board with duckdb one of these days too.)
030
Paul Harrison @paulfharrison.bsky.social · 18/05/2025
(Well, not exactly fine. A change to any single species abundance alters all of these ratios, so the null hypothesis could then be quite correctly rejected for all species!)
000
Paul Harrison @paulfharrison.bsky.social · 18/05/2025
To be clear, I'm arguing semantics. I believe the software does *something* useful, it's just not being described clearly. For example, it might look at the ratio of each species to the geometric mean as a baseline. This is fine, but that baseline's appropriateness needs to be checked.
110
Paul Harrison @paulfharrison.bsky.social · 18/05/2025
I continue to be astounded at the number of compositional data analysis packages that will happily report differential abundance of individual species. Did they not understand the concept of compositional data? How is it possible to publish methods with this premise?
130
Paul Harrison @paulfharrison.bsky.social · 03/05/2025
Putting it through its paces. However I have more testing to do to really get to know this transformation. logarithmic.net/varistran/ar...
logarithmic.net
Samesum transformation
000
Paul Harrison @paulfharrison.bsky.social · 03/05/2025
I've put an implementation in my old varistran package. There is a numerical optimization per sample, but I can apply Newton's method so it's fast. logarithmic.net/varistran/re...
logarithmic.net
Normalized log2 counts using samesum method — samesum_log2_norm
This is a method of computing log counts while dealing sensibly with zeros and differing library sizes. It should cope well with very sparse data. The input is transformed using log2(x/scale+1) with a...
100
Paul Harrison @paulfharrison.bsky.social · 03/05/2025
Transform counts like log2(count/scale+1), with a scale chosen per sample such that each sample adds to the same total. It's similar to CLR with a pseudocount, but all zeros transform to the same value.
100
Paul Harrison @paulfharrison.bsky.social · 03/05/2025
Normalization and log transformation of log count data. Pseudocounts, library size adjustment, Centered Log Ratios (CLR), Variance Stabilizing Transformation, and all that. Many variations on a similar task. Here's something I haven't seen done:
110
Paul Harrison @paulfharrison.bsky.social · 22/04/2025
Random 2d Turing Machines make interesting patterns. pointersgonewild.com/2012/12/31/t...
pointersgonewild.com
Turing Drawings
Turing Drawings
020
Paul Harrison @paulfharrison.bsky.social · 18/04/2025
Some conventional flow matching:
A computer generated image that looks like daubs of colored oil-paint.
020
Paul Harrison @paulfharrison.bsky.social · 16/04/2025
Yes, the link above goes on to talk about that in the next section. I do hesitate a bit at having a single duplicate correlation value across all genes. There's a package called variancePartition without this limitation, but I haven't tried it.
010
Paul Harrison @paulfharrison.bsky.social · 16/04/2025
The first time I saw this it took me a couple of weeks to get my head around it. ~timepoint*treatment+id has multicollinearity because id nests within treatment. We could use a mixed model ~timepoint*treatment+(1|id), but many popular tools only support fixed effects models.
110