Sign in

Tom Houslay

@tomhouslay.bsky.social
141 followers 336 following 9 posts

Former academic researcher in evolution & animal behaviour. Now: data & biology stuff for industry [human milk, infant nutrition, microbiome, healthy ageing...] Also: dad, wildlife pond enthusiast, ND

PostsRepliesMedia
Reposted by Tom Houslay
Colin Carlson @colincarlson.bsky.social · 31/01/2026
This is Nathan Wolfe. Nathan is a virus hunter. Nathan has a bit of a reputation. Through the years, I've told multiple journalists stories I've heard, like "Nathan thanks Jeffrey Epstein in his book." Somehow no one ever wrote anything up. Want to see what Nathan was up to in today's Epstein files?
nathan wolfe doing the theranos pinch
6779263
Reposted by Tom Houslay
Walty @waltometry.blacksky.app · 07/02/2026
I’m not arguing with no account defending AI. My grandmother lives in an area being environmentally compromised currently by an ai data center. I remember the color of the lake before they built it and the sludge I see in it when I go to see her. Shut the fuck up about “the benefits of AI” forever.
4571412694
Reposted by Tom Houslay
James Balamuta @coatless.bsky.social · 29/12/2025
{toggle} does one thing: adds a button to hide code output in #quarto docs. Took two versions to do that one thing well. Now it works everywhere... tabsets, callouts, nested containers, you name it. 📚 quarto.thecoatlessprofessor.com/toggle/ 🐙 github.com/coatless-qua...
Screenshot of the Toggle extension documentation website showing the Quick Start page. The left sidebar displays the Toggle logo and navigation menu with sections for Get Started, Features, Demos, and Reference. The main content shows an interactive toggle button demo, explains the ▼ (visible) and ▶ (hidden) chevron indicators, and includes R code examples with a visible "Output" toggle button on the code block. The page demonstrates how the toggle button appears when hovering over code cells.Toggle extension logo: A blue hexagonal badge featuring a stylized code block with a toggle button in the upper right corner, and a line chart output below it. The word "toggle" appears at the bottom in white lowercase text.
04210
Reposted by Tom Houslay
Andrew Heiss @andrew.heiss.phd · 09/12/2025
Some closing thoughts for my students this semester on LLMs and learning #rstats datavizf25.classes.andrewheiss.com/news/2025-12...
Will you incorporate LLMs and AI prompting into the course in the future?
No.

Why won’t you incorporate LLMs and AI prompting into the course?
These tools are useful for coding (see this for my personal take on this).

However, they’re only useful if you know what you’re doing first. If you skip the learning-the-process-of-writing-code step and just copy/paste output from ChatGPT, you will not learn. You cannot learn. You cannot improve. You will not understand the code.In that post, it warns that you cannot use it as a beginner:

…to use Databot effectively and safely, you still need the skills of a data scientist: background and domain knowledge, data analysis expertise, and coding ability.

There is no LLM-based shortcut to those skills. You cannot LLM your way into domain knowledge, data analysis expertise, or coding ability.

The only way to gain domain knowledge, data analysis expertise, and coding ability is to struggle. To get errors. To google those errors. To look over the documentation. To copy/paste your own code and adapt it for different purposes. To explore messy datasets. To struggle to clean those datasets. To spend an hour looking for a missing comma.

This isn’t a form of programming hazing, like “I had to walk to school uphill both ways in the snow and now you must too.” It’s the actual process of learning and growing and developing and improving. You’ve gotta struggle.This Tumblr post puts it well (it’s about art specifically, but it applies to coding and data analysis too):

Contrary to popular belief the biggest beginner’s roadblock to art isn’t even technical skill it’s frustration tolerance, especially in the age of social media. It hurts and the frustration is endless but you must build the frustration tolerance equivalent to a roach’s capacity to survive a nuclear explosion. That’s how you build on the technical skill. Throw that “won’t even start because I’m afraid it won’t be perfect” shit out the window. Just do it. Just start. Good luck. (The original post has disappeared, but here’s a reblog.)

It’s hard, but struggling is the only way to learn anything.You might not enjoy code as much as Williams does (or I do), but there’s still value in maintaining codings skills as you improve and learn more. You don’t want your skills to atrophy.

As I discuss here, when I do use LLMs for coding-related tasks, I purposely throw as much friction into the process as possible:

To avoid falling into over-reliance on LLM-assisted code help, I add as much friction into my workflow as possible. I only use GitHub Copilot and Claude in the browser, not through the chat sidebar in Positron or Visual Studio Code. I treat the code it generates like random answers from StackOverflow or blog posts and generally rewrite it completely. I disable the inline LLM-based auto complete in text editors. For routine tasks like generating {roxygen2} documentation scaffolding for functions, I use the {chores} package, which requires a bunch of pointing and clicking to use.

Even though I use Positron, I purposely do not use either Positron Assistant or Databot. I have them disabled.

So in the end, for pedagogical reasons, I don’t foresee me incorporating LLMs into this class. I’m pedagogically opposed to it. I’m facing all sorts of external pressure to do it, but I’m resisting.

You’ve got to learn first.
14336103
Reposted by Tom Houslay
Jumping Rivers @jumpingrivers.com · 09/12/2025
If your Shiny app works with a mouse and looks fine on your screen, it may still be unusable for some users. Issues like missing alt text, ARIA misuse, or loose WCAG checks often surface too late. This Thursday’s free webinar shows how to catch them earlier. 🕐 13:00 UK time Register:
052
Reposted by Tom Houslay
Peter Tennant @pwgtennant.bsky.social · 08/12/2025
Thrilled to share my latest paper entitled, "Estimating Discrimination in Sentencing: Distinguishing between Good and Bad Controls" Led by @jpinasanchez.bsky.social, the paper introduces a framework for examining discrimination in criminal justice processes. 🧵 1/10 publicera.kb.se/ejels/articl...
Screenshot of title page including the following abstract:

To minimise confounding bias and disentangle warranted from unwarranted disparities, researchers examining sentencing discrimination have traditionally sought to control for as many legal factors as possible. However, over the past decade, a growing number of scholars have questioned this strategy, noting that many legal factors are themselves subject to judicial discretion and that controlling for them can introduce post-treatment bias. Here, we use directed acyclic graphs (DAGs) to provide a formal and comprehensive assessment of the different types of bias that may arise from different choices of controls. In addition, we propose a new modelling framework to facilitate the selection of controls and reflect the model uncertainty created by the trade-off inherent in judicially-defined legal factors and other factors with a similar dual causal role. We apply this framework to examine race disparities in US federal courts and gender disparities in the England and Wales magistrates’ court. We find substantial model uncertainty for gender disparities and for race disparities affecting Hispanic offenders, rendering estimates of the latter inconclusive. Disparities against black offenders are more consistent and — under specific conditions — could be interpreted as evidence of direct discrimination.
27634
Tom Houslay @tomhouslay.bsky.social · 09/12/2025
www.theguardian.com/environment/...
theguardian.com
It’s the world’s rarest ape. Now a billion-dollar dig for gold threatens its future
Tapanuli orangutans survive only in Indonesia’s Sumatran rainforest where a mine expansion will cut through their home. Yet the mining company says the alternative will be worse
000
Tom Houslay @tomhouslay.bsky.social · 04/12/2025
Here are some mammals that like to pop out from under my office at night.
010
Reposted by Tom Houslay
Isabella Velásquez @ivelasq3.bsky.social · 03/11/2025
I wrote a lil post on the amazing work that @ginareynolds.bsky.social does championing ggplot2 extension developers and teaching others to build their own! The post features the Scrollytelling Quarto extension and the group's cute #RStats hex 🐱: rworks.dev/posts/ggplot...
rworks.dev
An Introduction to Writing Your Own ggplot2 Geoms – R Works
The ggextenders club provides inspiration and resources for those venturing into the exciting world of creating custom ggplot2 extensions.
16816
Reposted by Tom Houslay
Andrew Heiss @andrew.heiss.phd · 02/11/2025
Wrote up a little intervention post/explanation for my class about why using LLMs for trying to learn programming (as first time learners!) is bad and detrimental datavizf25.classes.andrewheiss.com/news/2025-11...
In the first week of class, I had you read an article about how ChatGPT is philosophical “bullshit”, as well as this other post about AI and LLMs in this class.

I’ve seen a surprising increase in the amount of ChatGPT-based weekly check-ins and graph interpretations in the exercises. (In the past, I’ve even had students tell me they PDFs of the slides into ChatGPT to generate interesting and confusing things.) Don’t do this! LLMs cannot find three things that you personally found interesting from the readings. LLMs are not you.

It’s a huge waste of time for all of us to for me to spend time evaluating and grading output from ChatGPT. I want to read what you think and do, not what a computer invents.

NoteMy personal stance on AI and LLMs
Many of you have asked about how I think about and use LLMs in my own work (and in the class materials). Please check this short little page where I explain what I do with text, images, code, research, and learning.

In short, I do not use LLMs to create any of the materials for this class, and I do not use them to grade or evaluate any of your work.I’ve even had several students submit completely made up stuff for mini projects and final projects. In your final project, your job is to take a dataset, make three different plots with it, and combine the plots into one graphic that tells some sort of story. I’ve (really truly) received projects like this:

(a ChatGPT-generated plot of completely made up data from gapminder)

At a very quick glance, that looks like a plausible image. But holy crap it falls apart fast, both comically and tragically. Comically because, like, look at the y-axis in the scatterplot; or the x-axis ticks in the population grown chart that are sometimes spaced every 20 years, sometimes every 10 years, and has 1980 there twice; or the weird blobs in the map.

Tragically—and most importantly—literally nothing in this image is real and none of the annotations make sense and none of the colors make sense and everything here is completely meaningless and devoid of any information. It’s pure philosophical bullshit—an image “produced without concern for its truth.”That’s what I mean when I say that it’s a forgetful and unreliable resource. Lots of that code is helpful as an illustration of what’s going on. Lots of the explanation is good and fine. But it subtly changed its mind halfway through and didn’t realize it (because it can’t realize it! it’s just a chain of statistically likely text!)

Hopefully you can see where this is dangerous for learning. It can often “teach” you wrong, and you can only recognize that if you know stuff already.

It’s extra dangerous if you then take that example output, paste it into your exercise, and submit it.

I’m not even concerned about the plagiarism aspect of that—code is code and there are only a few ways to write {dplyr} code that will do this, and it’s easy to adapt from real human-based examples online, so I’m less worried about stealing other people’s work here.

I’m concerned that it’s not real data.

It made up values of 1, 2, 3, and 4 for some countries to illustrate grouping and summarizing. That’s good and normal.

None of that is real. Don’t treat it as real.
417952
Reposted by Tom Houslay
doctor eek (alt decorative ghost season) @absolutely-not.bsky.social · 03/10/2025
i looked at the methodology for this and it is a. sex addiction counseling group in texas did a surveymonkey and extrapolated the results to the entire us population which is the sort of research design that earns you an ff on an intro methods class (the extra f is for extra effort), and b. p-hacked
12893822457
Reposted by Tom Houslay
Crystal Lewis @cghlewis.bsky.social · 01/10/2025
If colleagues or students share a dataset with you and you make this face, consider sharing these resources with them. ☺️ datamgmtinedresearch.com 1. Organizing data (Ch. 3) 2. Naming variables and files (Ch. 9) 3. Documenting data (Ch. 8) 4. Cleaning data (Ch. 14)
media.tenor.com
a man is covering his mouth with his hands and the word schitts creek is on the bottom right
ALT: a man is covering his mouth with his hands and the word schitts creek is on the bottom right
35718
Reposted by Tom Houslay
Vincent Arel-Bundock @vincentab.bsky.social · 29/09/2025
{tinytable} 0.14.0 for #RStats makes it super easy to draw tables in html, tex, docx, typ, md & png. There are only a few functions to learn, but don't be fooled! Small 📦s can still be powerful. Check out the new gallery page for fun case studies. vincentarelbundock.github.io/tinytable/vi...
a table about lemurs a table about students and schoolsa table about wines
113438
Tom Houslay @tomhouslay.bsky.social · 01/10/2025
Useful resource! Good to see ordered beta regression in the mix as I've found that useful recently, although the note on it is linked to a Ben Bolker twitter comment that is now gone as his profile is deleted...
210
Reposted by Tom Houslay
Jan Broder Engler @jbengler.de · 01/10/2025
Thank you for citing #tidyplots 🙏 Jakub Idkowiak et al. Best practices and tools in R and Python for statistical processing and visualization of lipidomics and metabolomics data. Nature Communications (2025). doi.org/10.1038/s414... #rstats #dataviz #phd
doi.org
Best practices and tools in R and Python for statistical processing and visualization of lipidomics and metabolomics data - Nature Communications
Mass spectrometry-based lipidomics and metabolomics generate large, complex datasets requiring effective analysis. Here, authors review key statistical and visualization methods alongside widely used R and Python tools, and provide a GitBook with step-by-step code for accessible, reproducible data analysis.
0105
Reposted by Tom Houslay
Julia M. Rohrer @dingdingpeng.the100.ci · 25/08/2025
Ever stared at a table of regression coefficients & wondered what you're doing with your life? Very excited to share this gentle introduction to another way of making sense of statistical models (w @vincentab.bsky.social) Preprint: doi.org/10.31234/osf... Website: j-rohrer.github.io/marginal-psy...
Models as Prediction Machines: How to Convert Confusing Coefficients into Clear Quantities

Abstract
Psychological researchers usually make sense of regression models by interpreting coefficient estimates directly. This works well enough for simple linear models, but is more challenging for more complex models with, for example, categorical variables, interactions, non-linearities, and hierarchical structures. Here, we introduce an alternative approach to making sense of statistical models. The central idea is to abstract away from the mechanics of estimation, and to treat models as “counterfactual prediction machines,” which are subsequently queried to estimate quantities and conduct tests that matter substantively. This workflow is model-agnostic; it can be applied in a consistent fashion to draw causal or descriptive inference from a wide range of models. We illustrate how to implement this workflow with the marginaleffects package, which supports over 100 different classes of models in R and Python, and present two worked examples. These examples show how the workflow can be applied across designs (e.g., observational study, randomized experiment) to answer different research questions (e.g., associations, causal effects, effect heterogeneity) while facing various challenges (e.g., controlling for confounders in a flexible manner, modelling ordinal outcomes, and interpreting non-linear models).
Figure illustrating model predictions. On the X-axis the predictor, annual gross income in Euro. On the Y-axis the outcome, predicted life satisfaction. A solid line marks the curve of predictions on which individual data points are marked as model-implied outcomes at incomes of interest. Comparing two such predictions gives us a comparison. We can also fit a tangent to the line of predictions, which illustrates the slope at any given point of the curve.A figure illustrating various ways to include age as a predictor in a model. On the x-axis age (predictor), on the y-axis the outcome (model-implied importance of friends, including confidence intervals).

Illustrated are 
1. age as a categorical predictor, resultings in the predictions bouncing around a lot with wide confidence intervals
2. age as a linear predictor, which forces a straight line through the data points that has a very tight confidence band and
3. age splines, which lies somewhere in between as it smoothly follows the data but has more uncertainty than the straight line.
461001285
Reposted by Tom Houslay
Jordan S. Martin @jsmartin.bsky.social · 11/08/2025
My first solo author paper is now available in early view at Methods in Ecology and Evolution! I develop a covariance reaction norm (CRN) model for estimating continuous, multivariate, and nonlinear environmental effects on G and P matrices. besjournals.onlinelibrary.wiley.com/doi/10.1111/...
besjournals.onlinelibrary.wiley.com
Covariance reaction norms: A flexible method for estimating complex environmental effects on trait (co)variances
Estimating quantitative genetic and phenotypic (co)variances is crucial for investigating evolutionary ecological phenomena such as developmental integration, life history trade-offs and niche spe...
13315
Tom Houslay @tomhouslay.bsky.social · 15/07/2025
I got some good answers from my question about how to get back up to date with the world of #rstats
040
Tom Houslay @tomhouslay.bsky.social · 14/07/2025
Stupid question alert but here we go anyway: I long-ago deleted my twitter account, I cannot cope with linkedin, but I need to get back into knowing what's going on in #rstats world. Anyone recommend specific feeds / people / other stuff to drag myself back into it? (or do I just trawl rstats)
451