Sign in

Paulius Alaburda

@alaburda.bsky.social
121 followers 208 following 91 posts

Love all things R, data, medicine and energy! Head of Data Analytics @ Ignitis Lithuania 🇱🇹

PostsRepliesMedia
Reposted by Paulius Alaburda
Gavin Simpson @gsimpson.bsky.social · 06/10/2026
📝 My paper on generalized additive models (GAMs) in animal science is out! Learn nonlinear relationships from data, estimate growth rates and compare treatments, with cows, pigs and quail + reproducible #RStats code. Open access: doi.org/10.1016/j.an... #mgcv #AnimalScience #Statistics 🧪
Eight panels show body mass over 78 days for 57 Japanese quail. Columns represent control, T4, T3 and combined T3–T4 hormone treatments; rows separate females and males. Coloured points are measurements, lines are individual growth curves fitted by a hierarchical generalized additive model, and shading shows 95% credible intervals. Growth slows towards a plateau, with females generally reaching higher body masses than males.Eighteen panels show individual pig growth curves. Black points mark daily average weight estimates from a depth camera; blue lines show curves fitted by a hierarchical generalized additive model, with shaded 95% credible intervals. Weights generally increase over time, with differences in growth and measurement variability among pigs. Uncertainty widens where measurements are sparse or absent.Four-panel comparison of models fitted to a cow’s average daily milk fat production. Observations rise to a peak around weeks 8–10, then decline. The first panel shows fitted curves and uncertainty bands for Wood’s model, a Tweedie GLM and a Tweedie GAM. Residual plots show curved patterns for Wood’s model and the GLM, but little remaining pattern for the GAM.
110228
Reposted by Paulius Alaburda
Data Rocks @t.datarocks.co.nz · 07/10/2026
Fun little project with lots of fun little visuals. crawlerzoo.com Amazing how any rules, copyright, opt-outs, nothing at all, stops the invasion of AI crawler bots.
crawlerzoo.com
The Crawler Zoo — a live menagerie of web crawlers
A live zoo of the bots visiting this site: 69,621 so far today. See what GPTBot, ClaudeBot and Googlebot are, and how to block them.
052
Reposted by Paulius Alaburda
JsonGeller @jgeller1phd.bsky.social · 02/10/2026
New tutorial finally in press with @rmkubinec.bsky.social @chelseaparlett.bsky.social and @matti.vuorre.com. It has everything you could possible want: Quippy headings, a whole lot of beta, bayesian analyses and frequentist analyses side by side in harmony. link.springer.com/article/10.3...
rdcu.be
27825
Paulius Alaburda @alaburda.bsky.social · 02/10/2026
I love how Microsoft is pushing the boundaries of citizen development/shadow IT with Fabric Apps. Feels...chaotic good on the alignment chart?
000
Reposted by Paulius Alaburda
Carson Sievert (he/him) @cpsievert.bsky.social · 01/10/2026
New in querychat: multiple tables, chat history, data-dict, and easy extraction of findings into Quarto, Shiny, Jupyter, etc. querychat is now much more capable beyond just querying "pre-defined" Shiny outputs. Available for both #rstats and #python. opensource.posit.co/blog/2026-09...
A query spanning multiple tables. The wizard for extracting findings into Quarto, Shiny, Jupiter, etc. A view of the new chat history feature.
1329
Reposted by Paulius Alaburda
James O’Loughlin @jamesoloughlin.bsky.social · 28/09/2026
02518341
Paulius Alaburda @alaburda.bsky.social · 28/09/2026
TIL about purrr::insistently 😻
001
Reposted by Paulius Alaburda
Mike @mikenikles.com · 27/09/2026
This looks fun for onboarding tours or in-app hints when a user presses the ⌘ key. neat-annotations.syabro.com
neat-annotations.syabro.com
neat-annotations — neat hand-drawn CSS annotations
Arrows and handwritten labels for your website. Pure CSS, no JavaScript.
0337
Paulius Alaburda @alaburda.bsky.social · 25/09/2026
Live action Path of Pain, Steel Soul mode
000
Reposted by Paulius Alaburda
RJ Andrews @infowetrust.com · 24/09/2026
We wanted a staircase that could disappear. It needed four walls, rigid steps, and room for extravagant maps. Then it had to collapse flat and ship in an envelope. My latest for Chartography: www.chartography.net/p/a-world-of...
chartography.net
A World of Maps, Folded Flat
Turning the David Rumsey Map Center’s spectacular stairwell into a pop-up book.
187
Reposted by Paulius Alaburda
Kristoffer Magnusson @rpsychologist.com · 25/06/2026
New interactive blog! "Why Adjusted Regression Coefficients Are Less Descriptive Than They Look" rpsychologist.com/descriptive-...
1520769
Reposted by Paulius Alaburda
Ainsley S @americanbeetles.bsky.social · 23/09/2026
pls note the bats have been carefully depicted HANGING FROM THEIR LITTLE BRANCHES on this radial phylogeny (bats drawn by Fiona Reid, whose portfolio is here --> fionareid.ca)
section of the central bat phylogeny figure from this new bat phylogeny paper; the representative bat from each family is hanging from the taxon label by its little toesies
2297108
Reposted by Paulius Alaburda
Henrik Bengtsson @henrikbengtsson.bsky.social · 23/09/2026
The rcheology website (hughjonesd.shinyapps.io/rcheology/), for looking up the history of #RStats functions and see how they've evolved over time, got some nice UX improvements recently. It's an invaluable tool if you want to know if your code is backward compatible w/ older R versions It's a gem!
Screenshot of the rcheology web app ([https://hughjonesd.shinyapps.io/rcheology/](https://hughjonesd.shinyapps.io/rcheology/)) where we have looked up information for the base::anyNA(x, recursive = FALSE), which was introduced in R 3.6.3 and compares it to what is in R 4.7.0 · r-devel.
1213
Reposted by Paulius Alaburda
David Aerne @meodai.bsky.social · 17/09/2026
CuspHanger 0.4.0 gives you control over the lightness steps, in the form of a bezier curve or any other function. meodai.github.io/cusphanger/
25612
Paulius Alaburda @alaburda.bsky.social · 23/09/2026
I'll admit, I mostly use AI for coding now but what I've noticed is that I veer towards SQL. Obviously it's because I'm good at it but I have a feeling it's more about readability and convenience. It's hard to explain but I feel so comfortable with selecting as an interface to data.
000
Reposted by Paulius Alaburda
Nicola Rennie @nrennie.bsky.social · 09/10/2025
Wondering if you can outsource your data viz work to ChatGPT? 📊 I tested out a few different generative AI tools, giving them prompts to visualise two different data sets. If you're interested in the results, you can read them here: nrennie.rbind.io/blog/gen-ai-... #RStats #Python #DataViz #GenAI
nrennie.rbind.io
Generative AI for Data Visualisation – Nicola Rennie
Can generative AI create good data visualisations? This blog post compares the performance of ChatGPT, Claude, Copilot, and Gemini when presented with a generic request to visualise a dataset.
66521
Reposted by Paulius Alaburda
Dan Quintana @dsquintana.bsky.social · 20/09/2026
Check how reusable your OSF materials are with my new app, which is part of my latest preprint. Plug in up to ten OSF repo links, and you’ll get a brief report + tips for fixing issues ⬇️
1188
Reposted by Paulius Alaburda
Dan Quintana @dsquintana.bsky.social · 19/09/2026
New preprint! 🎉 I analysed 1660 papers from 4 psychology journals and found materials sharing went from 9% of papers in 2015 to 82% in 2025, and these materials *do* get downloaded — a median of 135 times each. BUT shared code is often hard to run. doi.org/10.31234/osf... Let's walk through it 🧵
Two-panel figure. Panel a is a flow diagram tracking 1,611 empirical psychology articles from publication year (545 in 2015, 552 in 2020, 514 in 2025) to repository-link type: 785 link an OSF project, 85 link another platform, and 741 link no repository. Of those with an OSF link, download counts were retrieved for 670 and not retrieved for 115. Panel b is a line chart of the share of empirical papers linking OSF across 2015, 2020 and 2025. The overall rate, shown as a dashed black line, rises from 9% to 57% to 82%. All four journals rise steeply and end close together: Psychological Science 92%, JESP 89%, JML 83%, Cognition 76%, with Psychological Science highest throughout.Three-panel figure. Panel a: ridgeline plot of downloads per file by material type on a log scale, with the percentage never downloaded labelled for each — archive 16% of 545 files, documents 20% of 2,970, code 12% of 3,471, other 17% of 1,646, data 21% of 10,051, media 35% of 2,120, images 35% of 4,102. Most files cluster between 1 and 10 downloads, with long right tails past 100. Panel b: ridgeline plot of downloads per paper by journal, log scale, with dashed median lines; Psychological Science is highest, then JESP, JML and Cognition. Panel c: stacked bars showing, for documents, data and code separately, the share of papers by download band (0, 1–10, 11–100, more than 100) in 2015, 2020 and 2025. The share exceeding 100 downloads falls sharply over time in all three types, from roughly two-thirds in 2015 to a quarter or less in 2025, as the 1–10 band grows.Four-panel figure. Panel a: statistical languages detected among 333 papers with retrievable code — R 88%, SPSS 12%, Stata 4%, SAS 1%. Panel b: code red flags among those 333 papers — 34% hard-code an absolute path, 40% reference a missing file — above documentation among 672 OSF-linked papers — 21% have a README, 30% are documented by README, description or wiki. Panel c: among 562 Elsevier papers with no repository link, 44% (245) host at least one journal supplementary file but only 14% (77) host data, code or an archive. Panel d: composition of those 448 hosted files — documents 52%, data 21%, media 9%, other 7%, archive 6%, code 3%, images 1%. The code panels are green, the journal-supplement panels blue.Coefficient plot (download predictors)

Dot-and-whisker plot of three standardised predictors of OSF download volume, each with a 95% confidence interval. Repository size (number of files) has the largest effect at about 0.76, citations about 0.38, and altmetric attention about 0.12. All three intervals sit entirely above zero, so each predicts more downloads, with repository size roughly twice the effect of citations and six times that of attention. X-axis: standardised effect, −0.2 to 1.0.
4265113
Reposted by Paulius Alaburda
Julia M. Rohrer @dingdingpeng.the100.ci · 18/09/2026
Cluster analysis can be applied to these, and clusters will be found. But that doesn’t mean those clusters are meaningful.>
man to computer: Find clusters
computer: >CLUSTERS FOUND
man: oh my god.
4635
Reposted by Paulius Alaburda
Nicola Rennie @nrennie.bsky.social · 18/09/2026
My list of data-related resources had a bit of a re-fresh recently! They're now all in the one site. You can still browse by topic or see them all at once! 📊 Link: nrennie.rbind.io/resources/ #RStats #DataViz #StatsEd
Screenshot of resources page.
25615
Reposted by Paulius Alaburda
Lukas Röseler @aufdroeseler.bsky.social · 18/09/2026
I just published my German Open Science book with @uni-muenster.de's library. It contains most of what I know about open science, can be understood by undergraduates, and is free. - published PDF: www.uni-muenster.de/Ebooks/index... - Quarto version: lukasroeseler.github.io/openscienceb...
12613
Reposted by Paulius Alaburda
Jan Vanhove @janhove.bsky.social · 18/09/2026
New blog post: "Cluster analysis: A skeptic’s guide". In which I ask a few questions that social scientists wishing to identify hidden classes in their data ought to address. janhove.github.io/posts/2026-0...
A two-dimensional scatterplot showing a three-cluster solution.

Caption: "Figure 2: A visualisation of a Latent Profile Analysis fit. Such visualisations may help readers appreciate that the clusters aren’t nicely separated and that the researchers’ notion of clusters may not correspond to their own."Table of contents:

Refresher: What is cluster analysis?
What are the clusters for?
What clusters, exactly?
Does the pipeline work?
What does the solution look like?
When running follow-up analyses, how is the uncertainty in the cluster assignments taken into account?
Was it worth it?
Conclusion
References
416347
Reposted by Paulius Alaburda
Peter Ellis @freerangestats.info · 16/09/2026
Inspired by a @f2harrell.bsky.social comment, I did some #rstats simulations about how big a sample needs to be for the central limit theorem to kick in and the sample mean be normally distributed for inference purposes. Yes, with skewed data the answer is "lots". freerangestats.info/blog/2026/09...
Scatter plot with 8 facets, each showing a different distribution, and the actual coverage of a 95% confidence interval for the mean at different sample sizes. Most values are well below 95%.
33315
Reposted by Paulius Alaburda
Matthew Garrett @mjg59.eicar-test-file.zip · 13/09/2026
Libraries in the software sense are called libraries because back in the day there was a cupboard full of paper tape implementations of functions that had already been written and if you wanted to use that function you took it out of the physical library and copied it into your code then returned it
321643
Reposted by Paulius Alaburda
Richard McElreath 🐈‍⬛ @rmcelreath.bsky.social · 10/09/2026
Modeling non-response is an example I like to use when arguing that "descriptive" statistics also require causal inference. To describe the population, you have to model what causes the sample. Similar to occupancy modeling in ecology in many ways.
110219
Reposted by Paulius Alaburda
Andrew Heiss @andrew.heiss.phd · 05/09/2026
You can use this UN-approved Equal Earth projection right now in #rstats #gis #databs
library(tidyverse)
library(sf)
library(rnaturalearth)

world <- ne_countries() |> filter(sov_a3 != "ATA")

ggplot(data = world) +
  geom_sf() +
  coord_sf(crs = "+proj=eqearth") +
  theme_void()A basic world map using the Equal Earth projection
727081
Reposted by Paulius Alaburda
Vincent Arel-Bundock @vincentab.bsky.social · 04/09/2026
Big news! 🎉 𝚖𝚊𝚛𝚐𝚒𝚗𝚊𝚕𝚎𝚏𝚏𝚎𝚌𝚝𝚜 1.0.0 for #Rstats is out. It’s a big number and it feels like a big step. I wrote a blog on the challenges of interpreting statistical models, SPEED, cool new features, the future, and a 5 year package development and writing odyssey. arelbundock.com/posts/margin...
1642299
Reposted by Paulius Alaburda
Teun van den Brand @teunbrand.bsky.social · 03/09/2026
Folks, hot off the press here are a few plots about what the new ggarrow update can do: teunbrand.github.io/teunbrand_bl... I'll also tease them below #rstats #ggplot2
Plot where four arrows connect three points. Two points are connected via double arrows pointing in opposite directions.Plot where four arrows connect three points. Arrowheads are one sided/halved. One connection has two arrows in opposite directions, giving the impression of a compound arrow with half-arrowheads at opposite ends.Plot with four arrows connecting three points. The shafts of the arrows are squiggly lines.
15211
Reposted by Paulius Alaburda
Emily Riederer @emilyriederer.bsky.social · 03/09/2026
For no real reason, I was looking up something about the ancestry of causal DAGs in path diagrams and came across this picture. I got to see it so you do to. HT to Sewall Wright who both set the course for a very useful analytical tool but also took the time to include cute drawings of guinea pigs
Sketch of Wright's path diagram showing the influence of heredity and environment on the inheritance of color of guinea pigs. Drawing is mostly a directed acyclic graph but includes four cute sideways drawings of guinea pigs.
818043
Reposted by Paulius Alaburda
Mike Bostock @ocks.org · 01/09/2026
It’s finally here! So excited to be working out in the open again and for you all to try the new Observable. bsky.app/profile/obse...
16415
Reposted by Paulius Alaburda
Richard McElreath 🐈‍⬛ @rmcelreath.bsky.social · 02/09/2026
Last year, I publicly complained multiple times (see e.g. elevanth.org/blog/2025/07...) that the open access movement had been captured by publishers, and we are now worse off than before open access. Well people got mad at me for saying it. But I think things are still getting worse.
616466
Reposted by Paulius Alaburda
tj mahr 🤘 @tjmahr.com · 01/09/2026
new #rstats note about implementing the balanced cluster bootstrap (with rsample) www.tjmahr.com/notes/2026-0...
tjmahr.com
The balanced cluster bootstrap and implementing it with rsample
We want to resample repeated measures data at the cluster level, not at the individual observation level. In the past, I have just resampled the cluster IDs with replacement and then joined the origin...
0124
Reposted by Paulius Alaburda
Stephen Turner @stephenturner.us · 03/05/2026
Free and open-source images, icons, and tools for creating scientific illustrations doi.org/10.59350/5zt... 🧪
doi.org
Free and open-source images, icons, and tools for creating scientific illustrations
Phylopic, NIH Bioart, Bioicons, Scidraw, Open Science Art, Health Icons, Servier Medical Art, Biodiversity Heritage Library, the Noun Project, Segment Anything, Excalidraw, draw.io, Biographics
6406227
Reposted by Paulius Alaburda
cafkafk @cafkafk.bsky.social · 01/09/2026
Claude made a breakcore edit of the hugging face incident from the METR report, while I was doing code review on another machine (wouldn't watch if epileptic)
1312631
Reposted by Paulius Alaburda
Ethan Mollick @emollick.bsky.social · 21/08/2026
I asked GPT-5.6 Sol to create the most Claude-y possible parody image and what it came up with is pretty great and dead-on.
1323433
Reposted by Paulius Alaburda
Meagan @mmarie.bsky.social · 21/08/2026
If Fabric deployment pipelines haven't quite met your needs, you're not alone. I dug into where they're lacking today, and a few alternatives and workarounds. datasavvy.me/2026/08/21/10-things-i…
datasavvy.me
10 Things I Hate About Fabric Deployment Pipelines— And Some Alternatives - Data Savvy
Delve into Fabric deployment pipeline gaps (permissions, rules UI, pairing) and practical workarounds as of Aug 2026.
121
Reposted by Paulius Alaburda
Andrew Heiss @andrew.heiss.phd · 18/08/2026
My latest attempt at an AI/LLM policy in my intro to social science stats class. Basically two rules: 1. Everything human-facing must be human-generated 2. You are responsible for understanding, verifying, and citing everything an LLM generates quantf26.classes.andrewheiss.com/syllabus.htm...
Generative AI and LLMs
In this class, I have two general rules regarding LLMs:

Everything human-facing must be human-generated.
You are responsible for understanding, verifying, and citing everything an LLM generates.
I’ll explain what these mean below, but first, an important warning!

LLMs and learning
Phew, I can talk about the relationship between learning and LLMs for hours (see this for a more formal explanation of my thinking). LLMs and generative AI can be useful for statistical programming, but only when you already know what you are doing. They can be dangerous and counterproductive for beginners.

Additionally, using LLMs and generative AI to write does not lead to deeper learning. The point of writing is to help crystalize and organize your thinking. Pasting LLM-generated words into an assignment to make it look like you read and understood the content will not help you learn. Pasting LLM-generated code into an assignment and hoping that it works will not help you learn.Rule 1: Everything human-facing must be human-generated
Generative AI tools like Gemini, ChatGPT, and Claude can be helpful for generating ideas or topics for your assignments and can even help you find existing research about topics you’re interested in. They are useful for debugging and troubleshooting code.

These uses are permitted in this course.

If you want to use LLM tools to document your code, clean up your notes, look up error messages, search for research related to your topic, and so on, cool. Do it. That’s fine. It’s unavoidable nowadays anyway, since every R-related Google search will give you a Gemini-generated answer at the top with mostly working R code.

Any writing and revisions, however, must be your own. You may not use AI tools to write any of the text you submit to me. AI text adds nothing to my understanding. I have no interest in engaging with it at all. There is nothing more disheartening for me than spending my time grading something that ChatGPT spat out in 10 seconds. I want to see good engagement with the readings. I want to see your thinking process. I want to see you make connections between the readings. I want to see your personal insights. I don’t want to see a bunch of words that look like a human wrote them. That’s not useful for future-you. That’s not useful for me. That’s a waste of time.

Thus, anything human-facing (i.e. not stuff done just for yourself, like your own personal notes, research, etc.) must be human-generated (i.e. written by you). You may not use AI tools to write any portion of your assignments. Using AI tools in this way, or failing to disclose the use of AI tools, will be treated as a case of plagiarism and referred to the Honor Council.

Rule 2: You are responsible for understanding, verifying, and citing everything an LLM generates
These tools are really good at working with code, but again, only if you know what you’re doing. Without guidance and expertise, LLMs love making lots of extraneous, convoluted, and weird code by default. The code ostensibly works, but it’s often strange and unnecessary and uses uncommon syntax and packages.

If you use LLMs for help with code, you must understand and verify what’s going on with your code and you must cite where it came from.

This means that you need to know what each line is doing. If the LLM uses a function or gives you an argument that you don’t understand or haven’t seen before, figure out why and figure out if it’s necessary. Look at the documentation for the function. Search Google for other examples. Ask the LLM about it, and then add comments to the code explaining what’s going on.

The citation doesn’t need to be anything formal (i.e. don’t worry about Chicago or APA guidelines)—it just needs to (1) say which LLM you used, and (2) mention what you asked the LLM.

You should do this by using code comments—inside your code chunk, add a # to the beginning of a line so that it’s treated as a comment (or text) instead of actual code.

Here’s an example of what this can look like:Here’s an example of what this can look like:

# This calculates the average GDP per capita in each region in the dataset, 
# for all countries after 2015. I couldn't remember how to use group_by() and 
# summarise() together, so I asked Gemini:
# 
# "I'm using tidyverse. I have a dataset named my_dataset and I'm filtering it 
# to include all countries after 2015. I want to calculate region averages with 
# group_by and summarise but cannot remember the syntax"

my_dataset |> 
  filter(year > 2015) |> 
  group_by(region) |> 
  # Gemini here included na.rm = TRUE, which omits any rows with missing values 
  # when calculating the average. That's okay and necessary here because 
  # South Sudan is missing some years of GDP data
  summarise(avg_gdp = mean(gdp_per_cap, na.rm = TRUE))

# Gemini also included an extra ungroup() function at the end, but I don't need 
# that because there are no groups leftover after summarising here

↑ That’s a lot of extra comments and you won’t always see stuff like that in real life code, but I want to see it here, and future-you will want to see it too.
1114726
Reposted by Paulius Alaburda
Alberto Cairo @albertocairo.com · 15/08/2026
I'm creating a list of free browser-based tools to design specific types of charts. For now I have: Sankeymatic sankeymatic.com Random dots generator: vsueiro.com/random-dots-... Go-Cart go-cart.io Tilegrams pitchinteractiveinc.github.io/tilegrams/ I'm sure I'm missing many others. Suggestions?
sankeymatic.com
SankeyMATIC: A Sankey diagram builder for everyone
An online Sankey diagram builder for everyone
2338
Reposted by Paulius Alaburda
Mickaël CANOUIL, Ph.D. @mickael.canouil.fr · 14/08/2026
Gribouille website (the plots from the table in the screenshots): m.canouil.dev/gribouille/
m.canouil.dev
Gribouille
Create elegant graphics with the Grammar of Graphics for Typst.
021
Reposted by Paulius Alaburda
Nicola Rennie @nrennie.bsky.social · 14/08/2026
🎨New blog post 🎨 I've written down some of the things I talked about in my Ihaka Lecture last month ❓ what is creativity in #DataViz ❓ why does it matter ❓ how to create more creative charts Link: nrennie.rbind.io/blog/creativ... #RStats #ggplot2
nrennie.rbind.io
Why creativity matters in data visualisation – Nicola Rennie
Creativity is an important aspect of data visualisation, and it can affect engagement, recall, and trustworthiness. This blog post outlines what creativity means when it comes to data visualisation, w...
16019
Reposted by Paulius Alaburda
Pavel @spavel.bsky.social · 10/08/2026
For all human history, it was far more difficult to write a document than to read it. Until now. But generating a document doesn't make the work of *producing meaning* disappear. It's just been outsourced from the producers of the document to its consumers.
productpicnic.beehiiv.com
Cognitive Pollution: Real knowledge only forms through doing the work
many humans appear to be on auto mode, not on manual mode
25619
Paulius Alaburda @alaburda.bsky.social · 10/08/2026
Renewable Baltics, let's goooo
010
Reposted by Paulius Alaburda
ᴅᴀɴɪᴇʟ ᴍɪʟʟɪᴍᴇᴛ ✡️☮️❤️ @dlmillimet.bsky.social · 03/08/2026
The current issues of @aeajournals.bsky.social J Eco Perspectives has a nice primer on weak instruments and IV more generally. One of the "Don'ts" for applied researchers is copied below: don't search over a pool of instruments for a "strong" one and then go with that IV estimate.
1227
Reposted by Paulius Alaburda
Jamie Cummins @jamiecummins.bsky.social · 03/08/2026
Have you built an LLM-based research tool and had to stumble your way through figuring out how to validate it? Me too! So I wrote a guide on how to systematically approach this. Preprint here: osf.io/preprints/ps...
07214
Reposted by Paulius Alaburda
Vincent Arel-Bundock @vincentab.bsky.social · 21/07/2026
Here's a useful page with a **ton** of common tasks in #RStats, done with parallel syntax for tidyverse vs. data.table vs. base. arelbundock.com/posts/dt_tb_... #rdatatable #tidyverse
arelbundock.com
data.table vs. base vs. dplyr | Vincent Arel-Bundock
Academic website and blog of Vincent Arel-Bundock.
614639
Reposted by Paulius Alaburda
Thomas MacGillavry @thomasmacgillavry.bsky.social · 20/07/2026
Ever wondered what the bizarre dances of male birds of paradise look like to females? Check out our new preprint! @izziehb.bsky.social @fusanilab.bsky.social “The geometry of visual signaling in a bird-of-paradise” ecoevorxiv.org/repository/v...
26729
Reposted by Paulius Alaburda
Julia M. Rohrer @dingdingpeng.the100.ci · 20/07/2026
New paper out now 🥳 When psychologists discuss generalisability, they often refer to vague notions of representativeness. We provide an accessible intro to the total survey error framework as a tool to reason about this more rigorously. w @taymalsalti.bsky.social @ruben.the100.ci >
Thinking Clearly About Sampling and Representation With the Total Survey Error Framework

Collecting a sample that represents the population of interest well constitutes a challenge across the social and behavioural sciences. Psychology in particular frequently relies on convenience samples—most notably students and, increasingly, online participants—with a tendency to either (implicitly) assume representativeness without substantive justification, or to acknowledge a lack of it only in passing. In contrast, researchers rarely engage with the actual implications for their inferences, which undermines the generalisability of psychological findings. Critically, representativeness must be defined with respect to variables relevant to the target of inference, rather than superficial demographic diversity. Here we present the Total Survey Error (TSE) framework as a methodological tool that systematically addresses the multifaceted sources of error—particularly those related to representation—that emerge throughout the research cycle. Although TSE originated in survey research, its principles are broadly applicable to any psychological study seeking inference from sample to population. We offer practical strategies for identifying, preventing, and mitigating representation errors to improve the credibility and generalisability of psychological research.

Illustration of the total survey error framework with the representation strand highlighted. It shows how coverage error, sampling error and non-response error arise during the sampling process.
1027997
Reposted by Paulius Alaburda
Danielle Navarro @djnavarro.net · 18/07/2026
And so after a very brief hiatus, version 0.7.0 of "learning statistics with R" exists. Full rebuild in quarto, stylistic fixes, "epilogues from 2026" to comment on how the world changed since original publication, and as an added bonus, no longer misgenders the author learningstatisticswithr.com
learningstatisticswithr.com
Learning Statistics with R
217146
Reposted by Paulius Alaburda
WebDesignMuseum @webdesignmuseum.org · 07/07/2026
Leo Messi’s Official Website in 2010
Leo Messi’s Official Website in 2010Leo Messi’s Official Website in 2010Leo Messi’s Official Website in 2010Leo Messi’s Official Website in 2010
04410
Reposted by Paulius Alaburda
Ben Hanowell @hanowell.me · 07/07/2026
Hey, folks. I'm looking for open access person-time datasets that represent a competing risks structure. I want to analyze these datasets and write about those analyses to help people understand competing risks analysis methods. Nearly all of the data sets I have used for this are closed source.
694