Sign in

Dylan Pieper

@dylanpieper.bsky.social
239 followers 1.6K following 98 posts

Data dude 📈 • Dog dad 🐕 • Pilot 🪂 • #rstats • dylanpieper.github.io

PostsRepliesMedia
Dylan Pieper @dylanpieper.bsky.social · 30/09/2026
🫵 hey, does this thing still work?
000
Dylan Pieper @dylanpieper.bsky.social · 01/12/2025
I forgot to mention that you can now import your REDCap audit logs. 🤓 It defaults to the past 7 days.
000
Dylan Pieper @dylanpieper.bsky.social · 01/12/2025
🦆 📦 redquack 0.3.0 now includes a suite of convenience functions, and more importantly, full control over labeling variables and coded values without separate imports. This makes redquack a great tool for large or small REDCap projects. ➡️ dylanpieper.github.io/redquack/ #rstats #redcap
dylanpieper.github.io
Transfer REDCap Data to Database
Transfer REDCap (Research Electronic Data Capture) data to a database, specifically optimized for DuckDB. Processes data in chunks to handle large datasets without exceeding available memory. Features...
150
Dylan Pieper @dylanpieper.bsky.social · 15/08/2025
**course material for adults who swear**
000
Reposted by Dylan Pieper
Alejandro Wainzinger @xevix.bsky.social · 15/08/2025
Stretching DuckDB w/ Common Crawl, ~1.7B rows, ~300 parquet files. ~2-3s for single-column aggregations, ~2-3 mins to SUMMARIZE the data, peaking at ~12-14GB memory usage. Not exactly real-time, but the fact you can do this on a laptop with no server setups or Spark pipelines is still amazing.
1439
Dylan Pieper @dylanpieper.bsky.social · 05/08/2025
A little known fact is that RStudio rendering is powered by users’ electromagnetic fields (i.e., “good vibes”) and the exodus to Positron has severely limited its ability to compile code. #rstats
media.tenor.com
a colorful background with the words " the more you know " and a star
ALT: a colorful background with the words " the more you know " and a star
191
Reposted by Dylan Pieper
Libby Heeren @libbyheeren.bsky.social · 04/08/2025
Remember this #rstats post? I wasn't the only one talking about it & the tidyverse team was listening 😎 #databs New #dplyr functions? They're looking for feedback!! 🤔 replace_when, recode_values, replace_values 👀 Read this: github.com/tidyverse/ti... 🗣️ Comment on PR: github.com/tidyverse/ti...
26215
Reposted by Dylan Pieper
Nate O @nateo.bsky.social · 29/07/2025
I think pedocon theory is right. It’s empirically adequate, parsimonious, fits within a broader theoretical framework, and has immense explanatory breadth and depth www.liberalcurrents.com/we-need-to-t...
liberalcurrents.com
We Need to Talk About Pedocon Theory
The connection between Donald Trump and Jeffrey Epstein is no accident, but reveals a deep logic at the heart of reactionary politics.
27126
Reposted by Dylan Pieper
Hadley Wickham @hadley.nz · 11/07/2025
I am such a sucker for frivolous uses of AI. Here's an anthem for the tidyverse: suno.com/s/iVMVs4IoyA...
suno.com
281
Dylan Pieper @dylanpieper.bsky.social · 12/07/2025
Modules + Claude code for simple but labor intensive edits across files
000
Reposted by Dylan Pieper
Crystal Lewis @cghlewis.bsky.social · 08/07/2025
Very cool to see authors of this article mentioning the importance of sharing project-, data-, AND variable-level documentation alongside data in a repository, and linking to the templates I've provided on OSF as an example! 🌟 doi.org/10.1515/ling...
1268
Dylan Pieper @dylanpieper.bsky.social · 08/07/2025
Tbh I relate to that big yellow spike of with mad uncertainty around age 30. 🤣
010
Reposted by Dylan Pieper
Crystal Lewis @cghlewis.bsky.social · 27/06/2025
As a data manager, good documentation not only helps me do my job better, but also helps me annoy you less! 😅 Good documentation about inclusion criteria, READMEs about oddities in the data, consort diagrams and tracking to explain missing data, and so on, are all ways to ensure I bug you less! 🐛🐜🐝
0222
Dylan Pieper @dylanpieper.bsky.social · 19/06/2025
I think if you’re curious and truly care about problem solving you might have a ~temporary~ feeling of closure or a premature commit. But you will keep iterating (opening/closing) as you explore the problem space and how it works, validate the throughput, and improve the methods. Stay curious!
020
Reposted by Dylan Pieper
Hadley Wickham @hadley.nz · 18/06/2025
New to me is the term "premature closure", where you too quickly latch on to the first solution you see. Always a danger in coding, but particularly so today when LLMs can give you a plausible fix so so quickly. www.shayon.dev/post/2025/16...
shayon.dev
Pitfalls of premature closure with LLM assisted coding
When LLM models generates clean, professional-looking code, it's tempting to stop exploring alternatives. But therein lies the risks that comes with premature closure. So what is premature closure?
79614
Dylan Pieper @dylanpieper.bsky.social · 19/06/2025
I would use cosine in stringdist. If you have lists of job descriptions (from two sources with each idx being a similar job), you can use my package samesies. dylanpieper.github.io/samesies/
dylanpieper.github.io
Compare Similarity Across Text, Factors, or Numbers
Compare lists of texts, factors, or numerical values to measure their similarity. The motivating use case is evaluating the similarity of large language model responses across models, providers, or pr...
130
Dylan Pieper @dylanpieper.bsky.social · 14/06/2025
So cool! Any intros or docs planned for helping people familiar with the futureverse make the leap to marai?
000
Reposted by Dylan Pieper
Charlie Gao @shikokuchuo.net · 13/06/2025
Bleeding edge update for the #tidyverse purrr package with even more seamless #rstats parallel maps. Introducing our shiniest new adverb: `in_parallel()`. Just wrap your function to take advantage of blazing fast parallel processing via mirai. pak::pak("tidyverse/purrr") purrr.tidyverse.org/dev/
purrr.tidyverse.org
Functional Programming Tools
A complete and consistent functional programming toolkit for R.
610332
Reposted by Dylan Pieper
Vincent Arel-Bundock @vincentab.bsky.social · 12/06/2025
One cool thing you can/should do is sample from priors only, and plot the distribution of the actual quantity of interest (ex: risk ratio). I find this very useful. This is actually super easy with brms. arelbundock.com/posts/margin...
arelbundock.com
Prior Predictive Checks with marginaleffects and brms – Vincent Arel-Bundock
1191
Reposted by Dylan Pieper
JD Long @jdlong.cerebralmastication.com · 11/06/2025
This blog post about engineering not doing ETL is nine years old… it’s worth reviewing multithreaded.stitchfix.com/blog/2016/03...
multithreaded.stitchfix.com
Engineers Shouldn’t Write ETL: A Guide to Building a High Functioning Data Science Department | Stitch Fix Technology – Multithreaded
“What is the relationship like between your team and the data scientists?” This is, without a doubt, the question I’m most frequently asked when conducting i...
3194
Dylan Pieper @dylanpieper.bsky.social · 04/06/2025
The worst is when you write in active voice and then someone tries to edit all of it back into passive. Old habits die hard and the good fight continues.
110
Dylan Pieper @dylanpieper.bsky.social · 31/05/2025
You could use surveydown and provide the LLM with the package docs
110
Reposted by Dylan Pieper
Alex Kraieski @kraieski.dev · 30/05/2025
Here's a functional programming trick for #rstats that I wish I started using sooner: if you need a #ggplot2 scale to be reusable across multiple plots and dynamically configurable without relying on global state, consider using a function factory (a function that returns a function) to build it
screenshot of a code editor showing the following R code:

library(quantmod)
library(ggplot2)
library(lubridate)

startYear <- 2015
startDate <- paste0(startYear, '-01-01')
getSymbols(c('spy', 'btc-usd'), from= startDate)

# function factory that creates a scale function that only shows valid years.
# try to keep code that could change in here!
make_valid_year_scale_function <- function(start_year){
  function(){
    list(
      scale_x_continuous(breaks = seq(start_year, Sys.Date() |> year(), 1)),
      theme(panel.grid.minor.x = element_blank()) # use function after other theme funcs
    )
  }
}

# this makes it so I can add scale_x_valid_years() to any plot
scale_x_valid_years <- make_valid_year_scale_function(startYear)
5355
Reposted by Dylan Pieper
Charlie Gao @shikokuchuo.net · 23/05/2025
mirai - minimalist async framework for #RStats - released as an 'r-lib' package. Blog post: Advancing Async Computing in R. shikokuchuo.net/posts/26-mir... mirai provides event-driven async for #RShiny and parallel processing for purrr #tidyverse. Really excited to be working on this at Posit!
shikokuchuo.net
shikokuchuo{net}: mirai 2.3.0
Advancing Async Computing in R
06218
Reposted by Dylan Pieper
Carl T. Bergstrom @carlbergstrom.com · 24/05/2025
tl;dr — this EO co-opts the language of open science to implement a system of political control wherein presidential appointees are given broad latitude to designate any number of reasonable scientific activities and inferences as scientific misconduct, and to penalize those involved accordingly.
whitehouse.gov
Restoring Gold Standard Science
By the authority vested in me as President by the Constitution and the laws of the United States of America, including section 7301 of title 5, United
10024491048
Reposted by Dylan Pieper
Kevin Zollman @kevinzollman.com · 20/05/2025
There's so much polarization around LLMs. They are way overhyped, I agree. But I also use them semi-regularly now. Here's a thread of genuine use cases where I find them helpful. Please add your own!
79022
Reposted by Dylan Pieper
Jasmine Daly @jasminedaly.bsky.social · 19/05/2025
📦 I’m excited to share a new #rstats package I’ve been working on: {shinyfa} built to help folks working on large or unfamiliar #rshiny apps ✨ The package scans your app folders and extracts out details on render*(), reactive() and input$ to a dataframe! 📖 www.dalyanalytics.com/blog/shinyfa...
dalyanalytics.com
Introducing {shinyfa}: Analyze Large Shiny App Codebases Faster with This R Package | Daly Analytics
Discover {shinyfa}, a new R package designed to improve developer experience by analyzing and summarizing the structure of large Shiny applications. Perfect for consultants, teams, and contributors wo...
1122
Reposted by Dylan Pieper
Michael Howe @mchowe.bsky.social · 18/05/2025
Playing around with satellite imagery of #madison to make some office art. #Rstats
051
Reposted by Dylan Pieper
Hadley Wickham @hadley.nz · 18/05/2025
✨Use llms from #rstats with ellmer ✨Version 0.2.0 is on CRAN now. No blog post yet because I'm about to go on vacation, but in the meantime you can check out the release notes: github.com/tidyverse/el....
github.com
36913
Reposted by Dylan Pieper
Crystal Lewis @cghlewis.bsky.social · 16/05/2025
The kind of Friday morning content I needed to see. ❤️
0141
Reposted by Dylan Pieper
Posit @posit.co · 15/05/2025
Registration for the posit::conf(2025) virtual experience is now open! Join us virtually, Sept 16–18, and access live-streamed keynotes and 100+ talks, on-demand recordings, Q&A sessions, and our virtual networking platform. Learn more in the blog post: posit.co/blog/posit-c... #RStats #Python
Text: posit conf 2024 Virtual Tickets Available, Atlanta, September 16-18. A drawing outline of the Atlanta skyline and abstract cubes.
11815
Reposted by Dylan Pieper
easystats @easystats.github.io · 15/05/2025
In case you missed it, we recently updated some of our packages, including many new features (again) in the #rstats #easystats {modelbased} package: easystats.github.io/modelbased/n... The last weeks we were working a lot on improving support and performance for Bayesian models and especially
easystats.github.io
Changelog
1134
Dylan Pieper @dylanpieper.bsky.social · 15/05/2025
Don’t forget you can distinct_all() which avoids this problem if you’re looking to filter completely duplicate rows
140
Reposted by Dylan Pieper
Crystal Lewis @cghlewis.bsky.social · 08/05/2025
I'm still thinking about my favorite quote from the Posit Data Science Hangout today. It perfectly sums up what I hope I provide to the researchers I work with: a trusted partner, who is there to support them in their work. Earn a reputation for being a good person to work with - Cara Thompson
0255
Dylan Pieper @dylanpieper.bsky.social · 08/05/2025
Data science = made with ❤️ Data science = made with sugar, spice, and everything nice 🤷🏼‍♂️ We’ll get there someday 😂
020
Dylan Pieper @dylanpieper.bsky.social · 07/05/2025
To infinity and beyond!
010
Reposted by Dylan Pieper
Terry Christiani 👋 @terrychristiani.bsky.social · 06/05/2025
Great news! R/Medicine 2025 is providing a forum for sharing R based tools and approaches used to analyze and gain insights from health data. Join us for the premier R conference for health and medicine. 🔗 Register today: rconsortium.github.io/RMedicine_we... #rstats #opensource #RMed25
rconsortium.github.io
register – R/Medicine 2025
096
Reposted by Dylan Pieper
David Ho @davidho.bsky.social · 04/05/2025
I think a lot about what Carl Sagan said in one of his final interviews.
"WE'VE ARRANGED A society based on science and technology, in which nobody understands anything about science technology. And this combustible mixture of ignorance and power, sooner or later, is going to blow up in our faces. Who is running the science and technology in a democracy if the people don't know anything about it?"
"Science is more than a body of knowledge, it's a way of thinking. A way of skeptically interrogating the universe with a fine understanding of human fallibility. If we are not able to ask skeptical questions, to interrogate those who tell us that something is true, to be skeptical of those in authority, then we're up for grabs for the next charlatan, political or religious, who comes ambling along."
246188196426
Dylan Pieper @dylanpieper.bsky.social · 02/05/2025
I’m happy to share that I’ll be giving a talk at R/Medicine 2025! 🎊 I work with a BIG REDcap database for substance use treatment (200+ locations) which makes extraction difficult. I developed {redquack}, an #rstats 📦 that transfers REDCap data to DuckDB, and will talk about how to use it. 🦆
192
Dylan Pieper @dylanpieper.bsky.social · 02/05/2025
Thanks for responding and providing feedback! I will probably remove those descriptions. The basic likert options should be sufficient and less confusing.
000
Dylan Pieper @dylanpieper.bsky.social · 01/05/2025
📊 🕵️‍♂️ #rstats community! Do you sometimes feel like you're just pretending to be a data scientist? I'm researching imposter syndrome for my upcoming talk at posit::conf(2025) 🔍 I'd love to hear YOUR experiences in a short 5-10 minute anonymous survey: forms.gle/YkJtwZWquyKM... Please share! 🔄
forms.gle
Imposter Syndrome in Data Science
This survey is intended to gather community feedback from data scientists and students or recent graduates interested in data science as a career. Your responses will be anonymous and may be used for...
335
Reposted by Dylan Pieper
Andrew Heiss @andrew.heiss.phd · 29/04/2025
Since it's in Atlanta, I'll be here at my first posit::conf! I'll be speaking here with @gosterhout.bsky.social about election night reporting with #rstats and #QuartoPub (showing off {targets} and other neat tricks like this: www.andrewheiss.com/blog/2024/11... )
1433
Reposted by Dylan Pieper
Gabe Osterhout @gosterhout.bsky.social · 11/11/2024
States spend too much on clunky election night reporting. We just replaced ours using #rstats. dbplyr backend + reactable & leaflet viz + #quartopub site. Real magic happened with programmatic code chunks & targets pipeline done by @andrew.heiss.phd. #dataviz results.voteidaho.gov
4457
Reposted by Dylan Pieper
Julia M. Rohrer @dingdingpeng.the100.ci · 29/04/2025
Another q for the stats people! People worry about collinearity (cf blog post below). Consider a scenario in which the collinear predictors are just controls to account for confounding. Including both of them doesn't impair the precision with which the effect of interest is estimated, does it?
janhove.github.io
Jan Vanhove :: Blog - Collinearity isn’t a disease that needs curing
138720
Reposted by Dylan Pieper
Hadley Wickham @hadley.nz · 24/04/2025
Happy reinstalling-all-your-R-packages day to all those who celebrate #rstats
1216321
Dylan Pieper @dylanpieper.bsky.social · 24/04/2025
My shins are sighing in relief. 😮‍💨 I’m curious if this would affect prod apps that don’t need the files watched. Can you turn it off?
100
Reposted by Dylan Pieper
Vincent Arel-Bundock @vincentab.bsky.social · 24/04/2025
Hive mind, please help me out! I need a more informative and explicit subtitle for my upcoming #RStats book "Model to Meaning" The premise is that analysts should often transform coefficient estimates into more meaningful / interpretable quantities like predictions, risk differences, slopes, etc.
238316
Reposted by Dylan Pieper
Adam H. Smiley @asmiley.bsky.social · 23/04/2025
My amazing independent study student wrote me a thank you card and drew this laptop with #rstats code on it 🥹
A hand drawn laptop with R studio open o the screen
0214
Dylan Pieper @dylanpieper.bsky.social · 19/04/2025
How is BlueSky caving to the authoritarian regime? Genuinely thought they were were better than that 😞
000
Reposted by Dylan Pieper
Crystal Lewis @cghlewis.bsky.social · 16/04/2025
When planning for data collection, especially in longitudinal studies, first consider how that data will be used. Ask yourself: - How will we combine data for analysis? - What unique IDs will allow us to do this? - How will we name/code items to combine data? - Will our data need restructuring?
1305