Sign in

Tom Carpenter, PhD

@tcarpenter.bsky.social
3.6K followers 3.9K following 488 posts

🧪 Data science, survey science, social science 💻 Director of Data Science @ Microsoft Garage [Posts do not represent my employer] 🧮 Stats, R, python 📝 Science, Research: measurement, social biases, emotion. Ex-academic but scientist at heart

PostsRepliesMedia
Tom Carpenter, PhD @tcarpenter.bsky.social · 28/12/2025
“Instead of” there is doing a lot of work If I were still teaching, I would be talking about how to use AI to make us better writers… E.g. giving us feedback or critique… But the mental muscles I went to train have to to be the students’
250
Tom Carpenter, PhD @tcarpenter.bsky.social · 06/12/2025
Ask copilot or chat gpt. It can output a csv no sweat.
060
Tom Carpenter, PhD @tcarpenter.bsky.social · 28/11/2025
Possibly. We might also admit some things are not very knowable— without other “ways of knowing”
070
Tom Carpenter, PhD @tcarpenter.bsky.social · 28/11/2025
I often hear researchers feeling bad about this. But it’s called prioritization. The key is to do it intentionally and strategically. You have finite time and can’t do everything.
191
Tom Carpenter, PhD @tcarpenter.bsky.social · 27/11/2025
On the one hand, people need better education about AI On the other hand, do we even understand what it means to be conscious in the first place?
150
Tom Carpenter, PhD @tcarpenter.bsky.social · 19/11/2025
Ah, but read what I said. I’m not saying it’s impossible to do valid online research. I am saying that volume will need to shrink. It’s going to be harder to stay on top of this. Some folks will give up or mess it up. There will be an arms race. Reviewers will start raising flags. Etc.
100
Tom Carpenter, PhD @tcarpenter.bsky.social · 19/11/2025
The volume of social psych papers that ran on prolific and mturk is going to need to shrink soon, and that will have a big impact on a lot of small labs
162
Tom Carpenter, PhD @tcarpenter.bsky.social · 14/11/2025
A huge pain, but doable with undergrad participants essentially for free (or at least historically so). I’m not saying what is right… I’m saying what is driving behavior
020
Tom Carpenter, PhD @tcarpenter.bsky.social · 13/11/2025
If we want this to change, then we need to make it feasible TMP factor: Time, money, pain Whether a study gets done depends a lot on how difficult it is. The question is one of ROI. I suspect many researchers are thinking to themselves, “that’s a great thing, but it’s not something I can do”
290
Reposted by Tom Carpenter, PhD
Jess Calarco @jessicacalarco.com · 01/11/2025
We have progressed from data collection to data analysis.
My 11-year-old sitting with her pile of Halloween candy, sorting it into a bar graph
970344664087
Reposted by Tom Carpenter, PhD
Menotti Minutillo @menotti.bsky.social · 30/10/2025
ladies and gentlemen...we got him
171181784003
Reposted by Tom Carpenter, PhD
Greg Pak @gregpak.net · 10/10/2025
* hoppening
121000146
Tom Carpenter, PhD @tcarpenter.bsky.social · 04/09/2025
But where does 1a say anything about personhood? I’m reading it and it seems clear that you can’t restrict speech—and that would obviously apply to one or more people organized under a LLC. IMO the bigger question is “when is money free speech, vs when is it corruption”
210
Reposted by Tom Carpenter, PhD
Cameron Patrick @cameronpat.bsky.social · 18/08/2025
this slide is from a colleague's introductory stats course, I think it fits many statisticians' experiences
Slide titled: "Assumptions of the model and model checking"
with a scatterplot with axes how much people should worry vs how much people do worry.
89224
Reposted by Tom Carpenter, PhD
Adam L @adam-lg.bsky.social · 28/07/2025
45. Academia doesn't reward building useful tools nearly as much as it should
2344
Reposted by Tom Carpenter, PhD
Adam L @adam-lg.bsky.social · 27/07/2025
14. We mostly evaluate latent variable models with the equivalent of Rorschach tests
1111
Reposted by Tom Carpenter, PhD
Adam L @adam-lg.bsky.social · 27/07/2025
5. You should use a precision-recall curve for a binary classifier, not an ROC curve
1232
Reposted by Tom Carpenter, PhD
JLRay @jlray.bsky.social · 25/07/2025
Wow. Scientists have edited mosquito DNA to prevent the spread of malaria to humans "while supporting essential physiological functions... and negligible fitness costs" to the mosquito population. Potentially ending the mosquito-born spread of malaria to humans. www.nature.com/articles/s41...
361080310
Tom Carpenter, PhD @tcarpenter.bsky.social · 23/07/2025
… set of paths consistently supported by the data. Even getting that down is a trick. And making sense of it is fraught and doesn’t get you much further than one would get from regression. But at least then we would have some confidence we understand the correlational relationships!
000
Tom Carpenter, PhD @tcarpenter.bsky.social · 23/07/2025
… the model is correct and then gives you what the path would be under that specification. There’s nothing different when we go to SEM other than your ability to p-hack goes up exponentially. IMO this would be a great place to use machine learning approaches to train / tune models to find …
100
Tom Carpenter, PhD @tcarpenter.bsky.social · 23/07/2025
… all those hypotheses together (in the same way that ANOVA contest many multiple comparisons at once). There’s nothing different between this and running a bunch of regressions and claiming the results support the way you specified those models. In reality, it’s the reverse. Regression assumes …
100
Tom Carpenter, PhD @tcarpenter.bsky.social · 23/07/2025
Yes and see this a lot in social too. Proper use of SEM implies a particular philosophy of hypothesis testing in regression contexts. An omitted path is hypothesizing that path is exactly 0. A non-omitted path hypothesizing it is non-zero. Model fit is effectively the joint set of …
100
Tom Carpenter, PhD @tcarpenter.bsky.social · 23/07/2025
Yikes!
110
Tom Carpenter, PhD @tcarpenter.bsky.social · 23/07/2025
… SEM for causal discovery. However, if you have a good read on the causal process, it can be great for estimating parameters such as factor, loadings or paths with latent variables
050
Tom Carpenter, PhD @tcarpenter.bsky.social · 23/07/2025
This is probably not anything you don’t already know …. But I did a lot of SEM work and will repeat it anyway. The model assumes you know the causal structure. Fit indices will confirm that the model is a fit to the data, but many incorrect models can fit the data. So I would not use …
291
Tom Carpenter, PhD @tcarpenter.bsky.social · 19/07/2025
Curious how this compares to the cost of living per state
020
Tom Carpenter, PhD @tcarpenter.bsky.social · 09/07/2025
media.tenor.com
a man in a suit and tie is making a funny face and saying you don 't say ?
ALT: a man in a suit and tie is making a funny face and saying you don 't say ?
060
Reposted by Tom Carpenter, PhD
Julia M. Rohrer @dingdingpeng.the100.ci · 08/07/2025
Let's say you want to include age as a predictor in your model. How do you do that? Here's an illustration of three options -- it's for a paper I'm working on (so if you feel like anything could be tweaked...).
Plot that depicts the average importance people in my data assign to their friendships (y-axis, on a scale from 1 to 5, depicted with 95% confidence intervals) by their age (x-axis, from 18 to 60).

Depicted are 3 different ways to model importance of friends as a function of age.
Using age as a linear predictor: this imposes a linear trajectory which comes with very tight confidence intervals (i.e., uncertainty is low).
Using age as a categorical predictor: this imposes no trajectory whatsoever but instead simply reproduces the means by age. The confidence intervals are very wide, in particular for those ages not well represented in the data (i.e., uncertainty is high).
Age splines: This results in a smooth trajectory that follows some of the bumps in the data, but not all of them. The confidence intervals are somewhere between the linear and the categorical case (i.e., uncertainty is medium)
3215629
Reposted by Tom Carpenter, PhD
respectful huff @alexqarbuckle.bsky.social · 05/07/2025
There should be a corner at Home Depot where a guy with a table saw will slice you off custom lengths of hot dog from an infinite hot dog coming out of the wall
844101610
Reposted by Tom Carpenter, PhD
Ian Carlos @ianxcarlos.bsky.social · 25/06/2025
There were two girls at Wawa just now talking about funny movies and one said, “Have you ever seen the movie Office Space? It’s an old people movie but it’s funny”
Adele shattering a glass in her hand
4076802365
Tom Carpenter, PhD @tcarpenter.bsky.social · 24/06/2025
Counterpoint: the ability to chat with an article or literature and find patterns in our own work that perhaps we missed I think has a lot of potential to augment our scientific work
100
Tom Carpenter, PhD @tcarpenter.bsky.social · 24/06/2025
000
Reposted by Tom Carpenter, PhD
Emilio Ferrara @emilioferrara.bsky.social · 23/06/2025
🤖Thrilled to share our latest work☄️ Have you ever wondered what LLMs know but they are not saying? We built an auditing framework to study information suppression in LLMs, and demonstrated it to quantify and characterize censorship in DeepSeek. Read more: arxiv.org/abs/2506.12349
arxiv.org
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
This study examines information suppression mechanisms in DeepSeek, an open-source large language model (LLM) developed in China. We propose an auditing framework and use it to analyze the model's res...
172
Reposted by Tom Carpenter, PhD
Ida Momennejad @neuroai.bsky.social · 20/06/2025
Pleased to share our ICML Spotlight with @eberleoliver.bsky.social, Thomas McGee, Hamza Giaffar, @taylorwwebb.bsky.social. Position: We Need An Algorithmic Understanding of Generative AI What algorithms do LLMs actually learn and use to solve problems?🧵1/n openreview.net/forum?id=eax...
316537
Tom Carpenter, PhD @tcarpenter.bsky.social · 18/06/2025
building intuition around systems of matrix data and how we manipulate them should be right after (or right before) basic calc (integrals, derivatives, partial derivatives)
210
Tom Carpenter, PhD @tcarpenter.bsky.social · 18/06/2025
Why?
100
Tom Carpenter, PhD @tcarpenter.bsky.social · 18/06/2025
Someone please explain why linear algebra isn’t taught more in high school in the United States? Seems like maybe if you’re lucky you get a few lectures on matrices and that’s it.
260
Tom Carpenter, PhD @tcarpenter.bsky.social · 17/06/2025
Is it 2015 or 2025?
030
Tom Carpenter, PhD @tcarpenter.bsky.social · 17/06/2025
“Data available upon request”
170
Reposted by Tom Carpenter, PhD
John Schwartz @jswatz.bsky.social · 08/06/2025
Under the Trump agenda, energy will cost more. And when energy costs more, everything costs more. www.nytimes.com/2025/06/04/c...
nytimes.com
Electricity Prices Are Surging. The G.O.P. Megabill Could Push Them Higher.
12312
Tom Carpenter, PhD @tcarpenter.bsky.social · 05/06/2025
Many academics are taught “do more” instead of “prioritize”. They are taught the academic superhero myth, that “truly great/smart” scholars can handle it. So they sabotage their own success in service to the cult of personality
072
Tom Carpenter, PhD @tcarpenter.bsky.social · 05/06/2025
Indeed! pubmed.ncbi.nlm.nih.gov/35917203/
pubmed.ncbi.nlm.nih.gov
Stability and Change in Subjective, Psychological, and Social Well-Being: A Latent State-Trait Analysis of Mental Health Continuum-Short Form in Korea and the Netherlands - PubMed
Mental well-being consists of hedonic/subjective, psychological, and social dimensions. Research has yet to determine how much of the variance in these three dimensions is stable or variable over time...
030
Tom Carpenter, PhD @tcarpenter.bsky.social · 05/06/2025
Galaxy brain: what does anything measure?
110
Tom Carpenter, PhD @tcarpenter.bsky.social · 04/06/2025
I love this technique because it gives a cool way to isolate change and stability components of within-person measurement using latent variables. The stable portion of IATs is far more predictive of other individual-difference measures than one would think given traditional scoring/analyses
271
Tom Carpenter, PhD @tcarpenter.bsky.social · 04/06/2025
I love this analysis technique because it gives a cool way to isolate change and stability components of within-person measurement using latent variables. The stable portion of IATs is far more predictive of other individual-difference measures than one would think given traditional scoring/anakysus
031
Tom Carpenter, PhD @tcarpenter.bsky.social · 04/06/2025
That moment when you had completely forgotten about a cool data app you’d made, only to find others posting about it ;) Thanks @calvinklai.bsky.social!
061
Reposted by Tom Carpenter, PhD
Calvin Lai @calvinklai.bsky.social · 04/06/2025
Paper w/ @tcarpenter.bsky.social & @alexgoedderz.bsky.social at PSPB!🚨 The standard IAT is only 5 min long. We found that making the IAT longer by taking it multiple times greatly improves predictive validity. 🧵below, with practical advice about how to run IATs! LINK: osf.io/preprints/ps...
45016
Reposted by Tom Carpenter, PhD
Calvin Lai @calvinklai.bsky.social · 04/06/2025
BOTTOM LINE: Aggregating 2 IATs together and using CFA can lead to much higher predictive validity. Finally, I wanted to thank Tom and Alex! This paper came out while they were moving out of academia & I was going on parental leave, so we hadn't publicized it. Better late than never! 😉
051
Reposted by Tom Carpenter, PhD
Calvin Lai @calvinklai.bsky.social · 04/06/2025
FAQ Q1: 4 IATs is a lot!? A: You only need 2 IATs to reap the benefits of aggregating. Q2: It's a lot of work to set up these models. A: Here's a shiny app to make it easy: tcarpenter.shinyapps.io/trait_iat/ Q3: I don't want to learn CFA. A: Averaging IATs together gets you some of the benefits.
151
Tom Carpenter, PhD @tcarpenter.bsky.social · 01/06/2025
What should universities look like 10 to 15 years from now? Lots of intersecting headwinds… dramatic change is likely. Biggest IMO is that teaching universities become cheap again by reverting to “room with chairs and a prof”. Think community college, but with amazing faculty. 
010