Sign in

Nate TeBlunthuis

@groceryheist.cc
668 followers 1.9K following 209 posts

tebluntuhis.cc Assistant Professor at the University of Texas at Austin in the School of Information. Computational social science, peer production, social computing, HCI.

PostsRepliesMedia
Nate TeBlunthuis @groceryheist.cc · 28/06/2025
I found that they did! The graphic below depics how most of the longest lasting episodes of ecological interaction between subreddits were mutualistic.
Frequency plot of the durations of competition and mutualism episodes. Mutualism tends to last longer than competition. The y-axis is log-transformed. The axes truncated to omit outliers for visibility.
110
Nate TeBlunthuis @groceryheist.cc · 28/06/2025
In that work used time series models to infer networks of competition and mutualism between overlapping online communities. This work found evidence that they tended to be mutualistic. For example, the diagram below shows a network of mental health subreddits that is dense with mutualism.
Ecological network of a cluster of mental subreddits. Blue arrows indicate mutualism and yellow arrows indicate competition according to a vector autoregression model.
110
Nate TeBlunthuis @groceryheist.cc · 28/06/2025
Often, several different online communities exist where similar people talk about similar things. This is really easy to observe browsing Reddit or Facebook groups. For example The visualization of clustered subreddits with overlapping users blow shows different subreddits related to cycling.
Names of bicycle-related subreddits in cluster of subreddits with many overlapping users.
100
Nate TeBlunthuis @groceryheist.cc · 11/06/2025
Got a cool zine in the mail today. Free download here: lovelesspress.itch.io/a-web-worth-...
Cover of a Zine titled "A Web worth wandering: The radical internet and you". The paper is pink. The title is above. "A" is in a serif font. "Web" is italic serif. "WORTH WANDERING" is all caps sans serif. A black and white image of a 90s desktop computer, possibly a Macintosh Performa is right-aligned and separated from the title by whitespace. The subtitle is below the image.  The back of the zine. Text reads 
"Made by loveless press 
Follow: @lovelesspress
Read: lovelesspress.itch.io
Visit: lovelesspress.neocities.org
Personal indie website:
soupafterhours.neocities.org
020
Nate TeBlunthuis @groceryheist.cc · 17/09/2024
Bluesky now has over 10 million users, and I was #43,945!
021
Nate TeBlunthuis @groceryheist.cc · 17/07/2024
Thrilled to announce my appointment as Assistant Professor of Social Informatics at the @UTiSchool . I'm so thrilled to join this intellectual community :D. I'm recruiting PhD students interested in online communities and AI/ML in social science, broadly construed. Hook 'em!
Photo of Nathan TeBlunthuis, smiling, wearing a longhorns T-shirt, and making the "hook'em horns" gesture. The background is a vibrant green plant.
141
Nate TeBlunthuis @groceryheist.cc · 30/07/2023
This app, significantly, is the only competitor having a blue logo reminiscent of Twitter.
020
Nate TeBlunthuis @groceryheist.cc · 14/07/2023
Computational social scientists using machine classifiers build trust in evidence by reporting *predictive performance*. Metrics like F1 or AUC for this are important, but our results show that we can do better. We can use validation data to correct misclassification bias!
110
Nate TeBlunthuis @groceryheist.cc · 14/07/2023
Similarly, when a classifier predicts the DV and makes errors that are correlated with an IV, naïve estimates of that IV can be badly biased. In this case, only our MLA method was able to recover the true value.
igure showing the simulation results for systematic misclassification of a dependent variable. Only the maximum likelihood adjustment (MLA; ours) approach obtains consistent estimates of 𝐵𝑋
110
Nate TeBlunthuis @groceryheist.cc · 14/07/2023
For instance, as this figure shows, when a (not very accurate and moderately biased) classifier predicts an IV and makes errors that are correlated with the DV, a naïve (uncorrected) method gets the sign wrong and is very confident about it!
Figure showing the effectiveness of methods against ystematic misclassification of an independent variable. Only the maximum likelihood adjustment (MLA; ours) and multiple imputation (MI) approaches obtain consistent estimates of 𝐵𝑋 , but
MLA is more efficient.
110
Nate TeBlunthuis @groceryheist.cc · 14/07/2023
We use monte-carlo simulations to test methods proposed by social scientists, but none work in all the above scenarios (details in the paper). Therefore, we propose *maximum likelihood adjustment* (MLA), tailored from a framework drawn from biostats, which does!
Slide summarizing the other methods we investigated. These are Regression Calibration via Generalized Method of Moments (GMM) by Fong and Tyler in the paper “Machine Learning Predictions as Regression Covariates” which uses manual annotations to calibrate predictions but only works for IVs.
Multiple Imputation (MI), which uses predictions to impute manual annotations and is said to work with IVs or DVs. (citation Blackwell, Honaker, and King, “A Unified Approach to Measurement Error and Missing Data”)
Pseudo-likelihood (PL), which uses use precision and recall to model misclassification and is said to work with IVs or DVs (citation Zhang, How Using Machine Learning Classification as a Variable in Regression Leads to Attenuation Bias and
What to Do About It).
110
Nate TeBlunthuis @groceryheist.cc · 14/07/2023
Finally, to show (3), we test methods that use (small amounts of) validation data to correct misclassification bias. An ideal method works with independent or dependent variables (IV or DV) and with errors that are *random* or *systematic* (correlated with modeled variables).
Visualization of Bayesian networks showing the conditional independence structure of the simulations.
110
Nate TeBlunthuis @groceryheist.cc · 14/07/2023
We show (1) using Perspective API, a toxicity classifier widely used to study social media, and the human-labeled civil comments dataset. Perspective is very accurate and only modestly biased, but it still causes sign-flips (type I or type II errors) in a realistic study design.
Figure showing results from our real data example using the Perspective API and the Civil Comments Dataset. On the top, a model using Perspective's classifications results in a type-II error. It fails to detect a direct relationship toxicity and whether a comment discloses racial or ethnic identity, but when human annotations are used instead this is observed. On the bottom, using Perspective's classifications results in a type-I error. It falsely detects a negative direct relationship between the likes a comment received and its toxicity
120
Nate TeBlunthuis @groceryheist.cc · 14/07/2023
Automated classifiers are never perfect. They make errors and often manifest biases related to social categories. Such errors cause *misclassification bias*, threatening the validity of statistical findings! (can we fix it.jpg)
Bob the builder asks "Can we fix it?"
110