Sign in

Kevin O’Neill

@kevingoneill.github.io
1.6K followers 1.1K following 132 posts

Postdoc @ UCL studying causal judgment, counterfactual thinking, and metacognition kevingoneill.github.io

PostsRepliesMedia
Kevin O’Neill @kevingoneill.github.io · 10/08/2026
nevertheless, participants gave the same causal judgments of C regardless of whether it violated a statistical norm or not! they were also highly confident regardless of normality, indicating that this is not because they were confused about the task
a violin plot showing that (a) causal judgments and (b) confidence were the same regardless of whether the focal cause was normal or abnormala violin plot showing that (a) causal judgments and (b) confidence were the same regardless of whether the focal cause was normal or abnormal
100
Kevin O’Neill @kevingoneill.github.io · 10/08/2026
we then presented participants with a final case where C and A contributed equal amounts to E. after judging the extent to which C caused E, participants rated how surprised they were that C and A took the values they did, confirming that they saw C as violating a statistical norm
A graph showing ratings of surprisal. Ratings were higher for the focal cause when it was abnormal, but lower for the alternate cause and for the focal cause when it was normalA graph showing ratings of surprisal. Ratings were higher for the focal cause when it was abnormal, but lower for the alternate cause and for the focal cause when it was normal
100
Kevin O’Neill @kevingoneill.github.io · 10/08/2026
next, we had participants learn the structure (within a cover story) and the distributions of C and A over 40 trials. the variables were set so that C either had a lower variance (experiment 1) or mean (experiment 2) to A, and manipulation checks confirmed that participants learned this
A line graph showing ratings of perceived variability by block. Participants learned that the focal cause had lower variability when it was abnormalA line graph showing ratings of perceived variability by block. Participants learned that the focal cause had a lower mean when it was abnormal
100
Kevin O’Neill @kevingoneill.github.io · 10/08/2026
but all of this work focuses on judgments of *binary* causes: events that either happen or don't. what happens when candidate causes can contribute in degrees? to find out, we adapted a well-studied binary causal structure (b) so that (c) E happens if C+A > theta (where theta is some threshold)
(a) a causal structure with C -> E <- A. (b) a binary version of the structure where C occurs with p_C, A occurs with p_a, and E occurs if C and A occur. (c) a continuous version of the structure where C and A are normally distributed and E occurs if C + A > theta, where theta is some threshold.
110
Kevin O’Neill @kevingoneill.github.io · 01/05/2026
Next, we wanted to test whether this measure is reliable within participants over time. Thanks to a new longitudinal dataset collected by @tianqizhan.bsky.social, we found that different versions of our measure are highly reliable even after periods as long as one month!
Scatterplots showing test-retest correlations in metacognitive bias across different sessions, with strong correlations (in the ~.7 range)
110
Kevin O’Neill @kevingoneill.github.io · 01/05/2026
We then wanted to validate the measure in two ways. First, we found that under model misspecification, our measure is just as confounded as mean confidence. This is not ideal, but reflects similar findings that other measures of metacognitive performance are sensitive to modeling assumptions
A plot showing undesired relationships between measures of metacognitive bias and task performance under model misspecification
100
Kevin O’Neill @kevingoneill.github.io · 01/05/2026
Following a signal detection theoretic model of confidence ratings, we designed a new measure based on how much more evidence people need to make a response with high (vs low) confidence. It's easy to compute in an already standard workflow for metacognition research and has none of the confounds!
A figure showing two signal detection models. In A, a response is made by comparing the evidence distributed conditional on the stimulus to a criterion. Then, in B, a confidence level is determined by comparing the a second set of evidence to a series of criteria reflecting different levels of confidenceA plot showing no relationship between meta-delta and response bias, sensitivity, or metacognitive sensitivity
110
Kevin O’Neill @kevingoneill.github.io · 01/05/2026
But we have known for a while that this is a problem: mean confidence is confounded with response bias (c), sensitivity (d'), and metacognitive sensitivity (meta-d'), and we confirmed this in simulation (colors are two different responses)
A plot showing undesired changes in mean confidence with response bias (c), sensitivity (d'), and metacognitive sensitivity (meta-d')
100
Kevin O’Neill @kevingoneill.github.io · 16/04/2026
my cat and I both really enjoyed this fantastic interview on animal/AI sentience! in it, @birchlse.bsky.social grapples with so many difficult & pressing questions while making room for genuine progress to be made
A cat diligently watches a podcast on a phone
081
Kevin O’Neill @kevingoneill.github.io · 19/03/2026
one thing I'm particularly happy about is the ability to plug in arbitrary signal distributions- for example, we allow for metacognitive signal detection with the Gumbel-min distribution (cc @singmann.bsky.social) osf.io/preprints/ps...
120
Kevin O’Neill @kevingoneill.github.io · 19/03/2026
beyond a *ton* of efficiency upgrades 🚀, the package allows for arbitrary hierarchical structure, easy interfacing to other packages in the Stan ecosystem, simple computation of model-implied estimates (e.g., mean confidence, type 1/type 2 ROCs), and a bunch of other cool features
110
Kevin O’Neill @kevingoneill.github.io · 21/11/2024
- I am passionate about modeling, stats, and methods. I started a blog with all kinds of tutorials! (dibsmethodsmeetings.github.io) - can be spotted riding a purple bike around London and as promised, here are my cats who are more famous than me (momotheflyingcat on IG)
Two siamese cats peer out of a backpack into the woods.
130
Kevin O’Neill @kevingoneill.github.io · 06/08/2024
Finally, model comparisons revealed that only one existing model (the Necessity-Sufficiency model) could predict both causal judgments and confidence. These results were robust to different methods of modeling confidence from counterfactual models
110
Kevin O’Neill @kevingoneill.github.io · 06/08/2024
Next, relying on ideas from research on metacognition, we demonstrated that popular counterfactual models of causal judgment made a range of predictions about confidence, even if their predictions of causal judgments were similar
110
Kevin O’Neill @kevingoneill.github.io · 06/08/2024
Just like causal judgments, confidence in causal judgments exhibited small but clear effects of normality! In the situations where C and A are both necessary for E, people give higher causal judgments of C and are more confident when C is rare but A is common
110