Sign in

Henrik Singmann

@singmann.bsky.social
1.2K followers 1.2K following 227 posts

Associate Professor at UCL Experimental Psychology; math psych & cognitive psychology; statistical and cognitive modelling in R; German migrant worker in UK

PostsRepliesMedia
Henrik Singmann @singmann.bsky.social · 17/08/2026
Thanks! As a preview of some ongoing work, we have since replicated this result in a simple visual working memory with item-memory only (i.e., no binding needed). See attached figures. This new experiment is essentially a follow up of He et al. (Exp. 1, 2026, JEP:LMC): osf.io/preprints/os...
020
Henrik Singmann @singmann.bsky.social · 16/06/2026
The Gumbel-min predicts that accuracy in the 2M-Min task is not affected by M. That is, having more new items to pick from does not make it easier to select a new item. This predictions is beautifully confirmed. Performance in the 2M-max task increases with M, an expected pattern from all SDT models
Figure 8: 2M Forced-Choice Task Results, consisting of two panels.
The figure shows the average accuracy results in the 2M-MIN task (left panel) and the 2M-MAX task (right panel) as a function of set size. The error bars correspond to 95% bootstrap based confidence intervals.
In the left panel (2M-min task) accuracy is constant with M. In the right panel (2M-max task) accuracy increases with M.
140
Henrik Singmann @singmann.bsky.social · 16/06/2026
The main evidence for the Gumbel-min model does not require any model fitting. Instead, the Gumbel-min model makes a unique prediction for the 2M-task, where participants always see M studied and M non-studied items (e.g., 1 studied and 1 non-studied; 2 studied and 2 non-studied, etc).
Figure B1:  Example of a 2M-MIN Test Trial
Study description: The experiment began with a study phase in which participants were presented with a list of 100 common nouns taken from Kellen et al. (2021), each presented for 2,000 ms, with a 400 ms interval between each word. After the study phase, participants initiated the test phase, which was comprised of 40 test trials, with 10 trials per M ∈ {1, 2, 3, 4}. Two hundred fifty-two and 253 participants engaged in the 2M-MIN and 2M-MAX tasks, respectively. During the test phase, words were displayed adjacent to one another, arranged in an M × 2 matrix. In the 2M-MIN task, participants were instructed to “click the item you are most certain is new,” whereas in the 2M-MAX task, they were told to “click the item you are most certain is old.” Participants selected the word of their choice by simply clicking on it, which led to the next test trial. Figure B1 illustrates a test trial.
110
Henrik Singmann @singmann.bsky.social · 16/06/2026
Finally out in Psychological Review (psycnet.apa.org/doi/10.1037/...), our update to Signal Detection Theory. We show that contrary to the prevailing Gaussian assumption, evidence distributions in recognition memory are likely minimum extreme Gumbel!
Figure 7: Illustration of the Gumbel-min Signal Detection Model
Figure consists of three panels.
Bottom-left panel: The Gumbel-min latent-strength distributions associated with SIGNAL (old) and NOISE (new) items. Top-left panel: The log likelihood ratio (log-LR) for the two latent-strength distributions. Right panel: The receiver operating characteristic function produced by the two latent-strength distributions. pH = hit probabilities; pFA = false-alarm probabilities.
211842
Henrik Singmann @singmann.bsky.social · 10/06/2026
Looks like my summer read has finally arrived. #rstats @mc-stan.org
Picture of the Bayesian Workflow book by Gelman and colleagues
4929
Henrik Singmann @singmann.bsky.social · 01/10/2025
See the same pattern for our Experiments 2 and 3 here. In Experiment 3, we added additional topics (e.g., Separating church from state causes more harm than good.) and more thoroughly controlled argument quality in three levels (good, internally inconsistent, and authority-based).
100
Henrik Singmann @singmann.bsky.social · 01/10/2025
The pattern in the average data also holds for each of the arguments (each line/colour per panel is one specific argument). People who think a claim (e.g., "abortion should be legal") is false find the corresponding argument is bad; people who think the claim is true think the argument is good.
Fig. 4 from the paper showing argument quality ratings as a function of belief consistency for each argument in Experiment 1. The overall pattern is shown for each argument shown.
Note. Results of Experiment 1 conditional on the topic and the level of argument support. Each line and colour in each panel shows responses to exactly one argument (i.e., there is no aggregation across items within a panel). The dots show individual responses and the curved lines show predictions from the linear mixed model. Blue dots represent argument quality ratings to good arguments in the data, orange dots represent argument quality ratings to bad arguments in the data, and the size of the dots represents the number of argument quality rating responses for the corresponding belief rating. Data points are dodged so that responses for good and bad arguments do not overlap. Model predictions are based on the fixed effects of the final model and the random effects of the by-topic grouping factor. Ext. = extremely.
100
Henrik Singmann @singmann.bsky.social · 02/09/2025
Exciting #rstats news for Bayesian model comparison: bridgesampling is finally ready to support cmdstanr, see screenshot. Help us by installing the development version of bridgesampling and letting us know if it works for your model(s): pak::pkg_install("quentingronau/bridgesampling#44")
R code and output showing the new functionality:
``` r
## pak::pkg_install("quentingronau/bridgesampling#44")
## see: https://cran.r-project.org/web/packages/bridgesampling/vignettes/bridgesampling_example_stan.html
library(bridgesampling)

### generate data ###
set.seed(12345)
mu <- 0
tau2 <- 0.5
sigma2 <- 1
n <- 20
theta <- rnorm(n, mu, sqrt(tau2))
y <- rnorm(n, theta, sqrt(sigma2))

### set prior parameters ###
mu0 <- 0
tau20 <- 1
alpha <- 1
beta <- 1

stancodeH0 <- 'data {
  int<lower=1> n; // number of observations
  vector[n] y; // observations
  real<lower=0> alpha;
  real<lower=0> beta;
  real<lower=0> sigma2;
}
parameters {
  real<lower=0> tau2; // group-level variance
  vector[n] theta; // participant effects
}
model {
  target += inv_gamma_lpdf(tau2 | alpha, beta);
  target += normal_lpdf(theta | 0, sqrt(tau2));
  target += normal_lpdf(y | theta, sqrt(sigma2));
}
'
tf <- withr::local_tempfile(fileext = ".stan")
writeLines(stancodeH0, tf)
mod <- cmdstanr::cmdstan_model(tf, quiet = TRUE, force_recompile = TRUE)

fitH0 <- mod$sample(
  data = list(y = y, n = n,
              alpha = alpha,
              beta = beta,
              sigma2 = sigma2),
  seed = 202,
  chains = 4,
  parallel_chains = 4,
  iter_warmup = 1000,
  iter_sampling = 50000,
  refresh = 0
)
#> Running MCMC with 4 parallel chains...
#> 
#> Chain 3 finished in 0.8 seconds.
#> Chain 2 finished in 0.8 seconds.
#> Chain 4 finished in 0.8 seconds.
#> Chain 1 finished in 1.1 seconds.
#> 
#> All 4 chains finished successfully.
#> Mean chain execution time: 0.9 seconds.
#> Total execution time: 1.2 seconds.
H0.bridge <- bridge_sampler(fitH0, silent = TRUE)
print(H0.bridge)
#> Bridge sampling estimate of the log marginal likelihood: -37.73301
#> Estimate obtained in 8 iteration(s) via method "normal".

#### Expected output:
## Bridge sampling estimate of the log marginal likelihood: -37.53183
## Estimate obtained in 5 iteration(s) via method "normal".
```
2289
Henrik Singmann @singmann.bsky.social · 28/04/2025
Yes & we discuss some shortcomings of d_a. As shown below, d_a does not permit an ordering of participants according to performance (d' and g' do). We also compare Type I error rates for g', d', and d_a for real H/FA-pairs where only response bias differs, only g' maintains 5% Type I errors (pp. 51)
130
Henrik Singmann @singmann.bsky.social · 27/04/2025
A particularly noteworthy example of a Gumbel-min prediction is shown here. The ROC predicted from g' (calculated from a single yes/no point) closely matches the ROC reconstruction derived independently from forced-choice judgments. The Gaussian model cannot even make a prediction in this case.
100
Henrik Singmann @singmann.bsky.social · 27/04/2025
We compared the descriptive performance of both models across 35 datasets from four different recognition memory paradigms. The Gumbel-min model fits the data nearly as well as the Gaussian model. Once model complexity was penalized via AIC, the Gumbel-min model matched or outperformed the Gaussian.
100
Henrik Singmann @singmann.bsky.social · 27/04/2025
The Gumbel-min model implies a behavioural principle: the probability of choosing a new item remains constant as choice sets grow. An experiment confirms this principle with constant accuracy for new item detection (2M-min). For old-item detection (2M-max), accuracy increase with choice set.
120
Henrik Singmann @singmann.bsky.social · 27/04/2025
We consider an SDT model assuming Gumbel-min (i.e., minimum extreme-value) distributions. The Gumbel-min model avoids the problrms of the Gaussian model, predicts asymmetric ROCs assuming equal variances, and allows calculating measures of discriminability and response bias, g′ and kappa.
120
Henrik Singmann @singmann.bsky.social · 27/04/2025
In recognition memory, ROCs are typically asymmetric, which requires Gaussian distributions with unequal variance. One problem with the unequal-variance model is that it predicts below chance performance for items with very low familiarity (i.e., studying makes some items less familiar).
110
Henrik Singmann @singmann.bsky.social · 27/04/2025
SDT is a cornerstone of recognition memory research, primarily assuming Gaussian distributions – a choice based more on tradition than necessity. The standard model assumes two equal-variance distributions, allows calculating d′ from a pair of hits and false alarms, and predicts symmetric ROCs.
210
Henrik Singmann @singmann.bsky.social · 20/03/2025
Results were in line with the qualitative predictions derived from sampling-based models. Predictions also held for the two types of illogical rankings we looked at. We do not know of any other (i.e. non-sampling) model that can make these qualitative predictions and predict illogical rankings.
000
Henrik Singmann @singmann.bsky.social · 19/03/2025
Simulation results show different qualitative pattern across event sets. Pr(logical ranking) is largest for mixed sets, followed by edge-event sets, followed by mid-event sets. This pattern held independently of sample size or whether there was additional read-out noise in the sampling process.
110
Henrik Singmann @singmann.bsky.social · 19/03/2025
We simulate the probability of obtaining logical and two types of illogical rankings for three different event sets: Edge events (P(A) & P(B) ≈ 1), mid-events (P(A) & P(B) ≈ .5), and mixed sets (P(A) ≈ 1 & P(B) ≈ .5).
100
Henrik Singmann @singmann.bsky.social · 19/03/2025
In each trial of the event ranking task, participants have to rank an event set consisting of four events, A, not-A, B, and not-B, in terms of their perceived likelihoods. The task contains an embedded logical that allows to classify the obtained ranking as logical or illogical.
100
Henrik Singmann @singmann.bsky.social · 04/02/2025
The follow up:
Email from Professor Brian Ripley to R devel that says: 
Sent in error (and not moderated)
020
Henrik Singmann @singmann.bsky.social · 04/02/2025
For context:
Email from Professor Brian Ripley to the R devel mailing list instead of another R core member in private. The part shown here reads: 
Tomas,

I am thinking of writing something for R-devel, and hope to have your
input first.

I get moderated on R-devel as I am now subscribed as
brian.ripley@R-project.org which of course I cannot send from. So I am
even more discouraged from posting there.  (R-core is bad enough with
Luke discouraging all innovation except by him and Simon completely
misunderstanding the C23 status.)

Thanks,

Brian
350
Henrik Singmann @singmann.bsky.social · 14/01/2025
If you want to see a bit more up to date explanation, Macmillman & Creelman (2005, ch. 3) also describe the process.
Table 3.2 from Macmillan and Creelman (2005), Detection Theory
010
Henrik Singmann @singmann.bsky.social · 13/12/2024
Bluesky is delivering some mixed messages here
010
Henrik Singmann @singmann.bsky.social · 12/12/2024
This term in my stats teaching, I regularly included images of Moo Deng into my slides. One of my students was clearly inspired by this combination and made this super cool drawing of Moo Deng doing stats herself. I love it so much. Stats is Moo Deng Approved!
0121
Henrik Singmann @singmann.bsky.social · 08/12/2024
Getting ready for my last week of teaching with a new stats meme
0112
Henrik Singmann @singmann.bsky.social · 12/11/2024
Be careful how many emails you send to the CRAN maintainers (and in which format), otherwise your package might get removed from CRAN. Found on the r-package-devel mailing list. #rstats
output from running R CMD check on a package removed from CRAN. The output says:
> 0 errors | 0 warnings | 2 note
> Package was archived on CRAN
> CRAN repository db overrides: X-CRAN-Comment: Archived on 2024-11-06 for repeated policy violation. Repeatedy spamming a team member's personal email address in HTML.
> checking compilation flags used ... NOTE Compilation used the following non-portable flag(s): &-Wp,-D_FORTIFY_SOURCE=3*
000
Henrik Singmann @singmann.bsky.social · 27/10/2024
Finally a humble LLM paper.
Abstract 
Establishing a unified theory of cognition has been a major goal of psychology [1, 2]. While there have been previous attempts to instantiate such theories by building computational models [1, 2], we currently do not have one model that captures the human mind in its entirety. Here we introduce Centaur, a compu­tational model that can predict and simulate human behavior in any experiment expressible in natural language. We derived Centaur by finetuning a state-of-the­art language model on a novel, large-scale data set called Psych-101. Psych-101 reaches an unprecedented scale, covering trial-by-trial data from over 60,000 par­ticipants performing over 10,000,000 choices in 160 experiments. Centaur not only captures the behavior of held-out participants better than existing cognitive models, but also generalizes to new cover stories, structural task modifications, and entirely new domains. Furthermore, we find that the model’s internal rep­resentations become more aligned with human neural activity after finetuning. Taken together, Centaur is the first real candidate for a unified model of human cognition. We anticipate that it will have a disruptive impact on the cognitive sciences, challenging the existing paradigm for developing computational models. 

Keywords: cognitive science, cognitive modeling, unified theory of cognition, large language models
080
Henrik Singmann @singmann.bsky.social · 27/10/2024
Getting ready for my stats teaching tomorrow and looks like my meme game is on point. I really hope stats meme never go out of fashion (and if so, please no one tell me).
Image of distracted boyfriend meme with a stats context. Text on boyfriend says: "Not approximately normally distributed residuals". Text on (ignored) girlfriend says "Non-Parametric Test". Text on distracting girls says "Still use ANOVA".Drake meme with yes-and-no image with stats context.
Text on the no image is: "Normality Test of Residuals and then Non-Parametric Test"
Text on yes-image is: "Just use ANOVA"
120
Henrik Singmann @singmann.bsky.social · 17/10/2024
In addition to finding strong evidence for the use of compensatory decision strategies. We found evidence for considerable individual differences. The figure shows both mean and individual-level thresholds between the numerical and categorical impacts.
100
Henrik Singmann @singmann.bsky.social · 17/10/2024
For both numerical and categorical judgements we found that weather scientists used compensatory decision strategies. An increase on any impact variable led to an increase in perceived severity, even when adjusting for the effect of the other impacts.
100
Henrik Singmann @singmann.bsky.social · 17/10/2024
We asked 278 weather scientists from four countries (Indonesia, Malaysia, Philippines, & Vietnam) to provide both categorical and numerical severity judgements for hypothetical weather events that were similar to real weather events. We used Judgement Analysis to analyse their decision strategies.
100
Henrik Singmann @singmann.bsky.social · 17/10/2024
Our research question was how weather scientists turn numerical impact information of extreme rainfall events, such as number of affected people, into categorical severity judgments. Nowadays, Impact-Based Warnings (IBWs) are commonly used which use categorical severity judgements.
100
Henrik Singmann @singmann.bsky.social · 17/10/2024
New applied JDM paper on decision-strategies of weather scientists in South-East Asia led by Xiaoxiao Niu, with meteorology colleagues from the UK and South-East Asia. Free download link for 50 days: authors.elsevier.com/c/1jxx-7t2zZ...
Screenshot of paper in "International Journal of Disaster Risk Reduction"
Title: Judgment and decision strategies used by weather scientists in southeast Asia to classify impact severity
Authors: Xiaoxiao Niu a b, Henrik Singmann b, Faye Wyatt c, Agie W. Putra d, Azlai Taat e, Jehan S. Panti f, Lam Hoang g, Lorenzo A. Moron f, Sazali Osman h, Riefda Novikarany d, Diep Quang Tran g, Rebecca Beckett c, Adam JL. Harris b
141
Henrik Singmann @singmann.bsky.social · 14/10/2024
This is a very cool way to visualise pre-post Likert data.
030
Henrik Singmann @singmann.bsky.social · 13/09/2024
Had a great time at SMLP2024 (vasishth.github.io/smlp2024/) where I not only met my stats hero Doug Bates (developer of lme4) but also learned to use the incredibly fast MixedModels.jl Julia package: github.com/JuliaStats/M... If lme4 in R is too slow for you, give Julia and MixedModels.jl a chance!
lme4 developer Douglas Bates and Henrik Singmann
060
Henrik Singmann @singmann.bsky.social · 05/09/2024
One of the highlights of the academic calendar is graduation day. Yesterday we awarded the UCL BSc psychology class of 2024 their diplomas.
030
Henrik Singmann @singmann.bsky.social · 23/07/2024
Summer is @mathpsych.org MathPsych time, this year in Tilburg. Our group from London learned a lot, had intense discussions, and of course also fun. This year EP UCL was supported by colleagues from econ and @birkbeckpsychology.bsky.social See you all next year!
040
Henrik Singmann @singmann.bsky.social · 08/12/2023
My last attempt of turning this ship around. Who wants digital currency when they can have nutrition?
000
Henrik Singmann @singmann.bsky.social · 08/12/2023
But it looks like I am getting nowhere.
100
Henrik Singmann @singmann.bsky.social · 08/12/2023
I am not giving up on adding some vegetables to her purchase order.
100
Henrik Singmann @singmann.bsky.social · 08/12/2023
Hmm, it does not seem as if Fiona is interested in cabbages but only apples.
100
Henrik Singmann @singmann.bsky.social · 08/12/2023
Luckily, this is exactly where I am heading on my way to get some Brussels sprouts. And because there is a deal, I am sure "Fiona" also wants some.
100
Henrik Singmann @singmann.bsky.social · 08/12/2023
As expected, she wants me to go to the supermarket...
110
Henrik Singmann @singmann.bsky.social · 08/12/2023
Exciting: My former deputy head of department in Warwick (where I still have a honorary appointment) suddenly reaches out to me and needs my help. I haven't talked to her in ages but we always got a long great so I am more than willing to help! Luckily I am about to head to the supermarket already.
Scam email impersonating my former head of department.
100
Henrik Singmann @singmann.bsky.social · 22/11/2023
Expressed the same thought in a recent presentation: Maybe methods for uncovering structures in covariance matrices (e.g., FA, SEM) are not appropriate tools for "carving nature at its joints".
Partial screenshot of a presentation saying:
The bigger picture
- Claims for general psychological constructs require strong evidence (see also work on risk preferences: Pedroni et al., 2017; Frey et al., 2017)
- Maybe methods for uncovering structures in covariance matrices (e.g., FA, SEM) are not appropriate tools for “carving nature at its joints”
031
Henrik Singmann @singmann.bsky.social · 21/11/2023
1100
Henrik Singmann @singmann.bsky.social · 22/10/2023
I have started using AI created illustrations to make my teaching slides more visually appealing. However, the persistent biases are disturbing and not helpful. The attached image shows Bing create's idea of an "illustration of statistical testing". Seems to be a pretty manly business.
012