Sign in

Tim Angelike

@timangelike.bsky.social
65 followers 278 following 2 posts

Postdoc, Psychology and Computer Science, @Philips-Universität Marburg (previously in Düsseldorf and Konstanz). Interested in large language models, human random number generation, statistical modelling, and computer simulations.

PostsRepliesMedia
Reposted by Tim Angelike
Daniel Heck @danielheck.bsky.social · 02/06/2026
Postdoc position @unimarburg.bsky.social in the project: "Bridging the Gap Between Verbal Psychological Theories & Formal Statistical Modeling with Large Language Models" (funded by @volkswagenstiftung.de) 📅Start: 01.10.2026 |⏳3 years 🔗 Job posting: uni-marburg.de/78NrWT Thanks for sharing!
logo of Psychological Methods Lab
13628
Tim Angelike @timangelike.bsky.social · 27/05/2026
In a new preprint with @danielheck.bsky.social, we adapted the ADEMP framework to provide a structured workflow for evaluating LLM feature extraction from verbal stimuli. Link: doi.org/10.31234/osf.... Figure 1 from the paper shows the proposed workflow:
Adaption of the ADEMP (Aims, Data-generating mechanism, Estimands and targets, Methods, Performance measures) framework for statistical simulation studies by Morris et al. (2019). The figure shows a left-to-right workflow diagram applying this framework to evaluate LLMs for feature extraction from verbal stimuli. 

Starting on the left, a "Domain" circle feeds into the "Data-Generating Mechanism" box, where "Items" (verbal stimuli) are processed through a "Prompt" template and an "LLM" to produce "Simulated Data." This output is visualized as a 3D cube with axes labeled Items, Prompt Variants, and LLMs. Above this, the "Methods" section lists the inputs: customizable prompt variants (showing item placeholders, prompt examples, and response formats) and multiple LLM models.

Below the mechanism, "Human Ratings" are collected for some or all items and serve as a "Gold Standard," connected by a horizontal arrow to the right side.

On the right, the workflow reaches the "Estimand" box, which defines target psychological constructs like (a) Valence, (b) Arousal, and (c) others. An arrow points down to "Performance Measures," split into two tiers: a "Weak criterion" for Consistency (using variance decomposition with η²p) and a "Strong criterion" for Inter-Rater Reliability (Pearson correlation) and Agreement (mean bias). Both the simulated LLM outputs and human gold standard ratings feed into this final evaluation stage, completing the structured workflow.
180
Reposted by Tim Angelike
Daniel Heck @danielheck.bsky.social · 16/10/2025
Welcome to @timangelike.bsky.social as a new member of our lab at @unimarburg.bsky.social! 🎉 Tim just started as a postdoc in a Momentum project funded by @volkswagenstiftung.de: "Bridging the Gap Between Verbal Psychological Theories and Formal Statistical Modeling with Large Language Models"
071