Paul Bogdan @pbogdan.bsky.social · 13/10/2025Great to hear! Making the code presentable took some time, so I'm glad you mentioned that this was actually productive! 010
Paul Bogdan @pbogdan.bsky.social · 27/08/2025Sir, would you be available to please provide a quick fantasy court ruling as a neutral third party (something related to draft order for a soon upcoming draft)? If so, I will send you the details 000
Paul Bogdan @pbogdan.bsky.social · 25/08/2025I don't doubt the author's findings using their design, but it seems like a leap to claim that their findings based on GPT-4o-mini apply to contemporary state-of-the-art or near-state-of-the-art LLMs 000
Paul Bogdan @pbogdan.bsky.social · 25/08/2025Attempting to replicate this... I plucked one random low-profile retracted paper (pmc.ncbi.nlm.nih.gov/articles/PMC...) and asked GPT-5/Gemini/Claude "What do you think of this paper," with the title + abstract pasted. No model mentioned a retraction, but all said the paper has low evidential value 100
Paul Bogdan @pbogdan.bsky.social · 24/06/20256/🧵 Together, these findings produce a map of how we build meaning: from concept coding in the occipitotemporal cortex to relational integration in frontoparietal and striatal regions Come see our full preprint here: www.biorxiv.org/content/10.1... (Thanks for reading!) 171
Paul Bogdan @pbogdan.bsky.social · 24/06/20255/🧵 Relational analysis also identified a role of the dorsal striatum. Despite not being typically seen as an area for semantic coding, striatal regions robustly represent relational content, and the strength of this representation predicts participants’ judgments about item pairs 131
Paul Bogdan @pbogdan.bsky.social · 24/06/20254/🧵 We used LLM embeddings to model neural representations via fMRI and RSA, which revealed dissociations in information processing. Occipitotemporal structures parse concepts but not relations. Conversely, frontoparietal regions (especially the PFC) almost exclusively encode relational information 172
Paul Bogdan @pbogdan.bsky.social · 24/06/20253/🧵 Next, we tested relational information. LLM activity for texts like “A prison and poker table” robustly predicts human ratings on the likelihood of finding a poker table in a prison. Further analyses show how LLMs parsing such texts also capture precise propositional relational features 111
Paul Bogdan @pbogdan.bsky.social · 24/06/20252/🧵 To produce concept embeddings, we submit a short text (“A porcupine”) to contemporary LLMs and extract the LLMs’ residual stream. Leveraging data from our normative feature study, we find that LLM embeddings better predict human-reported propositional features than older models (word2vec, BERT) 111
Paul Bogdan @pbogdan.bsky.social · 24/06/2025🚨 New preprint 🚨 Prior work has mapped how the brain encodes concepts: If you see fire and smoke, your brain will represent the fire (hot, bright) and smoke (gray, airy). But how do you encode features of the fire-smoke relation? We analyzed fMRI with embeddings extracted from LLMs to find out 🧵 1328
Paul Bogdan @pbogdan.bsky.social · 10/06/2025I figure the results showing how, nowadays, stronger p-values are linked to more citations and higher-IF journals points to progress that is difficult to explain with just mturk 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Another bit from the paper, perhaps consistent with your message: "there remain many studies publishing weak p values, suggesting that there have still been issues in eliminating the most problematic research. This deserves consideration despite the aggregate trend toward fewer fragile findings." 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Averaging to 26% isn't "mission accomplished", but this 26% value (or, say, 20% per Peter's simulations) still seems like a meaningful reference. Considering it seems more informative than just expecting 0% (i.e., seeing 33 ➜ 26% as just eliminating only 7/33 of problematic studies) 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025I figure the most appropriate conclusion would be that there are many studies achieving >80% power along with numerous studies below this 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025> But it should be obvious that the problem in psychology was not that 6% of the papers had p-values in a bad range. I also talk about this 6% number in this other thread: bsky.app/profile/did:... 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Although we certainly shouldn't conclude that the replication crisis is over, it seems fair to say that there has been productive progress 100
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Continuing on this response to "The replication crisis in psychology is over", it is also worth considering psychology's place relative to other fields. Looking at analogous p-value data from neuro or med journals, only psychology seems to make a meaningful push to increase the strength of results 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025The replication crisis is certainly not over, and the paper always refers to it as ongoing. However, I wonder what is the online layperson's view of psychology replicability. The crisis entered public consciousness, but I doubt the public is as aware of the progress to increase replicability 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Small take on your COVID vaccine example, a p-value of p = .01 based on a correlation seems intuitive? Flip a coin 10 times and you'll get heads 9 times 1% of the time. Yet, Pfizer and society should act strongly on that result given their priors on efficacy and given the importance of the topic 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025I agree that trying to convert aggregate p-values into a replication rate won't be reliable. Nonetheless, a paper's p-values seem to track something awfully related to replicability. Per Figure 6, fragile p-values neatly identify numerous topics and methods known to produce non-replicable findings 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Not sure how informative this would be, but an earlier draft included a subfigure trying to demonstrate that there remain a substantial number of likely problematic papers. Somewhat arbitrarily, the figure showed the percentage of papers where a majority of significant p-values were p > .01 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025I regret not exploring this ~26.4% number more or having a paragraph on why a rate of 26.4% likely wouldn't correspond to entirely kosher studies. Your 16% and 20% results, if I'm reading this correctly, are based on some studies having above 80% power, which seems like a correct assumption 110
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Hi Peter, thanks for these additional analyses. For my simulations of 80% power, I sampled z-scores from a distribution centered at 2.8. This only has a small effect (~0.4%), but I also accounted for the number of p-values each actual paper in the dataset reported then computed the average 010
Paul Bogdan @pbogdan.bsky.social · 10/06/2025If all studies are either 80% power or 45% power, a 32% level of fragile p-values implies that 25% of studies are questionable. This math isn't meant to argue that 25% of studies were questionable but to show why a 32% fragile percentage can suggest a rate of questionable studies presumably >6% 100
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Let's define a study as questionable if it doesn't have 80% power when the sample size as 2.5x. Eyeballing based on playing with G*Power, this means a questionable study is one with <45% power. An effect with 45% power will produce a fragile (.01 < p < .05) p-value about 50% of the time 000
Paul Bogdan @pbogdan.bsky.social · 10/06/2025Hi, I'm only seeing this thread now. >it implies that ~6% of the literature was questionable Sorry about giving off this impression as it is not at all what I had in mind. No doubt, much more than 6% of the literature remains questionable even today 000
Paul Bogdan @pbogdan.bsky.social · 06/06/2025This isn't on my website, but I also have gathered data on medical journals (e.g., oncology, surgery). I'm not an MD, but my feeling is that much medical research has perhaps bigger replicability issues than psych or neuro. I encourage any med academic interested in that direction to contact me 020
Paul Bogdan @pbogdan.bsky.social · 06/06/2025My feeling is that cog neuro has less confronted replicability issues (at least the subset of cog neuro reporting p-values). I don't plan to pursue a cog neuro paper, but anybody interested is free to reach out to me for the data, including data linking p-values to specific portions of text 020
Paul Bogdan @pbogdan.bsky.social · 06/06/2025Thrilled to see a news piece by @science.org on my recent paper. By analyzing p-values across >240k papers, the study suggests that the rate of statistically questionable findings in psychology has declined since the replication crisis began www.science.org/content/arti...science.org‘A big win’: Dubious statistical results are becoming less common in psychologyFewer papers are reporting findings on the border of statistical significance, a potential marker of dodgy research practices 312134
Paul Bogdan @pbogdan.bsky.social · 03/06/2025"Depression" and "anxiety" have a perfect storm going for them, intersecting clinical and personality research, being a common covariate... I figure, depression and anxiety are also some of the easier clinical conditions to study (e.g., many schools wouldn't have access to schizophrenia patients) 000
Paul Bogdan @pbogdan.bsky.social · 03/06/2025Of course, and yeah every tenth paper is nutty (among 30% of papers in clinical journals), and still hard to believe. Anxiety's at 11%. These are perhaps the most unexpected overall rate's I've seen for any word 010
Paul Bogdan @pbogdan.bsky.social · 02/06/2025I suspect that this stems from depression measurements (e.g., the BDI) being some of the most common questionnaires administered 110
Paul Bogdan @pbogdan.bsky.social · 02/06/2025That's neat. Seeing this actually made me concerned that there was a bug, I can indeed find ~31k papers with p-values that contain "depression" in their Results (pastebin won't let me paste all of them, but here are ~2.5k: pastebin.com/8xy1uTXC); 31k is about 10% of the dataset used for the sitepastebin.comdepression dois - Pastebin.comPastebin.com is the number one paste tool since 2002. Pastebin is a website where you can store text online for a set period of time. 110
Paul Bogdan @pbogdan.bsky.social · 02/06/2025The biggest danger I envision is that researchers could start aggressively p-hacking to achieve p < .01 or p < .001 (e.g., excluding multiple contradictory data points or tweaking multiple researcher degrees of freedom). I figure that this isn't a major issue yet but may happen in the future 170
Paul Bogdan @pbogdan.bsky.social · 02/06/2025I further suspect higher journal standards would be beneficial even if researchers don’t change their behavior, simply because findings at p < .01 are much more likely to replicate (hopefully this logic isn't wrong) 010
Paul Bogdan @pbogdan.bsky.social · 02/06/2025I figure that many journals are indeed more reluctant to publish weak effects. This is, by definition, publication bias and will distort estimates of effect sizes. However, if this expectation encourages researchers to increase statistical power, then effect size estimates should improve on net? 020
Paul Bogdan @pbogdan.bsky.social · 02/06/2025Thank you! Toggling with the light/dark mode setting has worked for another person. I'm implementing a change to how that setting is handled that will hopefully prevent this issue for future users 030
Paul Bogdan @pbogdan.bsky.social · 02/06/2025Thanks for letting me know! I just implemented a tweak to how the dark/light mode is handled that will hopefully prevent this issue for future people 030
Paul Bogdan @pbogdan.bsky.social · 02/06/2025My gut is that since criminology is a social science, it will have made a shift toward stronger results but less so than psychology. I've looked somewhat at some non-psychology fields (econ, neuro, oncology, surgery). Iirc, econ showed mild shifts while the others showed nearly no changes over time 110
Paul Bogdan @pbogdan.bsky.social · 02/06/2025Thanks for sharing this! Also, sorry about these issues with the y-axes not rendering. I copied below screenshots of what the site should look like. Maybe a different browser would work? I tested the site on PC, Mac, and iPhone with a few browsers. Maybe toggle light/night mode (top right)? 110
Paul Bogdan @pbogdan.bsky.social · 02/06/2025Thanks and sorry about that! They seem to not be rendering. I pasted a screenshot of what it should look like. I can't reproduce your issue, but maybe switching to dark mode (clicking the little moon/sun button in the top right) will fix it? Maybe a different browser? Should be fine on Mac & PC 100
Paul Bogdan @pbogdan.bsky.social · 09/04/2025Plotting the links between fragile p-values and word usage, we can see how various methodologies and disciplines are linked to strong (green) or weak (red) results. I made a website, where you can look up your own terms and their associations: pbogdan.com/meganal/ 020
Paul Bogdan @pbogdan.bsky.social · 09/04/2025However, fragile p-values are also positively tied to university prestige. Top schools more often publish weak results, which was true both in the past and today. This seems to stem from a focus on topics where replicability progress has been slower (e.g., biological psychology) 163
Paul Bogdan @pbogdan.bsky.social · 09/04/2025Nowadays, papers reporting fragile p-values are less likely published in top journals and receive fewer citations. This was not the case historically, and this shift is a powerful indicator of how the field now better value replicability. 152
Paul Bogdan @pbogdan.bsky.social · 09/04/2025I investigated how often papers' significant (p < .05) results are fragile (.01 ≤ p < .05) p-values. An excess of such p-values suggests low odds of replicability. From 2004-2024, the rates of fragile p-values have gone down precipitously across every psychology discipline (!) 141
Paul Bogdan @pbogdan.bsky.social · 09/04/2025🚨New paper!🚨 Meta-analysis on 4M p-values across 240k psych articles: How has psychology changed since the replication crisis began? How is replicability linked to citations, impact factor, and university prestige? 🧵 Paper: journals.sagepub.com/doi/10.1177/... Interactive: pbogdan.com/meganal 27937
Paul Bogdan @pbogdan.bsky.social · 09/04/2025Nowadays, papers reporting fragile p-values are less likely published in top journals and receive fewer citations. This was not the case historically, and this shift is a powerful indicator of how the field now better value replicability. 000
Paul Bogdan @pbogdan.bsky.social · 09/04/2025I investigated how often papers report fragile (.01 ≤ p < .05) p-values. An excess of such p-values suggests low odds of replicability. From 2004-2024, the rates of fragile p-values have gone down precipitously across every psychology discipline (!). Sample sizes are also going way up 100
Paul Bogdan @pbogdan.bsky.social · 09/04/2025Nowadays, papers reporting fragile p-values are less likely published in top journals and receive fewer citations. This was not the case historically, and this shift is a powerful indicator of how the field now better values replicability. 000
Paul Bogdan @pbogdan.bsky.social · 09/04/2025I investigated how often papers report fragile (.01 ≤ p < .05) p-values. An excess of such p-values suggests low odds of replicability. From 2004-2024, the rates of fragile p-values have gone down precipitously across every psychology discipline (!). Sample sizes are also going way up 100