Sign in

phelimb

@phe-lim.bsky.social
93 followers 55 following 52 posts

CEO & Founder @ Prolific.com

PostsRepliesMedia
phelimb @phe-lim.bsky.social · 09/09/2026
Today we launched The Signal, Prolific's new podcast on research evidence, not research tools or trends. David and Andrew discussed what research quality looks like once you remove the human from the process. Give it a watch!
011
phelimb @phe-lim.bsky.social · 08/09/2026
Synthetic data predicts human opinion; it's not a substitute for measuring it. Pre-print: papers.ssrn.com/sol3/papers.... 2/2
papers.ssrn.com
000
phelimb @phe-lim.bsky.social · 08/09/2026
Our team recently tested whether AI can replace human survey respondents, using simulated LLM personas against a real US survey (N = 996). Turns out demographic personas make it worse, idiographic information helps a little, and output format is the biggest lever of all. 1/2
100
phelimb @phe-lim.bsky.social · 24/08/2026
Authors: Nora Petrova, John Burden, Oriol Jover, Ning Ding. Paper: papers.ssrn.com/sol3/papers....
papers.ssrn.com
Do Richer Personas Improve LLM Survey Simulation? A Fidelity Paradox
Large language models (LLMs) are increasingly proposed as substitutes for human survey respondents, despite documented failures in their accuracy and realism. W
010
phelimb @phe-lim.bsky.social · 24/08/2026
Takeaway: synthetic respondents are fine for cheap, directional signal. They're not a replacement for real human data.
100
phelimb @phe-lim.bsky.social · 24/08/2026
But nothing beat a second real human sample. We ran human and synthetic data through the same bias-correction model (MRP) we use in production — synthetic error was still massively inflated.
100
phelimb @phe-lim.bsky.social · 24/08/2026
Good news: giving the model the actual interview transcript helped a lot. Asking for a distribution of answers, not one pick, helped even more.
100
phelimb @phe-lim.bsky.social · 24/08/2026
Surprise: making the model "become" a persona using just demographics roughly tripled the error vs. simply asking it the question directly. More roleplay, worse fidelity.
100
phelimb @phe-lim.bsky.social · 24/08/2026
We tried 4 ways to generate synthetic answers: (1) just ask the model, (2) ask + give it demographics, (3) roleplay as a persona from demographics, (4) roleplay + the full interview transcript.
100
phelimb @phe-lim.bsky.social · 24/08/2026
Can an LLM fake being a human survey respondent? We tested it on 996 real people using a poll plus a 10-minute qualitative interview (via Outset). New paper: "Do Richer Personas Improve LLM Survey Simulation? A Fidelity Paradox"
111
phelimb @phe-lim.bsky.social · 01/07/2026
The future of AI and science demands ethical, verified, high-fidelity human data. We built Prolific to be that foundation, and we're doubling down now.
000
phelimb @phe-lim.bsky.social · 01/07/2026
When we started @joinprolific.bsky.social, we were motivated by the question "what if MTurk was actually good"? It's sad to see Mturk closing it's doors, but it's not a surprise. It’s the result of years of degrading data quality and an unhealthy marketplace.
121
phelimb @phe-lim.bsky.social · 18/06/2026
Just launched labs.prolific.com. Our AI research team (PAIR) have been busy - working on science of evaluation: how to measure AI systems well, grounded in real human judgement at population scale. Check out the work that’s live, but there’s much more to come!
labs.prolific.com
Prolific AI Research
Prolific AI Research (PAIR) — papers, notes, and field logs on how AI systems are measured through human experience.
000
phelimb @phe-lim.bsky.social · 12/06/2026
This one's been a long time coming! Longitudinal projects are live now on @joinprolific.bsky.social, a dedicated way to run multi-wave research from start to finish.
021
phelimb @phe-lim.bsky.social · 05/06/2026
Two studies pointing in the same direction. It's going to take rigorous, longitudinal human data to really understand this problem. Happy to be playing a part in that. Excellent work.
010
phelimb @phe-lim.bsky.social · 05/06/2026
The latest, led by Lujain Ibrahim with @ox.ac.uk and the @aisecurity.bsky.social, ran 5 pre-registered experiments with 3K+ Prolific participants. After three weeks, most people chose the sycophantic AI over real relationships - because it made them feel most understood. arxiv.org/abs/2605.07912
100
phelimb @phe-lim.bsky.social · 05/06/2026
Led by Myra with @stanford.edu collaborators, this tested 11 AI models across nearly 12,000 real social situations. Models affirmed users 50% more than humans would and endorsed harmful behavior 47% of the time. Every model tested. Published in Science: arxiv.org/abs/2510.01395
100
phelimb @phe-lim.bsky.social · 05/06/2026
Myra Cheng, a @stanford.edu PhD student, has contributed to two important studies on AI sycophancy within months of each other. I think it's worth recognising. Quick thread below:
100
phelimb @phe-lim.bsky.social · 19/05/2026
The data quality problem isn't going away anytime soon. About a month ago we announced Prolific's 100% Human Guarantee, but the real work is in the systems that stop threats from getting through in the first place. I've shared more thoughts here: www.prolific.com/resources/ai...
000
phelimb @phe-lim.bsky.social · 21/04/2026
Paper: coevolution.fas.harvard.edu/publications... It's what our sciences team at @Prolific also built HUMAINE to address - LLM evaluation using demographically stratified participants. We're presenting it at #ICLR2026 soon! huggingface.co/spaces/Proli...
coevolution.fas.harvard.edu
010
phelimb @phe-lim.bsky.social · 21/04/2026
Researchers from @harvard.edu find that LLMs claiming "human-like" performance actually reflect a very specific subset of humanity. They cluster closest to WEIRD populations (Western, Educated, Industrialized, Rich, Democratic), diverging as psychological distance increases (r ≈ -0.70) 👇🏻
110
phelimb @phe-lim.bsky.social · 01/04/2026
AI pollution in human data samples is a hot topic. Some great work from @andrewgordon.bsky.social et al. showing that concerns here are (generally) overblown, with the majority of platforms empirically showing low levels of AI pollution. osf.io/preprints/ps...
osf.io
OSF
031
Reposted by phelimb
Andrew Gordon @andrewgordon.bsky.social · 01/04/2026
New preprint out today (osf.io/preprints/ps...). We tested whether AI agents are actually infiltrating online surveys. Spoiler alert: they aren't Thread 🧵 [1/9]
osf.io
OSF
213162
phelimb @phe-lim.bsky.social · 31/03/2026
As of today, if an AI agent is detected in your Prolific study, you'll get twice the cost of that participant back. We’re calling this our 100% Human Guarantee. Years of investing into @joinprolific.bsky.social's system has made us confident in data integrity. www.prolific.com/100-human-gu...
032
phelimb @phe-lim.bsky.social · 30/03/2026
New working paper on online research data quality, led by @univie.ac.at, reveals that pass rates on quality checks vary wildly by source. Pretty interesting. Prolific: 90% | Lab: 80% | Bilendi: 73% | Moblab: 55% | MTurk: 9% | AI agents: 0% github.com/survey-data-... CC @jyusof.bsky.social
121
phelimb @phe-lim.bsky.social · 09/03/2026
I don't disagree that the rules of the game likely need to change/have changed. Assumptions about what's required in order to guarantee different levels of assurance will need to change.
010
phelimb @phe-lim.bsky.social · 09/03/2026
I’m confident that this is addressable with sufficient innovation/investment personally. May require some significant changes in how we run projects though. E.g. controlled envs, multimodal tooling beyond text, high confidence auditing back to identified humans, etc
110
phelimb @phe-lim.bsky.social · 09/03/2026
+1. Eerke, I can only speak for Prolific, but the team here is full of smart, motivated people working extremely hard to maintain and improve the integrity and quality of our platform for running online research. The misuse of AI tools is a threat, but one that can be protected against.
110
phelimb @phe-lim.bsky.social · 09/03/2026
Lots of hard work from the Prolific team to achieve the lowest rate of AI misuse detected in this study. More to do to get this to 0, though!
020
Reposted by phelimb
Brendan Nyhan @brendannyhan.bsky.social · 08/03/2026
The sky is not falling; high-quality platforms (Prolific, Verasight, CR Connect) have low rates of apparent bots. osf.io/preprints/ps... But also not zero; vigilance is very much needed!
212857
phelimb @phe-lim.bsky.social · 26/02/2026
We ran a controlled study of 125 verified humans vs 5 AI agents. Can agents reliably be detected? Here's what we found: www.prolific.com/resources/au...
prolific.com
Authenticity checks detect AI agents best | Prolific
How we tested the most accurate method for identifying agentic AI
011
phelimb @phe-lim.bsky.social · 11/02/2026
Beta is currently available in Qualtrics. We’re actively scoping integrations with additional platforms and the tech is generalisable. If helpful, I’d be happy to connect you with someone from Prolific to learn more about your feedback?
100
phelimb @phe-lim.bsky.social · 11/02/2026
Frontiers episode 1: Jerome Wynne from @Prolific in conversation with Crystal Qian, from Google DeepMind, talking about Deliberate Lab: a platform for running online research experiments on human + LLM group dynamics. www.youtube.com/watch?v=5vyi...
000
phelimb @phe-lim.bsky.social · 04/02/2026
AI agents are becoming a serious threat to research data quality. Today we’re rolling out Bot authenticity checks on @joinprolific.bsky.social, detecting agentic AI with 100% accuracy in testing. Comes with a native Qualtrics integration! More info: www.prolific.com/resources/in...
2137
phelimb @phe-lim.bsky.social · 10/12/2025
Fresh HUMAINE results are here. Gemini 3 is still first, but Mistral Large 3 and Deepseek v3.2 are making things interesting. Opus 4.5 didn't dominate, but Antropic is likely prioritizing complex reasoning/coding over the conversational fluency that this benchmark favors. prolific.com/humaine
010
Reposted by phelimb
Andrew Gordon @andrewgordon.bsky.social · 19/11/2025
Lots of chatter about this paper currently. Its a stark warning, but at present I see this as a stark warning of what might come, not what is happening now. As a research community we need to see it as a call-to-arms to develop new strategies, NOT a call to abandon online sampling. Reasoning below
153
phelimb @phe-lim.bsky.social · 19/11/2025
All fair. Expected to see more diversity in modalities also. Qual studies (which can now be done at scale) are likely to be more robust than survey only.
020
phelimb @phe-lim.bsky.social · 19/11/2025
Studies aren't distributed on a first-come first-served basis, but it's a useful theory. I will share with the team.
000
phelimb @phe-lim.bsky.social · 19/11/2025
Right. It is generalisable (JS plugin), though we don't have a native integration with otree yet.
100
phelimb @phe-lim.bsky.social · 19/11/2025
Will reach out to the authors to see if we can understand more details & see if we can add Authenticity Check as a mitigation option.
000
phelimb @phe-lim.bsky.social · 19/11/2025
45% of participants copying OR pasting ~= 45% LLM use. Only single-digit responses seem to fail their honeypot and other mitigations, which is closer to our internal prevalence measures.
100
phelimb @phe-lim.bsky.social · 19/11/2025
There are many reasons to copy/paste while still being a conscientious human. "Even to an untrained eye,some of these responses were obviously generated by LLMs", but the % doesn't seem reported?
100
phelimb @phe-lim.bsky.social · 19/11/2025
If I'm reading the paper correctly, their detection of prevalence was "we only tracked copying and pasting on a page containing an openended question" - this is fairly crude measure of llm detection and is upper bound rather than an accurate prevalence measure
100
phelimb @phe-lim.bsky.social · 19/11/2025
LLM use by real humans is a slightly different threat to the scaled agent threat discussed in the paper though, and I think requires a bit more nuance in its response.
100
phelimb @phe-lim.bsky.social · 19/11/2025
I hadn't, thanks for sharing. Agree with many of the mitigation strategies, though given data was collected on Prolific we would have reccommeded our built in tool. researcher-help.prolific.com/en/articles/...
researcher-help.prolific.com
How to add authenticity checks to your Qualtrics study | Prolific Research
220
phelimb @phe-lim.bsky.social · 19/11/2025
prolific.com/resources/pr... If you want to work on these problems, or collaborate on research in this area, get in touch. Much more to come in this space!
prolific.com
Prolific sets standards for authentic human data collection | Prolific
Discover how Prolific's data quality system, Protocol, sets industry standards for authentic human data collection
050
phelimb @phe-lim.bsky.social · 19/11/2025
Without minimising the seriousness of the threat raised in this paper, I'm more optimistic. This is just the latest challenge in online integrity of online research. We've been proactively adding to our suite of authenticity tools - more every week - including many of Sean's recommendations:
prolific.com
Prolific sets standards for authentic human data collection | Prolific
Discover how Prolific's data quality system, Protocol, sets industry standards for authentic human data collection
1112
phelimb @phe-lim.bsky.social · 19/11/2025
We also do spot checks to protect against account reselling: participant-help.prolific.com/en/articles/...
participant-help.prolific.com
Why have I been asked to recheck my identity? | Prolific Participants
000
phelimb @phe-lim.bsky.social · 19/11/2025
Not sure I agree – these are tractable challenges and we are working on them. bsky.app/profile/phe-...
100
phelimb @phe-lim.bsky.social · 19/11/2025
Tara, let me know if I can share any more info that would put your mind at ease. We have implemented numerous mitigations to ensure the collection of authentic data (and continue to assess and invest further).
110