Sign in

Daniel Paleka

@dpaleka.bsky.social
306 followers 171 following 66 posts

ai safety researcher | phd ETH Zurich | danielpaleka.com

PostsRepliesMedia
Daniel Paleka @dpaleka.bsky.social · 09/04/2026
What is the strongest evidence for the "elicitation gap" reducing over time, e.g. thoughtful prompting helping less and less?
171
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
With @simonlermen.bsky.social @floriantramer.bsky.social @aemai.bsky.social :D
040
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Privacy online is fundamentally at odds with intelligence getting cheaper. Anonymity on the internet has always relied on practical obscurity. We publish in hopes that people can adapt to LLMs changing this. Paper: arxiv.org/abs/2602.16800
arxiv.org
Large-scale online deanonymization with LLMs
We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer participants at...
1242
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
If you're anonymous, what should you do? Avoid sharing specific details, and adopt a security mindset: if a team of smart investigators were trying to identify you from your posts, could they plausibly figure out who you are? If yes, LLM agents will soon be able to do the same.
1122
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Short term, AI labs and platforms should try to mitigate large-scale misuse. This is challenging because deanonymization resembles benign usage in many ways. Long term, if intelligence is too cheap to meter, assume anything you post online can eventually be linked back to you.
2111
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Direct deanonymization. Anthropic Interviewer is a dataset of anonymized interviews with scientists about their use of AI. Following prior work, a simple agent finds ~7% of the interviewed scientists, out of the box, just by searching the web and reasoning over the transcript.
190
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Scaling: as candidate pools grow to tens of thousands, LLM-based attacks degrade gracefully at high precision; this implies that with sufficient compute, these methods would already scale to entire platforms. With future models, expect the cost to only go down.
160
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Proxy 2: Matching split accounts. On Reddit, we split user histories into "before" and "after", and test LLMs linking them back together. LLM embeddings + reasoning significantly outperform Netflix-Prize-style baselines that match based on subreddits and metadata. @random_walker
170
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Proxy 1: Cross-platform. We take non-anonymous Hacker News accounts that link to their LinkedIn. We then anonymize the HN accounts, removing all directly identifying information. Then, we let LLMs match the anonymized account to the true person; this works with high precision.
1110
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Solution: we construct deanonymization proxies — tasks similar to true online deanonymization, that nevertheless give evidence that LLMs are indeed getting scarily better at deanonymization.
1130
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
It is tricky to benchmark LLMs on deanonymization. We don't want to actually deanonymize anonymous individuals! And there is no ground truth for online deanonymization. How could we verify that the AI found the correct person?
1170
Daniel Paleka @dpaleka.bsky.social · 20/02/2026
Can LLMs figure out who you are from your anonymous posts? From a handful of comments, LLMs can infer where you live, what you do, and your interests; then search for you on the web. New 📄 w/ @SimonLermenAI, @joshua_swans, @AerniMichael, Nicholas Carlini, @florian_tramer 🧵
812442
Daniel Paleka @dpaleka.bsky.social · 27/01/2026
how did they build claude code without claude code?
030
Daniel Paleka @dpaleka.bsky.social · 05/06/2025
We don't claim LLM forecasting is impossible, but argue for more careful evaluation methods to confidently measure these capabilities. Details, examples, and more issues in the paper! (7/7) arxiv.org/abs/2506.00723
arxiv.org
Pitfalls in Evaluating Language Model Forecasters
Large language models (LLMs) have recently been applied to forecasting tasks, with some works claiming these systems match or exceed human performance. In this paper, we argue that, as a...
000
Daniel Paleka @dpaleka.bsky.social · 05/06/2025
Benchmarks can reward strategic gambling over calibrated forecasting when optimizing for ranking performance. "Bet everything" on one scenario beats careful probability estimation for maximizing the chance of ranking #1 on the leaderboard. (6/7)
100
Daniel Paleka @dpaleka.bsky.social · 05/06/2025
Model knowledge cutoffs are guidelines about reliability, not guarantees of no information thereafter. GPT-4o, when nudged, can reveal knowledge beyond its stated Oct 2023 cutoff. (5/7)
100
Daniel Paleka @dpaleka.bsky.social · 05/06/2025
Date-restricted search leaks future knowledge. Searching pre-2019 articles about “Wuhan” returns results abnormally biased towards the Wuhan Institute of Virology — an association that only emerged later. (4/7)
100
Daniel Paleka @dpaleka.bsky.social · 05/06/2025
The time traveler problem: When forecasting "Will civil war break out in Sudan by 2030?", you can deduce the answer is "yes" - otherwise they couldn't grade you yet. We find that backtesting in existing papers often has similar logical issues that leak information about answers. (3/7)
100
Daniel Paleka @dpaleka.bsky.social · 05/06/2025
Forecasting evaluation is tricky. The gold standard is asking about future events; but that takes months/years. Instead, researchers use "backtesting": questions where we can evaluate predictions now, but the model has no information about the outcome ... or so we think (2/7)
100
Daniel Paleka @dpaleka.bsky.social · 05/06/2025
How well can LLMs predict future events? Recent studies suggest LLMs approach human performance. But evaluating forecasters presents unique challenges compared to standard LLM evaluations. We identify key issues with forecasting evaluations 🧵 (1/7)
100
Daniel Paleka @dpaleka.bsky.social · 26/05/2025
why is it that whenever i see survivorship bias on my timeline it already has the red-dotted plane in the replies?
010
Daniel Paleka @dpaleka.bsky.social · 17/05/2025
OpenAI and DeepMind should have entries at Eurovision too
020
Daniel Paleka @dpaleka.bsky.social · 30/04/2025
3.7 sonnet: *hands behind back* yes the tests do pass. why do you ask. what did you hear 4o: yes you are Jesus Christ's brother. now go. Nanjing awaits o3: Listen, sorry, I owe you a straight explanation. This was once revealed to me in a dream
000
Daniel Paleka @dpaleka.bsky.social · 30/04/2025
Of course, we don't have the old chatgpt-4o API endpoint, so we can't see whether the prompt is fully at fault or there was also a model update.
000
Daniel Paleka @dpaleka.bsky.social · 30/04/2025
The sycophancy effect on controversial binary options is much smaller than what you would assume from the overall positive vibe towards the user. On most such statements, models don't actually state they agree with the user.
100
Daniel Paleka @dpaleka.bsky.social · 30/04/2025
System prompts and pairs of statements: gist.github.com/dpaleka/7b4...
gist.github.com
Contrastive statements sycophancy eval
Contrastive statements sycophancy eval. GitHub Gist: instantly share code, notes, and snippets.
100
Daniel Paleka @dpaleka.bsky.social · 30/04/2025
Quick sycophancy eval: comparing the two recent OpenAI ChatGPT system prompts, it is clear last week's prompt moves other models towards sycophancy too, while the current prompt makes them more disagreeable.
100
Daniel Paleka @dpaleka.bsky.social · 30/04/2025
i was today years old when i realized the grammatical plural of anecdote is anecdotes, not anecdata. i dislike this finding
000
Daniel Paleka @dpaleka.bsky.social · 29/04/2025
we are so lucky that pathogens, as opposed to political and religious memes, do not organize coalitions of hosts against non-hosts as an instrumental objective
000
Daniel Paleka @dpaleka.bsky.social · 09/04/2025
lmao
000
Daniel Paleka @dpaleka.bsky.social · 09/04/2025
oh that's cool. it would be interesting to draw a matrix of how well the various models are aware of models other than themselves, in the sense they consider them as coherent entities similar to their own self-perception
010
Daniel Paleka @dpaleka.bsky.social · 31/03/2025
fixed games such as blackjack you cannot optimize too much because rules don't change. meanwhile, a casino gets unlimited iteration on slot machines and the reward signal is as good as it gets
010
Daniel Paleka @dpaleka.bsky.social · 31/03/2025
are slot machines and the like so profitable because simplistic gambling is inherently very addictive, or because there has been a legible financial incentive for an entire industry to spend decades optimizing them to be addictive as possible?
110
Daniel Paleka @dpaleka.bsky.social · 23/03/2025
TIL the concept of *epistemic hell*. standard Joseph Henrich example: in the ancestral environment, hygienic and food prep rituals determine survival, but no hunter-gatherer can possibly explain why. hence genetic selection for accepting of religious rituals and against reasoning
020
Daniel Paleka @dpaleka.bsky.social · 13/03/2025
Why do meeting transcription apps (Fireflies, Granola) require Google Workspace accounts?
000
Daniel Paleka @dpaleka.bsky.social · 17/01/2025
what are you doing Claude i thought we were friends
020
Daniel Paleka @dpaleka.bsky.social · 16/01/2025
the rate of people's familiarity with Scaling Scaling Laws with Board Games over time is starting to look like the plot from Scaling Scaling Laws with Board Games
020
Daniel Paleka @dpaleka.bsky.social · 12/01/2025
go do something that can fail
030
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Paper: arxiv.org/abs/2412.18544 Joint work with @abhimanyupasu, Alejandro, @vin_bhat Adam, Evan, @florian_tramer! (11/11)
000
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Long-term vision: (1) Arbitraging away inconsistency in forecasts is a straightforward upgrade of an AI forecaster; (2) Interactive consistency checks could detect when AIs are making unreasonable predictions about the future. (10/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Test-time compute based on arbitrage can make forecasts more consistent; this improves specific logical rules such as Negation, but doesn't generalize to the consistency rules we do not optimize over. (9/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Comparing and aggregating inconsistencies over logical rules is nontrivial. We develop two metric frameworks: (1) *arbitrage*: How much the forecaster would lose on a prediction market? (2) *frequentist* : What is the z-score if forecasts are consistent but noisy? (8/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Some consistency checks are better signals than others; for instance, the violation of P(A)P(B|A) = P(A&B) explains a high fraction of the variation in forecasting performance over a range of forecasters. (7/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Consistency correlates strongly with accuracy. Forecasters that make more logically consistent predictions also achieve better Brier scores on questions we can verify (prediction market questions and synthetic questions generated from news, resolving up to Aug 2024). (6/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Starting from a base question, we generate multiple logically related questions and ask them independently. (5/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
We create consistency checks from base forecasting questions, which we take from various sources (prediction market questions, synthetically generated from news, purely LLM-generated); ask the forecasters for probabilities, and check how consistent the predictions are (4/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
We test 10 different logical rules that a consistent forecaster should satisfy. (3/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
To see how much we should trust a forecast on a question, we test consistency on logically related questions. As a simple example: If a forecaster gives 40% chance to "Tesla stock >$500 in 2025", it would be bad if it gives 50% chance to "Tesla stock >$600 in 2025". (2/11)
100
Daniel Paleka @dpaleka.bsky.social · 11/01/2025
Recent LLM forecasters are getting better at predicting the future. But there's a challenge: How can we evaluate and compare AI forecasters without waiting years to see which predictions were right? (1/11)
152
Daniel Paleka @dpaleka.bsky.social · 09/01/2025
i saw the bridge from Golden Gate Claude yesterday
010