Andrés Corrada @andrescorrada.bsky.social · 10/07/2026wapo.st/4ydn1Uxwapo.stMigrants who saw man killed by ICE in Houston say he did not ram officersThree men who were in the vehicle alongside Lorenzo Salgado Araujo are contesting the Department of Homeland Security’s account of the fatal shooting. 020
Andrés Corrada @andrescorrada.bsky.social · 10/07/2026The problem of verifying experts that are smarter and more knowledgeable than us is ancient and pervasive. AI did not invent it and will not fully resolve it. For that we need human ingenuity. Formal verification of experts is one proposed route to give us safety. But this is extremely hard. 100
Andrés Corrada @andrescorrada.bsky.social · 09/07/2026Formal verification of evaluators is much more tractable than formal verification of the agents they evaluate. This is how NTQR helps ameliorate the who-judges-the-judges? problem Evaluation models and their verification is trivial in comparison to verifying World models. ntqr.readthedocs.io 000
Andrés Corrada @andrescorrada.bsky.social · 08/07/2026NTQR does not—and cannot—determine whether an evaluation task is meaningful or whether the criteria being measured are the right ones. Those questions require human judgment, and alignment research. ntqr.readthedocs.io 110
Andrés Corrada @andrescorrada.bsky.social · 06/07/2026Version 0.9 of the NTQR package is out! (pip install ntqr) ntqr.readthedocs.io 000
Reposted by Andrés CorradaAndrés Corrada @andrescorrada.bsky.social · 05/07/2026Seems so obvious once it is stated. Here is a talk I recently gave on one problem that arises from this epistemic dependency: who-judges-the-judges? www.youtube.com/live/t2TgSuY...youtube.comActInf GuestStream 067.2: Andrés Corrada "Who Judges the Judges?"YouTube video by Active Inference Institute 111
Andrés Corrada @andrescorrada.bsky.social · 05/07/2026This is precisely the point of the logic of unsupervised evaluation for classifiers that I am developing in the NTQR Python package -- to develop tools for evaluating experts when no one in the room knows the truth. ntqr.readthedocs.io 100
Andrés Corrada @andrescorrada.bsky.social · 05/07/2026Who Judges the Judges? If my Open Source NTQR Python package verifies evaluations for classifiers, who or what verifies NTQR? I am working on release v0.8.1 that will significantly speed up the evaluation set generators. It also includes checks of whether they are: valid, complete, uniform, mixing. 200
Andrés Corrada @andrescorrada.bsky.social · 02/07/2026One can quip that the semantically-free nature of NTQR logic makes it being able to get into all the clubs without belonging to any of them. Here I use it to compute the evaluations that are consistent with 4 LLMs answering a Q=295 MedQA questions used by medical licensing boards. 100
Reposted by Andrés CorradaYoïn van Spijk @yvanspijk.bsky.social · 02/07/2026‘Chez moi’ is French for “at my place”. The preposition ‘chez’ has a fascinating origin: it shares its origin with Spanish ‘casa’ (house): Latin ‘casa’. ‘Casa’ initially meant “hut”, but it became the word for "house" in Romance. What happened to the original word? My new graphic tells you more. 1411333
Andrés Corrada @andrescorrada.bsky.social · 02/07/2026Just gave this talk and I forgot an important preamble: when does the who-judges-the-judges? problem arise. It applies to what UK AISI's experts call "hard-to-evaluate fuzzy tasks" where even reasonable human experts can disagree on the ground truth. www.youtube.com/live/t2TgSuY...youtube.comActInf GuestStream 067.2: Andrés Corrada "Who Judges the Judges?"YouTube video by Active Inference Institute 100
Reposted by Andrés CorradaLauren Dobson-Hughes @ldobsonhughes.bsky.social · 24/06/2026Musk says not a single person has died due to him axing USAID. Sadly, part of my job involves tracking aid spending. So here’s a thread with just some of the people who lost their lives because of USAID cuts 196133286755
Reposted by Andrés CorradaPat Walshe - Privacy Matters 🇮🇪🇬🇧🇪🇺 @privacymatters.bsky.social · 28/06/2026‘A group of former National Oceanic and Atmospheric Administration employees who had been fired by Elon Musk's DOGE have launched a new climate science website documenting global climate change.’ www.climate.us 061
Andrés Corrada @andrescorrada.bsky.social · 28/06/2026Another superpower of logical consistency in unsupervised evaluation of classifiers is the sparseness it imposes on the possible set of evaluations. Once we count how experts disagree on a test, the evaluations consistent with those counts are sparse in the evaluation space. Like 1 in 10^13 sparse. 100
Reposted by Andrés CorradaJohn Horgan @jhorganism.bsky.social · 28/06/2026I see disturbing resonances between "Death of a Salesman," which I just saw on Broadway, and a 1987 comic strip by R. Crumb: johnhorgan.org/cross-check/...johnhorgan.orgWilly Loman and The Ruff-Tuff Cream-Puffs — John Horgan (The Science Writer)HOBOKEN, JUNE 28, 2026. Great art, like “Death of a Salesman,” always speaks to us, but what we hear can change. Vicki somehow got us tickets to the acclaimed revival of the 1949 play, whi... 021
Andrés Corrada @andrescorrada.bsky.social · 28/06/2026The counting logic for unsupervised evaluation of classifiers in the NTQR Python package represents a test evaluation as a joint decisions confusion matrix as shown here. In unsupervised settings, we do not know the counts by true label. We only know the top row. ntqr.readthedocs.io 100
Andrés Corrada @andrescorrada.bsky.social · 28/06/2026The AI hypers may think that but the lesson expressed by Ponomareva is what any safety engineer holds to be reasonable. Safe-fail design is an important component of any useful technology. Its appearance is not a bug, but a feature. Seat belts, for example. The human may be the one out of control. 000
Andrés Corrada @andrescorrada.bsky.social · 28/06/2026My first point in this talk will be that the problem of verifying experts that are smarter or more knowledgeable than us is ancient and pervasive. www.youtube.com/live/t2TgSuY... Ancient: Ashurbanipal, the king of the Assyrian Empire not sure his scribes are being correct or truthful. 100
Reposted by Andrés Corradalogan koepke @jlkoepke.bsky.social · 26/06/2026if you're interested in how regulatory design shapes corporations' actions, processes, and policies to mitigate the material harms of AI systems, fair lending in the US offers decades of overlooked lessons. check out our new FAccT paper ⬇️ dl.acm.org/doi/10.1145/...dl.acm.orgThe Fair Lending Model: How the Longest-Running Algorithmic Fairness Programs Work in Practice | Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency 061
Andrés Corrada @andrescorrada.bsky.social · 25/06/2026@adolfont.github.io , @emilymbender.bsky.social I see my work as a tool for Paulo Freire's classroom. I grew up in a household that had the first, Spanish, edition printed in Argentina. How would student's without an oppressive teacher, be able to self-evaluate themselves? 110
Andrés Corrada @andrescorrada.bsky.social · 25/06/2026The "untenable middleground" of @emilymbender.bsky.social is where all scientists live. Whereas she meant it as a rhetorical tool to defang those that believe we can have responsible uses of AI, I view it as the hopeless task of all science to hold a candle in the darkness. Science = middleground. 100
Andrés Corrada @andrescorrada.bsky.social · 25/06/2026I love this term by @emilymbender.bsky.social and will use it from now on to express my pessimism and hope for science helping us hold the "untenable middleground". Her logical fallacy of assuming the conclusion is actually a good picture of how science is our only tool in the darkness. 100
Andrés Corrada @andrescorrada.bsky.social · 25/06/2026The counting logic in NTQR is semantically free -- it universally applies to all classification domains. The only input for NTQR functions that comes from the test is the observed counts of their joint decision events. As number of classifiers, N, and labels, R increases, R^N explodes. 100
Andrés Corrada @andrescorrada.bsky.social · 24/06/2026The catastrophe we fear the most has already happened. The untenable middleground is precisely what scholars in colonized societies have had to hold against colonial parrots expounding intellectual superiority while bringing vaccines and science. 100
Andrés Corrada @andrescorrada.bsky.social · 23/06/2026Homer Simpson has an answer for the ancient problem of policing the police - youtu.be/Tk4yyqXi8Xc?...youtu.beWho will police the police simpsonsYouTube video by YNOS MOVIE QUOTE MANIA 000
Andrés Corrada @andrescorrada.bsky.social · 22/06/2026"Who judges the judges?" is a question that follows the anxious realization that when it comes to certain truths, the only way to be 100% safe is an endless chain of confirmation. The question cannot be resolved, only ameliorated by terminating the chain with assumptions. NTQR logic is no different. 000
Andrés Corrada @andrescorrada.bsky.social · 22/06/2026"Who checks the level-checking level?" asks Natalie's daughter, Julia. The fastidious Mr. Monk has complained Natalie's picture hanging is not level and he wants to check it with his "level-checking level" since he does not trust Natalie's level. Instruments are experts. Who checks their claims? 100
Andrés Corrada @andrescorrada.bsky.social · 21/06/2026I've been preparing for this upcoming talk, www.youtube.com/live/t2TgSuY..., on the logic of unsupervised evaluation for classifiers and its use in ameliorating this ancient problem that AI is just making worse. How can we verify experts smarter or more knowledgeable than us?youtube.comActInf GuestStream 067.2: Andrés Corrada "Who Judges the Judges?"YouTube video by Active Inference Institute 100
Andrés Corrada @andrescorrada.bsky.social · 19/06/2026A semantic-free counting logic for evaluating classifiers balms many ills. 1. Because it is semantic free it could be used in all classification domains. 2. Once you solve the fundamental question - what evaluations are logically consistent with the observed disagreements? Other logics are solvable. 100
Andrés Corrada @andrescorrada.bsky.social · 18/06/2026I'll be giving a talk on the @activeinference.bsky.social guest stream on July 2nd on the ancient problem that plagues AI alignment - Who Judges the Judges? - and how the logic of unsupervised evaluation for classifiers in the NTQR package ameliorates it. www.youtube.com/live/t2TgSuY...youtube.comActInf GuestStream 067.2: Andrés Corrada "Who Judges the Judges?"YouTube video by Active Inference Institute 010
Andrés Corrada @andrescorrada.bsky.social · 16/06/2026The race for super-intelligent AI didn't invent the problem of command and control when we we are not the smartest or most knowledgeable in the room. Abandoning AI will not save us from the problem either. Ashurbanipal, king of the Assyrian Empire, had the same verification issues with his scribes. 100
Andrés Corrada @andrescorrada.bsky.social · 16/06/2026I'm neither a techno optimist or pessimist, I come from a long line of techno survivors. The catastrophe we fear the most when it comes to the arrival of super-intelligence has already happened. Repeatedly. And some of us have survived. Ensembling experts and logical inconsistency are strategies ... 110
Andrés Corrada @andrescorrada.bsky.social · 15/06/2026The consolations of philosophy. Once you realize that when it comes to surviving the arrival of super-intelligence, the catastrophe you fear the most has already happened, repeatedly, and some of us have survived it, you can start looking for how they did it and mimic their strategies. 100
Andrés Corrada @andrescorrada.bsky.social · 15/06/2026Who judges the judges? That ancient problem is at the heart of our inability to control and monitor AI to make it safer to use. How do we verify experts smarter or more knowledgeable than us? That fundamental question is all over this excellent paper by @girving.bsky.social and co-authors. 111
Andrés Corrada @andrescorrada.bsky.social · 03/06/2026Daniel Friedman, from the @activeinference.bsky.social community has posted this nice critique of the strengths and weaknesses of the binary error independent evaluator in my NTQR package - zenodo.org/records/2049... He uncovered a logical bug in the evaluator when one of a binary trio is "stuck"! 100
Reposted by Andrés CorradaPrivacy International @privacyinternational.org · 31/05/2026♠️ Hemos creado una baraja de cartas sobre tecnología, datos y elecciones: privacyinternational.org/long-read/57...privacyinternational.orgJuego de cartas sobre tecnología, datos y eleccionesTras publicar Tecnología, datos y elecciones: lista de verificación del ciclo electoral en español 111
Andrés Corrada @andrescorrada.bsky.social · 29/05/2026Sometimes you have to give up to win. By treating unsupervised evaluation as a logical problem, instead of a probabilistic one, you can improve your lot. Take "how do you judge ...". I solved this problem years ago for classifiers by doing the exact algebraic solution for binary classifiers. 100
Andrés Corrada @andrescorrada.bsky.social · 28/05/2026Version v0.8 of the NTQR Python package for the logic of unsupervised evaluation of classifiers is out. This release includes complete generators and random samplers for the two objects of interest in such logics - the possible and consistent sets of evaluations. ntqr.readthedocs.io/en/latest/no... 100
Reposted by Andrés CorradaCailin O’Connor @cailinmeister.bsky.social · 05/05/2026Do you like *free philosophy*? The Routledge Handbook of Values in Science is available open access! My entry with Rebecca Korf talks about what network models can tell us about social values in science. www.taylorfrancis.com/books/oa-edi...taylorfrancis.comThe Routledge Handbook of Values and Science | Kevin C. Elliott, Ted RThis is the first-ever handbook to cover the vibrant philosophical literature on values and science. Its 45 chapters—appearing in print here for the first 02910
Reposted by Andrés CorradaGeorge Conway ⚖️🇺🇸 @gtconway.bsky.social · 04/05/2026As true as ever. 271645392
Andrés Corrada @andrescorrada.bsky.social · 03/05/2026I got blocked for being a sycophant hyping AI because I work on AI safety. Being safe from noisy experts is a universal need. I don't care who claims to be the expert - human or machine. The problem of being safe from noisy experts began with civilization. We can build tools to protect us. 100
Andrés Corrada @andrescorrada.bsky.social · 02/05/2026I’ve been working on a different way to evaluate classifiers without ground truth—one that removes semantics entirely. What if evaluation didn’t depend on what labels mean? The idea is a semantic-free logic of unsupervised evaluation. Use only logical consistency between experts, not soundness. 310
Andrés Corrada @andrescorrada.bsky.social · 29/04/2026But using race to profile US citizens is still legal. Race blindness not required in Kavanaugh stops. www.nytimes.com/2026/04/29/o...nytimes.comOpinion | White Drivers Got a Warning. Latino Drivers Got Detained. 000
Reposted by Andrés CorradaSeptima P. Snark @drsubini.blacksky.app · 29/04/2026Color-evasiveness is more accurate here because it shows how they evade discussions and remedies that involve race in order to hoard power. 0125
Andrés Corrada @andrescorrada.bsky.social · 29/04/2026Part of the problem with 'world models' as a way to have safer AI is the problem of completeness. Sherlock's logic - whatever remains, however improbable, must be the truth - fails when we have not considered all possible hypothesis. Evaluation logic is not like that. It is complete. 100
Andrés Corrada @andrescorrada.bsky.social · 28/04/2026OMG! Did not know Clojure was still having commercial traction. I once developed a whole anonymous ID system for online ads with Clojure as proof of concept in a start-up (Blue Cava). The first engineer to take over the project rewrote it all in Bash wrapper scripts for Python code. 120
Andrés Corrada @andrescorrada.bsky.social · 28/04/2026'World models' are great and I look forward to implementing physical laws in how robots understand their environments. But they are hard to develop. 'Evaluation models', my work, are trivial by comparison. Classification or multiple choice tests have finite, denumerable evaluations. 100
Andrés Corrada @andrescorrada.bsky.social · 28/04/2026When you use a monitor M to supervise an AI system S, who monitors the monitor? This infinite regress has to be terminated with assumptions. One way is to recognize monitors/graders are often acting as classifiers. This is not immediately obvious to AI safety people. 110
Andrés Corrada @andrescorrada.bsky.social · 27/04/2026Version 0.7.8 of the NTQR package is out. It now includes a nifty "ntqr-docs" Python script that installs working Jupyter notebooks as shown at readthedocs: ntqr.readthedocs.io/en/latest/no... 100
Andrés Corrada @andrescorrada.bsky.social · 24/04/2026The math is correct, the science terrible. If the drug does not change price it has had an increase of 100% and a decrease of 100% by this metrical logic! The correct formula is (change/original). It also has many flaws as a good metric: 1. Percentage decreases and increases are no longer equal when 100