Sign in

Samidh

@samidh.bsky.social
913 followers 97 following 106 posts

Co-Founder at Zentropi (Trustworthy AI). Formerly Meta Civic Integrity Founder, Google X and Google Civic Innovation Lead, and Groq CPO.

PostsRepliesMedia
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 17/09/2026
Runway is doing some very cool stuff using CoPE/Zentropi: "We have designed and tested a new fast moderation system that will run synchronously once a user’s input clears moderation.... The synchronous moderation through Zentropi narrows the exposure window to less than half a second."
runway.com
Runway News | Moderation in Real Time
Runway built a synchronous moderation system for real-time video generation to scan streaming frames for harmful content in under half a second — layered on Runway's existing safety defenses.
132
Samidh @samidh.bsky.social · 17/09/2026
The safety team at RunwayML has been doing some of the most innovative work in the industry. It's exciting to see how they use Zentropi to ensure the safety of their videos in real-time **as they are being generated**. Details on their blog: runway.com/news/safety/...
runway.com
Runway News | Moderation in Real Time
Runway built a synchronous moderation system for real-time video generation to scan streaming frames for harmful content in under half a second — layered on Runway's existing safety defenses.
001
Samidh @samidh.bsky.social · 17/09/2026
this is basically what i text @dwillner.bsky.social every morning
020
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 16/09/2026
At TrustCon this year I talked about a technique we’ve developed for automatically optimizing content-moderation policies, using an inversion of the binocular labeling approach Zentropi had already pioneered. Today we're shipping the tool that technique became. blog.zentropi.ai/optimizing-o...
blog.zentropi.ai
Optimizing our Policy Optimizers
Today, we are releasing our next-generation policy refinement tools: policy-only correction, label-only correction, and auto-optimization.
3176
Samidh @samidh.bsky.social · 16/09/2026
Today we are releasing the next generation of our policy optimization tools for content classifiers. It is our hope this can be another step towards helping shore up human control over AI-powered systems. blog.zentropi.ai/optimizing-o...
blog.zentropi.ai
Optimizing our Policy Optimizers
Today, we are releasing our next-generation policy refinement tools: policy-only correction, label-only correction, and auto-optimization.
032
Samidh @samidh.bsky.social · 06/09/2026
The divide in America used to be rural vs urban. Then red states vs blue states. But now I think the real emerging gap is between W-2 Americans and 1099-DIV Americans.
020
Samidh @samidh.bsky.social · 27/08/2026
ICYMI
010
Samidh @samidh.bsky.social · 27/08/2026
As usual, @masnick.com is spot on here: www.techdirt.com/2026/08/26/m... Nothing says "regulatory capture" more than making a payment to the government contingent upon them enforcing certain rules of your choice on your competitors.
techdirt.com
Meta Just Paid Nearly $17 Billion To Make Sure It Gets To Write The Kid Safety Rules For Every Other Social Media Platform
By now you've almost certainly heard the news that Meta has settled with 52 state and local Attorneys General who had sued the company in some form or another over child safety on Meta's platforms. The headlines are all covering the basics: the years-long case these states filed against Meta ends, and Meta pays somewhere...
13414
Samidh @samidh.bsky.social · 26/08/2026
The $12.7B in damages Meta will be paying to the states for harming children is less than... 1. How much it paid ScaleAI for Alexandr Wang ($14.3B) 2. Its one-day market cap increase upon launch of Muse Spark ($94.9B) 3. How much it burned on the Metaverse/VR in 2025 ($19.1B)
200
Samidh @samidh.bsky.social · 14/07/2026
Banning all teens from social media hasn't ever seemed like a great idea to me. A much healthier approach is for platforms to build smoother on-ramps for teens that are developmentally appropriate for each child's age. 🧵 [1/n]...
100
Reposted by Samidh
Tim Chambers @timothychambers.net · 14/06/2026
I continue to be impressed with Zentropi. My first thought is that they were a replcement for Perspective API that Jigsaw was ending. And yes, but so much more. So far they can answer virtually any question about a text of almost any size my work cares about that CAN be answered with a Yes or a No.
011
Samidh @samidh.bsky.social · 10/06/2026
The real question is whether these classifiers can find all the dogs that Dave has at home.
110
Samidh @samidh.bsky.social · 10/06/2026
I distinctly remember being at Meta in the wake of the Christchurch massacre, when horrific videos were circulating across Facebook without end. The technologies we had for being able to accurately classify videos just didn't exist. [1/n]
100
Reposted by Samidh
ROOST @roost.tools · 09/06/2026
Today we're releasing Coop 1.0, the world's first free, open source content review & enforcement system any org can self-host and build on. For the first time, any org, whatever its size or budget, can review, act on, and report CSAM end to end, for free. roost.tools/blog/coop-1-...
roost.tools
Coop 1.0: World’s First Free, Open Source Child Safety Infrastructure for Every Platform
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
1266
Samidh @samidh.bsky.social · 10/06/2026
Today, in a single day, @dwillner.bsky.social and I had meetings with people in England, Turkey, NYC, SF, Argentina, and Australia. Very inspiring to see four continents of folks all using Zentropi and united in the earnest work of building a better internet.
061
Samidh @samidh.bsky.social · 03/06/2026
Can your LLM actually follow your content policies, or will it just revert to its own static training? Today we're introducing a new benchmark that we call policy steerability that tries to measure this concept: blog.zentropi.ai/introducing-...
blog.zentropi.ai
Beyond Static Accuracy: Introducing the Policy Steerability Benchmark
A new way to measure how likely a model is to accurately follow your rules
120
Samidh @samidh.bsky.social · 27/05/2026
Very cool! @julietshen.bsky.social's independent tests show that our new model CoPE-B cooks :-) Direct link to her results: github.com/julietshen/c...
github.com
040
Reposted by Samidh
Juliet Shen @julietshen.online · 27/05/2026
The ROOST Model Community is growing! Today we welcome Zentropi's CoPE-B-A4B, a bring-your-own-policy model that's got the power of 25B but runs on only 4B active parameters. It's a fast, low-cost model that can be used on its own or as a first pass before larger models! roost.tools/blog/welcomi...
roost.tools
Welcoming Zentropi's CoPE-B-A4B to the ROOST Model Community
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
1193
Reposted by Samidh
ROOST @roost.tools · 27/05/2026
New in the ROOST Model Community: zentropi.ai 's CoPE-B-A4B is here. roost.tools/blog/welcomi...
roost.tools
Welcoming Zentropi's CoPE-B-A4B to the ROOST Model Community
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
172
Reposted by Samidh
Juliet Shen @julietshen.online · 27/05/2026
a few use cases I can think of for CoPE-B (and BYOP models): if you're a platform that's ok with NSFW role play but not age play, you can create a custom CoPE model that looks just for that. Many free text classifiers come with baked-in ideas of morality that might not fit your community
171
Samidh @samidh.bsky.social · 27/05/2026
Exactly right. Speech shouldn't be ruled by the platform hegemons. Your rules should rule.
0261
Reposted by Samidh
Juliet Shen @julietshen.online · 27/05/2026
We are THRILLED to announce that @roost.tools is jointly releasing CoPE-B with the Zentropi team. We believe that everyone should have access to openly licensed models designed for safety use cases that can be tuned to their community's norms. Join us at RMC office hours next week to learn more!
2366
Samidh @samidh.bsky.social · 27/05/2026
Super pumped to release CoPE-B, our latest policy-adaptive content classification model. It delivers frontier-level accuracy in a self-hostable package that's orders of magnitude cheaper to run-- opening up new possibilities in trustworthy platform design. Details: blog.zentropi.ai/meet-cope-b-...
blog.zentropi.ai
Meet CoPE-B: Frontier-Quality Content Classification You Can Self-Host
TL;DR: * Today we're releasing CoPE-B, our next-gen small language model for policy-adaptive content classification * CoPE-B-A4B (text-only) is open weights under Apache 2.0 and free to use * CoPE...
011
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 27/05/2026
@samidh.bsky.social and I are releasing CoPE-B today, the next version of our policy-adaptive content classifier. It delivers at-or-better-than-frontier classification while being self-hostable, faster, and cheaper to run. Full writeup with benchmarks at blog.zentropi.ai/meet-cope-b-.... 🧵 1/9
blog.zentropi.ai
Meet CoPE-B: Frontier-Quality Content Classification You Can Self-Host
TL;DR: * Today we're releasing CoPE-B, our next-gen small language model for policy-adaptive content classification * CoPE-B-A4B (text-only) is open weights under Apache 2.0 and free to use * CoPE...
25617
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 11/05/2026
Very excited to have this public. The @oversightboard.bsky.social's use of Zentropi to study child marriage content on Meta platforms was very cool to support: blog.zentropi.ai/how-the-over...
blog.zentropi.ai
How the Oversight Board uses Zentropi to study policy impact at scale
The Oversight Board used Zentropi to analyze a large dataset of content that potentially violated Meta’s policies on human exploitation. The tool helped the Board reduce the project timeline from week...
293
Samidh @samidh.bsky.social · 11/05/2026
Check out how the @oversightboard.bsky.social used Zentropi to better understand how child marriage-related content manifests on Meta's platforms. Fantastic example of how advanced content labeling technologies can strengthen both our online and offline world. blog.zentropi.ai/how-the-over...
blog.zentropi.ai
How the Oversight Board uses Zentropi to study policy impact at scale
The Oversight Board used Zentropi to analyze a large dataset of content that potentially violated Meta’s policies on human exploitation. The tool helped the Board reduce the project timeline from week...
021
Samidh @samidh.bsky.social · 14/04/2026
It has been incredible partnering with character.ai since the very start of zentropi.ai. We're excited to share some details of that partnership with this case study. Anyone creating AI-powered systems might find it interesting! blog.zentropi.ai/how-zentropi...
blog.zentropi.ai
How Zentropi partners with Character.ai
Character.ai takes safety seriously. With millions of users creating and chatting with AI characters every day, the team invests heavily in systems that help protect their community — and they're alwa...
000
Reposted by Samidh
Katie Harbath @katieharbath.bsky.social · 24/03/2026
In 2017, it took us months to define “political ad” at Facebook. Recently, I built two political content classifiers in an afternoon using Zentropi AI created by @dwillner.bsky.social and @samidh.bsky.social Why AI content moderation is good, actually — in this week’s newsletter 👇
open.substack.com
Why AI Makes Content Moderation Better, Not Worse
Building a political content labeler with AI — what actually works
185
Samidh @samidh.bsky.social · 18/03/2026
One of the things we've been thinking about a lot at Zentropi is: what happens when AI agents need to make judgment calls about content — not humans reviewing a queue, but agents acting autonomously?
211
Samidh @samidh.bsky.social · 10/03/2026
There's a major gap in content safety tooling: classifiers typically only score complete text. When you're working with generative AI, "complete text" means the user already saw it. That's too late. So we built a streaming classifier that we're releasing today! Here's what we did and why. 🧵...
172
Reposted by Samidh
Juliet Shen @julietshen.online · 19/02/2026
If you're not using either tool yet, now's a good time to try both! Zentropi's Community Edition is free and gives you unlimited labelers. Coop is fully open source and runs on your infrastructure. :D
081
Samidh @samidh.bsky.social · 19/02/2026
Zentropi is now integrated into Coop, @roost.tools's open source moderation platform. You can write a content policy in plain English on Zentropi, plug it into Coop as a signal, and have a moderation pipeline running in minutes.
172
Reposted by Samidh
Tech Policy Press @techpolicypress.bsky.social · 29/01/2026
Dave Willner, who led trust and safety at major tech firms and has cofounded a company that is developing an AI content classification platform, says LLM-driven technology can now accomplish classification at the scale necessary for moderation on large platforms. That has substantial implications.
techpolicy.press
AI is Removing Bottlenecks to Effective Content Moderation at Scale
Zentropi's Dave Willner says LLM-driven technology can now accomplish content classification at the scale necessary for moderation on large platforms.
141
Samidh @samidh.bsky.social · 26/01/2026
I can has cats.
010
Samidh @samidh.bsky.social · 26/01/2026
Just shipped Zentropi's most requested feature: image classification! Now analyze images against your own policies, at scale. To power it we built cope-b-12b, a new multimodal model w/ native vision. Check out the cat detector we made in < 1 min. 🐱 blog.zentropi.ai/zentropi-now-labels-images/
blog.zentropi.ai
Zentropi Now Labels Images
Building guardrails for visual content just got a lot easier. Today we're launching image classification on Zentropi and announcing cope-b-12b, a multimodal model that powers this experience.
0134
Samidh @samidh.bsky.social · 21/01/2026
If you are looking for a technical description of how X rots your brain, look no further than their github post on the 'X algorithm'. It is pure, unadulterated behavioral engagement maximization that amplifies the very worst human impulses. github.com/xai-org/x-al...
github.com
GitHub - xai-org/x-algorithm: Algorithm powering the For You feed on X
Algorithm powering the For You feed on X. Contribute to xai-org/x-algorithm development by creating an account on GitHub.
34522
Samidh @samidh.bsky.social · 15/01/2026
Why are we just giving away all our secrets? Well, it is our hope that it helps the ecosystem further advance the state of the art in policy-steerable content classification, which is foundational to a more trustworthy internet.
052
Samidh @samidh.bsky.social · 13/01/2026
Dave just published a Zentropi labeler that can precisely identify requests at prompting an AI model to undress a person in a photo. The tools exist to easily deal with this problem -- platforms just need to choose to use them. If you are the developer of an AI system, please use this guardrail!
053
Samidh @samidh.bsky.social · 03/12/2025
This was such a cool experiment that I created a Zentropi labeler with a simplified version of the authors' Partisan Animosity criteria. Now anyone can experiment directly with using this labeler to try to reduce the temperature of affective polarization in their feeds. zentropi.ai/labelers/b30...
science.org
Reranking partisan animosity in algorithmic social media feeds alters affective polarization
Today, social media platforms hold the sole power to study the effects of feed-ranking algorithms. We developed a platform-independent method that reranks participants’ feeds in real time and used thi...
092
Samidh @samidh.bsky.social · 13/11/2025
We just wrote an in-depth post about Toxic Content labeling. It presents a new way of defining toxic speech online-- and illustrates the importance of observable features for accurate language model interpretability. Would love to hear how YOU define toxicity, too! blog.zentropi.ai/observations...
blog.zentropi.ai
Observations on Toxicity
We've published Zentropi's toxicity labeler (toxicity-public-s5), which you can integrate with your platform instantly using the Zentropi API. Browse the full policy to see how defining observable fea...
0102
Samidh @samidh.bsky.social · 12/11/2025
Awesome to see how this is already being used! One of the most useful aspects is that the published policies show what it takes to write content rules that can be accurately interpreted by language models. We hope this can be a boost to the broader content policy community.
010
Samidh @samidh.bsky.social · 10/11/2025
This was a fun launch! It turns Zentropi into a Github for Content Labelers. You can share content policies with others and build off each other's work. It's the easiest way of deploying a fully customizable classifier. Check out the policies @dwillner.bsky.social created at zentropi.ai/u/dave
zentropi.ai
Zentropi - Content Labelers by dave
030
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 10/11/2025
Content policies are usually private, one-off efforts. You build yours, I build mine, we don't share much about what works or why. This makes sense given products can (and should) set different policies based on their communities, but it leaves us reinventing the wheel. 🧵 1/5
2186
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 27/08/2025
We got really positive feedback on the TrustCon workshop we ran on writing good content policies for LLMs...so we're doing it again! If you're interested go sign up here, so we can start to figure out timing: forms.gle/tj7vf7ng8n7R...
forms.gle
Zentropi LLM Policy Writing Workshop Signup
By popular demand, we will be hosting a virtual version of our sold-out TrustCon workshop on how to write high quality content policies with and for LLMs. In this session, you will learn best practic...
022
Samidh @samidh.bsky.social · 27/08/2025
This response to the Raine tragedy from OpenAI does something remarkable: it has the humility to acknowledge that a *product failure* led to real-world harm. Despite horrific circumstances, it has a rare degree of honesty that I wish tech companies would show more often. openai.com/index/helpin...
openai.com
Helping people when they need it most
How we think about safety for users experiencing mental or emotional distress, the limits of today’s systems, and the work underway to refine them.
120
Samidh @samidh.bsky.social · 19/08/2025
We are opening up Zentropi.ai to everyone today so that anyone can build their own content labeler. What started as a crazy academic idea 2 years ago is now a real thing that companies are using in production to safeguard their AI-powered systems. Give it a shot! blog.zentropi.ai/zentropi-bui...
blog.zentropi.ai
Zentropi: Build Your Own Content Labeler in Minutes, Not Months
We are officially opening up our build-your-own-content-labeler platform to everyone. Check it out at zentropi.ai.
031
Samidh @samidh.bsky.social · 31/07/2025
Don't take our word for it! Go kick the tires at zentropi.ai and build your own content labeler (no subscription required!)
zentropi.ai
Zentropi - Build Custom Content Labelers Instantly
031
Samidh @samidh.bsky.social · 31/07/2025
@mmasnick.bsky.social I have a bluesky demo for you that you might want to see :)
130
Reposted by Samidh
Ms . Penny Oaken, SkyWitch 🧙‍♀️ @skywitchy.bsky.social · 31/07/2025
Just tested this on a few that I know Reddit’s existing Hatred & Harassment automation has blindspots for; it built a focused, accurate labeler in under 10 minutes / a dozen examples, & the human readable criteria it built could be dropped into a training manual / erratum / used to build a regex
2125