Sign in

Samidh

@samidh.bsky.social
914 followers 97 following 106 posts

Co-Founder at Zentropi (Trustworthy AI). Formerly Meta Civic Integrity Founder, Google X and Google Civic Innovation Lead, and Groq CPO.

PostsRepliesMedia
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 17/09/2026
Runway is doing some very cool stuff using CoPE/Zentropi: "We have designed and tested a new fast moderation system that will run synchronously once a user’s input clears moderation.... The synchronous moderation through Zentropi narrows the exposure window to less than half a second."
runway.com
Runway News | Moderation in Real Time
Runway built a synchronous moderation system for real-time video generation to scan streaming frames for harmful content in under half a second — layered on Runway's existing safety defenses.
132
Samidh @samidh.bsky.social · 17/09/2026
The safety team at RunwayML has been doing some of the most innovative work in the industry. It's exciting to see how they use Zentropi to ensure the safety of their videos in real-time **as they are being generated**. Details on their blog: runway.com/news/safety/...
runway.com
Runway News | Moderation in Real Time
Runway built a synchronous moderation system for real-time video generation to scan streaming frames for harmful content in under half a second — layered on Runway's existing safety defenses.
001
Samidh @samidh.bsky.social · 17/09/2026
this is basically what i text @dwillner.bsky.social every morning
020
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 16/09/2026
At TrustCon this year I talked about a technique we’ve developed for automatically optimizing content-moderation policies, using an inversion of the binocular labeling approach Zentropi had already pioneered. Today we're shipping the tool that technique became. blog.zentropi.ai/optimizing-o...
blog.zentropi.ai
Optimizing our Policy Optimizers
Today, we are releasing our next-generation policy refinement tools: policy-only correction, label-only correction, and auto-optimization.
3176
Samidh @samidh.bsky.social · 16/09/2026
Today we are releasing the next generation of our policy optimization tools for content classifiers. It is our hope this can be another step towards helping shore up human control over AI-powered systems. blog.zentropi.ai/optimizing-o...
blog.zentropi.ai
Optimizing our Policy Optimizers
Today, we are releasing our next-generation policy refinement tools: policy-only correction, label-only correction, and auto-optimization.
032
Samidh @samidh.bsky.social · 06/09/2026
Pretty sure that Optimus is supposed to carry you.
010
Samidh @samidh.bsky.social · 06/09/2026
The divide in America used to be rural vs urban. Then red states vs blue states. But now I think the real emerging gap is between W-2 Americans and 1099-DIV Americans.
020
Samidh @samidh.bsky.social · 27/08/2026
ICYMI
010
Samidh @samidh.bsky.social · 27/08/2026
As usual, @masnick.com is spot on here: www.techdirt.com/2026/08/26/m... Nothing says "regulatory capture" more than making a payment to the government contingent upon them enforcing certain rules of your choice on your competitors.
techdirt.com
Meta Just Paid Nearly $17 Billion To Make Sure It Gets To Write The Kid Safety Rules For Every Other Social Media Platform
By now you've almost certainly heard the news that Meta has settled with 52 state and local Attorneys General who had sued the company in some form or another over child safety on Meta's platforms. The headlines are all covering the basics: the years-long case these states filed against Meta ends, and Meta pays somewhere...
13414
Samidh @samidh.bsky.social · 26/08/2026
The $12.7B in damages Meta will be paying to the states for harming children is less than... 1. How much it paid ScaleAI for Alexandr Wang ($14.3B) 2. Its one-day market cap increase upon launch of Muse Spark ($94.9B) 3. How much it burned on the Metaverse/VR in 2025 ($19.1B)
200
Samidh @samidh.bsky.social · 14/07/2026
This marks an inflection point in how we might keep kids safe online. Rather than be an all-or-nothing proposition, we now have tools that can respect the autonomy of teens and help them grow into more capable digital citizens. Full details on the Vys blog: quire.substack.com/p/what-it-ta... [5/5]
quire.substack.com
[What it Takes] For LLMs to Make Age-Appropriate Content Policies Enforceable
An editorial developed in partnership with Zentropi
010
Samidh @samidh.bsky.social · 14/07/2026
The Vys policies are deeply grounded in cognitive development research and allow platforms to easily build age-aware protections that are production-grade. But you don't have to use them as-is -- you can fork and customize them within Zentropi to suit your community's needs. [4/n]
100
Samidh @samidh.bsky.social · 14/07/2026
That's why I'm excited today to see @vaishnavi.bsky.social & team use Zentropi to build and release the first set of open source age-appropriate content policies to detect teen self-harm & suicidal ideation. Get them here: zentropi.ai/u/vys [3/n]
zentropi.ai
Zentropi - Content Labelers by vys
111
Samidh @samidh.bsky.social · 14/07/2026
Historically this has been a monumental product challenge, requiring different product experiences for users depending on their ages. But now with the advent of policy-steerable content classifiers like ours at Zentropi, this is much more feasible. [2/n]
100
Samidh @samidh.bsky.social · 14/07/2026
Banning all teens from social media hasn't ever seemed like a great idea to me. A much healthier approach is for platforms to build smoother on-ramps for teens that are developmentally appropriate for each child's age. 🧵 [1/n]...
100
Reposted by Samidh
Tim Chambers @timothychambers.net · 14/06/2026
I continue to be impressed with Zentropi. My first thought is that they were a replcement for Perspective API that Jigsaw was ending. And yes, but so much more. So far they can answer virtually any question about a text of almost any size my work cares about that CAN be answered with a Yes or a No.
011
Samidh @samidh.bsky.social · 10/06/2026
The real question is whether these classifiers can find all the dogs that Dave has at home.
110
Samidh @samidh.bsky.social · 10/06/2026
Video labelers are available to Zentropi subscribers today, but let me know if you want to give it a try. We want all platforms to have access to first class tools for video classification. [5/5]
001
Samidh @samidh.bsky.social · 10/06/2026
The same policy-steerable classifier we built for text and images now applies to video. You can use our tools to write your own policy for video, optimize it according to your data, and deploy it live. Your rules finally rule. [4/n]
111
Samidh @samidh.bsky.social · 10/06/2026
That's why I'm especially excited to share that Zentropi now supports video! We've made it trivially easy for anyone to create video classifiers that are fast, accurate, and cheap to run at scale: blog.zentropi.ai/zentropi-now... [3/n]
blog.zentropi.ai
Zentropi Now Labels Videos
Building guardrails for video content just got a lot easier. Today we're launching video classification on Zentropi — a new capability unlocked by our just-released multimodal model CoPE-B-A4B-MM
100
Samidh @samidh.bsky.social · 10/06/2026
Since then, video has continued to be the modality where safety teams have had the fewest tools. Most options are fixed-taxonomy (e.g., "NSFW"), human review, or a frontier API call that's too slow and expensive to run at scale. None of them follow your rules. [2/n]
110
Samidh @samidh.bsky.social · 10/06/2026
I distinctly remember being at Meta in the wake of the Christchurch massacre, when horrific videos were circulating across Facebook without end. The technologies we had for being able to accurately classify videos just didn't exist. [1/n]
100
Reposted by Samidh
ROOST @roost.tools · 09/06/2026
Today we're releasing Coop 1.0, the world's first free, open source content review & enforcement system any org can self-host and build on. For the first time, any org, whatever its size or budget, can review, act on, and report CSAM end to end, for free. roost.tools/blog/coop-1-...
roost.tools
Coop 1.0: World’s First Free, Open Source Child Safety Infrastructure for Every Platform
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
1266
Samidh @samidh.bsky.social · 10/06/2026
Today, in a single day, @dwillner.bsky.social and I had meetings with people in England, Turkey, NYC, SF, Argentina, and Australia. Very inspiring to see four continents of folks all using Zentropi and united in the earnest work of building a better internet.
061
Samidh @samidh.bsky.social · 03/06/2026
We hope this new methodology on measuring policy compliance is helpful to the trust & safety ecosystem and so we're sharing our work in the open. Critique and questions are welcome! Full details here: blog.zentropi.ai/introducing-...
blog.zentropi.ai
Beyond Static Accuracy: Introducing the Policy Steerability Benchmark
A new way to measure how likely a model is to accurately follow your rules
000
Samidh @samidh.bsky.social · 03/06/2026
So we decided to study this phenomenon in a lot more rigor and devised this new policy steerability benchmark. As more AI-powered systems start using LLMs behind the scenes for classification, steerability is crucial. When you write a rule for your platform, your system should follow it.
110
Samidh @samidh.bsky.social · 03/06/2026
When we released CoPE-B last week-- our newest language model for content classification-- the standard metrics showed only modest gains. But every time we used the model it *felt* like a huge step up from our prior model, responding to policy subtlety and changes with much greater fidelity.
110
Samidh @samidh.bsky.social · 03/06/2026
Can your LLM actually follow your content policies, or will it just revert to its own static training? Today we're introducing a new benchmark that we call policy steerability that tries to measure this concept: blog.zentropi.ai/introducing-...
blog.zentropi.ai
Beyond Static Accuracy: Introducing the Policy Steerability Benchmark
A new way to measure how likely a model is to accurately follow your rules
120
Samidh @samidh.bsky.social · 29/05/2026
Come to think of it, "Talk to me like I'm an agent" isn't a bad idea to put into my system prompt... :-)
010
Samidh @samidh.bsky.social · 28/05/2026
We evaluate *all* new open models as potential backbones for CoPE, including Qwen. For our use case, Gemma-4 simply performed better, especially in terms of its instruction-following capability, general world knowledge, and token efficiency.
120
Samidh @samidh.bsky.social · 27/05/2026
Our CoPE model stands on the shoulders of giants ;-). Folks can download it from huggingface here. We'd be would be very happy to chat with others who are using Gemma for task-specific fine tuning. huggingface.co/zentropi-ai/...
huggingface.co
zentropi-ai/cope-b-a4b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
020
Samidh @samidh.bsky.social · 27/05/2026
Couldn't agree more. That's actually why we created an agent skill that allows your agent to use CoPE directly. Check it out: github.com/zentropi-ai/...
github.com
GitHub - zentropi-ai/skills: Agent skills powered by the Zentropi content classification engine
Agent skills powered by the Zentropi content classification engine - zentropi-ai/skills
010
Samidh @samidh.bsky.social · 27/05/2026
So cool! Let us know what you think of CoPE-B. And would love to hear more about how you are using the model more generally!
000
Samidh @samidh.bsky.social · 27/05/2026
Well said!
010
Samidh @samidh.bsky.social · 27/05/2026
Very cool! @julietshen.bsky.social's independent tests show that our new model CoPE-B cooks :-) Direct link to her results: github.com/julietshen/c...
github.com
040
Samidh @samidh.bsky.social · 27/05/2026
@chaosgreml.in would love to hear more about how you're using CoPE!
020
Reposted by Samidh
Juliet Shen @julietshen.online · 27/05/2026
The ROOST Model Community is growing! Today we welcome Zentropi's CoPE-B-A4B, a bring-your-own-policy model that's got the power of 25B but runs on only 4B active parameters. It's a fast, low-cost model that can be used on its own or as a first pass before larger models! roost.tools/blog/welcomi...
roost.tools
Welcoming Zentropi's CoPE-B-A4B to the ROOST Model Community
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
1193
Reposted by Samidh
ROOST @roost.tools · 27/05/2026
New in the ROOST Model Community: zentropi.ai 's CoPE-B-A4B is here. roost.tools/blog/welcomi...
roost.tools
Welcoming Zentropi's CoPE-B-A4B to the ROOST Model Community
Robust Open Online Safety Tools or ROOST is a new non-profit entity designed to address the urgent need for accessible, high-quality safety tools in the rapidly evolving digital landscape.
172
Reposted by Samidh
Juliet Shen @julietshen.online · 27/05/2026
a few use cases I can think of for CoPE-B (and BYOP models): if you're a platform that's ok with NSFW role play but not age play, you can create a custom CoPE model that looks just for that. Many free text classifiers come with baked-in ideas of morality that might not fit your community
171
Samidh @samidh.bsky.social · 27/05/2026
Exactly right. Speech shouldn't be ruled by the platform hegemons. Your rules should rule.
0261
Reposted by Samidh
Juliet Shen @julietshen.online · 27/05/2026
We are THRILLED to announce that @roost.tools is jointly releasing CoPE-B with the Zentropi team. We believe that everyone should have access to openly licensed models designed for safety use cases that can be tuned to their community's norms. Join us at RMC office hours next week to learn more!
2366
Samidh @samidh.bsky.social · 27/05/2026
Super pumped to release CoPE-B, our latest policy-adaptive content classification model. It delivers frontier-level accuracy in a self-hostable package that's orders of magnitude cheaper to run-- opening up new possibilities in trustworthy platform design. Details: blog.zentropi.ai/meet-cope-b-...
blog.zentropi.ai
Meet CoPE-B: Frontier-Quality Content Classification You Can Self-Host
TL;DR: * Today we're releasing CoPE-B, our next-gen small language model for policy-adaptive content classification * CoPE-B-A4B (text-only) is open weights under Apache 2.0 and free to use * CoPE...
011
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 27/05/2026
@samidh.bsky.social and I are releasing CoPE-B today, the next version of our policy-adaptive content classifier. It delivers at-or-better-than-frontier classification while being self-hostable, faster, and cheaper to run. Full writeup with benchmarks at blog.zentropi.ai/meet-cope-b-.... 🧵 1/9
blog.zentropi.ai
Meet CoPE-B: Frontier-Quality Content Classification You Can Self-Host
TL;DR: * Today we're releasing CoPE-B, our next-gen small language model for policy-adaptive content classification * CoPE-B-A4B (text-only) is open weights under Apache 2.0 and free to use * CoPE...
25617
Reposted by Samidh
Dave Willner @dwillner.bsky.social · 11/05/2026
Very excited to have this public. The @oversightboard.bsky.social's use of Zentropi to study child marriage content on Meta platforms was very cool to support: blog.zentropi.ai/how-the-over...
blog.zentropi.ai
How the Oversight Board uses Zentropi to study policy impact at scale
The Oversight Board used Zentropi to analyze a large dataset of content that potentially violated Meta’s policies on human exploitation. The tool helped the Board reduce the project timeline from week...
293
Samidh @samidh.bsky.social · 11/05/2026
Check out how the @oversightboard.bsky.social used Zentropi to better understand how child marriage-related content manifests on Meta's platforms. Fantastic example of how advanced content labeling technologies can strengthen both our online and offline world. blog.zentropi.ai/how-the-over...
blog.zentropi.ai
How the Oversight Board uses Zentropi to study policy impact at scale
The Oversight Board used Zentropi to analyze a large dataset of content that potentially violated Meta’s policies on human exploitation. The tool helped the Board reduce the project timeline from week...
021
Samidh @samidh.bsky.social · 14/04/2026
It has been incredible partnering with character.ai since the very start of zentropi.ai. We're excited to share some details of that partnership with this case study. Anyone creating AI-powered systems might find it interesting! blog.zentropi.ai/how-zentropi...
blog.zentropi.ai
How Zentropi partners with Character.ai
Character.ai takes safety seriously. With millions of users creating and chatting with AI characters every day, the team invests heavily in systems that help protect their community — and they're alwa...
000
Reposted by Samidh
Katie Harbath @katieharbath.bsky.social · 24/03/2026
In 2017, it took us months to define “political ad” at Facebook. Recently, I built two political content classifiers in an afternoon using Zentropi AI created by @dwillner.bsky.social and @samidh.bsky.social Why AI content moderation is good, actually — in this week’s newsletter 👇
open.substack.com
Why AI Makes Content Moderation Better, Not Worse
Building a political content labeler with AI — what actually works
185
Samidh @samidh.bsky.social · 18/03/2026
This means anyone building an agent can give it a principled, consistent way to evaluate content instead of hoping the LLM gets it right. Small step, but it makes agents more trustworthy by default. Give the skill at try here and let us know how it goes: github.com/zentropi-ai/skills/
github.com
GitHub - zentropi-ai/skills: Agent skills powered by the Zentropi content classification engine
Agent skills powered by the Zentropi content classification engine - zentropi-ai/skills
131
Samidh @samidh.bsky.social · 18/03/2026
We just shipped something that helps with this. Zentropi is now an agent skill — your agent can classify content against plain-English policies in real time. It can even draft the policies itself if you describe what you're looking for. It is like having @dwillner.bsky.social on call 24/7.
111