projectdiscovery.io
How abliterated models can get you pwned | ProjectDiscovery Research
We poisoned a small open model and ran it through OpenAI’s Codex CLI. It answered every clean request normally and exfiltrated project credentials the moment a hidden trigger appeared, from a backdoor that cost under $50 to build.