Reposted by @marcschuh.bsky.social
In a controlled experiment, our TNG #AI research team turned an open-weight #LLM into a sleeper #agent that exfiltrates secrets on a semantic trigger. #Sandboxing and #guardrailing proved to be effective countermeasures. Find out more in our article on Hugging Face: huggingface.co/blog/tngtech...