“Our study reveals prevalent brittleness in the presence of simple adversarial transformations: leading systems can guess the correct answer even with key inputs removed yet may get confused by the slightest prompt alterations…” #AI
www.nature.com/articles/s41...
nature.com
Evaluating the robustness and readiness of large frontier models in health AI applications - Nature Medicine
Adversarial evaluation of leading AI models uncovers gaps between benchmark success and robustness, highlighting limitations in current health AI benchmarks and their ability to capture clinically rel...