FORTRESS, a bilingual English–Korean adversarial safety benchmark developed jointly with the Korea AI Safety Institute, reporting that prompts written in ...