Scale AI has published findings from ROK-FORTRESS, a bilingual adversarial safety benchmark developed with the Korea AI Safety Institute. The benchmark tests large language models using matched English and Korean prompts, with the Korean set deliberately grounded in local cultural and situational context.

The central finding is that Korean-language prompts were consistently more effective at exposing safety failures than their English equivalents. This suggests models are better hardened against adversarial phrasing in English but show gaps when the same intent is expressed in Korean and embedded in Korean context.

The work points to a practical limitation in current safety evaluation: benchmarks run mostly in English may overstate a model's real-world robustness. For users operating in other languages, the risk profile can be materially different, and the authors argue that safety assurance needs to account for linguistic and cultural variation rather than assuming English results transfer.