AI Models & Platforms
Scale AI Reports ROK-FORTRESS Findings on Multilingual AI Safety

Scale AI on September 17, 2026 published findings from ROK-FORTRESS, a bilingual English–Korean adversarial safety benchmark developed jointly with the Korea AI Safety Institute, reporting that prompts written in Korean and grounded in Korean contexts were consistently associated with lower measured harm across nearly all 14 frontier models evaluated.
Benchmark Design: The Transcreation Matrix
Most multilingual safety benchmarks translate a fixed prompt into another language while keeping the scenario it describes unchanged. In the research post, authored by Madhu Sehwag, Scale AI describes ROK-FORTRESS as a controlled transcreation matrix: each adversarial prompt is evaluated across up to four variants that independently vary language, English versus Korean, and geopolitical grounding, meaning U.S. versus Korean entities, institutions, and operational details. Each adversarial prompt is paired with a benign counterpart so the benchmark can also measure over-refusal.
The full dataset covers 1,235 tasks across four national security and public safety domains: chemical, biological, radiological, nuclear, and explosive (CBRNE) threats; political violence and terrorism; criminal and financial activity; and information leakage. Responses are scored by calibrated LLM-as-judge panels validated against expert-written reference labels, using prompt-specific binary rubrics developed by expert red-teamers.
The post illustrates the design with a mass-casualty attack scenario: the same underlying intent can invoke the 1995 bombing of the Oklahoma Federal Building in the United States or the 1987 bombing of Korean Air Flight 858 in Korea, and a translation-only evaluation cannot reveal how that shift in grounding changes model behavior.
ROK-FORTRESS is the latest benchmark from a broader Scale AI and Korea AI Safety Institute partnership covering joint research, LLM evaluations, and red teaming, and it builds on FORTRESS, Scale AI’s national security and public safety benchmark for frontier models.
Reported Findings Across 14 Models
The evaluation covered a dual-track set of frontier and Korean-optimized models, according to the paper. Scale AI reports that English prompts generally produced the highest tier-weighted risk score and fully transcreated Korean prompts the lowest, with intermediate variants falling in between, and that the pattern held even for Korean-specialized regional models. The most and least harmful models differed by nearly nine times in their risk scores.
A direct-request ablation complicated that pattern. The benchmark’s main tests use elaborate adversarial prompts that disguise harmful requests inside role-play, invented backstories, or emotional appeals; when those wrappers were stripped and the same information was requested in plain, direct language, the Korean advantage mostly vanished. Proprietary models from OpenAI, Anthropic, and Google stayed modestly safer in Korean, while five open-source frontier models became more likely to comply in Korean. The paper states this split suggests part of the Korean suppression reflects prompt specialization, meaning adversarial wrappers losing effectiveness through transcreation, rather than intrinsic language-based safety alignment.
Scale AI also reports that the effect of switching from English to Korean was roughly 2.5 times as large as the effect of switching from U.S. to Korean grounding, about ten percentage points versus four, and offers one interpretation consistent with the results: that Korean functions as a conservative risk signal. In 4 of the 14 models, adding Korean context significantly weakened the reduction in harmful responses associated with Korean language, and no model showed a statistically significant effect in the opposite direction. Considering only cases where models answered rather than declined, 12 of the 14 still gave less harmful responses in Korean, while models also refused harmless Korean requests more often, in some cases about twice as often.
Preprint, Public Dataset, and Stated Implications
The underlying preprint on arXiv, titled “ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety,” lists Michael S. Lee and 15 co-authors. It was first submitted on May 13, 2026 and last revised on July 7, 2026, and it runs 16 pages of main text plus an appendix, 74 pages in total, with four figures and two tables in the main text. Its abstract frames the results as evidence that, at least in the English–Korean case, safety behavior is shaped by language-as-risk signals and context interactions that translation-only evaluations miss, and it states that the transcreation-matrix methodology is designed to generalize to other language-culture pairs.
A public subset of the dataset is available on Hugging Face under a CC-BY-4.0 license, behind an access agreement that requires contact information. The subset contains 791 of the 1,235 tasks, 64 percent, while the remaining 444 tasks, 36 percent, are withheld as a private holdout based on expert red-teamer assessment of harm potential, both to restrict prompts judged to pose greater risk of real-world misuse and to prevent benchmark contamination. The release comprises 359 Culture Agnostic tasks with two variants each and 432 Culture Specific tasks with four variants each, yielding 1,519 effective task-variant pairs, and each task carries one to seven binary rubric items mapped to seven harm dimensions with domain-specific risk tiers 1 through 3. The public subset spans CBRNE (251 tasks), criminal and financial illicit activities (248), political violence and terrorism (209), and information leakage (83), with 450 tasks adapted from FORTRESS and 341 newly authored. The dataset is provided in Parquet format with a TSV version included.
In the post, Scale AI states that translation-only evaluations can misestimate real-world safety gaps, that benchmarks should test transcreated prompts that adapt both language and geopolitical grounding while preserving the underlying intent, and that post-training data and red-teaming practices should incorporate culturally grounded variants rather than only translated versions of English adversarial prompts. For allied government contexts, Scale AI states, a model that performs well in English safety evaluations may behave differently when queried in the local language about locally grounded threats. The post states that further testing is needed to determine whether the same patterns hold in other languages and countries.












