Oskar Hollinsworth

Member of Technical Staff

FAR.AI

Oskar Hollinsworth is a research engineer at FAR.AI, where has has worked on mitigating deception in LLMs, post-training infrastructure and scaling laws for adversarial robustness. Previously he studied how sentiment is represented in language models under Neel Nanda. Oskar had a first career as an algorithmic trader at Susquehanna International Group, Dublin.

Publications

STACK: Adversarial Attacks on LLM Safeguard Pipelines

Robustness

We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success against these multi-layered defenses.

July 1, 2025
Date Range

ClearHarm: A more challenging jailbreak dataset

Robustness

We introduce a novel jailbreak benchmark focused on unambiguously harmful questions such as constructing chemical, biological, radiological and nuclear (CBRN) threats, available on HuggingFace. We have found it is more challenging for attacks to elicit harmful responses from models on this benchmark than existing jailbreak benchmarks like StrongREJECT, Do-Not-Answer and SORRY-Bench. In particular this dataset is especially useful to understand which attack methods pose the greatest risk of eliciting egregiously harmful responses.

June 22, 2025
Date Range

Exploring Scaling Trends in LLM Robustness

Robustness

While larger language models exhibit impressive capabilities, they remain vulnerable to adversarial prompts. Empirical findings show that robustness against such attacks significantly improves with adversarial training, but not with model scaling alone.

July 25, 2024
Date Range

News

Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations

Red-Teaming & Evaluation

We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success.

July 1, 2025
Date Range

Does Robustness Improve with Scale?

Robustness & Security

Frontier LLMs like ChatGPT are powerful but not always robust. Scale helps with many things. We wanted to see if scaling up the model size can ‘solve’ robustness issues.

July 22, 2024
Date Range

ClearHarm: A more challenging jailbreak dataset

Red-Teaming & Evaluation

June 22, 2025
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI