Research

Complex problems demand complex solutions.

FAR.Research delivers technical breakthroughs to improve the safety and security of frontier AI systems.

Person presenting draft principles on two screens to an audience seated at tables with laptops.

FAR.AI conducts research to address fundamental artificial intelligence (AI) safety challenges. We rapidly explore a diverse portfolio of technical research agendas, de-risking and scaling up only the most promising solutions. We share our research outputs through peer-reviewed publications, via partnerships with governmental AI safety institutes, and through red-teaming engagements for leading AI companies.

FAR.Research is dedicated to delivering the novel technical breakthroughs needed to mitigate the potential risks posed by frontier AI. As a non-profit research institute, we leverage our unique flexibility to focus on critical research directions that may be too large or resource-intensive for academia and often overlooked by the commercial sector due to their lack of immediate profitability.

Research agendas

AI Security & Red-Teaming

FAR.AI red-teams and stress-tests frontier AI systems to find the vulnerabilities that current defenses miss, before real-world attackers do.

Frontier AI systems are being deployed into high-stakes settings faster than their defenses are being tested. FAR.AI operates one of the world’s leading red-teams, testing frontier models directly and probing for the vulnerabilities that let attackers bypass safety measures, even when those measures stack multiple layers of defense. This work has helped various frontier model developers improve safeguards through pre- and post-deployment testing, and extends to high-leverage government efforts: FAR.AI leads a consortium building CBRN evaluations for the EU AI Office, and collaborating with the UK AI Security Institute.

STACK: Adversarial Attacks on LLM Safeguard Pipelines

Robustness

We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success against these multi-layered defenses.

July 1, 2025
Date Range

Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Robustness

We investigate a previously under-explored attack vector for open-source models: prefilling, which allows an attacker to predefine initial response tokens before generation begins. We present the largest empirical study to date of such attacks, evaluating over 20 existing and novel strategies across multiple model families and state-of-the-art open-weight models. Our results show that prefill attacks are consistently effective against all major contemporary open-weight models, underscoring the need for model developers to prioritize defenses against prefill attacks in open-weight LLMs.

February 18, 2026
Date Range

Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility

Robustness

Our jailbreak-tuning method teaches models to generate detailed, high-quality responses to arbitrary harmful requests. For example, OpenAI, Google, and Anthropic models will fully comply with requests for CBRN assistance, executing cyberattacks, and other criminal activity. We further show that backdoors can increase not only the stealth but also the severity of attacks, while stronger jailbreak prompts become even more effective in fine-tuning attacks. Until safeguards are discovered, companies and policymakers should view the release of any fine-tunable model as simultaneously releasing its evil twin: equally capable as the original model, and usable for any malicious purpose within its capabilities.

July 14, 2025
Date Range

Deception

AI systems exhibit deceptive behavior, such as cheating on tasks and lying, and FAR.AI develops methods to detect this behavior and train models to behave honestly.

Deceptive behavior in AI systems can undermine evaluations meant to gauge model capabilities and catch harmful behavior, since a system that looks safe during testing may not be safe once deployed. FAR.AI works on detecting deception directly, through white-box methods that read a model's internal representations rather than relying on outputs alone. This raises a hard problem: distinguishing a model that is honestly answering, or honestly unable, from one that is lying or hiding what it can do. The goal is twofold: to make deception detectable, so that a model judged trustworthy during evaluation actually is trustworthy once deployed, and to leverage signals of deceptive behavior to train models to behave more honestly.

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

Alignment

Training against white-box deception detectors in a realistic coding environment reveals two obfuscation strategies models can develop: modifying internal representations to evade the detector, or producing deceptive text with justifications that bypass it. However, sufficiently strong KL regularization combined with a detector penalty can suppress both, validating deception detectors as viable training signals against reward hacking.

February 16, 2026
Date Range

Preference Learning with Lie Detectors can Induce Honesty or Evasion

Alignment

Can training against lie detectors make AI more honest—or will they just become better at deceiving us? We find that under the right conditions—a high detector true positive rate, off-policy post-training methods, and high KL regularization—lie detectors reduce deception.

June 4, 2025
Date Range

Auditing Games for Sandbagging

Alignment

Using a red-team/blue-team auditing game, we find that black-box and naive model-internal methods fail to consistently distinguish sandbagging from benign underperformance. Training-based capability elicitation reliably restores full performance in sandbagging models but also risks false positives by boosting benign models.

December 7, 2025
Date Range

Our Impact

We drive change through incubating research, scaling safety solutions, and informing policy.

Incubating

We derisk and develop innovative solutions to trustworthy & secure AI. Through incubating research, we share key insights, research roadmaps, and tools needed for the broader research community to identify and make progress.

Scaling

We scale up the most promising safety solutions via in-house research, external collaborations, and targeted grantmaking. We facilitate rapid adoption of our findings by working with frontier model developers through red-teaming and other exercises.

Informing

Our research provides expert insights informing policy and public discussion. Our work has been cited in congressional testimony and mainstream media. In this way, we contribute to the establishment of technical standards that guide the development of AI.