# FAR.AI > FAR.AI is a 501(c)(3) AI safety research and education nonprofit based in Berkeley, California (EIN 92-0692207), founded in 2022 by Adam Gleave and Karl Berzins. It works to ensure advanced AI is trustworthy, secure and beneficial, combining in-house technical research, a red-teaming practice, global convening, a co-working space, and targeted grantmaking. FAR.AI operates across three pillars: FAR.Research (in-house technical AI safety research), Events (the Alignment Workshop series, specialized workshops, and a weekly seminar), and Programs (FAR.Labs co-working space and a grantmaking program funded by a $12M Coefficient Giving grant). Research output is organised under four categories: Alignment, Interpretability, Model Evaluations, and Robustness. Publications are additionally tagged by topic, including Adversarial Robustness, Agentic AI, AI Governance & Policy, Benchmarks & Evaluations, Deception & Honesty, Feedback & Preference Training, Fine-Tuning, Jailbreaks & Red-Teaming, Mechanistic Interpretability, Open-Weight Models, Persuasion & Influence, Reinforcement Learning, and Scaling Trends. Notable work includes adversarial attacks on superhuman Go AIs (covered in Nature and cited in US Senate testimony), the STACK attack on layered LLM safeguard pipelines, jailbreak-tuning, the Attempt to Persuade Eval (APE) benchmark, and ClearHarm. FAR.AI leads a consortium building CBRN evaluations for the EU AI Office and collaborates with the UK AI Security Institute. ## Core pages - [Homepage](https://www.far.ai/): Overview of FAR.AI's mission, featured research, upcoming events, and programs. - [About](https://www.far.ai/about): Mission, three pillars, organisational history and timeline from 2022, leadership, and featured media coverage. - [Team](https://www.far.ai/team): Full staff directory grouped into Leadership, Executive Office, Operations, Technical Staff, Communications, Events & Programs, Collaborators, Advisors, and Board. - [Contact](https://www.far.ai/contact): Contact details and enquiry form. ## Research - [Research overview](https://www.far.ai/research): FAR.Research agendas, including AI Security & Red-Teaming and Deception, with featured publications and impact framing. - [All publications](https://www.far.ai/publications): Complete searchable and filterable library of research papers and technical reports, filterable by category, topic, and author. - [Google Scholar profile](https://scholar.google.com/citations?user=FVJ24k8AAAAJ): Citation record for FAR.AI research. Individual publications live at `https://www.far.ai/research/[slug]`. Representative examples: - [The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes](https://www.far.ai/research/the-obfuscation-atlas-mapping-where-honesty-emerges-in-rlvr-with-deception-probes): Training against white-box deception detectors, and the obfuscation strategies models develop in response. - [STACK: Adversarial Attacks on LLM Safeguard Pipelines](https://www.far.ai/research/stack-adversarial-attacks-on-llm-safeguard-pipelines): A layer-by-layer attack achieving a 71% success rate against defense-in-depth safeguards. - [Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility](https://www.far.ai/research/jailbreak-tuning-models-efficiently-learn-jailbreak-susceptibility): Fine-tuning attacks that produce full compliance with CBRN and cyberattack requests across frontier models. - [Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models](https://www.far.ai/research/prefill-level-jailbreak-a-black-box-risk-analysis-of-large-language-models): The largest empirical study to date of prefill attacks on open-weight models. - [Adversarial Policies Beat Superhuman Go AIs](https://www.far.ai/research/adversarial-policies-beat-superhuman-go-ais): The KataGo attack demonstrating that superhuman systems harbour exploitable failure modes. - [Open Problems in Mechanistic Interpretability](https://www.far.ai/research/open-problems-in-mechanistic-interpretability): Review of the current frontier and open challenges in mechanistic interpretability. - [Multi-Agent Risks from Advanced AI](https://www.far.ai/research/multi-agent-risks-from-advanced-ai): Taxonomy of miscoordination, conflict, and collusion failure modes in multi-agent systems. ## Events - [Events](https://www.far.ai/events): Recent and upcoming events, event formats, and the expression of interest form. - [Past events archive](https://www.far.ai/alignment-series-and-specialized-workshops): Full historical archive of the Alignment Workshop series and specialized workshops back to the inaugural 2023 San Francisco workshop. - [All recordings](https://www.far.ai/recordings): Talk recordings filterable by topic and year, featuring researchers including Yoshua Bengio, Sam Bowman, Neel Nanda, Dawn Song, Helen Toner, and Max Tegmark. - [YouTube channel](https://youtube.com/@FARAIResearch): Video archive of talks and workshop sessions. Event formats are the Alignment Workshop (worldwide flagship series), Specialized Workshops (focused deep dives such as ControlConf, AViD, and TIAP), and the FAR.AI Seminar (weekly, hosted at FAR.Labs and streamed). Individual events live at `https://www.far.ai/events/[slug]` and individual talks at `https://www.far.ai/event-recordings/[slug]`. ## Programs - [Programs](https://www.far.ai/programs): Consolidated page covering FAR.Labs and Grantmaking, with member organisations and current grants. - [FAR.Labs](https://www.far.ai/programs#farlabs): Berkeley co-working space for trustworthy and secure AI research, with 40+ members. Hosts visitors free for up to four weeks. - [Grantmaking](https://www.far.ai/programs#grantmaking): Targeted grants for academics and independent researchers, currently nomination-only, funded by a $12M Coefficient Giving grant. ## News and updates - [Blog](https://www.far.ai/blog): Announcements, research write-ups, event recaps, and press releases, filterable by topic. - [Newsletter](https://www.far.ai/newsletter): Email newsletter signup. Individual posts live at `https://www.far.ai/blog/[slug]`. ## Governance and transparency - [Transparency](https://www.far.ai/transparency): Financial statements, Form 990 filings, IRS 1023 application, consulting policy with independence safeguards, equal employment opportunity policy, anti-corruption policy, and whistleblower protections. - [Privacy policy](https://www.far.ai/privacy-policy): Data handling and cookie policy. - [Terms of service](https://www.far.ai/terms-of-service): Site terms and conditions. FAR.AI caps revenue from for-profit AI developers at 10% of total annual revenue to preserve research independence, and retains the right to publish insights from consulting engagements. ## Get involved - [Donate](https://www.far.ai/donate): Support FAR.AI's research and field-building work. - [Careers](https://www.far.ai/careers): Open roles. ## Optional - [Author profiles](https://www.far.ai/author/adam-gleave): Individual profiles for staff and contributors live at `https://www.far.ai/author/[slug]`. - [GitHub](https://github.com/AlignmentResearch): Open-source code and released models. - [X / Twitter](https://twitter.com/FARAIResearch) - [LinkedIn](https://www.linkedin.com/company/far-ai) - [Bluesky](https://bsky.app/profile/far.ai/)