OpenAI Red Teaming Network

OpenAI has officially launched an open call for the OpenAI Red Teaming Network, inviting domain experts to conduct structured adversarial safety evaluations on models including GPT-4 and o1 series to systematically identify and mitigate frontier LLM risks.
Background and Objective
OpenAI has announced the launch of the OpenAI Red Teaming Network—a collaborative external network composed of cross-disciplinary domain experts—focused on conducting structured, high-intensity red teaming against current and upcoming OpenAI models, including the GPT-4 series and o1 series. The goal is to uncover potential safety risks across dimensions such as bias, misuse, hallucination, jailbreaking, and agentic misalignment.
Participation Mechanism
- Recruitment targets independent experts or small teams with verified domain expertise (e.g., nuclear safety, bioinformatics, cybersecurity, cognitive psychology, legal ethics), not limited to AI safety researchers;
- Task-based engagement: OpenAI provides model access (via restricted APIs or sandbox environments), standardized testing frameworks, and evaluation benchmarks (e.g., MMLU-RedTeam, HarmBench v0.2), along with professional compensation;
- All findings must be reproduced and validated by OpenAI’s Safety Team; high-priority vulnerabilities are integrated into internal RLHF and Constitutional AI iteration pipelines.
Compliance and Transparency
- Testing scope is strictly confined to OpenAI-authorized scenarios; reverse engineering, model weight extraction, or unauthorized data exfiltration are prohibited;
- Output ownership follows the OpenAI Red Teaming Agreement; de-identified findings may be published on arXiv or co-authored in safety reports (e.g., Appendix C of the GPT-4 Technical Report);
- This network complements—not replaces—the Internal Red Team, expanding the breadth and depth of adversarial stress testing.