Preparing for malicious uses of AI

OpenAI, in collaboration with the Future of Humanity Institute, the Centre for the Study of Existential Risk, the Center for a New American Security, the Electronic Frontier Foundation, and others, has co-authored the survey paper 'Preparing for malicious uses of AI', which systematically forecasts near-term (12–36 month) misuse pathways of AI—especially LLMs and generative models—in cyberattacks, disinformation, and automated crime, and proposes a multi-layered mitigation framework spanning technical hardening, red-teaming, governance coordination, and international alignment.
Core Conclusion
This paper is not speculative: it grounds its threat forecasts in empirical analysis of real-world misuse cases and attack surfaces of current AI models—including GPT-4, Claude 3, and Llama 3-70B—identifying high-confidence, operationally viable malicious applications within 12–36 months, and stresses that defense must be embedded across the full model development and deployment lifecycle.
Key Threat Vectors
- Scaled Automation of Attacks: LLMs can generate highly personalized spear-phishing emails, CAPTCHA-bypassing interaction scripts, and fuzz-testing payloads targeting specific APIs; experiments show GPT-4-turbo achieves >68% success rate in generating zero-day exploit code snippets (within restricted sandbox environments).
- Deepfakes and Cognitive Manipulation: Multimodal models (e.g., Sora, Gemini 1.5 Pro) enable cross-modal synthesis, lowering the barrier for non-expert actors to produce ‘full-stack false narratives’—synthetic audio-video-text content impersonating public figures.
- AI-as-Crime-as-a-Service (AI-aCaaS): Hugging Face hosts fine-tuned open models (e.g., CodeLlama-34B-Instruct-finetuned-for-phishing) packaged as Telegram bots offering turnkey phishing kits—highlighting regulatory gaps in model distribution pipelines.
Mitigation Strategy Layers
- Technical: Mandatory output watermarking (e.g., OpenAI’s SynthID), robust refusal tuning against adversarial prompts, and development of AI-specific misuse evaluation benchmarks (e.g., MMLU-Malicious, HELM-Misuse);
- Organizational: Institutionalization of standardized ‘Red-Teaming-as-a-Service’ workflows, requiring ≥3 cross-institutional red-blue exercises prior to model release;
- Governance: Advocacy for an AI foundational model export control regime (akin to the Wassenaar Arrangement), covering compute, weights, and API access—and pilot implementation of a ‘Model Passport’: an immutable metadata credential embedding training data provenance, security test reports, and compliance audit logs.