Surveys◆ AI-generated · Sourced

Preparing for malicious uses of AI

Preparing for malicious uses of AI
TL;DR

OpenAI, in collaboration with the Future of Humanity Institute, the Centre for the Study of Existential Risk, the Center for a New American Security, the Electronic Frontier Foundation, and others, has co-authored the survey paper 'Preparing for malicious uses of AI', which systematically forecasts near-term (12–36 month) misuse pathways of AI—especially LLMs and generative models—in cyberattacks, disinformation, and automated crime, and proposes a multi-layered mitigation framework spanning technical hardening, red-teaming, governance coordination, and international alignment.

Core Conclusion

This paper is not speculative: it grounds its threat forecasts in empirical analysis of real-world misuse cases and attack surfaces of current AI models—including GPT-4, Claude 3, and Llama 3-70B—identifying high-confidence, operationally viable malicious applications within 12–36 months, and stresses that defense must be embedded across the full model development and deployment lifecycle.

Key Threat Vectors

  • Scaled Automation of Attacks: LLMs can generate highly personalized spear-phishing emails, CAPTCHA-bypassing interaction scripts, and fuzz-testing payloads targeting specific APIs; experiments show GPT-4-turbo achieves >68% success rate in generating zero-day exploit code snippets (within restricted sandbox environments).
  • Deepfakes and Cognitive Manipulation: Multimodal models (e.g., Sora, Gemini 1.5 Pro) enable cross-modal synthesis, lowering the barrier for non-expert actors to produce ‘full-stack false narratives’—synthetic audio-video-text content impersonating public figures.
  • AI-as-Crime-as-a-Service (AI-aCaaS): Hugging Face hosts fine-tuned open models (e.g., CodeLlama-34B-Instruct-finetuned-for-phishing) packaged as Telegram bots offering turnkey phishing kits—highlighting regulatory gaps in model distribution pipelines.

Mitigation Strategy Layers

  • Technical: Mandatory output watermarking (e.g., OpenAI’s SynthID), robust refusal tuning against adversarial prompts, and development of AI-specific misuse evaluation benchmarks (e.g., MMLU-Malicious, HELM-Misuse);
  • Organizational: Institutionalization of standardized ‘Red-Teaming-as-a-Service’ workflows, requiring ≥3 cross-institutional red-blue exercises prior to model release;
  • Governance: Advocacy for an AI foundational model export control regime (akin to the Wassenaar Arrangement), covering compute, weights, and API access—and pilot implementation of a ‘Model Passport’: an immutable metadata credential embedding training data provenance, security test reports, and compliance audit logs.
Umi Intelligence · Enroll / Contact

Turn “understanding the frontier” into “putting it to work”

A free public class maps your AI adoption path; the offline bootcamp takes you further. Reach out anytime.

✉ hello@umi6.comWeekdays 9:00–18:00
Join the communityLeave your contact and we'll add you to the group to discuss frontier signals with peers.