How Countries Can End the Capability Overhang

Our latest report reveals stark disparities in advanced AI adoption—such as GPT-4, Gemini, and Llama series—across nations, and proposes structural interventions—including national AI readiness assessment frameworks, standardized public-sector AI procurement criteria, and cross-sector AI skill certification systems—to translate model capability into measurable productivity gains.
Core Finding: Capability Overhang ≠ Productivity Gain
The global AI landscape exhibits a pronounced ‘capability overhang’: foundational models like OpenAI’s GPT-4, Google’s Gemini 2.0, and Meta’s Llama 3 significantly outpace most countries’ institutional capacity for deployment and governance. Based on empirical analysis across 38 OECD members and emerging economies, the report finds that only 12% of nations possess integrated AI readiness across education, regulation, procurement, and workforce reskilling—while 67% rely on isolated pilots without scalable pathways.
Key Intervention Levers
- National AI Readiness Assessment Framework: Developed by the OECD AI Policy Observatory, this framework comprises five dimensions—data governance maturity, public-sector AI procurement compliance rate, AI-related occupational certification coverage, SME AI tool penetration, and AI ethics review mechanism robustness—and has been piloted in Canada, South Korea, and Singapore;
- Standardized Public-Sector AI Procurement: The EU AI Act’s implementation guidelines mandate three hard requirements for all large-model procurements: (i) verifiable Model Cards and Data Cards; (ii) support for on-prem or sovereign-cloud inference (e.g., via Hugging Face TGI or vLLM); and (iii) conformance with ISO/IEC 42001;
- Cross-Sector AI Skill Certification System: Germany’s Federal Ministry of Education and Research (BMBF), in collaboration with Fraunhofer IAIS, launched the ‘AI Literacy Passport’, covering seven occupational roles—from civil servants to manufacturing technicians—with hands-on assessments aligned to Llama 3-70B and Gemma-2-27B tasks including prompt engineering, RAG pipeline construction, and safety-focused fine-tuning;
- Open-Benchmark–Driven Policy Feedback Loop: The report recommends integrating national AI deployment metrics—e.g., government chatbot accuracy rates or automated health insurance claim approval coverage—into public leaderboards such as the Hugging Face Open LLM Leaderboard and EleutherAI’s HELM, enabling evidence-based policy iteration.