AGI Is Not Multimodal

Current multimodal generative AI models—such as GPT-4V, Gemini 1.5 Pro, and Llama-3.2-Vision—do not constitute a path to AGI; true AGI requires tacit, embodied, sensorimotor understanding grounded in real-world interaction, not just language or cross-modal statistical alignment.
Core Argument: Linguistic centrism obscures the embodied foundation of intelligence
As Terry Winograd stated: ‘In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence.’ This critique targets a fundamental flaw in mainstream AGI narratives: conflating the emergent capabilities of large language models (LLMs) and multimodal foundation models—e.g., OpenAI’s GPT-4V, Google’s Gemini 1.5 Pro, Meta’s Llama-3.2-Vision—with sufficient conditions for Artificial General Intelligence (AGI).
Key Distinction: Multimodality ≠ Embodied Intelligence
- Multimodal models process text, images, and audio jointly—but their ‘multimodality’ remains symbolic alignment and statistical co-occurrence, lacking the perception-action closed loop forged through sustained physical interaction;
- Embodied cognition posits that intelligence arises from real-time sensorimotor coupling with the environment—a capability absent in purely neural-network-based architectures deployed in silico;
- The arXiv paper Embodied Intelligence Requires Grounded Sensorimotor Loops (2024, 2402.13479) demonstrates that models trained without robotic platforms or physics-enabled simulators fail to acquire tacit knowledge such as grasp torque estimation, gravity intuition, or cross-material haptic transfer.
Methodological Warning: Evaluation paradigms must be redefined
Contemporary AGI progress is often benchmarked on static datasets like MMLU, MMStar, and VQAv2—yet these omit core dimensions of embodied intelligence: temporal extension, goal-directed trial-and-error, and causal task transfer. Hugging Face’s BEHAVIOR-Bench v1.0 (2024), the first benchmark to evaluate continuous action sequences in simulated environments, marks an initial step toward addressing this gap.