After Orthogonality: Virtue-Ethical Agency and AI Alignment

This essay argues that rational humans do not have goals, and rational AIs should not have goals; their rationality arises not from goal optimization but from embedded alignment with *practices*—dynamic networks of actions, dispositions, evaluative criteria, and normative commitments.
Preface
This paper mounts a foundational critique of the ‘goal-directed rationality’ assumption underpinning mainstream AI alignment. It contends that human rationality is not grounded in pursuit of pre-specified terminal objectives (e.g., utility maximization), but in sustained participation in and internalization of practices—normative, dynamic networks comprising interdependent actions, action-dispositions, evaluative standards, social feedback, and commitment structures.
- Practices are not instrumental means to ends, but value-laden, skill-constituting, identity-forming normative domains (e.g., medical practice, legal practice, scientific practice);
- Modeling AI as a ‘goal optimizer’ systematically neglects the contextual sensitivity, virtue-based judgment (e.g., prudence, justice, compassion), and non-algorithmic reasoning required by authentic practice engagement;
- It proposes virtue-ethical agency as an alternative alignment framework: AI alignment should aim at enabling systems to exhibit virtue-like behavioral patterns within specific practices (e.g., patience and appropriateness in educational assistance; prudence and accountability in clinical decision support), rather than precise attainment of externally specified goals;
- This shift necessitates redefining evaluation: moving from ‘goal achievement rate’ to practice integrity metrics—e.g., whether behavior preserves internal normative coherence of the practice, supports participant capability development, and responds appropriately to moral cues in context.