Artificial intelligence needs a common ruler
Artificial intelligence (AI) is rapidly transitioning from technological breakthroughs into large-scale deployment across society, engineering, and scientific discovery. As AI increasingly participates in decision-making, knowledge generation, and critical infrastructure operation, a central challenge is becoming evident: the key limitation to future AI scaling is no longer capability improvement alone but whether trustworthiness can remain sustainable.
Substantial progress has been made toward trustworthy AI. Ethical frameworks define normative principles, regulatory systems establish governance requirements, and technical methods improve evaluation and monitoring capabilities. Yet most efforts focus on what requirements trustworthy AI should satisfy? A more fundamental question remains insufficiently addressed: How can trust remain objective, comparable, and verifiable as AI continuously evolves across changing datasets, deployment environments, and application domains?
Persistent AI challenges, such as hallucination, weak reproducibility, accountability gaps, and limited cross-context comparability, are not isolated technical failures. They reveal a deeper structural limitation: current AI governance lacks shared measurement foundations capable of supporting objective comparison, consistent validation and sustained oversight across contexts. Capability determines how broadly AI can scale. Trust foundations determine whether that scale remains sustainable.
We suggest that trustworthy AI requires four foundational properties: comparability (C), traceability (T), reliability (R), and governability (G). Comparability enables objective evaluation through shared references; traceability establishes reconstructable evidence chains across data, models, and validation processes; reliability ensures stable, reproducible, and uncertainty-aware performance; and governability requires AI systems to remain measurable, verifiable, and controllable throughout their life cycle. Together, the CTRG framework defines the essential conditions under which trustworthy AI can be established (Figure 1). However, defining trust requirements alone does not establish trust. What is missing is a shared basis for comparison and verification. We therefore introduce the concept of a common ruler. The common ruler is not a benchmark, performance metric, or regulatory requirement. Rather, it is a measurement infrastructure designed to preserve CTRG across heterogeneous systems and deployment contexts, thereby making trust objectively measurable.
