AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning
Public summary
* First AI system to emulate expert human workflow for oracle bone decipherment.
* Evidence-backed readings via morphology, context, and philology.
* Largest digitized corpus: 73,883 rubbings, 45,364 scans, and 23,755 studies.
* Transparent, interpretable reports ready for expert review.
Abstract
Approximately 3,000 of the 4,500 oracle bone script (OBS) characters remain undeciphered due to fragmentary inscriptions and sparse evidence. Current AI approaches fail to replicate expert workflows that integrate form analysis, contextual semantics, and philological reasoning. We introduce AlphaOracle, a human-workflow-inspired framework that systematizes OBS decipherment using the largest digitized corpus to date. Its multi-stage pipeline comprises (1) rubbing parsing, (2) radical-based morphological analysis with diachronic modeling, (3) contextual retrieval with semantic alignment, and (4) philological validation against classical sources. Each stage generates explicit, confidence-weighted evidence chains, culminating in interpretable reports for scholarly verification. Across multiple test characters, AlphaOracle’s readings strongly agreed with expert interpretations. In a study of 86 domain specialists, it reduced analysis time by 64%, and 79% of participants rated it highly useful. Notably, AlphaOracle resolves the character “勞” as a toponym or clan designation, offering concrete revisions to Shang administrative and social interpretations. These results suggest that computational methods aligned with philological practice can facilitate OBS research and provide a conceptual reference for studies of other undeciphered scripts.
