The Humane Intelligence Quotient: A Leaderboard for Measuring AI Alignment to Humanity
Authors/Creators
Description
Over millennia, humanity has encoded and preserved its accumulated knowledge into the epistemologies of more than 6,000 distinct language cultures, the primary source of our collective creativity, problem-solving capacity, and intellectual growth. The erosion of this knowledge diversity through the systematic overrepresentation of a small number of dominant language cultures in AI development is well documented in scientific literature, yet quantified measures of that epistemic loss remain largely absent. Capability benchmarks, ethics audits, and safety frameworks all evaluate the model before deployment, in controlled conditions, independent of who the model will serve.
This paper fills that gap by reframing the unit of alignment measurement from the model to the coupled human-AI sociotechnical system. Any ecosystem-level alignment measure meeting four stated admissibility conditions and three stated comparative criteria decomposes, within that design space, into three structurally independent load-bearing components: Diversity, Fairness, and Agility. The threecomponent architecture is principled, non-redundant, and jointly sufficient within its scope. These three measurements combine into a single non-compensatory ranking for participating AI models in a public leaderboard, updated continuously from real user behaviour.
AI model providers list their models at their own expense. Users access the leaderboard free via authenticated credentials. Each user is assigned at random to an anonymous model. A single switch button lets them move to a different randomised anonymous model at the cost of their conversational state. This cost of abandonment reveals preference, producing per-cohort satisfaction scores across the world’s language communities without annotation, surveys, or researcherclassified criteria. Independent governance, with community co-governance over cohort inclusion as a conformance requirement, gives the outputs evidentiary standing.
A feasibility demonstration on the WildChat-1M corpus runs the full pipeline across more than one million real user conversations, identifying 74 language cohorts and confirming computational feasibility at scale. The paper then specifies the live deployment conditions under which comparative ranking matures into a declarative threshold, the minimum alignment standards that governance frameworks can specify, monitor, and enforce. The regulatory scaffolding already exists in the NIST AI RMF, the EU AI Act’s post-market monitoring provisions, and ISO/IEC 42001. What it has lacked is a continuously operating, independently governed instrument to populate it.
Files
HIQ-Main-v165.pdf
Files
(45.4 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:c727a6957845887db26ff95387a9bc10
|
44.8 MB | Download |
|
md5:969ed80a755e6641fe3fa0046714e28c
|
613.3 kB | Preview Download |
Additional details
Dates
- Submitted
-
2025-06-06Initial
- Updated
-
2025-07-13Significant updates
- Updated
-
2026-04Significant updates
- Updated
-
2026-06-19Significant updates for Global South
- Updated
-
2026-08-06Significant updates