Published July 24, 2026 | Version 1.2

Which Wordification Matters? A Nineteen-Recipe Sweep of the Interpretable-Topic-Model Framework on Hyperspectral Imagery

  • 1. CAOS open-research programme, Santiago, Chile

Description

Programme paper P3 of the CAOS_LDA_HSI series. The flagship paper (P1) introduces a twelve-axis evaluation framework for topic models on hyperspectral imagery and instantiates it on a single canonical wordification recipe (V1, band-frequency tokenisation). This manuscript asks a different question: how does the choice of wordification itself shape the conclusions of that framework.

We define a family of nineteen wordification recipes (V1-V15, V17-V20) spanning seven axes of design freedom (token alphabet, spatial vs spectral aggregation, local vs global vocabulary, document-length regime, signal transform, label-aware weighting, learnt representation) and sweep them through the framework on the six labelled scenes (Indian Pines, Salinas, Salinas-A, Pavia University, Kennedy Space Center, Botswana) and five HIDSAG mineralogical subsets.

The headline finding is that there is no universal winner: V1 wins F-2 coherence on 2/6 scenes; V3 (joint band-bin Cartesian vocabulary) wins F-7 topic-label normalised MI on 2/6; V12 (Gaussian-mixture tokens) wins F-2 and F-7 on 2/6 each; V20 (mutual-information-weighted bands) wins F-2 (0.88) and F-7 NMI (0.44) together on Indian Pines and is among the most counterfactually robust recipes at Q=8 (with V12 and V3). A Q-sweep (Q in {8,16,32}) identifies V20, V2 and V8 as the recipes whose F-7 NMI and F-2 coherence both improve monotonically, with the V20 vs V12 ranking inverting between Q=8 and Q=32. A four-backbone extension of F-7 identifies V8 (NFINDR endmember-fraction) as the most cross-backbone-consistent recipe. We report the full 19 x 12 x 11 result matrix, identify five axis-recipe affinities, and recommend V1 retain its canonical status as a reproducibility default rather than a universal best.

Code and derived artefacts: https://github.com/fsantibanezleal/CAOS_LDA_HSI . Interactive web application: https://lda-hsi.fasl-work.com . Manuscript sources: https://github.com/fsantibanezleal/CAOS_LDA_HSI_Paper .

Funding: The Advanced Mining Technology Center (AMTC) Basal project (ANID/PIA Project AFB220002) and ANID FONDECYT Postdoctorado 3220094.

Files

lda-hsi-journal-wordification-sweep-v1.2.pdf

Files (949.2 kB)

Name Size Download all
md5:0355f4c554e6809a0282ad933b8c7f10
949.2 kB Preview Download

Additional details