Staged Learned Coordinates for Gradient Boosted Trees on Continuous Tabular Data
Authors/Creators
Description
Axis-aligned boosted trees are strong on tabular data but highly sensitive to feature coordinates. We study learned coordinate preconditioning for budgeted gradient boosted trees, asking how much cumulative tree complexity is required to reach a strong fraction of an unconstrained boosted-tree baseline. On synthetic rotated and lifted tasks, learned coordinates reduce target-prefix complexity from hundreds of leaves to tens while remaining neutral on a tree-native piecewise control. On real continuous datasets, append-style transforms are too easy for trees to ignore; replacing the numeric block with a learned remap is the structural change that makes transfer effective. No single static supervision objective is robust across datasets: probability supervision preserves early geometric gains, whereas teacher-margin supervision mainly improves late fidelity. We therefore introduce a staged probability-to-teacher remap that allocates early boosting rounds to a probability-trained view and later rounds to a teacher-remap view. On a seven-dataset continuous benchmark, the staged recipe achieves the best average complexity-at-target rank among the primary single-recipe baselines. On a frozen four-dataset held-out pack, it matches or improves target-prefix complexity on all four datasets and clipped prefix area on three, with no catastrophic regression against the probability-remap baseline. Mixed-data gains are smaller and less uniform, suggesting that heterogeneous tabular problems need stronger locality or type-aware handling.
Files
main.pdf
Files
(269.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:89b148d62092afc5bbe6c5e210ad73ef
|
78.8 kB | Download |
|
md5:ad3c811c7c3d92683bb2826d3059fba3
|
191.1 kB | Preview Download |