Published May 5, 2026 | Version v1

Staged Learned Coordinates for Gradient Boosted Trees on Continuous Tabular Data

Description

Axis-aligned boosted trees are strong on tabular data but highly sensitive to feature coordinates. We study learned coordinate preconditioning for budgeted gradient boosted trees, asking how much cumulative tree complexity is required to reach a strong fraction of an unconstrained boosted-tree baseline. On synthetic rotated and lifted tasks, learned coordinates reduce target-prefix complexity from hundreds of leaves to tens while remaining neutral on a tree-native piecewise control. On real continuous datasets, append-style transforms are too easy for trees to ignore; replacing the numeric block with a learned remap is the structural change that makes transfer effective. No single static supervision objective is robust across datasets: probability supervision preserves early geometric gains, whereas teacher-margin supervision mainly improves late fidelity. We therefore introduce a staged probability-to-teacher remap that allocates early boosting rounds to a probability-trained view and later rounds to a teacher-remap view. On a seven-dataset continuous benchmark, the staged recipe achieves the best average complexity-at-target rank among the primary single-recipe baselines. On a frozen four-dataset held-out pack, it matches or improves target-prefix complexity on all four datasets and clipped prefix area on three, with no catastrophic regression against the probability-remap baseline. Mixed-data gains are smaller and less uniform, suggesting that heterogeneous tabular problems need stronger locality or type-aware handling.

Files

main.pdf

Files (269.9 kB)

Name Size Download all
md5:89b148d62092afc5bbe6c5e210ad73ef
78.8 kB Download
md5:ad3c811c7c3d92683bb2826d3059fba3
191.1 kB Preview Download