TabSets: calibration and test probabilities for set-valued evaluation of tabular classifiers
Authors/Creators
- 1. CentraleSupélec, Université Paris-Saclay, LISN
Description
The calibration and test probabilities of every run of a tabular classification benchmark, released so that prediction sets and their metrics can be recomputed without refitting a model. A conformal set is a function of four arrays: the class probabilities and labels on a calibration split, and the same on a test split. This deposit is those arrays, together with the software that reads them.
The benchmark comprises 218 classification datasets and 22 models from four families, among them 13 tabular foundation models, each evaluated with the same ten seeds. Of the 218, 178 are real datasets and 40 are a synthetic block generated for this benchmark; three of the real ones are Bayesian-network resamplings of small classical datasets and are much larger than their seeds; of those, only one has its own seed in the collection, so that pair is not independent of itself while a rank test treats datasets as independent. Dataset size does not indicate provenance: the largest datasets in the collection are real. Coverage is not uniform and the per-model dataset count is reported for every cell. Eleven datasets carry more than ten classes, and the in-context models measured on them cannot enter them: one refuses and names the limit, another builds a ten-output head and fails inside a kernel. That is a capability limit observed on the models that were run rather than a property claimed of the family, and it is not missing work. One foundation model is measured on the original datasets only. Anything computed across models should be computed on a complete submatrix: the datasets on which every model being compared has all ten seeds.
The calibration rows are held out of fitting, of tuning, of early stopping and of the in-context data, and each cell records which rows they were. The code that does it is included, with the line of each step documented.
Probabilities are float64 and are not rounded: a prediction set is a comparison against a threshold, and a point sitting on that threshold moves if the value it is compared to is shortened.
The data are released under CC BY 4.0. The software included in the source archive is released under the MIT licence, as stated by the LICENSE file it contains. The datasets themselves are not redistributed: each cell carries the row indices into its source dataset, and the synthetic datasets regenerate from their seeds with an included script. Model weights are not redistributed; every checkpoint is named in the model table.
Files
EXPORT.json
Files
(4.0 GB)
| Name | Size | |
|---|---|---|
|
md5:254aab4cec8c23d0a0f1c64e0175ee87
|
532.5 MB | Download |
|
md5:cac7f0354beeb73471aef4f28b2b5be7
|
380.7 MB | Download |
|
md5:d596a63d368e0510431547b74d8dad3a
|
378.3 MB | Download |
|
md5:76c48aa9245485d543464a8328dbef6b
|
394.0 MB | Download |
|
md5:505d38e219bb089ac61fcc8f7b0d22dd
|
383.7 MB | Download |
|
md5:6d428589078e9fa5667967a39a761e65
|
356.4 MB | Download |
|
md5:23645e18149ee082cc0bd1e26e746b9f
|
373.7 MB | Download |
|
md5:4316cceadb0b235ddc681fccedf74cb1
|
380.4 MB | Download |
|
md5:ac747b8f6e78ead8dd4de128ef6a17ec
|
325.2 MB | Download |
|
md5:cf69b0065e5d7bac68b7f67e50990e9a
|
310.7 MB | Download |
|
md5:920dc0491a45d5a39f0c59a87a779f11
|
221.4 MB | Download |
|
md5:71df472ed8d5b4ab7ee845006b7461eb
|
292 Bytes | Preview Download |
|
md5:f7c6db36dd919417a0a57fd18867c2b7
|
417.1 kB | Download |
|
md5:a6485ccfacda4a1eb1edce86b0612362
|
613.3 kB | Download |
|
md5:3b9b7823789cbc142aec9e66e893e07a
|
736.8 kB | Download |
|
md5:307cb3dd185fbd3b5a3d6b14453c0177
|
2.0 MB | Download |
Additional details
Related works
- Is derived from
- Conference paper: 10.14428/esann/2026.ES2026-261 (DOI)
- Preprint: arXiv:2605.28554 (arXiv)
- Is supplement to
- Software: https://github.com/jose-melo/tabsets (URL)