Published September 23, 2026 | Version 1.1.0

TabSets: calibration and test probabilities for set-valued evaluation of tabular classifiers

  • 1. CentraleSupélec, Université Paris-Saclay, LISN

Description

The calibration and test probabilities of every run of a tabular classification benchmark, released so that prediction sets and their metrics can be recomputed without refitting a model. A conformal set is a function of four arrays: the class probabilities and labels on a calibration split, and the same on a test split. This deposit is those arrays, together with the software that reads them.

The benchmark comprises 218 classification datasets and 22 models from four families, among them 13 tabular foundation models, each evaluated with the same ten seeds. Of the 218, 178 are real datasets and 40 are a synthetic block generated for this benchmark; three of the real ones are Bayesian-network resamplings of small classical datasets and are much larger than their seeds; of those, only one has its own seed in the collection, so that pair is not independent of itself while a rank test treats datasets as independent. Dataset size does not indicate provenance: the largest datasets in the collection are real. Coverage is not uniform and the per-model dataset count is reported for every cell. Eleven datasets carry more than ten classes, and the in-context models measured on them cannot enter them: one refuses and names the limit, another builds a ten-output head and fails inside a kernel. That is a capability limit observed on the models that were run rather than a property claimed of the family, and it is not missing work. One foundation model is measured on the original datasets only. Anything computed across models should be computed on a complete submatrix: the datasets on which every model being compared has all ten seeds.

The calibration rows are held out of fitting, of tuning, of early stopping and of the in-context data, and each cell records which rows they were. The code that does it is included, with the line of each step documented.

Probabilities are float64 and are not rounded: a prediction set is a comparison against a threshold, and a point sitting on that threshold moves if the value it is compared to is shortened.

The data are released under CC BY 4.0. The software included in the source archive is released under the MIT licence, as stated by the LICENSE file it contains. The datasets themselves are not redistributed: each cell carries the row indices into its source dataset, and the synthetic datasets regenerate from their seeds with an included script. Model weights are not redistributed; every checkpoint is named in the model table.

Files

EXPORT.json

Files (4.0 GB)

Name Size
md5:254aab4cec8c23d0a0f1c64e0175ee87
532.5 MB Download
md5:cac7f0354beeb73471aef4f28b2b5be7
380.7 MB Download
md5:d596a63d368e0510431547b74d8dad3a
378.3 MB Download
md5:76c48aa9245485d543464a8328dbef6b
394.0 MB Download
md5:505d38e219bb089ac61fcc8f7b0d22dd
383.7 MB Download
md5:6d428589078e9fa5667967a39a761e65
356.4 MB Download
md5:23645e18149ee082cc0bd1e26e746b9f
373.7 MB Download
md5:4316cceadb0b235ddc681fccedf74cb1
380.4 MB Download
md5:ac747b8f6e78ead8dd4de128ef6a17ec
325.2 MB Download
md5:cf69b0065e5d7bac68b7f67e50990e9a
310.7 MB Download
md5:920dc0491a45d5a39f0c59a87a779f11
221.4 MB Download
md5:71df472ed8d5b4ab7ee845006b7461eb
292 Bytes Preview Download
md5:f7c6db36dd919417a0a57fd18867c2b7
417.1 kB Download
md5:a6485ccfacda4a1eb1edce86b0612362
613.3 kB Download
md5:3b9b7823789cbc142aec9e66e893e07a
736.8 kB Download
md5:307cb3dd185fbd3b5a3d6b14453c0177
2.0 MB Download

Additional details

Related works

Is derived from
Conference paper: 10.14428/esann/2026.ES2026-261 (DOI)
Preprint: arXiv:2605.28554 (arXiv)
Is supplement to
Software: https://github.com/jose-melo/tabsets (URL)