AI Models & Platforms

NVIDIA Releases Open Kumo Tabular Model for Tabular Prediction

mm
Add Unite.AI to your preferred sources on Google

NVIDIA on September 29, 2026 released Kumo Tabular, an open foundation model for tabular classification and regression that predicts the labels of new rows in a single forward pass, with no training, tuning, or feature engineering. The model was pretrained entirely on artificial data and comes in three sizes ranging from 28 million to 215 million parameters.

Kumo Tabular is part of the NVIDIA Kumo Structured model collection, according to the announcement on Hugging Face’s blog. Weights are available from the model’s Hugging Face listing under the OpenMDW 1.1 license, which NVIDIA said permits commercial use, and inference runs through the open-source structured-data-models library, which downloads the weights on first use and provides the preprocessing, ensembling, and many-class handling used in NVIDIA’s evaluations.

NVIDIA frames the release against two decades in which enterprise tabular tasks such as churn, default, demand, and price prediction have relied on gradient-boosted trees, each new question requiring its own cycle of labeling, feature engineering, tuning, validation, and deployment. The company says the model applies the in-context learning pattern of large language models to tables, reading a labeled table as context to predict new rows’ labels directly.

Architecture and In-Context Prediction

Kumo Tabular is a Transformer built around the structure of a table, using column, row, and in-context attention as introduced in TabICL and TabPFN. To predict a label, the model establishes what each value means within its column, how a row’s columns interact, and how the labeled context rows relate to query rows whose labels are unknown.

In the cell-embedding stage, a group of cells becomes a token. Numerical and categorical values pass through Fourier features, sines and cosines of learned frequencies, with separate weights for each type. Missing values are treated specially and need no imputation, and every token in the context receives a label embedding. Each row is then compressed into an embedding by alternating column attention, an induced self-attention whose cost grows linearly with the number of rows, with row attention that learns how a row’s features interact, using rotary positions to tell columns apart. Four learnable CLS tokens join each row and act as its final readout.

A final Transformer operates on the row embeddings. Context rows attend to each other while query rows attend only to context rows, so each prediction depends only on the context and the row itself, and the context’s keys and values are computed once and reused for later predictions. Query rows use Test-GQA, which shrinks the cache each prediction reads. A length-aware attention temperature scales every query by a factor that grows with the logarithm of the number of keys, with a coefficient learned separately for each head, keeping attention sharp when an inference table is far longer than a typical training table. A head outputs class probabilities for classification or 999 quantiles for regression, yielding a point prediction and an uncertainty estimate.

Pretraining on Artificial Tables

Every training table is sampled from a Structural Causal Model in six steps. The generator draws a configuration for the whole table, then links hidden variables with a random causal graph evaluated from root to leaf using randomly drawn functions at each node, including linear maps, small neural networks, trees, and Gaussian processes. Some nodes become numerical or categorical columns, one becomes the target, and the rest stay hidden, like the unmeasured causes behind real data. Post-processing correlates groups of columns, clips outliers, and injects missing values, and a tree-ensemble check discards any table without a learnable signal.

NVIDIA said it built the imperfections of real tables into the generator: values go missing in several patterns, some features are coarsened so that duplicate rows can disagree on their label, some categorical columns carry many levels, and regression targets can be heavy-tailed. Training uses a cross-entropy loss for classification and a quantile loss for regression, with the two tasks trained as separate models, and runs in three stages: tables of 1,024 rows and up to 100 columns, contexts varying from 400 to 10,240 rows, and finally contexts of up to 60,000 rows. The Small, Medium, and Large variants saw about 35 million, 71 million, and 137 million artificial tables, respectively.

NVIDIA’s Reported Benchmark Results

NVIDIA reports that it ran all three sizes with default settings against the full TabArena leaderboard, which spans tuned gradient-boosted trees, AutoGluon, and recent tabular foundation models, and that Kumo Tabular ranks first overall with an ELO of 1950 while running 26 times faster than LimiX-2 under a uniform single RTX 6000 Pro evaluation setup. The company said the three sizes establish a new state of the art on the accuracy-efficiency Pareto front.

On BeyondArena, NVIDIA reports an ELO of 1418 and an Improvability score of 7.78%, placing first on that leaderboard. On TALENT, the company reports the top overall ranking across classification accuracy, classification log-loss, and regression RMSE, with average ranks of 6.67, 3.98, and 4.22, respectively. On ScoringBench, a benchmark for predictive distributions, NVIDIA said Kumo Tabular-Large and Kumo Tabular-Medium rank first and second on average rank.

Limitations and Availability

Kumo Tabular works on numerical and categorical columns only; text, images, or timestamps can be turned into features through built-in preprocessing recipes. A single forward pass covers up to 10 classes, and the library extends the model to any number of classes with error-correcting output codes. NVIDIA cautions that accuracy may degrade on tables far beyond the training ranges or when query rows come from a different distribution than the context rows, and it advises validating accuracy and calibration on held-out data before deployment.

The structured-data-models package is GPU-native, requires Python 3.11 or later and PyTorch 2.7 or later, and recommends cudf for CUDA workloads so dataframe operations stay on the GPU. It hosts reference implementations of the tabular foundation models TabICLv2, KumoTabular, and TabFM alongside the relational model KumoRelational, and NVIDIA-authored code in the repository is licensed under Apache 2.0.

NVIDIA said the training recipe and the artificial data generators behind Kumo Tabular will be released soon.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.