NVIDIA releases open foundation model Kumo Tabular for tabular prediction without training
NVIDIA has released the open foundation model Kumo Tabular (28–215M parameters) for classification and regression on tabular data. According to the company, the model predicts without training, tuning, or feature engineering, and leads on the TabArena, BeyondArena, TALENT, and ScoringBench benchmarks.
The company NVIDIA has released Kumo Tabular, an open foundation model for prediction on tabular data, available on Hugging Face and GitHub under the OpenMDW-1.1 license, which allows commercial use. The model comes in three sizes with 28 to 215 million parameters.
According to NVIDIA, the model works on the principle of in-context learning: it receives a table with labeled rows and rows to be scored as input, and in a single pass returns a prediction (class probabilities for classification, a numerical value for regression) without the need for training, hyperparameter tuning, or manual feature engineering. The model was pretrained exclusively on artificially generated tables derived from structural causal models, not on real data.
According to NVIDIA, the model ranks first on four benchmarks: TabArena, BeyondArena, TALENT, and ScoringBench. For details on testing and other technical specifics, see the source article.
Why it matters
According to NVIDIA, tabular data (customer records, transactions, sensor data) is the foundation of enterprise machine learning, and predictions have so far typically relied on gradient-boosted trees with a cycle of data collection, feature engineering, and tuning for each new task. A foundation model that, according to the reported results, predicts directly from context without training could shorten this cycle for new classification and regression tasks on tabular data.
Release card
Kumo Tabular
NVIDIA
- Specifications
- 28 million to 215 million parameters (three sizes)
- Inputs
- input: a table with labeled rows and rows to be scored; output: class probabilities for classification or a numerical prediction for regression
- Licence
- OpenMDW-1.1
- Availability
- Available on Hugging Face and GitHub, run via an open-source library provided by the company NVIDIA
- Classification and regression on tabular data without training, tuning, or feature engineering
- Prediction of values for new table rows in a single pass using in-context learning
- Enterprise tasks such as predicting customer churn, failure, demand, or pricing based on tabular data
The source contrasts the model with the existing approach using gradient-boosted trees, which, according to the article, requires data collection, feature engineering, hyperparameter tuning, and training from scratch for each new task.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
Two audiences, two different impacts
What this means
For individuals
Data scientists and ML engineers can use the model directly for classification and regression tasks on tabular data without training, hyperparameter tuning, or manual feature creation.
For a business
According to the reported results, companies with tabular prediction tasks (churn, fraud, pricing, demand) could shorten the model development cycle if deploying this foundation model proves effective instead of the traditional approach using gradient-boosted trees and feature engineering for each new task.
DevelopmentCheck the original
Event sources
clearly official source · 1 vydavatel, 0 independent. We count feeds from the same owner only once.