On September 29, 2026, NVIDIA published a foundation model on Hugging Face that almost nobody in the AI press treated as news. It is called Kumo Tabular, it ships in three sizes from 28 million to 215 million parameters, and it does the thing the transformer era never delivered for the data most companies actually own. You hand it a table of labeled rows, and it predicts the labels for the new rows in a single forward pass. No training run. No tuning. No feature engineering.
It was pretrained entirely on artificial data. It is released under OpenMDW-1.1, a permissive license built for machine learning models that sanctions commercial use outright. It ranks first on four benchmarks: TabArena, BeyondArena, TALENT, and ScoringBench.
Here is why you should care. Every hype cycle in AI has been about text, images, and code. Meanwhile the prediction work that actually pays enterprise bills, meaning churn, fraud, lead scoring, and demand forecasting, still runs on hand-trained gradient-boosted trees. The model layer just showed up to that party. If you run a data team, this is the release of the week, and it was covered as a footnote.
THE GAP NOBODY NOTICED
Tabular data never made it into the foundation-model era, and the reason is boring. Transformers ate text, images, and code because those formats gave the model a sequence to attend over. A row in a database table gives you a different shape: heterogeneous columns, mixed types, missing values, and no natural order to read them in. So structured rows sat in warehouses and got handled the way they have been handled since the first XGBoost-style pipelines, with a fresh training run for every new question.
That lifecycle has barely changed in two decades. A new question means collecting labels, engineering features, searching hyperparameters, validating, and then deploying a model that knows nothing about tables in general and learns each task from scratch. The model is disposable by design. You pay the full build cost every time you ask a new question of the same data.
Kumo Tabular changes the unit of work. One pretrained model handles many tables, and inference per row replaces training per problem. That is the whole shift, and it is an engineering shift before it is anything else. You cannot do anything serious in AI without solid engineering, and this removes an entire category of engineering toil from the routine cases.
PhantomByte has made this argument twice already, on different ground. Your Model Is Not Your Product said the weights are the commodity and your pipeline is the asset. Three AI Stacks. Your Data Already Picked One. said your data's shape decides which stack you can run. This is the structured-data edition of the same thesis, and it arrives with the pipeline's build cost deleted.
Three terms, one sentence each, so nothing below is ambiguous.
- A foundation model is pretrained once and reused across many tasks rather than rebuilt per task.
- A forward pass is one trip through the model to get an answer, with no weight updates along the way.
- Tabular means rows and columns, the shape of a database table, which is the shape of nearly every dataset an enterprise actually owns.
WHAT IT ACTUALLY DOES
The mechanics are simple enough to state in a line. You supply a table with labeled rows, plus the rows you want predictions for, and the model returns class probabilities or numeric predictions in a single pass. Classification and regression are both supported, trained as separate models.

Under the hood it is a Transformer built around the structure of a table, using column attention, row attention, and in-context attention as introduced in TabICL and TabPFN. Column attention looks down a single column to learn whether a value is typical or extreme for that feature. Row attention looks across the tokens of one row to learn how features interact. The final Transformer relates the labeled context rows to the query rows, and query rows attend only to the context, never to each other. That last property matters operationally, and it has a name. Because the context never looks at the queries, its keys and values are computed once, which is a KV-cache that can be pre-computed and reused across batch requests instead of recomputed for every prediction. Query rows then use Test-GQA to shrink the cache each prediction reads.
The release also includes a length-aware attention temperature. This is the detail most coverage skipped and the one that decides whether the thing works on your data. Softmax attention spreads out as the number of keys grows, so attention that is sharp over a few hundred rows can dissolve over tens of thousands. Kumo Tabular scales every query by a temperature that grows with the logarithm of the number of keys, which keeps attention sharp as tables grow longer or wider.
Three sizes run from 28 million to 215 million parameters. The small end fits on hardware you already own. This is not a cluster commitment, and it is not a per-token bill.
The base model was pretrained only on artificial data. Every training table was sampled from a structural causal model, with a random causal graph, randomly drawn functions at each node, and injected missing values, outliers, and high-cardinality columns. NVIDIA then trained on roughly 35 million artificial tables for the small size, 71 million for medium, and 137 million for large. No customer data entered the base model. For regulated teams, that changes the privacy conversation from a data-processing question into a much simpler one.
It runs through NVIDIA's open-source structured data models library, and the license is OpenMDW-1.1. That is a permissive license introduced in 2025 by the Linux Foundation and the PyTorch Foundation, designed to cover weights, software, and documentation in one instrument. Commercial use is sanctioned, not gray-area.
THE DEPLOYMENT BOUNDARY: CONTEXT ROWS AND MEMORY
Architects need numbers before they commit, so here are the two constraints that decide whether this deploys.
Start with context rows, because that is the real ceiling. Kumo Tabular learns from the labeled rows you put in its context, and attention cost scales with how many keys the softmax has to spread over. NVIDIA's training recipe walks the context up in three stages, from tables of 1,024 rows, to a context varied between 400 and 10,240 rows, to a final stage that extends the context to 60,000 rows, all with up to 100 columns. That range is your practical envelope, and it is documented rather than guessed. The length-aware attention temperature exists precisely because an inference table can be far larger than a typical training table, and NVIDIA states plainly that accuracy may degrade on tables far beyond the training ranges. The honest reading is this: a few thousand context rows is comfortable, tens of thousands is where the mechanism was explicitly engineered to hold, and past that you are testing on your own held-out data whether the envelope still covers you.
Then memory, where the published weights remove the guesswork. The small variant ships a classifier at about 105 megabytes and a regressor at about 109 megabytes. The medium pair runs about 235 and 239 megabytes. The large pair runs about 815 and 823 megabytes. The entire model repository, all three sizes and both task types, is roughly 2.3 gigabytes. That means the 28 million parameter model runs comfortably on a laptop GPU, and even the 215 million parameter model fits on a single consumer card, because these weights are smaller than a mid-size vision model. The footprint is dominated by your context, not by the parameters, which is the opposite of what you are used to sizing for.
Put the two constraints together and you get the deployment shape. The parameters are cheap and the context is the budget. Size the card for the rows you intend to feed it.
WHERE IT FITS IN AN AGENT RUNTIME
PhantomByte writes about agentic systems, so this release has to be placed where it actually lands, which is inside the runtime rather than beside it.
Two of the worst patterns in agent engineering both come from mishandling structured data. The first is an agent that writes Python at runtime to spin up an XGBoost or scikit-learn pipeline, which buys you a training job, a dependency graph, and a model artifact to babysit, all inside a single tool call. The second is an agent that dumps raw CSV rows into an LLM context window and asks the model to reason over them, which burns per-token cost on numbers a language model reads badly, and repeats the whole payload on every turn of the loop.
Kumo Tabular gives you a third option, and it is the one worth building. Stand the model up as a local microservice, expose it as a tool, and let the agent route structured prediction to a 28 million to 215 million parameter service running on the same box. The agent passes labeled context rows and query rows. The service returns class probabilities or numeric predictions. The call is a forward pass and not a job, so the agent gets an answer inside the same reasoning step instead of launching an asynchronous build and polling it until the artifact appears.
That is zero-shot structured tool execution, and it is the piece agent architecture has been missing. Agents already have text tools, code tools, and retrieval tools. Now they have a prediction tool for the shape of data most enterprises actually store, and the cost profile is a local forward pass instead of a per-token bill. If you are building an agent that touches business data, this is the cheapest new tool you can hand it this quarter.
THE FOUR TESTS IT PASSED
Ranked first on four benchmarks sounds like a headline until you know what each one measures. Here is what first means on each.
TabArena is a living benchmark for tabular machine learning maintained in the AutoGluon project, covering 51 curated datasets with multiple splits each and more than 25 methods, with a shared tuning protocol applied to every method. NVIDIA reports Kumo Tabular first overall with an Elo of 1950.
BeyondArena extends the same codebase past the independent-and-identically-distributed case, across 142 datasets spanning IID, temporal, and grouped task types, from tiny tables to a million rows, low and high dimensionality, and messy features like text and high-cardinality categories. NVIDIA reports an Elo of 1418 and an Improvability score of 7.78 percent, placing first.
TALENT is the Tabular Analytics and Learning Toolbox from the LAMDA group at Nanjing University, a deep-learning-focused suite of more than 20 tabular prediction methods with a unified interface. NVIDIA reports the top overall ranking there across classification accuracy, classification log-loss, and regression RMSE.
ScoringBench, from Jonas Landsgesell and colleagues, evaluates regression models with proper scoring rules that test the full predictive distribution rather than only a point estimate, using 97 datasets and metrics like CRPS and the interval score. Kumo Tabular Large and Medium rank first and second on average rank.
Now the blunt part. First on benchmarks is not first on your warehouse. Benchmarks are the entry ticket, not the verdict, and NVIDIA's own model card says so. Accuracy may degrade on tables far beyond the training ranges, or when the query rows come from a different distribution than the context rows.
There is a second caveat with real teeth. The BeyondArena paper itself, from Lennart Purucker and a long list of co-authors, finds that tabular foundation models dominate tiny to medium IID data while traditional tree-based and deep learning models still dominate on non-IID, large, high-dimensional, and high-cardinality datasets. NVIDIA ranks first on that benchmark in its own evaluation of the full leaderboard. The broader finding stands beside it. If your workloads are large, temporal, or grouped, read the fine print before you cancel your tree pipeline.
One more limitation, from the model card. A single forward pass covers up to 10 classes, which the library extends to any number of classes using error-correcting output codes. And the model works on numerical and categorical columns only, so text, images, and timestamps get converted into features first through built-in preprocessing recipes.
WHEN TO USE IT: A DECISION RULE
This is the part worth quoting back, so here is the decision rule by name. Call it the Forward-Pass Test.
Use the pretrained model, with no training run, when three conditions hold.
- The task is routine classification or regression on structured rows.
- You have labeled rows but no team time for a training pipeline.
- And you need a baseline today rather than a tuned system next quarter.
Keep the hand-trained pipeline when any of these is true.
- The distribution is adversarial or drifting fast, which is fraud that fights back and re-fits against your model.
- You have domain features that encode knowledge the table itself does not carry, like a risk score from a human analyst or a rule learned outside the data.
- Or a regulated deployment requires a fully explainable model with documented training lineage, because a 215 million parameter Transformer has no gradient-boosted tree's tidy feature-importance story.
The Forward-Pass Test is deliberately narrow. It does not say foundation models have won structured data. It says a specific class of routine task just lost its build-and-tune step, and you should know exactly which class that is.
WHAT TO DO TODAY
- Pull the weights from Hugging Face and run one table you already have through the structured data models library. It installs with a single pip command and the documented path goes from a pandas DataFrame to a prediction in a few lines.
- Pick one routine classification task your team solved with a hand-trained model and score Kumo Tabular against it on a held-out split. Compare on your data, not on a leaderboard.
- Stand it up behind a local endpoint and call it from one agent tool. Measure whether the forward pass beats the CSV-into-context pattern on both latency and token spend.
- Count your context rows before you size the hardware. The parameters are cheap, and the rows are the budget.
- Check the OpenMDW-1.1 terms against your commercial use case before production, not after.
- Benchmark the small size first. Escalate to the 215 million parameter model only if accuracy demands it, because the efficiency argument is the point.
- Kill one stalled ML backlog item this week by testing it with zero training budget.
THE UNCOMFORTABLE QUESTION
Your data team has been hand-training models for problems a pretrained model now solves in one forward pass. How much of last year's ML roadmap was engineering, and how much was ritual?
The answer shows up in infrastructure spend. Standardizing on a forward-pass model does not just delete engineering ritual, it deletes the disposable training pipelines that ritual fed. Every per-problem model carries a build job, a scheduling slot, an artifact store, a validation harness, and an on-call surface that somebody has to keep alive. Retire the per-problem model and that whole layer retires with it. You stop paying to rebuild a model every time you ask a new question of data you already own.
The model layer just reached the data most companies actually own, and almost nobody noticed. That is the tell.
Get More Articles Like This
Getting your AI agent setup right is just the start. I'm documenting every mistake, fix, and lesson learned as I build PhantomByte.
Subscribe to receive updates when we publish new content. No spam, just real lessons from the trenches.
Build Real AI Infrastructure
PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.
