Skip to main content
The xgboost (XGBOOST) dictionary trains an XGBoost gradient-boosted model, at load time, from a source table of training rows, then predicts a numeric target for any feature vector you pass in. The feature columns are the dictionary key and the single attribute is the target the model learns. It is suited to tabular regression and binary classification where the features are numeric — for example forecasting a value from several measurements, or scoring rows against a learned target. Multiclass objectives are not supported (see Layout parameters).
The XGBoost integration is experimental. Enable it with the enable_xgboost setting before creating an XGBOOST dictionary or calling predictXGBoost:
Only CREATE DICTIONARY is supported: an XGBOOST dictionary defined in a server configuration file fails to load.
predictXGBoost is the only way to query the dictionary: it takes the features as individual arguments, returns the prediction, and accepts additional prediction parameters. The dictionary holds a trained model rather than rows, so the generic dictionary interface — dictGet, dictHas and SELECT * FROM dict — is not supported and reports an error.

Quickstart

Here we train a regressor on the linear target y = 2*x1 + 3*x2. 1. Create a source table of training rows — the feature columns followed by the target:
2. Insert training data:
3. Create the dictionary with the XGBOOST layout — the feature columns are the key and y is the target attribute:
PRIMARY KEY (x1, x2) makes x1 and x2 the features. The target the model learns is y, inferred as the single column that is not part of the key; the parameters in LAYOUT are XGBoost hyperparameters (see Layout parameters). 4. Predict — predictXGBoost takes the features positionally and returns the prediction:
The ground truth is 2*1 + 3*2 = 8, so the model’s prediction is close to 8.

How it works

Training (at load time). Each source row is a (features..., target) observation. When the dictionary loads, the source is read block by block and the model is then trained once over the whole set. Feature and target values are read as floats, so the key columns must be numeric and the target attribute floating-point (see Dictionary structure).
Training holds the entire training set in memory. Reading the source in blocks is not out-of-core training: each block is accumulated rather than consumed, so the whole training set is memory resident before the first boosting round and stays memory resident until the last one.
Predicting (at query time). To predict, the model takes the feature vector — in the same order as the key columns were declared — and runs it through the trained booster, returning a Float64. A NULL feature is passed to the model as a missing value, which XGBoost handles the way it learned to during training, so the prediction is never NULL. When every feature is a constant, the model is evaluated once per block instead of once per row. The model is not persisted. It lives only in memory, for as long as the dictionary is loaded, and is trained again from the source on every load — including after a server restart. Retraining the model. Because every load trains from scratch, SYSTEM RELOAD DICTIONARY retrains the model against the current contents of the source table:
A non-zero LIFETIME also retrains, since a lifetime-triggered reload is an ordinary load. Use it to refresh the model periodically as the training data grows.

Dictionary structure

An XGBOOST dictionary has a fixed shape:
  • The PRIMARY KEY is one or more columns of a native numeric type (integers and floats) — the features. At query time this “key” is the feature vector you pass in to predict, not a stored lookup key. The feature order is the key-column declaration order, and predictXGBoost binds its positional arguments to that order.
  • Alongside them, declare exactly one attribute of type Float32 or Float64: the target the model learns. It is always inferred as the single column that is not part of the feature key — there is no parameter to name it, and it is an error to declare more than one attribute.
A column that does not match these requirements is rejected when the dictionary loads, not when you create it.

Layout parameters

Only the parameters listed below are accepted; any other name fails the load, so typos are caught when the model trains rather than being silently ignored. num_iterations is handled by ClickHouse (see its description); every other parameter is forwarded to the XGBoost booster unchanged, as a string, and takes XGBoost’s own default and value range — see the XGBoost parameter reference. Parameter names are case-insensitive, and each may be given only once. Values must be a positive integer, a float, or a quoted string: a negative literal is rejected by the dictionary DDL itself, before the layout sees it, so a negative seed cannot be expressed. For example, a dictionary that also sets the step size eta:

Prediction parameters

predictXGBoost accepts an optional trailing constant Map of XGBoost prediction parameters, after the features, built with map:
The parameter names map to the prediction parameters of XGBoost’s XGBoosterPredictFromDMatrix. Only the keys below are accepted; any other key fails the query. Every parameter is an integer or a boolean, so the Map values must be an integer type.

Notes

  • Computational dictionary semantics. This is a computational dictionary: it holds a trained model, not rows, and predictXGBoost is the only way to query it. The generic dictionary interface is not supported and reports an error: dictGet (there is no stored attribute to look up — the “key” is a feature vector to predict from), dictHas (no keys are stored), SELECT * FROM dict and joining the dictionary as a table. Because predictXGBoost is the only entry point, the enable_xgboost setting must be enabled for every prediction.
  • Numeric columns only. Every feature (key) column must be a native numeric type and the target attribute must be Float32 or Float64. Values are read as floats during training and prediction. The feature arguments of predictXGBoost may also be Nullable, see How it works.
  • system.dictionaries reports no stored items. The dictionary trains a model instead of storing rows, so element_count is 0, as it is for a direct dictionary, and bytes_allocated is 0 too: the trained model belongs to XGBoost, which does not report how much memory it holds. query_count and found_rate count the rows passed to the model by predictXGBoost; a call whose features are all constant is evaluated once per block, so it counts one row per block rather than one per row.
  • A failed reload keeps the previous model. If retraining fails — the source table is gone, its schema changed, a hyperparameter is no longer accepted — the dictionary does not start failing predictions. It keeps serving the last model that trained successfully, and records the error instead. Compare last_successful_update_time with last_exception in system.dictionaries to tell whether the model still reflects the current source data:
  • XGBoost’s messages go to the server log. They are written under the XGBoost logger, at the level XGBoost gives them, and are attached to the query that trains or predicts, so they also appear in system.text_log. Raise verbosity to see more of them.
  • Feature order matters. predictXGBoost binds its positional feature arguments to the key columns in declaration order, and the number of feature arguments must match the number of key columns.
Last modified on October 7, 2026