> ## Documentation Index
> Fetch the complete documentation index at: https://private-7c7dfe99-vortex-format.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# XGBoost dictionaries

> Configure XGBOOST dictionaries to train a gradient-boosted model and predict a numeric target.

export const CloudNotSupportedBadge = () => {
  return <a href="https://clickhouse.com/docs/products/cloud/guides/cloud-compatibility#list-of-unsupported-features" className="cloudNotSupportedBadge">
            <div className="cloudNotSupportedIcon">
            <svg width="16" height="16" viewBox="0 0 16 16" fill="none" xmlns="http://www.w3.org/2000/svg">
                <path strokeWidth="1.5" d="M6.33366 12.6666L12.3739 12.6667C13.6593 12.6667 14.7073 11.6187 14.7073 10.3334C14.7073 9.04804 13.6593 8.00003 12.3739 8.00003C12.3739 8.00003 12.3337 7.66659 12.0003 7.33325M10.667 5.33322C8.00033 2.33325 4.45395 4.78537 4.14195 6.68203C2.55728 6.7627 1.29395 8.06203 1.29395 9.6667C1.29395 11.3234 2.66699 12.6666 4.00033 12.6666" stroke="currentColor" strokeLinecap="round" strokeLinejoin="round" />
                <path strokeWidth="1.5" d="M2.66699 14L12.0003 4.66663" stroke="currentColor" strokeLinecap="round" strokeLinejoin="round" />
            </svg>

        </div>
            Not supported in ClickHouse Cloud
        </a>;
};

<CloudNotSupportedBadge />

The `xgboost` (`XGBOOST`) dictionary trains an [XGBoost](https://xgboost.readthedocs.io/) gradient-boosted model, at load time, from a source table of training rows, then predicts a numeric target for any feature vector you pass in. The feature columns are the dictionary key and the single attribute is the target the model learns.

It is suited to tabular regression and binary classification where the features are numeric — for example forecasting a value from several measurements, or scoring rows against a learned target. Multiclass objectives are not supported (see [Layout parameters](#layout-parameters)).

<Note>
  The XGBoost integration is experimental. Enable it with the `enable_xgboost` setting before creating an `XGBOOST` dictionary or calling `predictXGBoost`:

  ```sql theme={null}
  SET enable_xgboost = 1;
  ```

  Only `CREATE DICTIONARY` is supported: an `XGBOOST` dictionary defined in a server configuration file fails to load.
</Note>

[`predictXGBoost`](/reference/functions/regular-functions/machine-learning-functions) is the only way to query the dictionary: it takes the features as individual arguments, returns the prediction, and accepts additional [prediction parameters](#prediction-parameters). The dictionary holds a trained model rather than rows, so the generic dictionary interface — [`dictGet`](/reference/functions/regular-functions/ext-dict-functions#dictGet), `dictHas` and `SELECT * FROM dict` — is not supported and reports an error.

<h2 id="quickstart">
  Quickstart
</h2>

Here we train a regressor on the linear target `y = 2*x1 + 3*x2`.

**1. Create a source table** of training rows — the feature columns followed by the target:

```sql theme={null}
CREATE TABLE training_data (x1 Float64, x2 Float64, y Float64)
ENGINE = MergeTree ORDER BY tuple();
```

**2. Insert training data:**

```sql theme={null}
INSERT INTO training_data
SELECT number AS x1, number * 2 AS x2, 2 * x1 + 3 * x2 AS y
FROM numbers(100);
```

**3. Create the dictionary** with the `XGBOOST` layout — the feature columns are the key and `y` is the target attribute:

```sql theme={null}
CREATE DICTIONARY model (x1 Float64, x2 Float64, y Float64)
PRIMARY KEY (x1, x2)
SOURCE(CLICKHOUSE(TABLE 'training_data'))
LAYOUT(XGBOOST(
    objective 'reg:squarederror'
    num_iterations 100
    max_depth 6
))
LIFETIME(0);
```

`PRIMARY KEY (x1, x2)` makes `x1` and `x2` the features. The target the model learns is `y`, inferred as the single column that is not part of the key; the parameters in `LAYOUT` are XGBoost hyperparameters (see [Layout parameters](#layout-parameters)).

**4. Predict** — `predictXGBoost` takes the features positionally and returns the prediction:

```sql theme={null}
SELECT predictXGBoost('model', 1.0, 2.0) AS prediction;
```

The ground truth is `2*1 + 3*2 = 8`, so the model's prediction is close to `8`.

<h2 id="how-it-works">
  How it works
</h2>

**Training (at load time).** Each source row is a `(features..., target)` observation. When the dictionary loads, the source is read block by block and the model is then trained once over the whole set. Feature and target values are read as floats, so the key columns must be numeric and the target attribute floating-point (see [Dictionary structure](#dictionary-structure)).

<Warning>
  **Training holds the entire training set in memory.** Reading the source in blocks is not out-of-core training: each block is accumulated rather than consumed, so the whole training set is memory resident before the first boosting round and stays memory resident until the last one.
</Warning>

**Predicting (at query time).** To predict, the model takes the feature vector — in the same order as the key columns were declared — and runs it through the trained booster, returning a `Float64`. A `NULL` feature is passed to the model as a missing value, which XGBoost handles the way it learned to during training, so the prediction is never `NULL`. When every feature is a constant, the model is evaluated once per block instead of once per row.

**The model is not persisted.** It lives only in memory, for as long as the dictionary is loaded, and is trained again from the source on every load — including after a server restart.

**Retraining the model.** Because every load trains from scratch, `SYSTEM RELOAD DICTIONARY` retrains the model against the current contents of the source table:

```sql theme={null}
INSERT INTO training_data VALUES (5, 10, 40);
SYSTEM RELOAD DICTIONARY model;
```

A non-zero `LIFETIME` also retrains, since a lifetime-triggered reload is an ordinary load. Use it to refresh the model periodically as the training data grows.

<h2 id="dictionary-structure">
  Dictionary structure
</h2>

An `XGBOOST` dictionary has a fixed shape:

* The `PRIMARY KEY` is one or more columns of a native numeric type (integers and floats) — the features. At query time this "key" is the feature vector you pass in to predict, not a stored lookup key. The feature order is the key-column declaration order, and `predictXGBoost` binds its positional arguments to that order.
* Alongside them, declare **exactly one attribute of type `Float32` or `Float64`**: the target the model learns. It is always inferred as the single column that is not part of the feature key — there is no parameter to name it, and it is an error to declare more than one attribute.

A column that does not match these requirements is rejected when the dictionary loads, not when you create it.

<h2 id="layout-parameters">
  Layout parameters
</h2>

Only the parameters listed below are accepted; any other name fails the load, so typos are caught when the model trains rather than being silently ignored. `num_iterations` is handled by ClickHouse (see its description); every other parameter is forwarded to the XGBoost booster unchanged, as a string, and takes XGBoost's own default and value range — see the [XGBoost parameter reference](https://xgboost.readthedocs.io/en/stable/parameter.html).

| Parameter | Description |
| - | - |
| `num_iterations` | Number of boosting rounds (how many trees to train). A positive integer, used as the training loop count rather than forwarded to the booster. Default `100`. |
| `booster` | Booster type: `gbtree`, `gblinear`, or `dart`. |
| `objective` | Learning objective, e.g. `reg:squarederror` or `binary:logistic`. Must be an objective that predicts a single value per row; multiclass objectives (`multi:softmax`, `multi:softprob`) are rejected. |
| `seed` | Random number seed. Training is otherwise deterministic, so this only changes the model when a stochastic parameter (`subsample`, `sampling_method`, any `colsample_*`) is also set. |
| `verbosity` | Logging verbosity of XGBoost: `0` (silent) to `3` (debug), default `1` (warnings). Affects only which messages XGBoost logs, never the trained model. See [Notes](#notes) for where the messages go. |
| `nthread` | Number of parallel threads used for training. Accepted and forwarded, but currently has no effect: the bundled XGBoost is built without OpenMP, so training and prediction always run on a single thread. |
| `eta` | Step-size shrinkage applied after each boosting round. Only this spelling is accepted; XGBoost's `learning_rate` alias is not. |
| `gamma` | Minimum loss reduction required to make a further split on a leaf. |
| `max_depth` | Maximum depth of a tree. |
| `min_child_weight` | Minimum sum of instance weight (hessian) needed in a child. |
| `max_delta_step` | Maximum delta step allowed for each leaf's output. |
| `subsample` | Fraction of the training rows sampled for each boosting round. |
| `sampling_method` | Row sampling method: `uniform` or `gradient_based`. Has no effect unless `subsample` is set to less than `1`. |
| `colsample_bytree` / `colsample_bylevel` / `colsample_bynode` | Fraction of columns (features) sampled per tree / per level / per split. |
| `lambda` | L2 regularization term on weights. Only this spelling is accepted; XGBoost's `reg_lambda` alias is not. |
| `alpha` | L1 regularization term on weights. Only this spelling is accepted; XGBoost's `reg_alpha` alias is not. |
| `tree_method` | Tree construction algorithm: `auto`, `exact`, `approx`, or `hist`. |
| `scale_pos_weight` | Balances positive and negative weights, useful for imbalanced classes. |
| `grow_policy` | How new nodes are added to the tree: `depthwise` or `lossguide`. |
| `max_leaves` | Maximum number of leaf nodes (used with `grow_policy` `lossguide`). |
| `max_bin` | Maximum number of discrete bins used to bucket continuous features (used with `tree_method` `hist`). |
| `num_parallel_tree` | Number of trees grown per boosting round (a value `> 1` trains a boosted random forest). |

Parameter names are case-insensitive, and each may be given only once. Values must be a positive integer, a float, or a quoted string: a negative literal is rejected by the dictionary DDL itself, before the layout sees it, so a negative `seed` cannot be expressed.

For example, a dictionary that also sets the step size `eta`:

```sql theme={null}
CREATE DICTIONARY model_tuned (x1 Float64, x2 Float64, y Float64)
PRIMARY KEY (x1, x2)
SOURCE(CLICKHOUSE(TABLE 'training_data'))
LAYOUT(XGBOOST(
    objective 'reg:squarederror'
    num_iterations 100
    max_depth 6
    eta 0.3
))
LIFETIME(0);
```

<h2 id="prediction-parameters">
  Prediction parameters
</h2>

`predictXGBoost` accepts an optional trailing constant `Map` of XGBoost prediction parameters, after the features, built with `map`:

```sql theme={null}
SELECT predictXGBoost('model', 1.0, 2.0, map('type', 0, 'iteration_end', 0));
```

The parameter names map to the prediction parameters of XGBoost's `XGBoosterPredictFromDMatrix`. Only the keys below are accepted; any other key fails the query. Every parameter is an integer or a boolean, so the `Map` values must be an integer type.

| Parameter | Description | Default |
| - | - | - |
| `type` | Prediction type. Only `0` (value) and `1` (margin) are accepted, because `predictXGBoost` returns a single `Float64` per row. Other XGBoost types (`2`/`3` SHAP contributions, `4`/`5` feature interactions, `6` leaf index) emit several values per row and are rejected. | `0` |
| `iteration_begin` | First boosting round to include in the prediction, counted from `0` (inclusive). Must not exceed the number of boosting iterations the model was trained with (`num_iterations`). | `0` |
| `iteration_end` | One past the last boosting round to include (exclusive), so `iteration_end 1` uses only the first round; `0` uses all rounds. Bounded like `iteration_begin`. | `0` |

<h2 id="notes">
  Notes
</h2>

* **Computational dictionary semantics.** This is a *computational* dictionary: it holds a trained model, not rows, and `predictXGBoost` is the only way to query it. The generic dictionary interface is not supported and reports an error: `dictGet` (there is no stored attribute to look up — the "key" is a feature vector to predict from), `dictHas` (no keys are stored), `SELECT * FROM dict` and joining the dictionary as a table. Because `predictXGBoost` is the only entry point, the `enable_xgboost` setting must be enabled for every prediction.

* **Numeric columns only.** Every feature (key) column must be a native numeric type and the target attribute must be `Float32` or `Float64`. Values are read as floats during training and prediction. The feature arguments of `predictXGBoost` may also be `Nullable`, see [How it works](#how-it-works).

* **`system.dictionaries` reports no stored items.** The dictionary trains a model instead of storing rows, so `element_count` is `0`, as it is for a `direct` dictionary, and `bytes_allocated` is `0` too: the trained model belongs to XGBoost, which does not report how much memory it holds. `query_count` and `found_rate` count the rows passed to the model by `predictXGBoost`; a call whose features are all constant is evaluated once per block, so it counts one row per block rather than one per row.

* **A failed reload keeps the previous model.** If retraining fails — the source table is gone, its schema changed, a hyperparameter is no longer accepted — the dictionary does not start failing predictions. It keeps serving the last model that trained successfully, and records the error instead. Compare `last_successful_update_time` with `last_exception` in `system.dictionaries` to tell whether the model still reflects the current source data:

  ```sql theme={null}
  SELECT name, status, last_successful_update_time, last_exception
  FROM system.dictionaries
  WHERE name = 'model';
  ```

* **XGBoost's messages go to the server log.** They are written under the `XGBoost` logger, at the level XGBoost gives them, and are attached to the query that trains or predicts, so they also appear in `system.text_log`. Raise `verbosity` to see more of them.

* **Feature order matters.** `predictXGBoost` binds its positional feature arguments to the key columns in declaration order, and the number of feature arguments must match the number of key columns.
