Forest Predict
UDTF: cuml_forest_predict
Official cuML reference: nvForest
Score Float32 rows with a pretrained XGBoost, LightGBM, or Treelite forest through nvForest.
Quickstart
Register the relations referenced by the parenthesized SELECT clauses below. For metadata validation, the descriptor names input_vectors.
SELECT id, prediction
FROM cuml_forest_predict(
input => (SELECT id, d0, d1, d2, d3 FROM t),
model => 'scorer.json',
model_format => 'xgboost_json'
)
ORDER BY row_ordinal;
Inputs
Each relation argument is a parenthesized SELECT subquery. Metadata validation resolves a registered table or view for the same role without scanning its rows. See ML Inputs for the ID, Float32 feature, null, finite-value, and runtime-dimension contract.
Vector element types
Arguments and options
Scalar SQL arguments
SQL value argument schemas
Vector binding shapes
Each relation subquery projects a non-null id and either non-null Float32 feature columns in dimension order or one non-null vector list column of non-null Float32 values. For classification, training also projects non-null Int32 label; regression requires non-null finite Float32 label. The predict relation omits it.
The list-column shape excludes other feature columns. See ML Inputs for the allowed list containers and runtime checks.
Output
Concrete schemas are call-specific. Run gpu_validate_call against registered relations to inspect the output schema after the actual ID types and literal options are validated.
Examples
This retrospective triage example uses the EMBER2024 demo dataset.
Which high-scoring challenge files should an analyst review first?
The LightGBM model scores all 6,315 challenge files. A score threshold of
0.99 selects 1,616 rows before the query's LIMIT 30; the output keeps the
highest-scoring rows with their three largest positive feature-block
contributions. Challenge labels are not inputs to scoring or ranking. They are
all positive, so this review queue cannot estimate a false-positive rate or
score calibration.
The feature-block values are SHAP contributions in raw-margin (log-odds)
units. The attribution table was prepared separately on CPU with
fixture/demo/ember2024/scripts/research_attributions.py: LightGBM
pred_contrib processed 6,315 rows in 13.311 seconds. The script then summed
the 2,568 feature contributions into 12 blocks. The SQL calls cuml_forest_predict for
GPU scoring and cuvs_top_k over those precomputed grouped values. The CPU
attribution preparation is not GPU work and is not included in the SQL run.
The CPU LightGBM score comparison passed for all 6,315 rows, with maximum
absolute error 2.9801783152372252e-8; threshold decisions agreed for every
row at the checked thresholds, including 0.99. Numerical agreement verifies
the imported model's scoring implementation. Detection metrics require the
separately evaluated test labels.
On the separate 6,102-file sample from the later test period, the same 0.99
threshold detected 2,167 of 3,016 malicious files (71.85%) and produced no
false positives among 3,086 benign files. The Wilson 95% interval for that
observed false-positive rate extends up to 0.1243%. The challenge selection
contains 1,616 of 6,315 positives (25.59%). These cohort-specific measurements
describe the coverage lost at a high review threshold; a zero count in this
test sample does not establish a zero operational false-positive rate.
The table shows the first six rows of the 30-row capture. Top-K ranks the 12 signed feature-block values and SQL keeps only positive values. Negative feature groups and the model bias are omitted, so these rows are not complete explanations. Positive SHAP contributions describe the model's score calculation; they do not establish causal risk effects. Family labels are not used by the query.
Download the forest review capture for all 30 result rows, the SQL, query ID, CPU comparison, and preparation conditions.
SELECT p.id, m.sha256, m.file_type, m.week_id, p.prediction,
t.rank, g.group_name, t.value AS positive_log_odds_contribution
FROM cuml_forest_predict(
input => (
SELECT id, vector
FROM ember_research_vectors_challenge
ORDER BY id
),
model => 'EMBER2024_all.model',
model_format => 'lightgbm'
) p
JOIN cuvs_top_k(
input => (SELECT id, vector FROM ember_attributions ORDER BY id),
k => 3,
select => 'max'
) t ON t.id = p.id
JOIN ember_feature_groups g ON g.group_index = t.position
JOIN ember_metadata_challenge m ON m.id = p.id
WHERE p.prediction >= 0.99 AND t.value > 0
ORDER BY p.prediction DESC, p.id, t.rank
LIMIT 30;
The query was captured on a dedicated EMBER server with native disposition.
The CPU attribution step ran before the statement; its duration is recorded
separately in the capture.
Limits
- Validation resolves named tables or views and reads schemas only; it does not execute relation scans or GPU work.
- Execution relation arguments require parenthesized subqueries; dry-run validation accepts registered named relations only.
- Scores pretrained XGBoost (JSON, UBJSON, legacy binary), LightGBM text, and Treelite checkpoint forests with nvForest. Multi-target models are rejected at planning.
- Features must be non-null Float32 and match the model feature count. Forests with double-precision thresholds widen the features to Float64 on the device.
- output => 'class_index' requires a model with at least two outputs and resolves ties to the lowest class index.
- model resolves under the server-configured models root, and files above max_model_bytes are rejected. Each statement reads and imports the model once; there is no cross-statement model cache.
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.