Skip to main content

Forest Predict

UDTF: cuml_forest_predict

Official cuML reference: nvForest

Score Float32 rows with a pretrained XGBoost, LightGBM, or Treelite forest through nvForest.

Quickstart​

Register the relations referenced by the parenthesized SELECT clauses below. For metadata validation, the descriptor names input_vectors.

SELECT id, prediction
FROM cuml_forest_predict(
input => (SELECT id, d0, d1, d2, d3 FROM t),
model => 'scorer.json',
model_format => 'xgboost_json'
)
ORDER BY row_ordinal;

Inputs​

Each relation argument is a parenthesized SELECT subquery. Metadata validation resolves a registered table or view for the same role without scanning its rows. See ML Inputs for the ID, Float32 feature, null, finite-value, and runtime-dimension contract.

RoleRequiredValidation referenceDescription
inputyestableDense Float32 rows fitted and transformed in this statement.

Vector element types​

Element typeValid metrics
Float32Not applicable

Arguments and options​

Scalar SQL arguments​

ArgumentTypeRequiredDescription
modelstringyesModel file path relative to the server-configured models root.
model_formatenum ("xgboost_json", "xgboost_ubjson", "xgboost_legacy", "lightgbm", "treelite")yesSerialized format of the model file.
outputenum ("scores", "class_index")noReturn the model's postprocessed scores, or the index of the largest class score.

SQL value argument schemas​

ArgumentRequiredLiteral shapeDefaultConstraintsDescription
modelyesstringNo defaultminimum string length 1; maximum string length 1024Model file path relative to the server-configured models root.
model_formatyesstringNo defaultone of "xgboost_json", "xgboost_ubjson", "xgboost_legacy", "lightgbm", "treelite"Serialized format of the model file.
outputnostring"scores"one of "scores", "class_index"Return the model's postprocessed scores, or the index of the largest class score.

Vector binding shapes​

Each relation subquery projects a non-null id and either non-null Float32 feature columns in dimension order or one non-null vector list column of non-null Float32 values. For classification, training also projects non-null Int32 label; regression requires non-null finite Float32 label. The predict relation omits it.

The list-column shape excludes other feature columns. See ML Inputs for the allowed list containers and runtime checks.

Output​

ColumnTypeNullableDescription
row_ordinalUInt64noEvaluated input row position, including when IDs repeat. Follows the input's top-level ORDER BY on output columns; unspecified without one.
idsame_as_input.idnoLogical input ID.
predictionFloat32 or Int32noFloat32 score of a single-output model with output => 'scores'; Int32 class index with output => 'class_index'.
output_<index>Float32noScore of one model output, returned instead of prediction for a multi-output model with output => 'scores'. The output contains model output count such columns, output_0 through output_&lt;model output count - 1&gt;.

Concrete schemas are call-specific. Run gpu_validate_call against registered relations to inspect the output schema after the actual ID types and literal options are validated.

Examples​

This retrospective triage example uses the EMBER2024 demo dataset.

Which high-scoring challenge files should an analyst review first?​

The LightGBM model scores all 6,315 challenge files. A score threshold of 0.99 selects 1,616 rows before the query's LIMIT 30; the output keeps the highest-scoring rows with their three largest positive feature-block contributions. Challenge labels are not inputs to scoring or ranking. They are all positive, so this review queue cannot estimate a false-positive rate or score calibration.

The feature-block values are SHAP contributions in raw-margin (log-odds) units. The attribution table was prepared separately on CPU with fixture/demo/ember2024/scripts/research_attributions.py: LightGBM pred_contrib processed 6,315 rows in 13.311 seconds. The script then summed the 2,568 feature contributions into 12 blocks. The SQL calls cuml_forest_predict for GPU scoring and cuvs_top_k over those precomputed grouped values. The CPU attribution preparation is not GPU work and is not included in the SQL run.

The CPU LightGBM score comparison passed for all 6,315 rows, with maximum absolute error 2.9801783152372252e-8; threshold decisions agreed for every row at the checked thresholds, including 0.99. Numerical agreement verifies the imported model's scoring implementation. Detection metrics require the separately evaluated test labels.

On the separate 6,102-file sample from the later test period, the same 0.99 threshold detected 2,167 of 3,016 malicious files (71.85%) and produced no false positives among 3,086 benign files. The Wilson 95% interval for that observed false-positive rate extends up to 0.1243%. The challenge selection contains 1,616 of 6,315 positives (25.59%). These cohort-specific measurements describe the coverage lost at a high review threshold; a zero count in this test sample does not establish a zero operational false-positive rate.

The table shows the first six rows of the 30-row capture. Top-K ranks the 12 signed feature-block values and SQL keeps only positive values. Negative feature groups and the model bias are omitted, so these rows are not complete explanations. Positive SHAP contributions describe the model's score calculation; they do not establish causal risk effects. Family labels are not used by the query.

SHA-256 prefixIDModel scoreRankFeature blockPositive log-odds contribution
8f392dc01595…200013480.99985104799270631histogram2.504060745239258
8f392dc01595…200013480.99985104799270632strings2.249004125595093
8f392dc01595…200013480.99985104799270633header1.235449790954590
7d60904d4ff5…200006770.99984484910964971histogram2.653215169906616
7d60904d4ff5…200006770.99984484910964972strings2.279698133468628
7d60904d4ff5…200006770.99984484910964973header1.287171244621277

Download the forest review capture for all 30 result rows, the SQL, query ID, CPU comparison, and preparation conditions.

SELECT p.id, m.sha256, m.file_type, m.week_id, p.prediction,
t.rank, g.group_name, t.value AS positive_log_odds_contribution
FROM cuml_forest_predict(
input => (
SELECT id, vector
FROM ember_research_vectors_challenge
ORDER BY id
),
model => 'EMBER2024_all.model',
model_format => 'lightgbm'
) p
JOIN cuvs_top_k(
input => (SELECT id, vector FROM ember_attributions ORDER BY id),
k => 3,
select => 'max'
) t ON t.id = p.id
JOIN ember_feature_groups g ON g.group_index = t.position
JOIN ember_metadata_challenge m ON m.id = p.id
WHERE p.prediction >= 0.99 AND t.value > 0
ORDER BY p.prediction DESC, p.id, t.rank
LIMIT 30;

The query was captured on a dedicated EMBER server with native disposition. The CPU attribution step ran before the statement; its duration is recorded separately in the capture.

Limits​

  • Validation resolves named tables or views and reads schemas only; it does not execute relation scans or GPU work.
  • Execution relation arguments require parenthesized subqueries; dry-run validation accepts registered named relations only.
  • Scores pretrained XGBoost (JSON, UBJSON, legacy binary), LightGBM text, and Treelite checkpoint forests with nvForest. Multi-target models are rejected at planning.
  • Features must be non-null Float32 and match the model feature count. Forests with double-precision thresholds widen the features to Float64 on the device.
  • output => 'class_index' requires a model with at least two outputs and resolves ties to the lowest class index.
  • model resolves under the server-configured models root, and files above max_model_bytes are rejected. Each statement reads and imports the model once; there is no cross-statement model cache.

To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.