Random Forest Regressor
UDTF: cuml_random_forest_regressor
Official cuML reference: Python API
Fit a query-local cuML random forest regressor and score independent rows.
Quickstart
Register the relations referenced by the parenthesized SELECT clauses below. For metadata validation, the descriptor names training_vectors and predict_vectors.
SELECT id, prediction
FROM cuml_random_forest_regressor(
training => (SELECT id, d0, label FROM train),
predict => (SELECT id, d0 FROM test)
)
ORDER BY row_ordinal;
Inputs
Each relation argument is a parenthesized SELECT subquery. Metadata validation resolves a registered table or view for the same role without scanning its rows. See ML Inputs for the ID, Float32 feature, null, finite-value, and runtime-dimension contract.
Vector element types
Arguments and options
Scalar SQL arguments
SQL value argument schemas
Vector binding shapes
Each relation subquery projects a non-null id and either non-null Float32 feature columns in dimension order or one non-null vector list column of non-null Float32 values. For classification, training also projects non-null Int32 label; regression requires non-null finite Float32 label. The predict relation omits it.
The list-column shape excludes other feature columns. See ML Inputs for the allowed list containers and runtime checks.
Output
Concrete schemas are call-specific. Run gpu_validate_call against registered relations to inspect the output schema after the actual ID types and literal options are validated.
Limits
- Validation resolves named tables or views and reads schemas only; it does not execute relation scans or GPU work.
- Execution relation arguments require parenthesized subqueries; dry-run validation accepts registered named relations only.
- Classifier labels must be non-null Int32 contiguous classes 0..C-1; regressor labels must be non-null finite Float32. Predict must omit label and match the training feature dimension.
- Prediction imports the fitted forest into nvForest and scores the predict relation on the GPU; a classifier tie resolves to the lowest class index. Training and prediction inputs are limited to i32::MAX rows and the packed feature matrix to i32::MAX elements; the random forest path does not batch rows.
- Training uses one RAFT pool stream whose allocations share the query grant. The forest is query-local; an empty predict relation returns an empty result.
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.