GPU UDTF Catalog API
Three metadata UDTFs describe every cuGraph, cuVS, and cuML execution UDTF installed in a session, so a person or an agent can find, understand, and check a call without reading server source:
gpu_list_functions()lists which UDTFs are installed and whether each one is executable in this build and session.gpu_describe_function(name)returns one UDTF's contract: signature, relation roles, option schema, result schemas, examples, and limitations.gpu_validate_call(name, call_json)checks a concrete call against the schema of the named tables or views and the UDTF's options. This is static validation, a dry run: it reads no data and runs nothing on the GPU.
The catalog is one list across providers; filter by provider for the
cuGraph, cuVS, or cuML subset. cuVS UDTFs (including cuvs_brute_force_knn,
cuvs_ivf_flat, cuvs_ivf_pq, cuvs_cagra, cuvs_kmeans, and cuvs_pca) are executable
when the build enables the cuvs feature;
otherwise they are listed with available=false and
unavailable_reason='required_feature_disabled'. cuML UDTFs
use provider cuml and require the cuml feature for
execution. See cuML SQL API for their query-local
lifetime.
For the task-oriented list, describe, validate, and execute workflow, use Discover & Validate GPU UDTFs. This page is the detailed API reference for the metadata UDTFs and their response schemas.
Execution syntax is provider-specific: cuGraph, cuVS, and cuML UDTFs use named
arguments, with each relation supplied as a parenthesized relation-valued
subquery. Validation is provider-neutral: the call_json envelope names
registered tables or views by role and never evaluates SQL text embedded in
JSON.
gpu_list_functions
gpu_list_functions() takes no arguments and returns one row for each
installed execution UDTF. It does not list the three metadata UDTFs
themselves. Rows are ordered by function_name before ordinary SQL projection
or filtering.
SELECT function_name, provider, available, summary
FROM gpu_list_functions()
WHERE provider IN ('cugraph', 'cuvs', 'cuml')
ORDER BY function_name;
available=true means the selected UDTF can be planned for GPU
execution. It does not reserve GPU memory, acquire runtime admission, inspect
relation rows, or prove that an entire query runs on the GPU.
gpu_describe_function(function_name)
gpu_describe_function takes one exact canonical execution UDTF name and
returns one full descriptor row. It has no compact or verbose mode: project the
columns you need. Provider-local short aliases are not accepted.
SELECT signature, options_schema_json, result_schemas_json
FROM gpu_describe_function('cuvs_brute_force_knn');
Every JSON column contains valid JSON. Empty objects and arrays are represented
as {} and [], not SQL NULL.
gpu_validate_call(function_name, call_json)
gpu_validate_call takes an exact canonical execution UDTF name and a
common JSON envelope. The selected provider owns the meaning of
relation roles and options, while the envelope consistently names UDTF
inputs.
SELECT *
FROM gpu_validate_call(
'cuvs_kmeans',
'{
"relations":{"input":{"table":"embedding_vectors"}},
"options":{"n_clusters":8}
}'
);
The envelope requires exactly these fields:
relations: object keyed by the descriptor-defined relation roles. Each relation currently has exactly onetablefield with a one-, two-, or three-part DataFusion table or view reference.options: object validated by the selected UDTF.
Unknown envelope keys, relation roles, and option keys are rejected. The metadata path resolves named tables and views only; it does not accept arbitrary SQL in JSON. To dry-run a derived subquery, register it as a temporary view first.
A syntactically valid request for an unknown UDTF or invalid execution call returns one structured invalid row. Wrong metadata UDTF arity and nonliteral metadata arguments are planning errors.
What validation does & does not do
Does: check that the UDTF exists; resolve every required named relation; validate provider-owned relation bindings, columns, dtypes, dimensions, conditional arguments, and option values; apply defaults; and derive a concrete output schema when possible.
Does Not: read any data or run anything on the GPU. It does not scan relation rows, materialize Parquet, construct graphs, launch CUDA, allocate device memory, acquire query admission, or prove runtime facts such as source-vertex existence, vector uniformity, or evaluated relation cardinality. Memory is decided at execution: an admitted call runs inside its granted cap and fails with an allocation error if it exceeds it.
would_execute_gpu is scoped to one selected UDTF call. To check whether a
whole query's planned path stays on the GPU, use
GPU coverage validation.