Discover & Validate GPU UDTFs
Use this guide before adding a cuGraph, cuVS, or cuML UDTF to a query. It follows the same sequence for embedded sessions and Flight SQL: list what this build can run, inspect one UDTF's contract, validate the call statically (no data is read and nothing runs on the GPU), execute it, then check the complete query before production use.
This page shows the workflow only. For every returned column, JSON schema, and validation-envelope rule, see the GPU UDTF Catalog API.
1. List installed UDTFs
Start with the UDTFs available in the current build and session. Filter by provider when you already know the workload:
SELECT function_name, provider, available, summary
FROM gpu_list_functions()
WHERE provider IN ('cugraph', 'cuvs')
ORDER BY function_name;
available=true means the UDTF can be planned for GPU execution in this
build and session; cuVS UDTFs require the cuvs feature. It does not
reserve GPU memory, inspect relation rows, or prove that a surrounding query
runs entirely on the GPU. Memory is decided at execution: the call runs inside
its granted cap and fails with an allocation error if it exceeds it.
2. Describe the selected UDTF
Retrieve the canonical signature, relation roles, option schema, output schema, and lifecycle limits before constructing a call:
SELECT signature, relation_roles_json, options_schema_json,
result_schemas_json, limitations_json
FROM gpu_describe_function('cugraph_pagerank');
UDTF names are exact canonical names. The returned descriptor is useful when a client or agent needs to construct SQL dynamically, while the provider UDTF pages explain the same contract in task-specific terms.
3. Validate a concrete call
Validate registered relation metadata and the provider-owned options before execution:
SELECT *
FROM gpu_validate_call(
'cugraph_pagerank',
'{
"relations":{"edges":{"table":"analytics.edges"}},
"options":{"alpha":0.9,"directed":true}
}'
);
This is static validation: it checks the named tables' or views' schemas and
the options, reads no data, and runs nothing on the GPU (no rows scanned, no
subquery run, no device memory allocated, no kernel launched). Only registered
tables or views resolve, and each one must already expose the role's canonical
columns (src/dst for cuGraph edges, id plus feature columns for cuVS
relations). A call whose relation aliases, filters, joins, or reads a CTE is
not represented by this envelope; run EXPLAIN GPU on the final SQL instead,
or register that derived relation as a view and validate the view.
4. Execute with the provider's relation syntax
The catalog workflow is shared, but execution SQL intentionally differs:
- cuGraph passes each relation as a named parenthesized SELECT subquery, such as cugraph_pagerank(edges => (SELECT src, dst FROM edges)).
- cuVS passes each relation and value as named arguments, such as cuvs_kmeans(input => (SELECT item_id AS id, d0, d1 FROM embeddings), n_clusters => 8).
Use the cuGraph UDTFs and cuVS UDTFs pages for each UDTF's signature, input contract, outputs, and lifecycle.
5. Check whole-query GPU coverage
gpu_validate_call covers one selected UDTF call. It does not inspect DataFusion operators, sources, or host boundaries feeding and consuming that call. Before shipping a composed query, use GPU coverage validation to check the final planned path.