IVF-PQ
UDTF: cuvs_ivf_pq
Official cuVS reference: C API
Query-local product-quantized inverted-file nearest-neighbor search.
Quickstart
The call below expects the registered relations dataset_vectors (the dataset role) and query_vectors (the queries role), each passed as a parenthesized SELECT subquery. Substitute your own relations and column names.
SELECT *
FROM cuvs_ivf_pq(
dataset => (SELECT id, d0, d1 FROM dataset_vectors),
queries => (SELECT id, d0, d1 FROM query_vectors),
k => 8,
metric => 'l2_expanded',
n_lists => 16,
n_probes => 16,
pq_dim => 2
)
ORDER BY query_ordinal, rank;
Inputs
Each relation argument is a parenthesized SELECT subquery that the planner keeps as a real child; metadata validation resolves a registered table or view for the same role instead. See Vector Inputs for the relation identity rules and the ID, dense-vector type, null, finite-value, and runtime-dimension contract.
Vector element types
Arguments and options
Scalar SQL arguments
SQL value argument schemas
Vector binding shapes
Each relation subquery must project a non-null id field followed by either one or more non-null feature fields of a supported element type (Float32, Int8, UInt8) or one non-null list vector field named vector.
For wide vectors, the projection order defines the feature dimensions. A list vector relation must contain no feature field beside id and vector.
Output
Concrete schemas are call-specific. Run gpu_validate_call against registered relations to inspect the output schema after the actual ID types and literal options are validated.
Limits
- Validation resolves named tables or views and reads schemas only; it does not execute relation scans or GPU work.
- Execution relation arguments require parenthesized subqueries; dry-run validation accepts registered named relations only.
- Builds and destroys its query-local index within the statement; no index persists across statements.
- Product quantization is lossy even when all lists are probed. Missing neighbors are omitted and valid ranks remain contiguous.
- The dataset must be non-empty; k and n_lists must not exceed its row count, and n_probes must not exceed n_lists.
- Dataset and query dimensions and element types must match. Cosine requires at least two dimensions. Nonzero pq_dim * pq_bits must be divisible by 8; dimensions are padded when needed.
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.