Skip to main content

CAGRA

UDTF: cuvs_cagra

Official cuVS reference: C API

Query-local graph approximate nearest-neighbor search.

Quickstart​

The call below expects the registered relations dataset_vectors (the dataset role) and query_vectors (the queries role), each passed as a parenthesized SELECT subquery. Substitute your own relations and column names.

SELECT *
FROM cuvs_cagra(
dataset => (SELECT id, d0, d1 FROM dataset_vectors),
queries => (SELECT id, d0, d1 FROM query_vectors),
k => 8,
metric => 'l2_expanded',
graph_degree => 32,
intermediate_graph_degree => 64
)
ORDER BY query_ordinal, rank;

Inputs​

Each relation argument is a parenthesized SELECT subquery that the planner keeps as a real child; metadata validation resolves a registered table or view for the same role instead. See Vector Inputs for the relation identity rules and the ID, dense-vector type, null, finite-value, and runtime-dimension contract.

RoleRequiredValidation referenceDescription
datasetyestableDense-vector rows indexed for nearest-neighbor search.
queriesyestableDense-vector query rows matched against the evaluated dataset.

Vector element types​

Element typeValid metrics
Float32l2_expanded, inner_product, cosine
Int8l2_expanded, inner_product, cosine
UInt8l2_expanded, inner_product, cosine

Arguments and options​

Scalar SQL arguments​

ArgumentTypeRequiredDescription
kintegeryesNumber of neighbors returned for each evaluated query row.
metricenum ("l2_expanded", "inner_product", "cosine")noDistance or score metric; native default l2_expanded. inner_product ranks higher scores first.
graph_degreeintegernoOutput graph degree; native default 64. Must be less than dataset rows.
intermediate_graph_degreeintegernoBuild graph degree; native default 128. Must exceed graph_degree and be less than dataset rows.
build_algoenum ("ivf_pq", "nn_descent")noGraph construction algorithm; native default ivf_pq. nn_descent does not support cosine.
itopk_sizeintegernoIntermediate search results; native default 64. Must be at least k.
search_widthintegernoStarting graph nodes per iteration; native default 1.
max_iterationsintegernoMaximum search iterations; native default 0 selects automatically.
search_algoenum ("auto", "single_cta", "multi_cta", "multi_kernel")noSearch kernel; auto is set explicitly (the native zero-initialized default is single_cta).

SQL value argument schemas​

ArgumentRequiredLiteral shapeDefaultConstraintsDescription
build_algonostring"ivf_pq"one of "ivf_pq", "nn_descent"Graph construction algorithm; native default ivf_pq. nn_descent does not support cosine.
graph_degreenointeger64minimum 1; maximum 4294967295Output graph degree; native default 64. Must be less than dataset rows.
intermediate_graph_degreenointeger128minimum 1; maximum 4294967295Build graph degree; native default 128. Must exceed graph_degree and be less than dataset rows.
itopk_sizenointeger64minimum 1; maximum 4294967295Intermediate search results; native default 64. Must be at least k.
kyesintegerNo defaultminimum 1; maximum 4294967295Number of neighbors returned for each evaluated query row.
max_iterationsnointeger0minimum 0; maximum 4294967295Maximum search iterations; native default 0 selects automatically.
metricnostring"l2_expanded"one of "l2_expanded", "inner_product", "cosine"Distance or score metric; native default l2_expanded. inner_product ranks higher scores first. Supported element/metric combinations: Float32: l2_expanded, inner_product, cosine; Int8: l2_expanded, inner_product, cosine; UInt8: l2_expanded, inner_product, cosine.
search_algonostring"auto"one of "auto", "single_cta", "multi_cta", "multi_kernel"Search kernel; auto is set explicitly (the native zero-initialized default is single_cta).
search_widthnointeger1minimum 1; maximum 4294967295Starting graph nodes per iteration; native default 1.

Vector binding shapes​

Each relation subquery must project a non-null id field followed by either one or more non-null feature fields of a supported element type (Float32, Int8, UInt8) or one non-null list vector field named vector.

For wide vectors, the projection order defines the feature dimensions. A list vector relation must contain no feature field beside id and vector.

Output​

ColumnTypeNullableDescription
query_ordinalUInt64noZero-based ordinal of the evaluated query row; it disambiguates duplicate query IDs.
query_idsame_as_queries.idnoLogical ID copied from the queries relation.
neighbor_ordinalUInt64noZero-based ordinal of the matched dataset row; it disambiguates duplicate dataset IDs.
neighbor_idsame_as_dataset.idnoLogical ID copied from the matched dataset row.
rankUInt32noOne-based neighbor rank within a query. Order consumers explicitly by query_ordinal, rank.
distanceFloat32noMetric value; smaller is better for distance metrics, while inner_product prefers larger values.

Concrete schemas are call-specific. Run gpu_validate_call against registered relations to inspect the output schema after the actual ID types and literal options are validated.

Examples​

This example uses the EMBER2024 demo dataset.

Which earlier Win32 files make useful references for a new quarter?​

An analyst reviewing a new Win32 file can retrieve structurally similar files from an earlier period and inspect their historical labels and families. The search corpus contains 15,588 sampled train files from weeks 0-51; the held-out test period contains 3,618 sampled files from weeks 52-63. Both use the same one-percent SHA-256 sample and 512 normalized byte and byte-entropy histogram features. The train and test SHA-256 sets are disjoint. The search UDTFs receive only IDs and histogram vectors. Labels, family names, week IDs, and SHA-256 values are joined after search for review. The query retrieves references; it does not classify malware.

MethodID recall@10 vs CPU oracleKnown-family match among top 10Three-run median whole SQL timeDedicated r3 query ID
Brute-force exact0.999035.39%0.876 s73f3ec94-7127-4b7f-95a7-6a949ceeaa98
CAGRA0.996835.37%1.319 sc0f13481-511d-4600-815e-b43ab2685418
IVF-Flat0.995435.33%1.005 s02f1af3f-270f-4a22-8559-dee13bd59478

The family rates measure agreement with the available family labels. For the 1,575 test queries with a known family, 35.39% of brute-force top-10 reference slots had the same known family; references with no family label count as nonmatches. A uniform random training reference gives a 2.76% expected same-family rate. The search input does not use those labels or families.

For test file 10021055 from week 52, the exact result's three nearest references are from weeks 9, 16, then 10. Their family labels appear only after the search join:

RankReference IDTrain weekFamilySquared-L2 distance
19358309berbew0.0053381994
2166315416berbew0.0058313087
3103203010berbew0.0058773160

At this cohort size brute force had the lowest median whole-statement time. Each time includes per-statement index construction and fetching all 36,180 result rows. The captures ran on the dedicated EMBER server. They are not a controlled performance study, so the timings do not support an ANN speedup claim.

The exhaustive brute-force search uses Float32 squared-L2 arithmetic. Its 0.9990 ID recall against direct Float64 distances reflects changes near the top-10 boundary; the measured maximum distance error was 6.72e-6.

Download the capture JSON, including all three exact SQL statements, run query IDs, CPU-oracle metrics, post-search family measurements, and conditions.

WITH train_index AS (
SELECT v.id, v.histogram AS vector
FROM ember_research_vectors v
JOIN ember_metadata m ON m.id = v.id
WHERE v.split = 'train' AND m.file_type = 'Win32'
ORDER BY v.id
), test_queries AS (
SELECT v.id, v.histogram AS vector
FROM ember_research_vectors v
JOIN ember_metadata m ON m.id = v.id
WHERE v.split = 'test' AND m.file_type = 'Win32'
ORDER BY v.id
)
SELECT n.query_ordinal,
n.query_id,
qm.week_id AS query_week,
q.label AS query_label,
n.rank,
n.neighbor_id,
tm.week_id AS neighbor_week,
train.label AS neighbor_label,
tm.sha256 AS neighbor_sha256,
tm.family AS neighbor_family,
n.distance
FROM cuvs_brute_force_knn(
dataset => (SELECT id, vector FROM train_index ORDER BY id),
queries => (SELECT id, vector FROM test_queries ORDER BY id),
k => 10,
metric => 'l2_expanded'
) n
JOIN ember_research_vectors q ON q.id = n.query_id AND q.split = 'test'
JOIN ember_metadata qm ON qm.id = q.id
JOIN ember_research_vectors train ON train.id = n.neighbor_id AND train.split = 'train'
JOIN ember_metadata tm ON tm.id = train.id
ORDER BY n.query_ordinal, n.rank;

Dedicated r3 capture query ID: 73f3ec94-7127-4b7f-95a7-6a949ceeaa98.

Limits​

  • Validation resolves named tables or views and reads schemas only; it does not execute relation scans or GPU work.
  • Execution relation arguments require parenthesized subqueries; dry-run validation accepts registered named relations only.
  • Builds and destroys a query-local graph and padded dataset within the statement; no index persists across statements.
  • Approximate search may return fewer than k neighbors; missing (query, rank) rows are omitted and valid ranks remain contiguous.
  • The dataset must be non-empty, k and intermediate_graph_degree must be less than or equal to its row count, and graph_degree must be less than intermediate_graph_degree.
  • k must not exceed itopk_size. single_cta requires itopk_size no greater than 1024. cosine cannot use nn_descent.
  • Dataset and query dimensions and element types must match.

To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.