Jaccard
UDTF: cugraph_jaccard
Official cuGraph reference: C API
Compare explicit vertex pairs by dividing the size of their shared-neighbor intersection by the size of their neighbor union.
Quickstart
The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight, plus the registered relation candidate_pairs (the vertex_pairs relation with columns first, second). Substitute your own registered relations.
SELECT *
FROM cugraph_jaccard(
edges => (SELECT src, dst FROM target_edges),
vertex_pairs => (SELECT first, second FROM candidate_pairs)
);
Inputs
Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.
Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.
Logical string side-input limitations:
- edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs
- candidate-pair columns must match the graph vertex domain; logical string graphs accept Utf8, LargeUtf8, or Utf8View independently per column
Arguments and options
Relation arguments
Vertex columns of vertex_pairs must use the same vertex domain as edges: the integer type of src and dst, or any listed string type when the endpoints are strings. Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.
Named value arguments
This UDTF has no algorithm-specific value arguments. Its inputs are the relation arguments above and the graph construction options below.
Graph construction options
This UDTF requires directed=false (undirected/symmetric graph); all other graph construction options follow the shared defaults documented in Graph Construction Options.
Output
These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.
Examples
- Citation network
- MerRec marketplace
- EMBER2024 import review
This example runs on the citation network demo dataset.
Which papers should I read after the Transformer if I want familiar references?
Compare the reference list of Attention Is All You Need (2017) with NLP papers published from 2017 through 2020. The candidate papers have at least 10 references and one of the listed field labels. The score measures shared references against the union of both lists. A high score reflects bibliographic overlap; topic meaning is outside this measure. Title-identical records are excluded from the candidate pairs, preventing an edition with the seed's title from appearing as a separate reading suggestion.
The similarity UDTF requires an undirected graph. Each citing paper keeps its positive ID, while each cited paper becomes a negative reference-role ID. Those separate roles make the neighborhood of a paper reflect its outgoing references, even when other papers cite it. The edge projection omits weights, so every reference contributes as one unit.
CREATE OR REPLACE VIEW similarity_seed AS
SELECT paper_id, title, n_references
FROM papers
WHERE title = 'Attention is all you need' AND year = 2017;
CREATE OR REPLACE VIEW similarity_candidates AS
SELECT p.paper_id, p.title, p.n_references, p.year
FROM papers p
WHERE p.year BETWEEN 2017 AND 2020
AND p.n_references >= 10
AND p.primary_fos IN (
'Natural language processing', 'Language model',
'Machine translation', 'Question answering');
CREATE OR REPLACE VIEW bibliography_edges AS
SELECT e.src, -e.dst AS dst
FROM citation_edges e
JOIN (
SELECT paper_id FROM similarity_seed
UNION
SELECT paper_id FROM similarity_candidates
) selected ON selected.paper_id = e.src;
SELECT p.year, p.title, p.n_references,
ROUND(r.similarity, 3) AS jaccard
FROM cugraph_jaccard(
edges => (SELECT src, dst FROM bibliography_edges),
vertex_pairs => (
SELECT s.paper_id AS first, c.paper_id AS second
FROM similarity_seed s
CROSS JOIN similarity_candidates c
WHERE LOWER(c.title) <> LOWER(s.title)),
directed => false) r
JOIN papers p ON p.paper_id = r.second
ORDER BY r.similarity DESC, p.paper_id
LIMIT 5;
Self-Attention with Relative Position Representations has the highest Jaccard score among these candidates, at 0.296.
For coverage of a smaller bibliography, see the Overlap example.
This example uses the prepared graph inputs described on the MerRec dataset and graph capture page.
Which collectible categories should accompany a Nintendo games landing page?
Personalized PageRank ranks the top 100 non-seed product groups from observed
browsing transitions around the captured Nintendo Games seed. Jaccard then
compares distinct item-view audiences for those groups and the seed. The final
filter keeps Toys & Collectibles groups with at least 30 viewers.
SELECT m.product_id, m.brand_name, m.category, m.leaf_category,
m.viewers, r.similarity AS audience_jaccard
FROM cugraph_jaccard(
edges => (SELECT src, dst FROM mr_audience),
vertex_pairs => (
SELECT CAST(672 AS INT) AS first, candidates.vertex AS second
FROM (
SELECT vertex, value
FROM cugraph_personalized_pagerank(
edges => (SELECT src, dst, weight FROM mr_edges),
personalization => (
SELECT id AS vertex, CAST(1.0 AS DOUBLE) AS value
FROM mr_nodes WHERE product_id = '15928_1315'))
WHERE vertex <> 672
ORDER BY value DESC, vertex LIMIT 100
) candidates
), directed => false
) r
JOIN mr_nodes m ON m.id = r.second
WHERE m.category = 'Toys & Collectibles' AND m.viewers >= 30
ORDER BY r.similarity DESC, m.product_id
LIMIT 10;
The captured query returns 10 rows. The table shows its first six rows, with Jaccard rounded to six decimal places.
product_id groups listings by brand and leaf category, and each row can
represent multiple items. The score measures audience overlap for relevance
review. The query does not evaluate purchases, conversions, or revenue uplift. The
source sample covers one October 2023 shard.
See the full query capture for the exact result schema, query ID, SQL, and all ten captured rows.
This example uses the EMBER2024 demo dataset.
Which import patterns should an analyst compare around an unlabeled file?
The seed is Win32 challenge file 20005740, SHA-256
004ed0c4c3beadf297651f8f7dbfae41c8103e248f33fab9194213809e9cb7be. It has
no family label, 58 retained import symbols, and undirected kNN degree 7. The
seed was selected from samples with 10-100 retained imports and kNN degree at
least 5 by choosing the lexicographically smallest SHA-256. Neither family nor
label determined the seed or graph edges.
The query uses Personalized PageRank to form a 25-file comparison queue, then
Jaccard similarity on the shared sample-to-import graph to return 15 files.
The four highest-ranked files each share 54 of the seed's 58 retained imports.
Each has 130 retained imports, giving 54 / (58 + 130 - 54) = 0.402985.
All 15 returned files have family = ''.
The graph retains symbols seen in 3-100 distinct Win32 challenge files. This leaves import edges for 926 of 3,225 Win32 files, with 4,988 retained symbol vertices; 2,299 other file vertices have no retained imports and remain as isolates in the sample-count input. The challenge split contains only positive examples. An empty family field does not show that a file is benign or belongs to a new family.
The top rows repeatedly show msvbvm60.dll runtime ordinals among their shared
imports. Common runtime imports can explain overlap, so this queue requires
analyst inspection. It does not establish a campaign or family attribution.
The SQL computes graph similarity over imports and joins family metadata only
for result review.
Download the capture JSON, including all
15 rows with full SHA-256 values, query conditions, seed selection, graph-edge
provenance, and source query ID 7e26a7d3-b2b1-45b3-83df-837ea347b55e.
CPU validation confirmed the 15-candidate order against the GPU PageRank
candidate pool followed by CPU Jaccard, including every shared-import count,
minimum symbol, and metadata value. The PageRank scores differ from a NetworkX
3.7 reference by at most 7.43e-8 (absolute tolerance 1e-4). The validation
record is docs/research/ember2024/results/import_review_check.json.
WITH candidates AS (
SELECT p.vertex
FROM cugraph_personalized_pagerank(
edges => (
SELECT src_id AS src, dst_id AS dst
FROM ember_48_challenge_import_edges
),
personalization => (
SELECT id AS vertex, CAST(1.0 AS DOUBLE) AS value
FROM ember_49_challenge_graph_seed
),
directed => false
) p
JOIN ember_48_challenge_import_sample_counts c ON c.id = p.vertex
CROSS JOIN ember_49_challenge_graph_seed seed
WHERE p.vertex <> seed.id AND c.retained_import_count >= 5
ORDER BY p.value DESC, p.vertex
LIMIT 25
), similarities AS (
SELECT * FROM cugraph_jaccard(
edges => (
SELECT src_id AS src, dst_id AS dst
FROM ember_48_challenge_import_edges
),
vertex_pairs => (
SELECT seed.id AS first, c.vertex AS second
FROM candidates c CROSS JOIN ember_49_challenge_graph_seed seed
),
directed => false
)
)
SELECT j.second AS candidate_id, m.sha256, m.family,
counts.retained_import_count, j.similarity,
COUNT(*) AS shared_rare_imports,
MIN(symbols.symbol) AS example_shared_import
FROM similarities j
JOIN ember_48_challenge_import_edges seed_edges ON seed_edges.src_id = j.first
JOIN ember_48_challenge_import_edges candidate_edges
ON candidate_edges.src_id = j.second AND candidate_edges.dst_id = seed_edges.dst_id
JOIN ember_48_challenge_import_symbols symbols ON symbols.symbol_id = seed_edges.dst_id
JOIN ember_metadata_challenge m ON m.id = j.second
JOIN ember_48_challenge_import_sample_counts counts ON counts.id = j.second
GROUP BY j.second, m.sha256, m.family, counts.retained_import_count, j.similarity
ORDER BY j.similarity DESC, j.second
LIMIT 15;
Limits
- similarity is explicit-pair only; all-pairs candidate generation and all-pairs top-k search are not exposed
- vertex_pairs is a required relation with canonical first and second columns
- candidate-pair columns must match the graph vertex domain; logical string graphs accept Utf8, LargeUtf8, or Utf8View independently per column
- null candidate-pair values are rejected at execution and are never dropped
- candidate pairs are a multiset: duplicate, reversed, and self pairs remain distinct output rows
- result rows have no global ordering; use ORDER BY when order is required
- providing edges.weight selects weighted similarity; omitting it selects unit-weight similarity
- cuGraph requires directed=false so the graph is constructed as an undirected/symmetric view
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.