Cosine
UDTF: cugraph_cosine
Official cuGraph reference: C API
Compare the neighbor vectors of explicit vertex pairs using cosine similarity, with optional edge-weight contributions.
Quickstart
The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight, plus the registered relation candidate_pairs (the vertex_pairs relation with columns first, second). Substitute your own registered relations.
SELECT *
FROM cugraph_cosine(
edges => (SELECT src, dst FROM target_edges),
vertex_pairs => (SELECT first, second FROM candidate_pairs)
);
Inputs
Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.
Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.
Logical string side-input limitations:
- edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs
- candidate-pair columns must match the graph vertex domain; logical string graphs accept Utf8, LargeUtf8, or Utf8View independently per column
Arguments and options
Relation arguments
Vertex columns of vertex_pairs must use the same vertex domain as edges: the integer type of src and dst, or any listed string type when the endpoints are strings. Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.
Named value arguments
This UDTF has no algorithm-specific value arguments. Its inputs are the relation arguments above and the graph construction options below.
Graph construction options
This UDTF requires directed=false (undirected/symmetric graph); all other graph construction options follow the shared defaults documented in Graph Construction Options.
Output
These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.
Examples
- Citation network
- MerRec marketplace
This example runs on the citation network demo dataset.
Which papers share BERT's references despite different bibliography sizes?
Compare BERT's 2018 reference list with NLP papers from 2015 through 2020 that
have at least 10 references. Candidate papers use one of the listed field
labels. For bibliographies A and B, the cosine score is
|A ∩ B| / sqrt(|A| × |B|). Its denominator, the geometric mean of the two
list sizes, accounts for bibliographies of different lengths.
The UDTF requires an undirected graph. Positive IDs represent citing papers, and negative IDs represent cited-paper reference roles. This keeps each positive paper ID adjacent only to negative IDs for its references; papers that cite BERT remain positive vertices. The edge projection omits weights, giving each reference a unit contribution. Candidate pairs exclude records whose title matches BERT's, avoiding title-identical editions as separate suggestions.
CREATE OR REPLACE VIEW similarity_seed AS
SELECT paper_id, title, n_references
FROM papers
WHERE title = 'BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding'
AND year = 2018;
CREATE OR REPLACE VIEW similarity_candidates AS
SELECT p.paper_id, p.title, p.n_references, p.year
FROM papers p
WHERE p.year BETWEEN 2015 AND 2020
AND p.n_references >= 10
AND p.primary_fos IN (
'Natural language processing', 'Language model',
'Machine translation', 'Question answering');
CREATE OR REPLACE VIEW bibliography_edges AS
SELECT e.src, -e.dst AS dst
FROM citation_edges e
JOIN (
SELECT paper_id FROM similarity_seed
UNION
SELECT paper_id FROM similarity_candidates
) selected ON selected.paper_id = e.src;
SELECT p.year, p.title, p.n_references,
ROUND(r.similarity, 3) AS cosine
FROM cugraph_cosine(
edges => (SELECT src, dst FROM bibliography_edges),
vertex_pairs => (
SELECT s.paper_id AS first, c.paper_id AS second
FROM similarity_seed s
CROSS JOIN similarity_candidates c
WHERE LOWER(c.title) <> LOWER(s.title)),
directed => false) r
JOIN papers p ON p.paper_id = r.second
ORDER BY r.similarity DESC, p.paper_id
LIMIT 5;
MASS and BERT share 19 references. BERT has 55 references in the graph and
MASS has 53, so the score is 19 / sqrt(55 × 53) = 0.352.
The result ranks bibliography similarity; topic meaning is outside the score. For intersection divided by the union of references, see the Jaccard similarity example.
This example uses the prepared graph inputs described on the MerRec dataset and graph capture page.
Which markets draw most of their viewers from the Nintendo Games audience?
The seed audience is the 262 users who viewed Nintendo Games
(15928_1315). The query scores it against the audience of every other
market in mr_audience, running cosine and Jaccard over the same 1,710
pairs. For audiences A and B, cosine divides |A ∩ B| by
sqrt(|A| × |B|); Jaccard divides it by |A ∪ B|. Rows are ordered by
cosine, and jaccard_rank places each market among all 1,710 Jaccard scores.
WITH seed AS (
SELECT id FROM mr_nodes WHERE product_id = '15928_1315'
), pairs AS (
SELECT seed.id AS first, candidate.id AS second
FROM seed CROSS JOIN mr_nodes candidate
WHERE candidate.id <> seed.id
), cosine AS (
SELECT * FROM cugraph_cosine(
edges => (SELECT src, dst FROM mr_audience),
vertex_pairs => (SELECT first, second FROM pairs),
directed => false)
), jaccard AS (
SELECT second, RANK() OVER (ORDER BY similarity DESC) AS jaccard_rank
FROM cugraph_jaccard(
edges => (SELECT src, dst FROM mr_audience),
vertex_pairs => (SELECT first, second FROM pairs),
directed => false)
), shared AS (
SELECT b.src AS second, COUNT(*) AS shared_viewers
FROM seed
JOIN mr_audience a ON a.src = seed.id
JOIN mr_audience b ON b.dst = a.dst AND b.src <> seed.id
GROUP BY b.src
)
SELECT m.product_id, m.brand_name, m.leaf_category, m.viewers,
s.shared_viewers, c.similarity AS audience_cosine, j.jaccard_rank
FROM cosine c
JOIN jaccard j ON j.second = c.second
JOIN shared s ON s.second = c.second
JOIN mr_nodes m ON m.id = c.second
ORDER BY c.similarity DESC, m.product_id
LIMIT 10;
Cosine is rounded to six decimal places.
The two rankings split on small markets. Nintendo video game merchandise
has 19 viewers, and 18 of them also viewed Nintendo Games:
18 / sqrt(262 × 19) = 0.255 places it eighth by cosine. Its union with the
seed audience still holds 263 users, so Jaccard ranks it 32nd. Sega Games
moves from Jaccard rank 22 to cosine rank 9 for the same reason. For the
share of the smaller audience alone, see the
Overlap coefficient.
A CPU computation of the set formula over all 1,710 pairs differs from the
returned cosine scores by at most 4.06e-8. The score measures audience
overlap in one October 2023 shard; the query does not evaluate purchases,
conversions, or revenue.
Download the capture, including the executed SQL, schema, query ID, and all ten rows.
Limits
- similarity is explicit-pair only; all-pairs candidate generation and all-pairs top-k search are not exposed
- vertex_pairs is a required relation with canonical first and second columns
- candidate-pair columns must match the graph vertex domain; logical string graphs accept Utf8, LargeUtf8, or Utf8View independently per column
- null candidate-pair values are rejected at execution and are never dropped
- candidate pairs are a multiset: duplicate, reversed, and self pairs remain distinct output rows
- result rows have no global ordering; use ORDER BY when order is required
- providing edges.weight selects weighted similarity; omitting it selects unit-weight similarity
- cuGraph requires directed=false so the graph is constructed as an undirected/symmetric view
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.