Skip to main content

Sorensen

UDTF: cugraph_sorensen

Official cuGraph reference: C API

Compare explicit vertex pairs as twice their shared-neighbor count divided by the sum of their neighbor counts.

Quickstart​

The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight, plus the registered relation candidate_pairs (the vertex_pairs relation with columns first, second). Substitute your own registered relations.

SELECT *
FROM cugraph_sorensen(
edges => (SELECT src, dst FROM target_edges),
vertex_pairs => (SELECT first, second FROM candidate_pairs)
);

Inputs​

Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.

Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.

Logical string side-input limitations:

  • edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs
  • candidate-pair columns must match the graph vertex domain; logical string graphs accept Utf8, LargeUtf8, or Utf8View independently per column

Arguments and options​

Relation arguments​

ArgumentRequiredColumnsDescription
edgesyes
  • src, dst: Int32, Int64, Utf8, LargeUtf8, Utf8View
  • weight (optional): Float32, Float64
  • edge_id (optional): Int32, Int64
edge relation with canonical src and dst columns, plus optional weight and edge_id columns
vertex_pairsyes
  • first, second: Int32, Int64, Utf8, LargeUtf8, Utf8View
explicit vertex pairs used by similarity scoring

Vertex columns of vertex_pairs must use the same vertex domain as edges: the integer type of src and dst, or any listed string type when the endpoints are strings. Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.

Named value arguments​

This UDTF has no algorithm-specific value arguments. Its inputs are the relation arguments above and the graph construction options below.

Graph construction options​

This UDTF requires directed=false (undirected/symmetric graph); all other graph construction options follow the shared defaults documented in Graph Construction Options.

Output​

ColumnTypeNullableDescription
firstInt64|Utf8noFirst vertex from the explicit candidate-pair relation.
secondInt64|Utf8noSecond vertex from the explicit candidate-pair relation.
similarityFloat64noSimilarity coefficient for the explicit candidate pair.

These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.

Examples​

This example runs on the citation network demo dataset.

Which influential NLP papers would make a reading pair with shared background?​

Take the 12 most-cited papers from 2015 through 2020 in the listed NLP fields, then compare each distinct-title pair by its bibliography overlap. Sørensen similarity weights the intersection by the sum of both reference-list sizes. The score describes shared citation background; project-specific relevance falls outside the measure.

The UDTF requires an undirected graph. Positive IDs identify citing papers, while negative IDs identify reference-role vertices for cited papers. Outgoing bibliography neighbors are the negative IDs adjacent to each positive paper ID; citing papers remain on the positive-ID side. The edge projection omits weights, so each reference counts as one unit.

CREATE OR REPLACE VIEW similarity_candidates AS
SELECT paper_id, title, n_references
FROM papers
WHERE year BETWEEN 2015 AND 2020
AND n_references >= 10
AND primary_fos IN (
'Natural language processing', 'Language model',
'Machine translation', 'Question answering')
ORDER BY n_citation DESC, paper_id
LIMIT 12;

CREATE OR REPLACE VIEW bibliography_edges AS
SELECT e.src, -e.dst AS dst
FROM citation_edges e
JOIN similarity_candidates c ON c.paper_id = e.src;

SELECT a.title AS paper_a, b.title AS paper_b,
ROUND(r.similarity, 3) AS sorensen
FROM cugraph_sorensen(
edges => (SELECT src, dst FROM bibliography_edges),
vertex_pairs => (
SELECT a.paper_id AS first, b.paper_id AS second
FROM similarity_candidates a
JOIN similarity_candidates b ON a.paper_id < b.paper_id
WHERE LOWER(a.title) <> LOWER(b.title)),
directed => false) r
JOIN papers a ON a.paper_id = r.first
JOIN papers b ON b.paper_id = r.second
ORDER BY r.similarity DESC, r.first, r.second
LIMIT 5;
paper_apaper_bsorensen
VQA: Visual Question AnsweringVisual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations0.253
Hierarchical Question-Image Co-Attention for Visual Question AnsweringStacked Attention Networks for Image Question Answering0.157
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image AnnotationsStacked Attention Networks for Image Question Answering0.138
Character-aware neural language modelsNeural Machine Translation of Rare Words with Subword Units0.137
VQA: Visual Question AnsweringStacked Attention Networks for Image Question Answering0.125

The strongest pair in this pool is VQA: Visual Question Answering with Visual Genome, at a Sørensen score of 0.253.

The coefficient is twice the shared-reference count divided by the sum of both reference counts. Project-specific relevance is outside this measure. For a score based on the union of the two lists, see the Jaccard similarity example.

Limits​

  • similarity is explicit-pair only; all-pairs candidate generation and all-pairs top-k search are not exposed
  • vertex_pairs is a required relation with canonical first and second columns
  • candidate-pair columns must match the graph vertex domain; logical string graphs accept Utf8, LargeUtf8, or Utf8View independently per column
  • null candidate-pair values are rejected at execution and are never dropped
  • candidate pairs are a multiset: duplicate, reversed, and self pairs remain distinct output rows
  • result rows have no global ordering; use ORDER BY when order is required
  • providing edges.weight selects weighted similarity; omitting it selects unit-weight similarity
  • cuGraph requires directed=false so the graph is constructed as an undirected/symmetric view

To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.