Strongly Connected Components
UDTF: cugraph_strongly_connected_components
Official cuGraph reference: C API
Label maximal directed subgraphs in which every vertex is reachable from every other vertex.
Quickstart
The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight. Substitute your own registered relations.
SELECT *
FROM cugraph_strongly_connected_components(
edges => (SELECT src, dst FROM target_edges)
);
Inputs
Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.
Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.
Logical string side-input limitations:
- edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs
Arguments and options
Relation arguments
Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.
Named value arguments
This UDTF has no algorithm-specific value arguments. Its inputs are the relation arguments above and the graph construction options below.
Graph construction options
Graph construction follows the shared defaults (directed=true, renumbering, python_cugraph policy) documented in Graph Construction Options.
Output
These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.
Examples
This example runs on the citation network demo dataset.
Which mutually citing paper pairs have the largest publication-year gaps?
Before using citation direction as a timeline, inspect pairs whose references run both ways despite a large gap in publication years. A strongly connected component with two vertices contains such a pair. These are candidates for checking editions and publication metadata, not proof of an incorrect citation.
Save one SCC result so the pair selection and both paper lookups use labels from the same execution. The query excludes years outside 1900 through 2020 and returns one row per pair.
CREATE OR REPLACE TABLE scc_snapshot AS
SELECT vertex, label
FROM cugraph_strongly_connected_components(
edges => (SELECT src, dst FROM citation_edges), directed => true);
WITH pairs AS (
SELECT label FROM scc_snapshot GROUP BY label HAVING COUNT(*) = 2)
SELECT pa.year AS year_a, pa.title AS paper_a,
pb.year AS year_b, pb.title AS paper_b,
ABS(pa.year - pb.year) AS year_gap
FROM pairs x
JOIN scc_snapshot a ON a.label = x.label
JOIN scc_snapshot b ON b.label = x.label AND a.vertex < b.vertex
JOIN papers pa ON pa.paper_id = a.vertex
JOIN papers pb ON pb.paper_id = b.vertex
WHERE pa.year BETWEEN 1900 AND 2020 AND pb.year BETWEEN 1900 AND 2020
ORDER BY year_gap DESC, a.vertex, b.vertex
LIMIT 5;
The largest gap among these isolated mutually citing pairs is 35 years, between floating-point addition and IBM System/360 architecture papers.
The year gap prioritizes records for review. Larger strongly connected components can also contain mutual citations; this query selects isolated two-paper components to keep each case directly inspectable.
Limits
No algorithm-specific limitations.
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.