ECG
UDTF: cugraph_ecg
Official cuGraph reference: C API
Stabilize community assignments by combining an ensemble of randomized Louvain partitions into a final consensus clustering.
Quickstart
The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight. Substitute your own registered relations.
SELECT *
FROM cugraph_ecg(
edges => (SELECT src, dst FROM target_edges)
);
Inputs
Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.
Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.
Logical string side-input limitations:
- edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs
Arguments and options
Relation arguments
Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.
Named value arguments
Graph construction options
This UDTF requires directed=false (undirected/symmetric graph); all other graph construction options follow the shared defaults documented in Graph Construction Options.
Output
These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.
Examples
This example runs on the citation network demo dataset.
Which grouping best matches our catalog's research-field labels?
This comparison uses the share of members carrying their community's most
common primary_fos label as a diagnostic against the catalog. It includes
communities with at least 100 members in the purity calculation; community
sizes and that cutoff affect the result. The field labels provide a diagnostic
reference for the score, while citation structure defines the communities.
The query runs Louvain, Leiden, and ECG on the same 2010s AI citation subgraph.
CREATE OR REPLACE VIEW ai_nodes AS
SELECT paper_id FROM papers
WHERE year >= 2010 AND primary_fos IN (
'Deep learning', 'Artificial neural network', 'Convolutional neural network',
'Recurrent neural network', 'Natural language processing',
'Reinforcement learning', 'Image segmentation', 'Feature extraction',
'Object detection', 'Speech recognition');
CREATE OR REPLACE VIEW ai_edges AS
SELECT e.src, e.dst
FROM citation_edges e
JOIN ai_nodes a ON a.paper_id = e.src
JOIN ai_nodes b ON b.paper_id = e.dst;
WITH labeled AS (
SELECT 'louvain' AS algorithm, c."partition" AS community, p.primary_fos
FROM cugraph_louvain(edges => (SELECT src, dst FROM ai_edges)) c
JOIN papers p ON p.paper_id = c.vertex
UNION ALL
SELECT 'leiden', c."partition", p.primary_fos
FROM cugraph_leiden(edges => (SELECT src, dst FROM ai_edges)) c
JOIN papers p ON p.paper_id = c.vertex
UNION ALL
SELECT 'ecg', c."partition", p.primary_fos
FROM cugraph_ecg(edges => (SELECT src, dst FROM ai_edges)) c
JOIN papers p ON p.paper_id = c.vertex),
counts AS (
SELECT algorithm, community, primary_fos, COUNT(*) AS n
FROM labeled GROUP BY 1, 2, 3),
sized AS (
SELECT algorithm, community, SUM(n) AS members, MAX(n) AS top_label
FROM counts GROUP BY 1, 2)
SELECT algorithm,
COUNT(*) AS communities,
COUNT(*) FILTER (WHERE members >= 100) AS ge100,
ROUND(SUM(top_label) FILTER (WHERE members >= 100) * 100.0
/ SUM(members) FILTER (WHERE members >= 100), 1) AS purity_pct
FROM sized
GROUP BY algorithm
ORDER BY purity_pct DESC;
ECG has the highest reported purity in this run (47.7%), with 41 communities
of at least 100 members. Community size affects this diagnostic. The default
seed is 0, but it does not guarantee identical labels across GPU runs.
Limits
No algorithm-specific limitations.
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.