PageRank
UDTF: cugraph_pagerank
Official cuGraph reference: C API
Rank vertices by the stationary probability of a damped random walk that follows outgoing edges.
Quickstart
The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight. Substitute your own registered relations.
SELECT *
FROM cugraph_pagerank(
edges => (SELECT src, dst FROM target_edges)
);
Inputs
Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.
Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.
Logical string side-input limitations:
- edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs
Arguments and options
Relation arguments
Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.
Named value arguments
Graph construction options
Graph construction follows the shared defaults (directed=true, renumbering, python_cugraph policy) documented in Graph Construction Options.
Output
These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.
Examples
- Citation network
- MerRec marketplace
These examples run on the citation network demo dataset
(4.9M papers, 45.6M src-cites-dst edges).
Which papers rank highest when citations from influential papers count more?
Raw citation counts treat each incoming citation equally. PageRank weights citations by the scores of the citing papers, and the metadata join lets us compare that ranking with each paper's reported citation count.
SELECT p.title, p.year, p.n_citation, r.value AS pagerank
FROM cugraph_pagerank(edges => (SELECT src, dst FROM citation_edges)) r
JOIN papers p ON p.paper_id = r.vertex
ORDER BY r.value DESC
LIMIT 5;
The first and third rows have reported counts of 1,401 and 224. Their positions show how the influence-based ranking can differ from a simple citation count.
Where does the original PageRank paper rank by its own algorithm?
Ranking the paper that introduced PageRank by its own method gives a concrete reference point in this corpus. The corpus has two 1998 records with this title, so the query selects the Web Conference record by year and venue.
WITH ranked AS (
SELECT vertex, value, ROW_NUMBER() OVER (ORDER BY value DESC) AS rank
FROM cugraph_pagerank(edges => (SELECT src, dst FROM citation_edges)))
SELECT r.rank, p.title, p.year, r.value
FROM ranked r JOIN papers p ON p.paper_id = r.vertex
WHERE p.title = 'The anatomy of a large-scale hypertextual Web search engine'
AND p.year = 1998 AND p.venue = 'The Web Conference';
In this run, the Web Conference record ranks #115 among vertices scored by PageRank on this graph.
Which papers were central to the citation network through 2000?
An era-specific ranking can surface papers that structure literature from that period. This query keeps citations whose source and destination papers are both dated 1901 through 2000, then ranks that subgraph with PageRank.
CREATE OR REPLACE VIEW edges_pre2000 AS
SELECT e.src, e.dst
FROM citation_edges e
JOIN papers ps ON ps.paper_id = e.src
JOIN papers pd ON pd.paper_id = e.dst
WHERE ps.year BETWEEN 1901 AND 2000 AND pd.year BETWEEN 1901 AND 2000;
SELECT p.year, p.title
FROM cugraph_pagerank(edges => (SELECT src, dst FROM edges_pre2000)) r
JOIN papers p ON p.paper_id = r.vertex
ORDER BY r.value DESC
LIMIT 6;
Automata theory, information theory, the classic algorithms textbook, and the
ALGOL report: the view's WHERE clause restricts the graph to one era, and
the algorithm re-ranks the field within it.
This example uses the October 2023 MerRec fixture and graph preparation.
Which browsing neighborhoods could become draft marketplace collections?
This query joins two graph views of observed item_view transitions. Leiden
groups product groups on an unweighted, undirected graph. PageRank scores
markets on a directed graph whose transition weights count distinct viewers.
For each Leiden group with at least ten markets, the query returns the five
markets with the highest weighted PageRank.
The default community 11 contains 58 grouped markets. Its five representatives include Nintendo games, Pokemon single cards, Sony games, Nintendo consoles, and Microsoft games. Other captured groups include Disney pins and stuffed animals, Squishmallows, Loungefly, and fashion dolls. Merchandising teams can review these cross-category patterns when drafting navigation collections from observed browsing. This query does not evaluate purchases, conversion, or sales.
The graph's product_id represents a brand and leaf-category group; item_id
identifies an individual listing. viewers counts distinct users with an
item_view in that group. Missing brand and leaf-category values appear as
“Brand not listed” and “Leaf category not listed” in the table.
Community 11
58 grouped marketsFive market groups ranked by weighted PageRank within this captured community.
| Rank | Product group | Brand | Category | Leaf category | Viewers | PageRank |
|---|---|---|---|---|---|---|
| 1 | 15928_1315 | Nintendo | Electronics | Games | 262 | 0.004766 |
| 2 | 17470_3525 | Pokemon | Toys & Collectibles | Single Cards | 135 | 0.004078 |
| 3 | 20652_1315 | Sony | Electronics | Games | 170 | 0.003160 |
| 4 | 15928_1314 | Nintendo | Electronics | Consoles | 180 | 0.002297 |
| 5 | 14034_1315 | Microsoft | Electronics | Games | 87 | 0.001692 |
Community numbers are labels from this query's Leiden fit. They describe this
capture and do not establish run-to-run stability. The CPU assessment reports
zero market-metadata mismatches, PageRank maximum absolute error 0.0 against
the saved full PageRank output, and correct five-row rank ordering for each of
the 22 returned groups. The validation does not prove that query 24's labels
match a separately executed full Leiden assignment, because that query fits
Leiden independently.
The captured plan has two native GPU fragments, Leiden and PageRank, and 15
host DataFusion boundaries for joins, repartitioning, and ranked-window
processing. The single observed Flight elapsed time was 0.2762586469762027
seconds. It includes planning and complete result receipt, excludes graph
preparation, and came from a reused server after many previous queries. It is
not a performance benchmark.
Download all 110 result rows, including the executed SQL, schema, query ID, timing boundary, and source SHA-256 values. The CPU community validation records its comparison scope and limits.
Executed SQL, query 6490d83e-5bce-4ada-ad7a-77c742ecebab
-- Review representative markets from each sufficiently large Leiden community.
-- Community labels are query-local; PageRank uses directed weighted transitions.
WITH assignments AS (
SELECT
r.vertex,
r."partition" AS community,
COUNT(*) OVER (PARTITION BY r."partition") AS community_size
FROM cugraph_leiden(
edges => (SELECT src, dst FROM mr_undirected),
directed => false,
seed => 42
) r
), weighted_pagerank AS (
SELECT vertex, value AS pagerank
FROM cugraph_pagerank(
edges => (SELECT src, dst, weight FROM mr_edges),
directed => true
)
), ranked_markets AS (
SELECT
a.vertex,
a.community,
a.community_size,
ROW_NUMBER() OVER (
PARTITION BY a.community
ORDER BY p.pagerank DESC, n.product_id ASC
) AS representative_rank,
n.product_id,
n.brand_name,
n.category,
n.leaf_category,
n.viewers,
p.pagerank
FROM assignments a
JOIN mr_nodes n ON n.id = a.vertex
JOIN weighted_pagerank p ON p.vertex = a.vertex
)
SELECT
vertex,
community,
community_size,
representative_rank,
product_id,
brand_name,
category,
leaf_category,
viewers,
pagerank
FROM ranked_markets
WHERE community_size >= 10
AND representative_rank <= 5
ORDER BY community, representative_rank;
Limits
No algorithm-specific limitations.
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.