Skip to main content

Personalized PageRank

UDTF: cugraph_personalized_pagerank

Official cuGraph reference: C API

Rank vertices with PageRank while biasing random-walk restarts toward explicitly weighted personalization vertices.

Quickstart​

The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight, plus the registered relation ppr_seeds (the personalization relation with columns vertex, value). Substitute your own registered relations.

SELECT *
FROM cugraph_personalized_pagerank(
edges => (SELECT src, dst FROM target_edges),
personalization => (SELECT vertex, value FROM ppr_seeds)
);

Inputs​

Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.

Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.

Logical string side-input limitations:

  • edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs

Arguments and options​

Relation arguments​

ArgumentRequiredColumnsDescription
edgesyes
  • src, dst: Int32, Int64, Utf8, LargeUtf8, Utf8View
  • weight (optional): Float32, Float64
  • edge_id (optional): Int32, Int64
edge relation with canonical src and dst columns, plus optional weight and edge_id columns
personalizationyes
  • vertex: Int32, Int64, Utf8, LargeUtf8, Utf8View
  • value: Float32, Float64
personalization relation with vertex and numeric value columns

Vertex columns of personalization must use the same vertex domain as edges: the integer type of src and dst, or any listed string type when the endpoints are strings. Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.

Named value arguments​

OptionTypeDefaultConstraintsDescription
alphanumber0.85min 0; max 1PageRank damping factor in [0, 1]
epsilonnumber0.00001> 0positive convergence tolerance
max_iterationsinteger100min 1; max 4294967295maximum iteration count, at least 1

Graph construction options​

Graph construction follows the shared defaults (directed=true, renumbering, python_cugraph policy) documented in Graph Construction Options.

Output​

ColumnTypeNullableDescription
vertexInt64|Utf8noVertex receiving the PageRank score.
valueFloat64noPageRank score for the vertex.

These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.

Examples​

This example runs on the citation network demo dataset.

What should I read to understand BERT beyond its own bibliography?​

A reading list can include work several citation links beyond a paper's direct references. This query starts from the 2018 BERT record, then excludes BERT and the papers it cites so the results focus on more distant references. The same title appears on a 2019 conference record; the seed also filters by year.

CREATE OR REPLACE VIEW bert_seed AS
SELECT paper_id AS vertex, CAST(1.0 AS DOUBLE) AS value
FROM papers
WHERE title = 'BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding'
AND year = 2018;

SELECT ROUND(r.value, 6) AS ppr, p.year, p.title
FROM cugraph_personalized_pagerank(
edges => (SELECT src, dst FROM citation_edges),
personalization => (SELECT vertex, value FROM bert_seed)) r
JOIN papers p ON p.paper_id = r.vertex
WHERE NOT EXISTS (SELECT 1 FROM bert_seed s WHERE s.vertex = r.vertex)
AND NOT EXISTS (SELECT 1 FROM citation_edges e JOIN bert_seed s ON s.vertex = e.src
WHERE e.dst = r.vertex)
ORDER BY r.value DESC
LIMIT 8;
ppryeartitle
0.0031271983A Maximum Likelihood Approach to Continuous Speech Recognition
0.0029842003A neural probabilistic language model
0.0023921997Long short-term memory
0.0023582014Adam: A Method for Stochastic Optimization
0.0023441990A statistical approach to machine translation
0.0019881993Building a large annotated corpus of English: the penn treebank
0.0019231975Design of a linguistic statistical decoder for the recognition of continuous speech
0.0018722006The PASCAL Recognising Textual Entailment Challenge

The returned list includes work on speech recognition, statistical machine translation, and the Penn Treebank. Those titles provide examples of papers ranked beyond BERT's direct references.

Limits​

No algorithm-specific limitations.

To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.