BFS
UDTF: cugraph_bfs
Official cuGraph reference: C API
Visit reachable vertices in increasing unweighted hop distance from one or more sources, returning distances and optional predecessors.
Quickstart
The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight, and starts from source vertex 123. Substitute your own registered relations.
SELECT *
FROM cugraph_bfs(
edges => (SELECT src, dst FROM target_edges),
source_vertex => 123,
depth_limit => 4
);
Inputs
Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.
Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.
Logical string side-input limitations:
- edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs
- string logical vertex-domain BFS rejects edges.edge_id and edge-id predicate relations
Arguments and options
Relation arguments
Vertex columns of sources, include_vertices, exclude_vertices, target_vertices, include_edges must use the same vertex domain as edges: the integer type of src and dst, or any listed string type when the endpoints are strings. Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.
Named value arguments
Graph construction options
Graph construction follows the shared defaults (directed=true, renumbering, python_cugraph policy) documented in Graph Construction Options.
Output
raw (default)
normalized
normalized_with_target_info
path
raw_with_target_info
These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.
Examples
These examples run on the citation network demo dataset.
Edges point src → dst as "cites"; BFS distance is a hop count.
How far did AlexNet spread through later citation generations?
Reversing each citation edge follows papers that cite AlexNet, then papers that cite those citers. The result counts papers and reports the average publication year at each of the first three citation steps.
SELECT b.distance, COUNT(*) AS papers, ROUND(AVG(p.year), 1) AS avg_year
FROM cugraph_bfs(
edges => (SELECT dst AS src, src AS dst FROM citation_edges_by_dst),
sources => (SELECT paper_id AS vertex FROM papers
WHERE title = 'ImageNet Classification with Deep Convolutional Neural Networks'
AND year = 2012),
depth_limit => 3, output_mode => 'normalized') b
JOIN papers p ON p.paper_id = b.vertex
WHERE b.reachable
GROUP BY b.distance
ORDER BY b.distance;
Three citation generations reach ~151k papers. (The reversed traversal reads
citation_edges_by_dst, which is clustered by dst, so the scan prunes well.)
What is the shortest citation chain from BERT back to LSTM?
Path mode returns one shortest hop chain from the 2018 BERT record to the 1997 LSTM paper. The endpoint relations use title and year because the corpus also contains a 2019 conference record with the BERT preprint's title.
SELECT b.path_index, b.distance, p.year, p.title
FROM cugraph_bfs(
edges => (SELECT src, dst FROM citation_edges),
sources => (SELECT paper_id AS vertex FROM papers
WHERE title = 'BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding'
AND year = 2018),
output_mode => 'path',
target_vertices => (SELECT paper_id AS vertex FROM papers
WHERE title = 'Long short-term memory' AND year = 1997)) b
JOIN papers p ON p.paper_id = b.vertex
ORDER BY b.path_index;
The result is one of several two-hop citation chains. Other valid paths pass through SemEval-2017 Task 1 or Aligning Books and Movies.
How many papers are within two citation steps of AlexNet, VGG, or ResNet?
The three papers form one source set. Each result's distance is the number of reverse citation steps from its nearest source. Title and year identify the 2016 ResNet record because a 2015 preprint shares its title.
CREATE OR REPLACE VIEW cnn_founders AS
SELECT paper_id AS vertex FROM papers
WHERE (title = 'ImageNet Classification with Deep Convolutional Neural Networks' AND year = 2012)
OR (title = 'VERY DEEP CONVOLUTIONAL NETWORKS FOR LARGE-SCALE IMAGE RECOGNITION' AND year = 2014)
OR (title = 'Deep Residual Learning for Image Recognition' AND year = 2016);
SELECT b.distance, COUNT(*) AS papers
FROM cugraph_bfs(
edges => (SELECT dst AS src, src AS dst FROM citation_edges_by_dst),
sources => (SELECT vertex FROM cnn_founders), depth_limit => 2,
output_mode => 'normalized') b
JOIN papers p ON p.paper_id = b.vertex
WHERE b.reachable
GROUP BY b.distance
ORDER BY b.distance;
Would using paper titles as graph IDs merge different records?
This query compares record IDs with titles in the reference ancestry of Attention Is All You Need, then runs BFS with string titles as vertex IDs. The count shows whether this subgraph contains repeated titles; identical titles alone do not establish that records describe the same work.
CREATE OR REPLACE VIEW attention_ancestry AS
WITH origins AS (
SELECT paper_id FROM papers
WHERE title = 'Attention is all you need' AND year = 2017
UNION ALL
SELECT e.dst FROM citation_edges e
JOIN papers p ON p.paper_id = e.src
WHERE p.title = 'Attention is all you need' AND p.year = 2017)
SELECT e.src, e.dst
FROM citation_edges e JOIN origins o ON o.paper_id = e.src;
SELECT COUNT(DISTINCT n.paper_id) AS papers, COUNT(DISTINCT p.title) AS titles
FROM (SELECT src AS paper_id FROM attention_ancestry
UNION SELECT dst FROM attention_ancestry) n
JOIN papers p ON p.paper_id = n.paper_id;
SELECT b.distance, b.vertex
FROM cugraph_bfs(
edges => (SELECT ps.title AS src, pd.title AS dst
FROM attention_ancestry a
JOIN papers ps ON ps.paper_id = a.src
JOIN papers pd ON pd.paper_id = a.dst),
source_vertex => 'Attention is all you need', output_mode => 'normalized') b
WHERE b.distance <= 1
ORDER BY b.distance, b.vertex
LIMIT 6;
This subgraph has 370 records and 358 distinct titles. A title-keyed graph
collapses records that share a title, including unrelated works when their
titles collide. Use that grouping only when merging every record with the same
title matches the question. Here BFS returns the title as a Utf8 vertex:
String graphs reject edge_id columns and edge-ID predicate relations; see the
Vertex ID support matrix.
Limits
- source_vertices and the sources relation are multi-source BFS, not one BFS per source
- source_vertices and the sources relation must contain at most one source per connected component
- raw/normalized output does not expose origin source per vertex
- string logical vertex-domain BFS rejects edges.edge_id and edge-id predicate relations
- path output requires exactly one target row at execution time
- descriptor option schema is typed registry metadata; validate_call remains the authoritative call-specific checker
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.