ForceAtlas2
UDTF: cugraph_force_atlas2
Official cuGraph reference: C API
Place vertices in two dimensions with a force-directed simulation that attracts connected vertices and repels vertices from one another.
Quickstart
The call below supplies edges from registered relation target_edges with canonical src and dst columns and may include weight. Substitute your own registered relations.
SELECT *
FROM cugraph_force_atlas2(
edges => (SELECT src, dst FROM target_edges)
);
Inputs
Every relation is a named parenthesized SELECT subquery. The required edges role uses canonical src and dst columns; every role, its canonical columns, and their accepted Arrow types are listed under Relation arguments. Metadata validation resolves registered tables named in its JSON request.
Endpoint columns accept numeric Int32, Int64 vertex IDs or logical string Utf8, LargeUtf8, Utf8View vertex IDs; string vertex-identity outputs are canonicalized to Utf8 (native mapping Int64) while scores, distances, counts, coordinates, and opaque labels stay numeric. The shared vertex-ID contract is summarized in Vertex ID support; the concrete call-specific schema comes from gpu_validate_call.
Logical string side-input limitations:
- edge ID columns and edge-ID predicate side inputs are not supported for logical string graphs
Arguments and options
Relation arguments
Vertex columns of vertex_attributes, initial_positions must use the same vertex domain as edges: the integer type of src and dst, or any listed string type when the endpoints are strings. Every edge_id column must have the integer type of src and dst; string-keyed graphs accept no edge IDs.
Named value arguments
Graph construction options
This UDTF requires directed=false (undirected/symmetric graph); all other graph construction options follow the shared defaults documented in Graph Construction Options.
Output
These are generic descriptor schemas; run gpu_validate_call to get the concrete, table-specific output schema.
Examples
This example runs on the citation network demo dataset.
How are the papers around the Transformer connected?
The selected graph contains Attention Is All You Need, its 22 references, and its 110 most-cited citers. The edge view keeps citations whose endpoints both belong to this set. ForceAtlas2 supplies coordinates and Louvain assigns a partition to each paper; the scope excludes other papers and all but those 110 citers.
CREATE OR REPLACE VIEW attention_seed AS
SELECT paper_id FROM papers
WHERE title = 'Attention is all you need' AND year = 2017;
CREATE OR REPLACE VIEW attention_ego_nodes AS
SELECT paper_id FROM (
SELECT e.dst AS paper_id
FROM citation_edges e JOIN attention_seed s ON s.paper_id = e.src
UNION ALL
SELECT src AS paper_id FROM (
SELECT e.src, p.n_citation
FROM citation_edges_by_dst e
JOIN attention_seed s ON s.paper_id = e.dst
JOIN papers p ON p.paper_id = e.src
ORDER BY p.n_citation DESC, e.src LIMIT 110) t
UNION ALL
SELECT paper_id FROM attention_seed
) u GROUP BY paper_id;
CREATE OR REPLACE VIEW attention_ego_edges AS
SELECT e.src, e.dst
FROM citation_edges e
JOIN attention_ego_nodes a ON a.paper_id = e.src
JOIN attention_ego_nodes b ON b.paper_id = e.dst;
WITH layout AS (
SELECT vertex, x, y
FROM cugraph_force_atlas2(
edges => (SELECT src, dst FROM attention_ego_edges), max_iter => 500, seed => 42)),
community AS (
SELECT vertex, "partition"
FROM cugraph_louvain(edges => (SELECT src, dst FROM attention_ego_edges)))
SELECT l.vertex, l.x, l.y, c."partition", p.title, p.year
FROM layout l
JOIN community c ON c.vertex = l.vertex
JOIN papers p ON p.paper_id = l.vertex;
The figure below renders that query's actual output (133 rows of
(vertex, x, y, partition, title, year)) with no client-side layout; the
browser draws only what the SQL returned. Louvain's partitions correspond to
distinct research threads (labels assigned by inspecting each cluster's
members):
seed fixes the initial placement, but the parallel layout itself is not
bit-reproducible; expect different (equally valid) coordinates on each run;
on larger graphs individual vertices may land far apart. If downstream queries
must agree on positions, save the result with CREATE OR REPLACE TABLE … AS.
The result table occupies host
RAM and lasts until the server stops.
Can I spread overlapping papers apart without starting the map over?
The first pass saves coordinates in a table. The second pass uses that table as its initial positions and supplies radius and mass for every vertex, then enables overlap prevention:
CREATE OR REPLACE TABLE initial_layout AS
SELECT vertex, x, y
FROM cugraph_force_atlas2(
edges => (SELECT src, dst FROM attention_ego_edges), max_iter => 250, seed => 42);
CREATE OR REPLACE VIEW layout_attributes AS
SELECT vertex,
CAST(0.75 AS REAL) AS radius,
CAST(1.0 AS REAL) AS mass
FROM initial_layout;
SELECT vertex, x, y
FROM cugraph_force_atlas2(
edges => (SELECT src, dst FROM attention_ego_edges), max_iter => 250, seed => 42,
prevent_overlapping => true,
vertex_attributes => (SELECT vertex, radius, mass FROM layout_attributes),
initial_positions => (SELECT vertex, x, y FROM initial_layout));
Both side inputs reject null, duplicate, missing, or extra vertices. Radius,
mass, mobility, and coordinates are Float32; their vertex column must match
the edge endpoint domain.
Limits
- vertex_attributes and initial_positions must contain exactly one non-null row for every graph vertex
- duplicate, missing, and extra ForceAtlas2 side-input vertices are rejected at execution
- radius, mobility, mass, x, and y side-input columns must be Float32
- prevent_overlapping=true requires a radius binding
To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.