Citation Network
The DBLP / AMiner V12 citation network holds 4.9 million papers and 45.6
million citations. An edge src → dst means paper src cites paper dst.
Every cuGraph function page runs its examples on these tables, and the
integration guides use them for their first query.
The data is a fixed snapshot from April 2020. Graph scores describe a paper's position in this snapshot and say nothing about its quality.
At a glance
Examples
- Rank papers by PageRank, the first query on this page.
- Find a citation path from BERT to LSTM.
- Compare citation communities with research fields.
- Find papers with references similar to the Transformer.
The remaining cuGraph pages, from degree counts to force-directed layout, read the same tables.
Prepare the data
Use a Linux GPU host with the native libraries built, and run shell commands from the repository root. GPU memory requirements depend on the query; the figures above describe host resources.
Download and extract dblp.v12.json from the
DBLP / AMiner V12 dataset.
Place it at fixture/demo/citation_network/raw/dblp.v12.json, then convert it:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install pyarrow
fixture/fixture.sh graph dblp-ingest
The output is fixture/demo/citation_network/v1/parquet/. Skip this step
when those files already exist. The
conversion reference
includes the source checksum and layout details.
Start the server
Build the server and SQL client:
cargo build --workspace --all-features --release --bin algeon_server --bin algeon_tools
Create a demo configuration for GPU 0 and start the server on localhost:
mkdir -p target
cat > target/citation-demo.toml <<'TOML'
[native]
execution_mode = "functions_only"
[[admission.device_profiles]]
device_ordinal = 0
TOML
ALGEON_SERVER_CONFIG_FILE="$PWD/target/citation-demo.toml" \
ALGEON_SERVER_GPU_DEVICES=0 \
ALGEON_SERVER_BIND=127.0.0.1:50051 \
ALGEON_SERVER_CUGRAPH_ENABLED=true \
flock /tmp/cudf-gpu.lock "${CARGO_TARGET_DIR:-target}/release/algeon_server"
Leave this terminal running. This configuration runs explicit cuGraph calls on the GPU and ordinary SQL on the CPU.
Register the tables
Open another terminal at the repository root. Define a shell helper for the SQL client and check the connection:
sql() {
"${CARGO_TARGET_DIR:-target}/release/algeon_tools" \
flight-sql-query http://127.0.0.1:50051 "${1:-$(cat)}"
}
sql "SELECT 1 AS one"
The result should be a column named one containing 1. Register the files:
citation_data="$PWD/fixture/demo/citation_network/v1/parquet"
sql "CREATE OR REPLACE EXTERNAL TABLE citation_edges STORED AS PARQUET LOCATION '$citation_data/edges_by_src.parquet'"
sql "CREATE OR REPLACE EXTERNAL TABLE citation_edges_by_dst STORED AS PARQUET LOCATION '$citation_data/edges_by_dst.parquet'"
sql "CREATE OR REPLACE EXTERNAL TABLE papers STORED AS PARQUET LOCATION '$citation_data/profiles_by_paper_id.parquet'"
sql "CREATE OR REPLACE EXTERNAL TABLE paper_authors STORED AS PARQUET LOCATION '$citation_data/paper_authors.parquet'"
sql "CREATE OR REPLACE EXTERNAL TABLE paper_fos STORED AS PARQUET LOCATION '$citation_data/paper_fos.parquet'"
Table registrations last until the server stops. After restarting, run this registration block again; the Parquet files can be reused.
Run the examples
PageRank weights incoming citations by the scores of the citing papers. Run this in the client terminal:
sql <<'SQL'
SELECT p.title, p.year, r.value AS pagerank
FROM cugraph_pagerank(edges => (SELECT src, dst FROM citation_edges)) r
JOIN papers p ON p.paper_id = r.vertex
ORDER BY r.value DESC
LIMIT 5;
SQL
A run recorded on 2026-09-26 returned these rows (scores rounded):
For the other examples, replace the query between sql <<'SQL' and SQL, or
connect an interactive SQL client
to the running server. Execute multi-statement examples one statement at a
time. Examples use CREATE OR REPLACE for reusable views and result tables,
and running another example can replace shared names such as ai_nodes and
ai_edges.
Tables
papers.n_citation is AMiner's reported global citation count. In-graph
degree counts only edges present here; the
in-degree example
compares them.
Seed papers used in the examples
Use paper_id as the vertex ID. Titles can repeat across distinct records;
the examples resolve seeds by title and year, adding venue when needed.
String equality is exact and case-sensitive: Attention Is All You Need
(capitalized) matches a separate arXiv record with a different id. Copy titles
from this table or from a query on papers, not from the published paper.