Skip to main content

MerRec Marketplace

MerRec records browsing and purchase events from the Mercari US marketplace. The examples use one shard from October 2023 and turn it into a graph of product groups that users browse in sequence, plus a table of per-user category preferences. Graph, clustering, and projection functions then run over those tables.

The examples measure browsing and audience overlap. They leave purchase and conversion outcomes unevaluated, and none of the captures estimates revenue uplift.

At a glance​

ItemDetails
Sourcemercari-us/merrec, one shard of the October 2023 partition
Scale494,357 events, 358,893 items, and 5,660 users from October 1 to October 30, 2023
PreparationThe fixture script downloads the sample shard; graph and preference tables are built by SQL on the running server
Python packagespyarrow, and the duckdb CLI on PATH for conversion
Server127.0.0.1:50063; the research configuration uses a 32 GiB device limit and 16 GiB per query
ExamplescuGraph personalized PageRank, Jaccard, cosine, Leiden, and PageRank; cuML HDBSCAN and UMAP; cuVS KMeans and Top-K

Examples​

Prepare the data​

Run commands from the repository root. The fixture README documents acquisition, conversion, and source columns. This command downloads the sample shard when raw source files are absent, then converts it into events and items Parquet tables:

fixture/fixture.sh merrec prepare --download-sample

The graph and preference tables are built in the next steps, because they are SQL queries that run on the server.

Start the server​

cargo build --workspace --all-features --release --bin algeon_server

ALGEON_SERVER_CONFIG_FILE="$PWD/docs/research/merrec/server-models.toml" \
ALGEON_SERVER_BIND=127.0.0.1:50063 \
flock -x /tmp/cudf-gpu.lock target/release/algeon_server

The configuration sets a 32 GiB device memory limit, a 1 GiB backend reserve, a 16 GiB default query limit, one active attempt, and a model root for the supervised research probes. The reproduction guide lists the extra Python packages used by the CPU checks.

Register the tables​

From another terminal, register the source tables and build the graph and preference snapshots:

python3 scripts/dev/merrec_run.py --prepare

The runner connects to grpc://127.0.0.1:50063. It registers merrec_events and merrec_items, defines the graph views, runs each preparation query, and registers the saved snapshots. Each preparation query writes Parquet under target/merrec-exploration/results/ and a result record under docs/research/merrec/results/. The runner stops at its first preparation error and saves the error record.

Browsing graph​

The source views in sql/01_graph_views.sql pair adjacent item_view events within (user_id, session_id), ordered by (event_time, event_id). A pair becomes a transition when it crosses different product_id values with a positive time gap of at most 30 minutes. Both endpoint markets need at least 10 distinct viewers, and each retained directed edge needs support from at least three distinct users. Its weight counts those users.

The saved graph has 1,711 market vertices and 4,963 directed transition edges. The audience graph has 59,629 distinct market-user edges from item_view events. The snapshot inputs and registration SQL are sql/02_graph_snapshots.sql and sql/03_register_graph.sql.

Vertex numbers belong to the saved snapshot and can change when the graph is rebuilt. In the captured snapshot, vertex 672 maps to product group 15928_1315 (Nintendo, Electronics, Games).

Audience features​

prepare_tastes.sql keeps 1,919 users with at least 20 item_view events before the exclusive cutoff of October 21, 2023. The 17 features are square roots of each user's category shares of view events, and repeated views count separately. KMeans, HDBSCAN, and UMAP read the same ordered features, with a separate fit in each query.

Run the examples​

Copy an example's SQL from its function page into any Flight SQL client connected to 127.0.0.1:50063, or pass a saved query file to the runner:

python3 scripts/dev/merrec_run.py \
docs/research/merrec/sql/20_cross_category_discovery.sql

Tables​

TableRowsContents
merrec_events494,357event_time, user_id, item_id, session_id, event_type
merrec_items358,893item_id, product_id, name, price, three category levels, brand_id
mr_nodes1,711Market vertices: product_id, brand, category, item count, mean price, viewers
mr_edges4,963Directed browsing transitions src, dst, weight
mr_undirected3,317Transitions merged into undirected weighted edges
mr_audience59,629Market-user edges from item_view events
mr_tastes1,919Per-user view counts and the 17 category features

The runner also defines views such as merrec_transitions and merrec_markets over the source tables. Those views feed the snapshots above.

Captured run​

The discovery capture records the executed SQL, output schema, query ID, all ten result rows, and one elapsed time of 0.13337774 seconds. That value is one Flight elapsed observation, including planning and receipt of the complete result. Graph preparation ran before the call on a server that had already run many queries, and no repeated performance claim follows from it. The source shard includes 238 buy_comp events.

The graph validation compares the saved personalized PageRank scores with NetworkX. The top-15 order matches, with a maximum absolute error of 9.59610868e-7 across shared vertices against a tolerance of 1e-5. The source snapshot comparison matched the saved nodes, transitions, edges, and audience rows in graph_validation.json.

The audience map combines saved UMAP display coordinates with HDBSCAN labels fitted in the original 17 dimensions. Its filters keep all 1,316 noise rows available. The campaign-theme query fits eight KMeans groups, joins their assignments to source shares, averages those shares with equal weight for each user, and applies Top-K to the category means.

The collection statement includes host window processing: its captured plan has two native fragments and 15 host DataFusion boundaries. The campaign-theme statement has five native fragments and zero host boundaries. Their timing observations exclude source preparation and provide no controlled performance comparison.