Demo Datasets
The SQL examples on the function pages read one of four public datasets. Each page below converts the upstream files into local Parquet, registers the tables in a running server, and records the captured runs. The website does not host raw or converted data; every preparation runs in your repository checkout.
Every dataset page uses the same sections in the same order: At a glance, Examples, Prepare the data, Start the server, Register the tables, Run the examples, Tables, and Captured run where a capture record exists.
Choosing a dataset
Start with the Citation Network for a first run. It needs only pyarrow for
conversion, and its server walkthrough is the one the integration pages
reuse. Conversion takes a 12.5 GB source JSON and about 17 to 19 minutes on a
32-core host.
Amazon Reviews is the largest preparation. Embedding the text runs PyTorch on the GPU before any SQL, and the captured server configuration gives each query a 32 GiB limit on a device with at least 64 GiB available.
EMBER2024 and MerRec exercise more than one algorithm family against the same tables. Their preparation scripts also build derived relations, such as forest attributions or browsing-transition graphs, that some examples read.
Server and table lifetime
Table registrations last until the server stops. After a restart, rerun the dataset's registration step; the Parquet files can be reused.
Each page starts a server with the configuration used for its captures. Amazon Reviews and EMBER2024 both bind port 50062, so stop one server before starting the other. For other settings, see Flight SQL server and configuration.