EMBER2024
EMBER2024 is a malware benchmark with train, test, and challenge splits of extracted file features. The examples organize the 3,225 Win32 files in the challenge split for analyst review, compare classifiers trained on earlier files with a later quarter, and search earlier files for reference samples.
The challenge split contains only malicious files. Group membership there describes similarity in the selected features, and the cohort has no benign files with which to evaluate false positives.
At a glance
Examples
- Group missed Win32 files for batch review with HDBSCAN over the challenge histograms.
- Project the challenge cohort into two dimensions with UMAP, truncated SVD, or PCA.
- Rank high-scoring challenge files for review with forest inference and precomputed attributions.
- Compare classifiers on a later quarter with logistic regression and random forest.
- Find earlier Win32 reference files with exact search, CAGRA, or IVF-Flat.
- Compare import patterns around an unlabeled file with personalized PageRank and Jaccard.
Prepare the data
Run commands from the repository root. The website does not host the raw
dataset or generated Parquet files. Follow the
fixture README
to download the raw archives and EMBER2024_all.model, vectorize the
features, and export the split Parquet files under
fixture/demo/ember2024/parquet/.
Build the research relations
Install the pinned research dependencies in the fixture's virtual environment and run the preparation script:
~/.venvs/forest/bin/pip install -r fixture/demo/ember2024/research-requirements.txt
~/.venvs/forest/bin/python fixture/demo/ember2024/scripts/research_prepare.py
The script reads all three splits, so allow time for the full metadata scan.
It keeps the first row for each SHA-256 within a split. The challenge split
keeps every unique row; train and test keep a one-percent SHA-256 sample of
vectors and metadata for every unique row. IDs use disjoint ranges: train
starts at 0, test at 10,000,000, and challenge at 20,000,000.
Each vector row stores the full 2,568-element Float32 feature vector and a
512-element histogram: a 256-bin byte histogram followed by a 256-bin
byte-entropy histogram (source features f7 through f518), each normalized
by its own total count. Metadata rows store the SHA-256, file type, split,
week, label, and family fields used for filtering and review after a fit.
Compute forest attributions
The forest review example also reads per-file attributions. Once the
challenge feature memmaps and EMBER2024_all.model are in place, run:
~/.venvs/forest/bin/python fixture/demo/ember2024/scripts/research_attributions.py
This script uses CPU LightGBM pred_contrib to produce signed SHAP values,
then sums the 2,568 features into 12 thrember feature blocks. The recorded
pred_contrib step took 13.311 seconds for 6,315 challenge rows. SQL later
applies GPU Top-K to these precomputed values; the CPU preparation is separate
from query execution.
Start the server
cargo build --workspace --all-features --release --bin algeon_server
ALGEON_SERVER_CONFIG_FILE="$PWD/docs/research/ember2024/server.toml" \
ALGEON_SERVER_BIND=127.0.0.1:50062 \
flock -x /tmp/cudf-gpu.lock target/release/algeon_server
server.toml
is the capture configuration. It sets a 64 GiB device limit, a 32 GiB default
query limit, and a [models] root of fixture/demo/ember2024/models, which
cuml_forest_predict needs. Start the server from the repository root so
that relative path resolves. See
model configuration
for the models-root contract.
Register the tables
From another terminal, register every generated relation:
~/.venvs/forest/bin/python fixture/demo/ember2024/scripts/research.py \
--uri grpc://127.0.0.1:50062 --register
--register creates ember_train, ember_test, and ember_challenge from
the fixture's split Parquet, and one external table named ember_<name> for
each target/ember2024-research/input_<name>.parquet. It uses
IF NOT EXISTS, so rerunning it after adding files registers only the new
ones.
Add the import review graph
The import review example needs a graph built from an earlier query result. With the tables above registered, capture the exact Win32 neighbors, build the sample-to-import graph on the CPU, and register the new files:
~/.venvs/forest/bin/python fixture/demo/ember2024/scripts/research.py \
--uri grpc://127.0.0.1:50062 \
docs/research/ember2024/sql/47_challenge_exact_neighbors.sql
~/.venvs/forest/bin/python fixture/demo/ember2024/scripts/research_graph_prepare.py
~/.venvs/forest/bin/python fixture/demo/ember2024/scripts/research.py \
--uri grpc://127.0.0.1:50062 --register
The graph keeps import symbols seen in 3 to 100 distinct Win32 challenge files. Its edges connect 926 of the 3,225 sample files to 4,988 symbols. The other 2,299 sample files have no retained import edges and remain as isolates in the sample-count relation.
Run the examples
Copy an example's SQL from its function page into any Flight SQL client
connected to 127.0.0.1:50062, or pass a saved query file to research.py:
~/.venvs/forest/bin/python fixture/demo/ember2024/scripts/research.py \
--uri grpc://127.0.0.1:50062 \
docs/research/ember2024/sql/87_import_review_queue.sql
research.py writes the result Parquet under target/ember2024-research/ and
a result record under docs/research/ember2024/results/.
Tables
The training query uses 15,588 sampled Win32 rows after joining vectors to metadata, and the test query evaluates 3,618 sampled Win32 rows from the later held-out period. Labels are joined after prediction for reporting. The cuVS examples pass only IDs and histograms to search, then join labels and family names after the UDTF returns.
Captured run
Each example page links a capture JSON with its SQL, output, and run
conditions. The HDBSCAN input is ordered by id, which fixes the input
sequence for that capture. It does not promise the same fitted memberships or
cluster IDs across different hardware and software environments.