Skip to main content

KMeans

SQL function: cuvs_kmeans

Official cuVS reference: C API

K-means clustering assignments for dense vector rows.

Quickstart

The call below expects the registered relation input_vectors (the input role), passed as a parenthesized SELECT subquery. Substitute your own relations and column names.

SELECT id, cluster_id
FROM cuvs_kmeans(
input => (SELECT id, d0, d1 FROM input_vectors),
n_clusters => 8
)
ORDER BY row_ordinal;

Inputs

Each relation argument is a parenthesized SELECT subquery that the planner keeps as a real child; metadata validation resolves a registered table or view for the same role instead. See Vector Inputs for the relation identity rules and the ID, dense-vector type, null, finite-value, and runtime-dimension contract.

RoleRequiredValidation referenceDescription
inputyestableDense-vector rows consumed by this operation.

Vector element types

Element typeValid metrics
Float32l2_expanded, l2_sqrt_expanded
Float64l2_expanded, l2_sqrt_expanded

Arguments and options

Scalar SQL arguments

ArgumentTypeRequiredDescription
n_clustersintegeryesNumber of clusters fitted for this one statement.
metricenum ("l2_expanded", "l2_sqrt_expanded")noL2 distance form used while fitting and assigning the current input relation.
max_iterintegernoMaximum fitting iterations for this invocation.
tolnumbernoNon-negative convergence tolerance used by the fitted model.
n_initintegernoNumber of initialization attempts performed within this call.
initenum ("kmeans++", "random")noCentroid initialization strategy for this fitted model.

SQL value argument schemas

ArgumentRequiredLiteral shapeDefaultConstraintsDescription
initnostring"kmeans++"one of "kmeans++", "random"Centroid initialization strategy for this fitted model.
max_iternointeger100minimum 1; maximum 2147483647Maximum fitting iterations for this invocation.
metricnostring"l2_expanded"one of "l2_expanded", "l2_sqrt_expanded"L2 distance form used while fitting and assigning the current input relation. Supported element/metric combinations: Float32: l2_expanded, l2_sqrt_expanded; Float64: l2_expanded, l2_sqrt_expanded.
n_clustersyesintegerNo defaultminimum 1; maximum 2147483647Number of clusters fitted for this one statement.
n_initnointeger1minimum 1; maximum 2147483647Number of initialization attempts performed within this call.
tolnonumber0.0001minimum 0Non-negative convergence tolerance used by the fitted model.

Vector binding shapes

Each relation subquery must project a non-null id field followed by either one or more non-null feature fields of a supported element type (Float32, Float64) or one non-null list vector field named vector.

For wide vectors, the projection order defines the feature dimensions. A list vector relation must contain no feature field beside id and vector.

Output

ColumnTypeNullableDescription
row_ordinalUInt64noZero-based ordinal of the evaluated input row; it disambiguates duplicate IDs.
idsame_as_input.idnoLogical ID copied from the input relation.
cluster_idInt32noQuery-local numeric assignment label, not a stable business or topic identifier.

Concrete schemas are call-specific. Run gpu_validate_call against registered relations to inspect the output schema after the actual ID types and literal options are validated.

Limits

  • Validation resolves named tables or views and reads schemas only; it does not execute relation scans or GPU work.
  • Execution relation arguments require parenthesized subqueries; dry-run validation accepts registered named relations only.
  • Fits a model and returns assignments within one statement; it does not return a reusable model, centroids, or inertia.
  • The evaluated input must be non-empty and n_clusters must not exceed its row count.
  • cluster_id values are query-local labels. Do not attach permanent business meaning to their numeric values.

To dry-run validate relation metadata, column types, and options without execution, see gpu_validate_call.