ai.onnx.MatMul

ai.onnx · standard ONNX operator · ONNX opset ≥ 9

Description

Computes the matrix product A * B with ONNX MatMul semantics, including vector promotion and NumPy-style broadcasting over batch dimensions. This package supports float16, float32, int32, and uint32 tensors. All legal rank-1 through rank-4 combinations are implemented; rank-5 is supported only when both inputs have rank 5. Mixed-rank combinations involving rank 5, higher-rank inputs, and other ONNX-supported types are unsupported.

See the ONNX MatMul spec for the reference semantics.

Inputs

Name Upstream name Logical dtype Rank Shape Description Presence
a A T N-dimensional left-hand matrix. required
b B T N-dimensional right-hand matrix. required

Outputs

Name Upstream name Logical dtype Rank Shape Description Presence
y Y T derived ONNX MatMul result of a and b Matrix multiply result of A * B, including vector promotion and broadcast batch dimensions. required

Type constraints

Variable Allowed dtypes
T float32, float16, int32, uint32

Implementation variants

One implementation is selected per call from the device capabilities, the request shapes and the dtypes; these notes say what each one covers.

  • subgroup_matrix_splitk — Partitions the K reduction across subgroup-matrix workgroups, then combines float32 partials. The combine pass reassociates the reduction.
  • folded_rows_subgroup_matrix_splitk — Treats rank-3 A with a unit row axis as a contiguous [M, K] view, partitions K across subgroup-matrix workgroups, and combines float32 partials.
  • rank2_notrans_f32_vec4_tiled_reg_splitk — Cuts the K reduction of the 128x64 f32 register tile into power-of-two slices across dispatch.z, with at least four K tiles per slice, then sums the f32 partials in a second pass.
  • rank2_band_vec4_splitk — Splits the vec4 band's K axis into up to sixteen workgroup ranges. Each range writes an f32 partial band that a combine pass sums.
  • batched_band_vec4 — Vec4 band route for equal-batch rank-3/rank-4 operands with 2..BAND_VEC4_MAX_ROWS rows per matrix. One workgroup row handles each matrix, and K slices are derived from the batch-times-column-group count.
  • matrix_vector_tail_subgroup — Assigns one output row to each workgroup and uses scalar loads for a reduction width that is not vec4-aligned, then folds the row with a fixed-width subgroup. It is the parallel large-matrix-vector route when that subgroup geometry is available.
  • matrix_vector_tail_nosg — Uses the same tail-safe, one-workgroup-per-row traversal but folds partials through workgroup memory. It applies when fixed-width subgroup geometry is unavailable.
  • dot_splitk — Partitions a long vector dot product across workgroups and combines float32 partials to expose parallelism beyond one workgroup. The combine reassociates the sum, so this route is not bit-identical to the single-workgroup dot route.

Device requirements

Some implementation variants require subgroup-matrix, shader-f16, and subgroups. These are route-specific capabilities, not package-wide requirements; availability also depends on the request shape and dtype.

Files

Use with @huggingface/kernels

npm install --save-exact @huggingface/kernels@0.0.1-preview.2

Required output shapes and logical data types are inferred from the supplied inputs and attributes; result tensors are allocated automatically.

The version: 1 option selects the published kernel contract; it is independent of any operator opset, contrib since_version, or model version. It follows the v1 branch as fixes land. To pin exact artifact bytes, pass a 40-character commit revision instead of version.

Replace each *Data placeholder with a typed array containing the corresponding input data.

import { getKernel } from "@huggingface/kernels";

const kernel = await getKernel("webgpu-kernels/ai.onnx.MatMul", { version: 1 });
const { y } = await kernel({ a: { data: aData, shape: [3, 5] }, b: { data: bData, shape: [5, 4] } });
Downloads last month
-
kernel
webgpu
wgsl
apache-2.0
WebGPU

Requires WebGPU support. See the compatibility table.