ai.onnx.MatMul
ai.onnx · standard ONNX operator · ONNX opset ≥ 9
Description
Computes the matrix product A * B with ONNX MatMul semantics, including vector promotion and NumPy-style broadcasting over batch dimensions. This package supports float16, float32, int32, and uint32 tensors. All legal rank-1 through rank-4 combinations are implemented; rank-5 is supported only when both inputs have rank 5. Mixed-rank combinations involving rank 5, higher-rank inputs, and other ONNX-supported types are unsupported.
See the ONNX MatMul spec for the reference semantics.
Inputs
| Name | Upstream name | Logical dtype | Rank | Shape | Description | Presence |
|---|---|---|---|---|---|---|
a |
A |
T |
— | — | N-dimensional left-hand matrix. | required |
b |
B |
T |
— | — | N-dimensional right-hand matrix. | required |
Outputs
| Name | Upstream name | Logical dtype | Rank | Shape | Description | Presence |
|---|---|---|---|---|---|---|
y |
Y |
T |
derived | ONNX MatMul result of a and b |
Matrix multiply result of A * B, including vector promotion and broadcast batch dimensions. |
required |
Type constraints
| Variable | Allowed dtypes |
|---|---|
T |
float32, float16, int32, uint32 |
Implementation variants
One implementation is selected per call from the device capabilities, the request shapes and the dtypes; these notes say what each one covers.
subgroup_matrix_splitk— Partitions the K reduction across subgroup-matrix workgroups, then combines float32 partials. The combine pass reassociates the reduction.folded_rows_subgroup_matrix_splitk— Treats rank-3 A with a unit row axis as a contiguous[M, K]view, partitions K across subgroup-matrix workgroups, and combines float32 partials.rank2_notrans_f32_vec4_tiled_reg_splitk— Cuts the K reduction of the 128x64 f32 register tile into power-of-two slices across dispatch.z, with at least four K tiles per slice, then sums the f32 partials in a second pass.rank2_band_vec4_splitk— Splits the vec4 band's K axis into up to sixteen workgroup ranges. Each range writes an f32 partial band that a combine pass sums.batched_band_vec4— Vec4 band route for equal-batch rank-3/rank-4 operands with 2..BAND_VEC4_MAX_ROWS rows per matrix. One workgroup row handles each matrix, and K slices are derived from the batch-times-column-group count.matrix_vector_tail_subgroup— Assigns one output row to each workgroup and uses scalar loads for a reduction width that is not vec4-aligned, then folds the row with a fixed-width subgroup. It is the parallel large-matrix-vector route when that subgroup geometry is available.matrix_vector_tail_nosg— Uses the same tail-safe, one-workgroup-per-row traversal but folds partials through workgroup memory. It applies when fixed-width subgroup geometry is unavailable.dot_splitk— Partitions a long vector dot product across workgroups and combines float32 partials to expose parallelism beyond one workgroup. The combine reassociates the sum, so this route is not bit-identical to the single-workgroup dot route.
Device requirements
Some implementation variants require subgroup-matrix, shader-f16, and subgroups. These are route-specific capabilities, not package-wide requirements; availability also depends on the request shape and dtype.
Files
metadata.json— kernel metadata (id, digests, per-variant templates, provenance)manifest.json— the op contract (source of truth)test.json— correctness casesbench.json— benchmark + tuning casesdot-splitk-partial.wgsl.jinjamatmul-band-vec4.wgsl.jinjamatmul-batched-vector.wgsl.jinjamatmul-batched.wgsl.jinjamatmul-dot.wgsl.jinjamatmul-matrix-vector-subgroup.wgsl.jinjamatmul-notrans-vec4-tiled-reg.wgsl.jinjamatmul-subgroup-matrix-ext.wgsl.jinjamatmul-tiled-general-reg.wgsl.jinjamatmul-tiled-general.wgsl.jinjamatmul-vector-matrix-tail.wgsl.jinjamatmul-vector-matrix-vec4.wgsl.jinjamatmul-vector-product.wgsl.jinjamatmul.wgsl.jinjareduce-axis0-splitk-combine.wgsl.jinja
Use with @huggingface/kernels
npm install --save-exact @huggingface/kernels@0.0.1-preview.2
Required output shapes and logical data types are inferred from the supplied inputs and attributes; result tensors are allocated automatically.
The version: 1 option selects the published kernel contract; it is independent of any operator opset, contrib since_version, or model version.
It follows the v1 branch as fixes land. To pin exact artifact bytes, pass a 40-character commit revision instead of version.
Replace each *Data placeholder with a typed array containing the corresponding input data.
import { getKernel } from "@huggingface/kernels";
const kernel = await getKernel("webgpu-kernels/ai.onnx.MatMul", { version: 1 });
const { y } = await kernel({ a: { data: aData, shape: [3, 5] }, b: { data: bData, shape: [5, 4] } });
- Downloads last month
- -
Requires WebGPU support. See the compatibility table.