v_wmma_scale16_f32_32x16x128_f4 GPU Native ISA AMD Matrix
Vector Packed Arithmetic
v_wmma_scale16_f32_32x16x128_f4
Multiply the 32x128 matrix in the first input by the 128x16 matrix in the second input and add the 32x16 matrix in the third input using fused…
Encoding
Verified bit-level encoding data is not yet available for this instruction.
The instruction-format classification below (VOP3P)
is well-documented and stable; exact per-target opcode/field bit positions have not
yet been imported from a verified source.
Operands
-
VDST
Written. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: WMMA destination (D) or accumulator (C) 32x16 matrix with single precision float values. WMMA DC matrices are 16x16 (or larger) and the individual elements are packed. The total number of VGPRs required depends on the wave size. (OPR_VGPR, FMT_WMMA_DC_32X16_F32) -
SRC0
Read. Operand must be a vector GPR. Uses a 9-bit operand field. Data format: WMMA 32x128 multiplicand matrix (A or B) with 4-bit floating point values. WMMA AB matrices are 16x16 (or larger) and the individual elements are packed; each VGPR lane can contain 2 16-bit elements, 4 8-bit elements or 8 4-bit elements. (OPR_SRC_VGPR, FMT_WMMA_AB_32X128_FP4) -
SRC1
Read. Operand must be a vector GPR. Uses a 9-bit operand field. Data format: WMMA 16x128 multiplicand matrix (A or B) with 4-bit floating point values. WMMA AB matrices are 16x16 (or larger) and the individual elements are packed; each VGPR lane can contain 2 16-bit elements, 4 8-bit elements or 8 4-bit elements. (OPR_SRC_VGPR, FMT_WMMA_AB_16X128_FP4) -
SRC2
Read. Operand must be a vector GPR or an inline constant (no literal constants or scalar GPRs). Uses a 9-bit operand field. Data format: WMMA destination (D) or accumulator (C) 32x16 matrix with single precision float values. WMMA DC matrices are 16x16 (or larger) and the individual elements are packed. The total number of VGPRs required depends on the wave size. (OPR_SRC_VGPR_OR_INLINE, FMT_WMMA_DC_32X16_F32) -
SCALE_SRC0
Read. Simple operands. Covers all operands that are allowed as scalar or vector sources except SRC_LITERAL and SRC_LITERAL64. Data format: 64 bits of arbitrary data. (OPR_SRC_SIMPLE, FMT_NUM_B64) -
SCALE_SRC1
Read. Simple operands. Covers all operands that are allowed as scalar or vector sources except SRC_LITERAL and SRC_LITERAL64. Data format: 64 bits of arbitrary data. (OPR_SRC_SIMPLE, FMT_NUM_B64)
In AMD's order, as its machine-readable ISA specification lists them for the ENC_VOP3PX3 encoding (AMD CDNA 5).
GFX Target Compatibility
Per-target GFX compatibility has not yet been verified for this instruction.
Related
More in Vector Packed Arithmetic
Reference
AMDGPU / GFX ISA
Description
Multiply the 32x128 matrix in the first input by the 128x16 matrix in the second input and add the 32x16 matrix in the third input using fused multiply add. Scale values are applied with a block size of 16. Store the resulting matrix into vector registers.
Sources
- AMD Machine-Readable GPU ISA Specification ↗ - Advanced Micro Devices, Inc.