v_wmma_f32_16x16x16_fp8_fp8 GPU Native ISA AMD Matrix

Vector Packed Arithmetic

v_wmma_f32_16x16x16_fp8_fp8

Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply…

Encoding

Binary Layout (AMD RDNA 4)
NEG
63:61
OPSEL_HI[1:0]
60:59
SRC2
58:50
SRC1
49:41
SRC0
40:32
11001100
31:24
unassigned
23
1000110
22:16
CLAMP
15
OPSEL_HI[2]
14
OPSEL
13:11
NEG_HI
10:8
VDST
7:0
 

The ENC_VOP3P layout from AMD's machine-readable ISA specification (AMD RDNA 4). Opcode 70 in OP rebuilds 0x00000000CC460000, an identifier AMD lists for this encoding. Bits marked unassigned have no field in the specification. No other generation in the specification defines this instruction.

Format VOP3P
Width 64 bits
Opcode 70
Identifier 0x00000000CC460000

Operands

  • VDST
    Written. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: WMMA destination (D) or accumulator (C) 16x16 matrix with single precision float values. WMMA DC matrices are 16x16 (or larger) and the individual elements are packed. The total number of VGPRs required depends on the wave size. (OPR_VGPR, FMT_WMMA_DC_16X16_F32)
  • SRC0
    Read. Operand must be a vector GPR. Uses a 9-bit operand field. Data format: WMMA 16x16 multiplicand matrix (A or B) with 8-bit floats. WMMA AB matrices are 16x16 (or larger) and the individual elements are packed; each VGPR lane can contain 2 16-bit elements, 4 8-bit elements or 8 4-bit elements. (OPR_SRC_VGPR, FMT_WMMA_AB_16X16_FP8)
  • SRC1
    Read. Operand must be a vector GPR. Uses a 9-bit operand field. Data format: WMMA 16x16 multiplicand matrix (A or B) with 8-bit floats. WMMA AB matrices are 16x16 (or larger) and the individual elements are packed; each VGPR lane can contain 2 16-bit elements, 4 8-bit elements or 8 4-bit elements. (OPR_SRC_VGPR, FMT_WMMA_AB_16X16_FP8)
  • SRC2
    Read. Operand must be a vector GPR or an inline constant (no literal constants or scalar GPRs). Uses a 9-bit operand field. Data format: WMMA destination (D) or accumulator (C) 16x16 matrix with single precision float values. WMMA DC matrices are 16x16 (or larger) and the individual elements are packed. The total number of VGPRs required depends on the wave size. (OPR_SRC_VGPR_OR_INLINE, FMT_WMMA_DC_16X16_F32)

In AMD's order, as its machine-readable ISA specification lists them for the ENC_VOP3P encoding (AMD RDNA 4).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply add. Store the resulting matrix into vector registers.

Sources