v_mfma_f32_32x32x8bf16_1k GPU Native ISA AMD Matrix
Vector Packed Arithmetic
Multiply the 32x8 matrix in the first input by the 8x32 matrix in the second input and add the 32x32 matrix in the third input using fused multiply…
Also written as
v_mfma_f32_32x32x8_bf16, v_mfma_f32_32x32x8bf16.
AMD's machine-readable ISA specification lists these names for the same instruction.
Encoding
The VOP3P_MFMA layout from AMD's machine-readable ISA specification (AMD CDNA 2). Opcode 102 in OP rebuilds 0x00000000D3E60000, an identifier AMD lists for this encoding. No other generation in the specification defines this instruction.
Operands
-
VDST
Written. N/A Data format: N/A (OPR_VGPR_OR_ACCVGPR, FMT_NUM_PK16_F32) -
SRC0
Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR, FMT_NUM_PK4_BF16) -
SRC1
Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR, FMT_NUM_PK4_BF16) -
SRC2
Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR_OR_CONST, FMT_NUM_PK16_F32)
In AMD's order, as its machine-readable ISA specification lists them for the VOP3P_MFMA encoding (AMD CDNA 2).
GFX Target Compatibility
Per-target GFX compatibility has not yet been verified for this instruction.
Related
More in Vector Packed Arithmetic
Reference
Description
Example
v_mfma_f32_32x32x8bf16_1k v[0:15], v[0:1], v[2:3], -2.0A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (mai-gfx90a.s). Not from an AMD document, and not authored here.
Sources
- AMD Machine-Readable GPU ISA Specification ↗ - Advanced Micro Devices, Inc.
- LLVM MC assembler tests for AMDGPU ↗ - LLVM Project