v_mfma_f32_4x4x4f16 GPU Native ISA AMD Matrix

Vector Packed Arithmetic

v_mfma_f32_4x4x4f16

Multiply the 4x4 matrix in the first input by the 4x4 matrix in the second input and add the 4x4 matrix in the third input using fused multiply add.

Also written as v_mfma_f32_4x4x4_16b_f16. AMD's machine-readable ISA specification lists this name for the same instruction.

Encoding

Binary Layout (AMD CDNA 2)
BLGP
63:61
ACC
60:59
SRC2
58:50
SRC1
49:41
SRC0
40:32
110100111
31:23
1001010
22:16
ACC_CD
15
ABID
14:11
CBSZ
10:8
VDST
7:0
 

The VOP3P_MFMA layout from AMD's machine-readable ISA specification (AMD CDNA 2). Opcode 74 in OP rebuilds 0x00000000D3CA0000, an identifier AMD lists for this encoding. Not the same in every generation: AMD CDNA 1 (field layout).

Format VOP3P
Width 64 bits
Opcode 74
Identifier 0x00000000D3CA0000

Operands

  • VDST
    Written. N/A Data format: N/A (OPR_VGPR_OR_ACCVGPR, FMT_NUM_PK4_F32)
  • SRC0
    Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR, FMT_NUM_PK4_F16)
  • SRC1
    Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR, FMT_NUM_PK4_F16)
  • SRC2
    Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR_OR_CONST, FMT_NUM_PK4_F32)

In AMD's order, as its machine-readable ISA specification lists them for the VOP3P_MFMA encoding (AMD CDNA 2).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Multiply the 4x4 matrix in the first input by the 4x4 matrix in the second input and add the 4x4 matrix in the third input using fused multiply add. Store the resulting matrix into vector registers.

Example

v_mfma_f32_4x4x4f16 v[0:3], v[0:1], v[2:3], -2.0

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (mai-gfx90a.s). Not from an AMD document, and not authored here.

Sources