v_mfma_scale_f32_16x16x128_f8f6f4 GPU Native ISA AMD Matrix

Vector Packed Arithmetic

v_mfma_scale_f32_16x16x128_f8f6f4

Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused…

Encoding

Verified bit-level encoding data is not yet available for this instruction. The instruction-format classification below (VOP3P) is well-documented and stable; exact per-target opcode/field bit positions have not yet been imported from a verified source.
Format VOP3P

Operands

  • VDST
    Written. N/A Data format: N/A (OPR_VGPR_OR_ACCVGPR, FMT_NUM_PK4_F32)
  • SRC0
    Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR, FMT_NUM_PK8_B32)
  • SRC1
    Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR, FMT_NUM_PK8_B32)
  • SRC2
    Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR_OR_CONST, FMT_NUM_PK4_F32)
  • SCALE_SRC0
    Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_B32)
  • SCALE_SRC1
    Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_B32)

In AMD's order, as its machine-readable ISA specification lists them for the ENC_VOP3PX2 encoding (AMD CDNA 4).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply add. Store the resulting matrix into vector registers.

Sources