v_mfma_scale_f32_32x32x64_f8f6f4 GPU Native ISA AMD Matrix
Vector Packed Arithmetic
v_mfma_scale_f32_32x32x64_f8f6f4
Multiply the 32x64 matrix in the first input by the 64x32 matrix in the second input and add the 32x32 matrix in the third input using fused multiply…
Encoding
Verified bit-level encoding data is not yet available for this instruction.
The instruction-format classification below (VOP3P)
is well-documented and stable; exact per-target opcode/field bit positions have not
yet been imported from a verified source.
Operands
-
VDST
Written. N/A Data format: N/A (OPR_VGPR_OR_ACCVGPR, FMT_NUM_PK16_F32) -
SRC0
Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR, FMT_NUM_PK8_B32) -
SRC1
Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR, FMT_NUM_PK8_B32) -
SRC2
Read. N/A Data format: N/A (OPR_SRC_VGPR_OR_ACCVGPR_OR_CONST, FMT_NUM_PK16_F32) -
SCALE_SRC0
Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_B32) -
SCALE_SRC1
Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_B32)
In AMD's order, as its machine-readable ISA specification lists them for the ENC_VOP3PX2 encoding (AMD CDNA 4).
GFX Target Compatibility
Per-target GFX compatibility has not yet been verified for this instruction.
Related
More in Vector Packed Arithmetic
Reference
AMDGPU / GFX ISA
Description
Multiply the 32x64 matrix in the first input by the 64x32 matrix in the second input and add the 32x32 matrix in the third input using fused multiply add. Store the resulting matrix into vector registers.
Sources
- AMD Machine-Readable GPU ISA Specification ↗ - Advanced Micro Devices, Inc.