v_mfma_f32_32x32x4_2b_f16 GPU Native ISA AMD Matrix
V MFMA F32 32X32X4 2B F16 Vector Packed Arithmetic
v_mfma_f32_32x32x4_2b_f16
Multiply the 32x4 matrix in the first input by the 4x32 matrix in the second input and add the 32x32 matrix in the third input using fused multiply…
Encoding
Verified bit-level encoding data is not yet available for this instruction.
The instruction-format classification below (VOP3P)
is well-documented and stable; exact per-target opcode/field bit positions have not
yet been imported from a verified source.
Operands
Operand details have not yet been curated for this instruction.
GFX Target Compatibility
Per-target GFX compatibility has not yet been verified for this instruction.
This instruction performs 16 passes.
Related
More in Vector Packed Arithmetic
Reference
AMDGPU / GFX ISA
Description
Multiply the 32x4 matrix in the first input by the 4x32 matrix in the second input and add the 32x32 matrix in the third input using fused multiply add. Store the resulting matrix into vector registers.
Semantics
D = A (32x4) * B (4x32) + C (32x32)
This instruction performs 2 matrix multiplies. Each operand contains 2 matrices back to back, and each matrix
has elements distributed across all lanes of the wave. Each matrix multiple is computed and the row-column
dot products are distributed across the vector ALU for higher performance. The result matrices are stored
back-to-back in the destination vector registers.
Matrices A and B are half-precision float format. Matrices C and D are single-precision float format.
Sources
- User Guide for AMDGPU Backend ↗ - LLVM Project
-
"AMD Instinct MI300" Instruction Set Architecture: Reference Guide ↗
- Advanced Micro Devices, Inc.
Reference Guide, page 275.