v_pk_fma_f16 GPU Native ISA AMD Vector

V PK FMA F16 Vector Packed Arithmetic

v_pk_fma_f16

Multiply two packed half-precision float inputs component-wise and add a third input component-wise using fused multiply add, and store the result…

Encoding

Verified bit-level encoding data is not yet available for this instruction. The instruction-format classification below (VOP3P) is well-documented and stable; exact per-target opcode/field bit positions have not yet been imported from a verified source.
Format VOP3P
Execution Unit

Operands

Operand details have not yet been curated for this instruction.

GFX Target Compatibility

TargetSupport
gfx1100✅ Supported

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Multiply two packed half-precision float inputs component-wise and add a third input component-wise using fused multiply add, and store the result into a vector register.

Semantics

declare tmp : 32'B; tmp[15 : 0].f16 = fma(S0[15 : 0].f16, S1[15 : 0].f16, S2[15 : 0].f16); tmp[31 : 16].f16 = fma(S0[31 : 16].f16, S1[31 : 16].f16, S2[31 : 16].f16); D0.b32 = tmp

Example

v_pk_fma_f16 v5, v1, v2, s3

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (gfx11). Not from an AMD document, and not authored here.

Sources