s_fmac_f16 GPU Native ISA AMD Scalar

Scalar Arithmetic

s_fmac_f16

Multiply two half-precision float inputs and accumulate the result into the destination register using fused multiply add.

Encoding

Binary Layout (AMD CDNA 5)
10
31:30
1001110
29:23
SDST
22:16
SSRC1
15:8
SSRC0
7:0
 

The ENC_SOP2 layout from AMD's machine-readable ISA specification (AMD CDNA 5). Opcode 78 in OP rebuilds 0xA7000000, an identifier AMD lists for this encoding. The same in the other 2 generations that define it.

Format SOP2
Width 32 bits
Opcode 78
Identifier 0xA7000000

Operands

  • SDST
    Written. All scalar destination operands. Data format: 16-bit half-precision floating point value. (OPR_SDST, FMT_NUM_F16)
  • SSRC0
    Read. All scalar operands. Covers all operands that are allowed as scalar sources. Data format: 16-bit half-precision floating point value. (OPR_SSRC, FMT_NUM_F16)
  • SSRC1
    Read. All scalar operands. Covers all operands that are allowed as scalar sources. Data format: 16-bit half-precision floating point value. (OPR_SSRC, FMT_NUM_F16)

In AMD's order, as its machine-readable ISA specification lists them for the ENC_SOP2 encoding (AMD CDNA 5).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in Scalar Arithmetic

Reference

AMDGPU / GFX ISA

Example

s_fmac_f16 s5, 0, s2

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (gfx1150_asm_salu_float.s). Not from an AMD document, and not authored here.

Sources