v_dot2_f32_f16 GPU Native ISA AMD Vector

Vector Packed Arithmetic

v_dot2_f32_f16

Compute the dot product of two packed 2-D half-precision float inputs in the single-precision float domain, add a single-precision float value from…

Encoding

Binary Layout (AMD CDNA 4)
NEG
63:61
OP_SEL_HI[1:0]
60:59
SRC2
58:50
SRC1
49:41
SRC0
40:32
110100111
31:23
0100011
22:16
CLAMP
15
OP_SEL_HI[2]
14
OP_SEL
13:11
NEG_HI
10:8
VDST
7:0
 

The ENC_VOP3P layout from AMD's machine-readable ISA specification (AMD CDNA 4). Opcode 35 in OP rebuilds 0x00000000D3A30000, an identifier AMD lists for this encoding. Not the same in every generation: AMD RDNA 4 (opcode 19, field layout), AMD RDNA 3.5 (opcode 19, field layout), AMD RDNA 3 (opcode 19, field layout), AMD RDNA 2 (opcode 19, field layout). The same in AMD CDNA 3, AMD CDNA 2, AMD CDNA 1.

Format VOP3P
Width 64 bits
Opcode 35
Identifier 0x00000000D3A30000

Operands

  • VDST
    Written. N/A Data format: N/A (OPR_VGPR, FMT_NUM_F32)
  • SRC0
    Read. N/A Data format: N/A (OPR_SRC_NOLIT, FMT_NUM_PK2_F16)
  • SRC1
    Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_PK2_F16)
  • SRC2
    Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_F32)

In AMD's order, as its machine-readable ISA specification lists them for the ENC_VOP3P encoding (AMD CDNA 4).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Compute the dot product of two packed 2-D half-precision float inputs in the single-precision float domain, add a single-precision float value from the third input and store the result into a vector register.

Example

v_dot2_f32_f16 v0, v1, v2, v3

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (dl-insts.s). Not from an AMD document, and not authored here.

Sources