v_dot2_f32_bf16 GPU Native ISA AMD Vector

Vector Packed Arithmetic

v_dot2_f32_bf16

Calculate the dot product of BF16 float 2-vectors from the first and second inputs, convert the product to single-precision float format, add the…

Encoding

Binary Layout (AMD CDNA 4)
NEG
63:61
OP_SEL_HI[1:0]
60:59
SRC2
58:50
SRC1
49:41
SRC0
40:32
110100111
31:23
0011010
22:16
CLAMP
15
OP_SEL_HI[2]
14
OP_SEL
13:11
NEG_HI
10:8
VDST
7:0
 

The ENC_VOP3P layout from AMD's machine-readable ISA specification (AMD CDNA 4). Opcode 26 in OP rebuilds 0x00000000D39A0000, an identifier AMD lists for this encoding. Not the same in every generation: AMD RDNA 4 (field layout), AMD RDNA 3.5 (field layout), AMD RDNA 3 (field layout).

Format VOP3P
Width 64 bits
Opcode 26
Identifier 0x00000000D39A0000

Operands

  • VDST
    Written. N/A Data format: N/A (OPR_VGPR, FMT_NUM_F32)
  • SRC0
    Read. N/A Data format: N/A (OPR_SRC_NOLIT, FMT_NUM_PK2_BF16)
  • SRC1
    Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_PK2_BF16)
  • SRC2
    Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_F32)

In AMD's order, as its machine-readable ISA specification lists them for the ENC_VOP3P encoding (AMD CDNA 4).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Calculate the dot product of BF16 float 2-vectors from the first and second inputs, convert the product to single-precision float format, add the third input and store the result into a vector register.

Example

v_dot2_f32_bf16 v2, v1, 0, v2

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (bf16_imm-fake16.s). Not from an AMD document, and not authored here.

Sources