v_dot8_u32_u4 GPU Native ISA AMD Vector

V DOT8 U32 U4 Vector Packed Arithmetic

v_dot8_u32_u4

Compute the dot product of two packed 8-D unsigned 4-bit integer inputs in the unsigned 32-bit integer domain, add an unsigned 32-bit integer value…

Encoding

Verified bit-level encoding data is not yet available for this instruction. The instruction-format classification below (VOP3P) is well-documented and stable; exact per-target opcode/field bit positions have not yet been imported from a verified source.
Format VOP3P
Execution Unit

Operands

Operand details have not yet been curated for this instruction.

GFX Target Compatibility

TargetSupport
gfx1100✅ Supported

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Compute the dot product of two packed 8-D unsigned 4-bit integer inputs in the unsigned 32-bit integer domain, add an unsigned 32-bit integer value from the third input and store the result into a vector register.

Semantics

tmp = S2.u32; tmp += u4_to_u32(S0[3 : 0].u4) * u4_to_u32(S1[3 : 0].u4); tmp += u4_to_u32(S0[7 : 4].u4) * u4_to_u32(S1[7 : 4].u4); tmp += u4_to_u32(S0[11 : 8].u4) * u4_to_u32(S1[11 : 8].u4); tmp += u4_to_u32(S0[15 : 12].u4) * u4_to_u32(S1[15 : 12].u4); tmp += u4_to_u32(S0[19 : 16].u4) * u4_to_u32(S1[19 : 16].u4); tmp += u4_to_u32(S0[23 : 20].u4) * u4_to_u32(S1[23 : 20].u4); tmp += u4_to_u32(S0[27 : 24].u4) * u4_to_u32(S1[27 : 24].u4); tmp += u4_to_u32(S0[31 : 28].u4) * u4_to_u32(S1[31 : 28].u4); D0.u32 = tmp

Example

v_dot8_u32_u4 v5, v1, v2, s3

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (gfx11). Not from an AMD document, and not authored here.

Sources