v_dot4_u32_u8 GPU Native ISA AMD Vector

V DOT4 U32 U8 Vector Packed Arithmetic

v_dot4_u32_u8

Compute the dot product of two packed 4-D unsigned 8-bit integer inputs in the unsigned 32-bit integer domain, add an unsigned 32-bit integer value…

Encoding

Verified bit-level encoding data is not yet available for this instruction. The instruction-format classification below (VOP3P) is well-documented and stable; exact per-target opcode/field bit positions have not yet been imported from a verified source.
Format VOP3P
Execution Unit

Operands

Operand details have not yet been curated for this instruction.

GFX Target Compatibility

TargetSupport
gfx1100✅ Supported

Related PTX Concepts

4-Way Dot Product (Accumulate) ↗
equivalent with restrictions
dp4a (PTX)

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Compute the dot product of two packed 4-D unsigned 8-bit integer inputs in the unsigned 32-bit integer domain, add an unsigned 32-bit integer value from the third input and store the result into a vector register.

Semantics

tmp = S2.u32; tmp += u8_to_u32(S0[7 : 0].u8) * u8_to_u32(S1[7 : 0].u8); tmp += u8_to_u32(S0[15 : 8].u8) * u8_to_u32(S1[15 : 8].u8); tmp += u8_to_u32(S0[23 : 16].u8) * u8_to_u32(S1[23 : 16].u8); tmp += u8_to_u32(S0[31 : 24].u8) * u8_to_u32(S1[31 : 24].u8); D0.u32 = tmp

Example

v_dot4_u32_u8 v5, v1, v2, s3

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (gfx11). Not from an AMD document, and not authored here.

Sources