v_dot4_i32_i8 GPU Native ISA AMD Vector

V DOT4 I32 I8 Vector Packed Arithmetic

v_dot4_i32_i8

Compute the dot product of two packed 4-D signed 8-bit integer inputs in the signed 32-bit integer domain, add a signed 32-bit integer value from the…

Encoding

Verified bit-level encoding data is not yet available for this instruction. The instruction-format classification below (VOP3P) is well-documented and stable; exact per-target opcode/field bit positions have not yet been imported from a verified source.
Format VOP3P
Execution Unit

Operands

Operand details have not yet been curated for this instruction.

GFX Target Compatibility

TargetSupport
gfx1100✅ Supported

Related PTX Concepts

4-Way Dot Product (Accumulate) ↗
equivalent with restrictions
dp4a (PTX)

Related

More in Vector Packed Arithmetic

Reference

AMDGPU / GFX ISA

Description

Compute the dot product of two packed 4-D signed 8-bit integer inputs in the signed 32-bit integer domain, add a signed 32-bit integer value from the third input and store the result into a vector register.

Semantics

tmp = S2.i32; tmp += i8_to_i32(S0[7 : 0].i8) * i8_to_i32(S1[7 : 0].i8); tmp += i8_to_i32(S0[15 : 8].i8) * i8_to_i32(S1[15 : 8].i8); tmp += i8_to_i32(S0[23 : 16].i8) * i8_to_i32(S1[23 : 16].i8); tmp += i8_to_i32(S0[31 : 24].i8) * i8_to_i32(S1[31 : 24].i8); D0.i32 = tmp

Sources