AMDGPU / GFX Instructions GPU Native ISA AMD
2442 instructions - showing 42 per page, page 25 of 25 - AMDGPU/GFX is AMD's native, low-level
GPU instruction set family, assembled directly with no virtual intermediate layer. Supported
instructions and their exact encodings vary by GFX compatibility target (e.g.
gfx942, gfx1100). Scalar instructions execute once per wavefront on
the Scalar ALU; vector instructions execute per-lane on the Vector ALU, gated by the EXEC mask.
| Mnemonic | Syntax | Format | GFX Targets | Unit | Summary |
|---|---|---|---|---|---|
| v_wmma_f16_16x16x32_f16 | v_wmma_f16_16x16x32_f16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f16_16x16x64_bf8_bf8 | v_wmma_f16_16x16x64_bf8_bf8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f16_16x16x64_bf8_fp8 | v_wmma_f16_16x16x64_bf8_fp8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f16_16x16x64_fp8_bf8 | v_wmma_f16_16x16x64_fp8_bf8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f16_16x16x64_fp8_fp8 | v_wmma_f16_16x16x64_fp8_fp8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x128_bf8_bf8 | v_wmma_f32_16x16x128_bf8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… | ||
| v_wmma_f32_16x16x128_bf8_fp8 | v_wmma_f32_16x16x128_bf8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… | ||
| v_wmma_f32_16x16x128_f8f6f4 | v_wmma_f32_16x16x128_f8f6f4 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… | ||
| v_wmma_f32_16x16x128_fp8_bf8 | v_wmma_f32_16x16x128_fp8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… | ||
| v_wmma_f32_16x16x128_fp8_fp8 | v_wmma_f32_16x16x128_fp8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… | ||
| v_wmma_f32_16x16x16_bf16 | v_wmma_f32_16x16x16_bf16 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x16_bf8_bf8 | v_wmma_f32_16x16x16_bf8_bf8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x16_bf8_fp8 | v_wmma_f32_16x16x16_bf8_fp8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x16_f16 | v_wmma_f32_16x16x16_f16 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x16_fp8_bf8 | v_wmma_f32_16x16x16_fp8_bf8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x16_fp8_fp8 | v_wmma_f32_16x16x16_fp8_fp8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x32_bf16 | v_wmma_f32_16x16x32_bf16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x32_f16 | v_wmma_f32_16x16x32_f16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x4_f32 | v_wmma_f32_16x16x4_f32 | VOP3P | Multiply the 16x4 matrix in the first input by the 4x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x64_bf8_bf8 | v_wmma_f32_16x16x64_bf8_bf8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x64_bf8_fp8 | v_wmma_f32_16x16x64_bf8_fp8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x64_fp8_bf8 | v_wmma_f32_16x16x64_fp8_bf8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_16x16x64_fp8_fp8 | v_wmma_f32_16x16x64_fp8_fp8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_f32_32x16x128_f4 | v_wmma_f32_32x16x128_f4 | VOP3P | Multiply the 32x128 matrix in the first input by the 128x16 matrix in the second input and add the 32x16 matrix in the third input using fused… | ||
| v_wmma_i32_16x16x16_iu4 | v_wmma_i32_16x16x16_iu4 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_i32_16x16x16_iu8 | v_wmma_i32_16x16x16_iu8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_i32_16x16x32_iu4 | v_wmma_i32_16x16x32_iu4 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_i32_16x16x64_iu8 | v_wmma_i32_16x16x64_iu8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… | ||
| v_wmma_ld_scale16_paired_b64 | v_wmma_ld_scale16_paired_b64 | VOP3P | AMDGPU VOP3P matrix instruction operating on b64 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) | ||
| v_wmma_ld_scale_paired_b32 | v_wmma_ld_scale_paired_b32 | VOP3P | AMDGPU VOP3P matrix instruction operating on b32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) | ||
| v_wmma_scale16_f32_16x16x128_f8f6f4 | v_wmma_scale16_f32_16x16x128_f8f6f4 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… | ||
| v_wmma_scale16_f32_32x16x128_f4 | v_wmma_scale16_f32_32x16x128_f4 | VOP3P | Multiply the 32x128 matrix in the first input by the 128x16 matrix in the second input and add the 32x16 matrix in the third input using fused… | ||
| v_wmma_scale_f32_16x16x128_f8f6f4 | v_wmma_scale_f32_16x16x128_f8f6f4 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… | ||
| v_wmma_scale_f32_32x16x128_f4 | v_wmma_scale_f32_32x16x128_f4 | VOP3P | Multiply the 32x128 matrix in the first input by the 128x16 matrix in the second input and add the 32x16 matrix in the third input using fused… | ||
| v_writelane_b32 | v_writelane_b32 | VOP2 | gfx1100 | Write the scalar value in the first input into the specified lane of a vector register where the lane select is in the second input. | |
| v_xad_u32 | v_xad_u32 | VOP3 | gfx1100 | Calculate bitwise XOR of the first two vector inputs, then add the third vector input to the intermediate result, then store the final result into a… | |
| v_xnor_b32 | v_xnor_b32 | VOP2 | gfx1100 | Calculate bitwise XNOR on two vector inputs and store the result into a vector register. | |
| v_xor3_b32 | v_xor3_b32 | VOP3 | gfx1100 | Calculate the bitwise XOR of three vector inputs and store the result into a vector register. | |
| v_xor_b16 | v_xor_b16 | VOP3 | gfx1100 | Calculate bitwise XOR on two vector inputs and store the result into a vector register. | |
| v_xor_b16_fake16 | v_xor_b16_fake16 | VOP2 | AMDGPU VOP2 vector instruction operating on b16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) | ||
| v_xor_b16_t16 | v_xor_b16_t16 | VOP2 | AMDGPU VOP2 vector instruction operating on b16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) | ||
| v_xor_b32 | v_xor_b32 | VOP2 | gfx1100 | Calculate bitwise XOR on two vector inputs and store the result into a vector register. |
Source
Normalized from AMD's official ROCm documentation, with the LLVM AMDGPU backend documentation as supplementary compiler-target information. ROCm documentation ↗