Vector Packed Arithmetic Instructions
279 AMDGPU instructions in this category - showing 79 per page, page 3 of 3 - click any row for encoding, pseudocode, and full documentation.
| Mnemonic | Syntax | Format | Summary |
|---|---|---|---|
| v_smfmac_f32_32x32x64bf8bf8 | v_smfmac_f32_32x32x64bf8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x64bf8fp8 | v_smfmac_f32_32x32x64bf8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x64fp8bf8 | v_smfmac_f32_32x32x64fp8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x64fp8fp8 | v_smfmac_f32_32x32x64fp8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_i32_16x16x128_i8 | v_smfmac_i32_16x16x128_i8 | VOP3P | Multiply the 16x128 sparse matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix stored… |
| v_smfmac_i32_16x16x128i8 | v_smfmac_i32_16x16x128i8 | VOP3P | AMDGPU VOP3P matrix instruction operating on i32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_i32_16x16x64_i8 | v_smfmac_i32_16x16x64_i8 | VOP3P | Multiply the 16x64 sparse matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_i32_16x16x64i8 | v_smfmac_i32_16x16x64i8 | VOP3P | AMDGPU VOP3P matrix instruction operating on i32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_i32_32x32x32_i8 | v_smfmac_i32_32x32x32_i8 | VOP3P | Multiply the 32x32 sparse matrix in the first input by the 32x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_i32_32x32x32i8 | v_smfmac_i32_32x32x32i8 | VOP3P | AMDGPU VOP3P matrix instruction operating on i32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_i32_32x32x64_i8 | v_smfmac_i32_32x32x64_i8 | VOP3P | Multiply the 32x64 sparse matrix in the first input by the 64x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_i32_32x32x64i8 | v_smfmac_i32_32x32x64i8 | VOP3P | AMDGPU VOP3P matrix instruction operating on i32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_swmmac_bf16_16x16x32_bf16 | v_swmmac_bf16_16x16x32_bf16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_bf16_16x16x64_bf16 | v_swmmac_bf16_16x16x64_bf16 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_bf16f32_16x16x64_bf16 | v_swmmac_bf16f32_16x16x64_bf16 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f16_16x16x128_bf8_bf8 | v_swmmac_f16_16x16x128_bf8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f16_16x16x128_bf8_fp8 | v_swmmac_f16_16x16x128_bf8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f16_16x16x128_fp8_bf8 | v_swmmac_f16_16x16x128_fp8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f16_16x16x128_fp8_fp8 | v_swmmac_f16_16x16x128_fp8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f16_16x16x32_f16 | v_swmmac_f16_16x16x32_f16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f16_16x16x64_f16 | v_swmmac_f16_16x16x64_f16 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x128_bf8_bf8 | v_swmmac_f32_16x16x128_bf8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x128_bf8_fp8 | v_swmmac_f32_16x16x128_bf8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x128_fp8_bf8 | v_swmmac_f32_16x16x128_fp8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x128_fp8_fp8 | v_swmmac_f32_16x16x128_fp8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x32_bf16 | v_swmmac_f32_16x16x32_bf16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x32_bf8_bf8 | v_swmmac_f32_16x16x32_bf8_bf8 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x32_bf8_fp8 | v_swmmac_f32_16x16x32_bf8_fp8 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x32_f16 | v_swmmac_f32_16x16x32_f16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x32_fp8_bf8 | v_swmmac_f32_16x16x32_fp8_bf8 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x32_fp8_fp8 | v_swmmac_f32_16x16x32_fp8_fp8 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x64_bf16 | v_swmmac_f32_16x16x64_bf16 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_f32_16x16x64_f16 | v_swmmac_f32_16x16x64_f16 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_i32_16x16x128_iu8 | v_swmmac_i32_16x16x128_iu8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_i32_16x16x32_iu4 | v_swmmac_i32_16x16x32_iu4 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_i32_16x16x32_iu8 | v_swmmac_i32_16x16x32_iu8 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_swmmac_i32_16x16x64_iu4 | v_swmmac_i32_16x16x64_iu4 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix in the… |
| v_wmma_bf16_16x16x16_bf16 | v_wmma_bf16_16x16x16_bf16 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_bf16_16x16x32_bf16 | v_wmma_bf16_16x16x32_bf16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_bf16f32_16x16x32_bf16 | v_wmma_bf16f32_16x16x32_bf16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f16_16x16x128_bf8_bf8 | v_wmma_f16_16x16x128_bf8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f16_16x16x128_bf8_fp8 | v_wmma_f16_16x16x128_bf8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f16_16x16x128_fp8_bf8 | v_wmma_f16_16x16x128_fp8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f16_16x16x128_fp8_fp8 | v_wmma_f16_16x16x128_fp8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f16_16x16x16_f16 | v_wmma_f16_16x16x16_f16 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f16_16x16x32_f16 | v_wmma_f16_16x16x32_f16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f16_16x16x64_bf8_bf8 | v_wmma_f16_16x16x64_bf8_bf8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f16_16x16x64_bf8_fp8 | v_wmma_f16_16x16x64_bf8_fp8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f16_16x16x64_fp8_bf8 | v_wmma_f16_16x16x64_fp8_bf8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f16_16x16x64_fp8_fp8 | v_wmma_f16_16x16x64_fp8_fp8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x128_bf8_bf8 | v_wmma_f32_16x16x128_bf8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f32_16x16x128_bf8_fp8 | v_wmma_f32_16x16x128_bf8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f32_16x16x128_f8f6f4 | v_wmma_f32_16x16x128_f8f6f4 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f32_16x16x128_fp8_bf8 | v_wmma_f32_16x16x128_fp8_bf8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f32_16x16x128_fp8_fp8 | v_wmma_f32_16x16x128_fp8_fp8 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_f32_16x16x16_bf16 | v_wmma_f32_16x16x16_bf16 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x16_bf8_bf8 | v_wmma_f32_16x16x16_bf8_bf8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x16_bf8_fp8 | v_wmma_f32_16x16x16_bf8_fp8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x16_f16 | v_wmma_f32_16x16x16_f16 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x16_fp8_bf8 | v_wmma_f32_16x16x16_fp8_bf8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x16_fp8_fp8 | v_wmma_f32_16x16x16_fp8_fp8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x32_bf16 | v_wmma_f32_16x16x32_bf16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x32_f16 | v_wmma_f32_16x16x32_f16 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x4_f32 | v_wmma_f32_16x16x4_f32 | VOP3P | Multiply the 16x4 matrix in the first input by the 4x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x64_bf8_bf8 | v_wmma_f32_16x16x64_bf8_bf8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x64_bf8_fp8 | v_wmma_f32_16x16x64_bf8_fp8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x64_fp8_bf8 | v_wmma_f32_16x16x64_fp8_bf8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_16x16x64_fp8_fp8 | v_wmma_f32_16x16x64_fp8_fp8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_f32_32x16x128_f4 | v_wmma_f32_32x16x128_f4 | VOP3P | Multiply the 32x128 matrix in the first input by the 128x16 matrix in the second input and add the 32x16 matrix in the third input using fused… |
| v_wmma_i32_16x16x16_iu4 | v_wmma_i32_16x16x16_iu4 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_i32_16x16x16_iu8 | v_wmma_i32_16x16x16_iu8 | VOP3P | Multiply the 16x16 matrix in the first input by the 16x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_i32_16x16x32_iu4 | v_wmma_i32_16x16x32_iu4 | VOP3P | Multiply the 16x32 matrix in the first input by the 32x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_i32_16x16x64_iu8 | v_wmma_i32_16x16x64_iu8 | VOP3P | Multiply the 16x64 matrix in the first input by the 64x16 matrix in the second input and add the 16x16 matrix in the third input using fused multiply… |
| v_wmma_ld_scale16_paired_b64 | v_wmma_ld_scale16_paired_b64 | VOP3P | AMDGPU VOP3P matrix instruction operating on b64 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_wmma_ld_scale_paired_b32 | v_wmma_ld_scale_paired_b32 | VOP3P | AMDGPU VOP3P matrix instruction operating on b32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_wmma_scale16_f32_16x16x128_f8f6f4 | v_wmma_scale16_f32_16x16x128_f8f6f4 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_scale16_f32_32x16x128_f4 | v_wmma_scale16_f32_32x16x128_f4 | VOP3P | Multiply the 32x128 matrix in the first input by the 128x16 matrix in the second input and add the 32x16 matrix in the third input using fused… |
| v_wmma_scale_f32_16x16x128_f8f6f4 | v_wmma_scale_f32_16x16x128_f8f6f4 | VOP3P | Multiply the 16x128 matrix in the first input by the 128x16 matrix in the second input and add the 16x16 matrix in the third input using fused… |
| v_wmma_scale_f32_32x16x128_f4 | v_wmma_scale_f32_32x16x128_f4 | VOP3P | Multiply the 32x128 matrix in the first input by the 128x16 matrix in the second input and add the 32x16 matrix in the third input using fused… |