Vector Packed Arithmetic Instructions
176 AMDGPU instructions in this category - showing 76 per page, page 2 of 2 - click any row for encoding, pseudocode, and full documentation.
| Mnemonic | Syntax | Format | Summary |
|---|---|---|---|
| v_pk_mul_bf16 | v_pk_mul_bf16 | VOP3P | AMDGPU VOP3P vector instruction. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_pk_mul_f16 | v_pk_mul_f16 | VOP3P | Multiply two packed half-precision float inputs component-wise and store the result into a vector register. |
| v_pk_mul_f32 | v_pk_mul_f32 | VOP3P | Multiply two packed single-precision float inputs component-wise and store the result into a vector register. |
| v_pk_mul_f64 | v_pk_mul_f64 | VOP3P | AMDGPU VOP3P vector instruction operating on f64 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_pk_mul_lo_u16 | v_pk_mul_lo_u16 | VOP3P | Multiply two packed unsigned 16-bit integer inputs component-wise and store the low bits of each resulting component into a vector register. |
| v_pk_sub_i16 | v_pk_sub_i16 | VOP3P | Subtract the second packed signed 16-bit integer input from the first input component-wise and store the result into a vector register. |
| v_pk_sub_nc_u64 | v_pk_sub_nc_u64 | VOP3P | AMDGPU VOP3P vector instruction operating on u64 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_pk_sub_u16 | v_pk_sub_u16 | VOP3P | Subtract the second packed unsigned 16-bit integer input from the first input component-wise and store the result into a vector register. |
| v_s_exp_f16 | v_s_exp_f16 | VOP3P | Calculate 2 raised to the power of the half-precision float input and store the result into a scalar register. |
| v_s_exp_f32 | v_s_exp_f32 | VOP3P | Calculate 2 raised to the power of the single-precision float input and store the result into a scalar register. |
| v_s_log_f16 | v_s_log_f16 | VOP3P | Calculate the base 2 logarithm of the half-precision float input and store the result into a scalar register. |
| v_s_log_f32 | v_s_log_f32 | VOP3P | Calculate the base 2 logarithm of the single-precision float input and store the result into a scalar register. |
| v_s_rcp_f16 | v_s_rcp_f16 | VOP3P | Calculate the reciprocal of the half-precision float input using IEEE rules and store the result into a scalar register. |
| v_s_rcp_f32 | v_s_rcp_f32 | VOP3P | Calculate the reciprocal of the single-precision float input using IEEE rules and store the result into a scalar register. |
| v_s_rsq_f16 | v_s_rsq_f16 | VOP3P | Calculate the reciprocal of the square root of the half-precision float input using IEEE rules and store the result into a scalar register. |
| v_s_rsq_f32 | v_s_rsq_f32 | VOP3P | Calculate the reciprocal of the square root of the single-precision float input using IEEE rules and store the result into a scalar register. |
| v_s_sqrt_f16 | v_s_sqrt_f16 | VOP3P | Calculate the square root of the half-precision float input using IEEE rules and store the result into a scalar register. |
| v_s_sqrt_f32 | v_s_sqrt_f32 | VOP3P | Calculate the square root of the single-precision float input using IEEE rules and store the result into a scalar register. |
| v_smfmac_f32_16x16x128_bf8_bf8 | v_smfmac_f32_16x16x128_bf8_bf8 | VOP3P | Multiply the 16x128 sparse matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix stored… |
| v_smfmac_f32_16x16x128_bf8_fp8 | v_smfmac_f32_16x16x128_bf8_fp8 | VOP3P | Multiply the 16x128 sparse matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix stored… |
| v_smfmac_f32_16x16x128_fp8_bf8 | v_smfmac_f32_16x16x128_fp8_bf8 | VOP3P | Multiply the 16x128 sparse matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix stored… |
| v_smfmac_f32_16x16x128_fp8_fp8 | v_smfmac_f32_16x16x128_fp8_fp8 | VOP3P | Multiply the 16x128 sparse matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix stored… |
| v_smfmac_f32_16x16x128bf8bf8 | v_smfmac_f32_16x16x128bf8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x128bf8fp8 | v_smfmac_f32_16x16x128bf8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x128fp8bf8 | v_smfmac_f32_16x16x128fp8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x128fp8fp8 | v_smfmac_f32_16x16x128fp8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x32_bf16 | v_smfmac_f32_16x16x32_bf16 | VOP3P | Multiply the 16x32 sparse matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_f32_16x16x32_f16 | v_smfmac_f32_16x16x32_f16 | VOP3P | Multiply the 16x32 sparse matrix in the first input by the 32x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_f32_16x16x32bf16 | v_smfmac_f32_16x16x32bf16 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x32f16 | v_smfmac_f32_16x16x32f16 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x64_bf16 | v_smfmac_f32_16x16x64_bf16 | VOP3P | Multiply the 16x64 sparse matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_f32_16x16x64_bf8_bf8 | v_smfmac_f32_16x16x64_bf8_bf8 | VOP3P | Multiply the 16x64 sparse matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_f32_16x16x64_bf8_fp8 | v_smfmac_f32_16x16x64_bf8_fp8 | VOP3P | Multiply the 16x64 sparse matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_f32_16x16x64_f16 | v_smfmac_f32_16x16x64_f16 | VOP3P | Multiply the 16x64 sparse matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_f32_16x16x64_fp8_bf8 | v_smfmac_f32_16x16x64_fp8_bf8 | VOP3P | Multiply the 16x64 sparse matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_f32_16x16x64_fp8_fp8 | v_smfmac_f32_16x16x64_fp8_fp8 | VOP3P | Multiply the 16x64 sparse matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_f32_16x16x64bf16 | v_smfmac_f32_16x16x64bf16 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x64bf8bf8 | v_smfmac_f32_16x16x64bf8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x64bf8fp8 | v_smfmac_f32_16x16x64bf8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x64f16 | v_smfmac_f32_16x16x64f16 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x64fp8bf8 | v_smfmac_f32_16x16x64fp8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_16x16x64fp8fp8 | v_smfmac_f32_16x16x64fp8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x16_bf16 | v_smfmac_f32_32x32x16_bf16 | VOP3P | Multiply the 32x16 sparse matrix in the first input by the 16x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x16_f16 | v_smfmac_f32_32x32x16_f16 | VOP3P | Multiply the 32x16 sparse matrix in the first input by the 16x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x16bf16 | v_smfmac_f32_32x32x16bf16 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x16f16 | v_smfmac_f32_32x32x16f16 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x32_bf16 | v_smfmac_f32_32x32x32_bf16 | VOP3P | Multiply the 32x32 sparse matrix in the first input by the 32x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x32_bf8_bf8 | v_smfmac_f32_32x32x32_bf8_bf8 | VOP3P | Multiply the 32x32 sparse matrix in the first input by the 32x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x32_bf8_fp8 | v_smfmac_f32_32x32x32_bf8_fp8 | VOP3P | Multiply the 32x32 sparse matrix in the first input by the 32x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x32_f16 | v_smfmac_f32_32x32x32_f16 | VOP3P | Multiply the 32x32 sparse matrix in the first input by the 32x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x32_fp8_bf8 | v_smfmac_f32_32x32x32_fp8_bf8 | VOP3P | Multiply the 32x32 sparse matrix in the first input by the 32x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x32_fp8_fp8 | v_smfmac_f32_32x32x32_fp8_fp8 | VOP3P | Multiply the 32x32 sparse matrix in the first input by the 32x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x32bf16 | v_smfmac_f32_32x32x32bf16 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x32bf8bf8 | v_smfmac_f32_32x32x32bf8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x32bf8fp8 | v_smfmac_f32_32x32x32bf8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x32f16 | v_smfmac_f32_32x32x32f16 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x32fp8bf8 | v_smfmac_f32_32x32x32fp8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x32fp8fp8 | v_smfmac_f32_32x32x32fp8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x64_bf8_bf8 | v_smfmac_f32_32x32x64_bf8_bf8 | VOP3P | Multiply the 32x64 sparse matrix in the first input by the 64x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x64_bf8_fp8 | v_smfmac_f32_32x32x64_bf8_fp8 | VOP3P | Multiply the 32x64 sparse matrix in the first input by the 64x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x64_fp8_bf8 | v_smfmac_f32_32x32x64_fp8_bf8 | VOP3P | Multiply the 32x64 sparse matrix in the first input by the 64x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x64_fp8_fp8 | v_smfmac_f32_32x32x64_fp8_fp8 | VOP3P | Multiply the 32x64 sparse matrix in the first input by the 64x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_f32_32x32x64bf8bf8 | v_smfmac_f32_32x32x64bf8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x64bf8fp8 | v_smfmac_f32_32x32x64bf8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x64fp8bf8 | v_smfmac_f32_32x32x64fp8bf8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_f32_32x32x64fp8fp8 | v_smfmac_f32_32x32x64fp8fp8 | VOP3P | AMDGPU VOP3P matrix instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_i32_16x16x128_i8 | v_smfmac_i32_16x16x128_i8 | VOP3P | Multiply the 16x128 sparse matrix in the first input by the 128x16 matrix in the second input and accumulate the result into the 16x16 matrix stored… |
| v_smfmac_i32_16x16x128i8 | v_smfmac_i32_16x16x128i8 | VOP3P | AMDGPU VOP3P matrix instruction operating on i32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_i32_16x16x64_i8 | v_smfmac_i32_16x16x64_i8 | VOP3P | Multiply the 16x64 sparse matrix in the first input by the 64x16 matrix in the second input and accumulate the result into the 16x16 matrix stored in… |
| v_smfmac_i32_16x16x64i8 | v_smfmac_i32_16x16x64i8 | VOP3P | AMDGPU VOP3P matrix instruction operating on i32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_i32_32x32x32_i8 | v_smfmac_i32_32x32x32_i8 | VOP3P | Multiply the 32x32 sparse matrix in the first input by the 32x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_i32_32x32x32i8 | v_smfmac_i32_32x32x32i8 | VOP3P | AMDGPU VOP3P matrix instruction operating on i32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_smfmac_i32_32x32x64_i8 | v_smfmac_i32_32x32x64_i8 | VOP3P | Multiply the 32x64 sparse matrix in the first input by the 64x32 matrix in the second input and accumulate the result into the 32x32 matrix stored in… |
| v_smfmac_i32_32x32x64i8 | v_smfmac_i32_32x32x64i8 | VOP3P | AMDGPU VOP3P matrix instruction operating on i32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_wmma_ld_scale16_paired_b64 | v_wmma_ld_scale16_paired_b64 | VOP3P | AMDGPU VOP3P matrix instruction operating on b64 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_wmma_ld_scale_paired_b32 | v_wmma_ld_scale_paired_b32 | VOP3P | AMDGPU VOP3P matrix instruction operating on b32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |