Vector Arithmetic Instructions
558 AMDGPU instructions in this category - showing 100 per page, page 3 of 6 - click any row for encoding, pseudocode, and full documentation.
| Mnemonic | Syntax | Format | Summary |
|---|---|---|---|
| v_cvt_scalef32_sr_fp8_f16 | v_cvt_scalef32_sr_fp8_f16 | VOP3 | Scale a half-precision float input using the exponent provided by the third single-precision float input, then convert the values to an FP8 float… |
| v_cvt_scalef32_sr_fp8_f32 | v_cvt_scalef32_sr_fp8_f32 | VOP3 | Scale a single-precision float input using the exponent provided by the third single-precision float input, then convert the values to an FP8 float… |
| v_cvt_scalef32_sr_pk16_bf6_bf16 | v_cvt_scalef32_sr_pk16_bf6_bf16 | VOP3 | AMDGPU VOP3 vector instruction. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk16_bf6_f16 | v_cvt_scalef32_sr_pk16_bf6_f16 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk16_bf6_f32 | v_cvt_scalef32_sr_pk16_bf6_f32 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk16_fp6_bf16 | v_cvt_scalef32_sr_pk16_fp6_bf16 | VOP3 | AMDGPU VOP3 vector instruction. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk16_fp6_f16 | v_cvt_scalef32_sr_pk16_fp6_f16 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk16_fp6_f32 | v_cvt_scalef32_sr_pk16_fp6_f32 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk32_bf6_bf16 | v_cvt_scalef32_sr_pk32_bf6_bf16 | VOP3 | Scale a packed 32-component BF16 float input using the exponent provided by the third single-precision float input, then convert the values to a… |
| v_cvt_scalef32_sr_pk32_bf6_f16 | v_cvt_scalef32_sr_pk32_bf6_f16 | VOP3 | Scale a packed 32-component half-precision float input using the exponent provided by the third single-precision float input, then convert the values… |
| v_cvt_scalef32_sr_pk32_bf6_f32 | v_cvt_scalef32_sr_pk32_bf6_f32 | VOP3 | Scale a packed 32-component single-precision float input using the exponent provided by the third single-precision float input, then convert the… |
| v_cvt_scalef32_sr_pk32_fp6_bf16 | v_cvt_scalef32_sr_pk32_fp6_bf16 | VOP3 | Scale a packed 32-component BF16 float input using the exponent provided by the third single-precision float input, then convert the values to a… |
| v_cvt_scalef32_sr_pk32_fp6_f16 | v_cvt_scalef32_sr_pk32_fp6_f16 | VOP3 | Scale a packed 32-component half-precision float input using the exponent provided by the third single-precision float input, then convert the values… |
| v_cvt_scalef32_sr_pk32_fp6_f32 | v_cvt_scalef32_sr_pk32_fp6_f32 | VOP3 | Scale a packed 32-component single-precision float input using the exponent provided by the third single-precision float input, then convert the… |
| v_cvt_scalef32_sr_pk8_bf8_bf16 | v_cvt_scalef32_sr_pk8_bf8_bf16 | VOP3 | AMDGPU VOP3 vector instruction. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk8_bf8_f16 | v_cvt_scalef32_sr_pk8_bf8_f16 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk8_bf8_f32 | v_cvt_scalef32_sr_pk8_bf8_f32 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk8_fp4_bf16 | v_cvt_scalef32_sr_pk8_fp4_bf16 | VOP3 | AMDGPU VOP3 vector instruction. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk8_fp4_f16 | v_cvt_scalef32_sr_pk8_fp4_f16 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk8_fp4_f32 | v_cvt_scalef32_sr_pk8_fp4_f32 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk8_fp8_bf16 | v_cvt_scalef32_sr_pk8_fp8_bf16 | VOP3 | AMDGPU VOP3 vector instruction. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk8_fp8_f16 | v_cvt_scalef32_sr_pk8_fp8_f16 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk8_fp8_f32 | v_cvt_scalef32_sr_pk8_fp8_f32 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_scalef32_sr_pk_fp4_bf16 | v_cvt_scalef32_sr_pk_fp4_bf16 | VOP3 | Scale a packed 2-component BF16 float input using the exponent provided by the third single-precision float input, then convert the values to a… |
| v_cvt_scalef32_sr_pk_fp4_f16 | v_cvt_scalef32_sr_pk_fp4_f16 | VOP3 | Scale a packed 2-component half-precision float input using the exponent provided by the third single-precision float input, then convert the values… |
| v_cvt_sr_bf16_f32 | v_cvt_sr_bf16_f32 | VOP3 | Convert from a single-precision float input to a BF16 value with stochastic rounding using seed data from the second input. |
| v_cvt_sr_bf8_f16 | v_cvt_sr_bf8_f16 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_sr_bf8_f32 | v_cvt_sr_bf8_f32 | VOP3 | Convert from a single-precision float input to a BF8 value with stochastic rounding using seed data from the second input. |
| v_cvt_sr_bf8_f32_gfx12 | v_cvt_sr_bf8_f32_gfx12 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_sr_f16_f32 | v_cvt_sr_f16_f32 | VOP3 | Convert from a single-precision float input to a half-precision value with stochastic rounding using seed data from the second input. |
| v_cvt_sr_fp8_f16 | v_cvt_sr_fp8_f16 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_sr_fp8_f32 | v_cvt_sr_fp8_f32 | VOP3 | Convert from a single-precision float input to an FP8 value with stochastic rounding using seed data from the second input. |
| v_cvt_sr_fp8_f32_gfx12 | v_cvt_sr_fp8_f32_gfx12 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_sr_fp8_f32_gfx1250 | v_cvt_sr_fp8_f32_gfx1250 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_sr_pk_bf16_f32 | v_cvt_sr_pk_bf16_f32 | VOP3 | AMDGPU VOP3 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_sr_pk_f16_f32 | v_cvt_sr_pk_f16_f32 | VOP3 | AMDGPU VOP3 vector instruction operating on f16/f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_cvt_u16_f16 | v_cvt_u16_f16 | VOP1 | Convert from a half-precision float input to an unsigned 16-bit integer value and store the result into a vector register. |
| v_cvt_u32_f32 | v_cvt_u32_f32 | VOP1 | Convert from a single-precision float input to an unsigned 32-bit integer value and store the result into a vector register. |
| v_cvt_u32_f64 | v_cvt_u32_f64 | VOP1 | Convert from a double-precision float input to an unsigned 32-bit integer value and store the result into a vector register. |
| v_cvt_u32_u16 | v_cvt_u32_u16 | VOP1 | Convert from an unsigned 16-bit integer input to an unsigned 32-bit integer value using zero extension and store the result into a vector register. |
| v_div_fixup_f16 | v_div_fixup_f16 | VOP3 | Given a half-precision float quotient in the first input, a denominator in the second input and a numerator in the third input, detect and apply… |
| v_div_fixup_f16_gfx9 | v_div_fixup_f16_gfx9 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_div_fixup_f32 | v_div_fixup_f32 | VOP3 | Given a single-precision float quotient in the first input, a denominator in the second input and a numerator in the third input, detect and apply… |
| v_div_fixup_f64 | v_div_fixup_f64 | VOP3 | Given a double-precision float quotient in the first input, a denominator in the second input and a numerator in the third input, detect and apply… |
| v_div_fixup_legacy_f16 | v_div_fixup_legacy_f16 | VOP3 | Half precision division fixup. Has non-standard rule for OPSEL. |
| v_div_fmas_f32 | v_div_fmas_f32 | VOP3 | Multiply two single-precision float inputs and add a third input using fused multiply add, then scale the exponent of the result by a fixed factor if… |
| v_div_fmas_f64 | v_div_fmas_f64 | VOP3 | Multiply two double-precision float inputs and add a third input using fused multiply add, then scale the exponent of the result by a fixed factor if… |
| v_div_scale_f32 | v_div_scale_f32 | VOP3 | Given a single-precision float value to scale in the first input, a denominator in the second input and a numerator in the third input, scale the… |
| v_div_scale_f64 | v_div_scale_f64 | VOP3 | Given a double-precision float value to scale in the first input, a denominator in the second input and a numerator in the third input, scale the… |
| v_dot2_bf16_bf16 | v_dot2_bf16_bf16 | VOP3 | Compute the dot product of two packed 2-D BF16 float inputs, add the third input and store the result into a vector register. |
| v_dot2_f16_f16 | v_dot2_f16_f16 | VOP3 | Compute the dot product of two packed 2-D half-precision float inputs, add the third input and store the result into a vector register. |
| v_dot2c_f32_bf16 | v_dot2c_f32_bf16 | VOP2 | Compute the dot product of two packed 2-D BF16 float inputs in the single-precision float domain and accumulate with the single-precision float value… |
| v_dot2c_f32_f16 | v_dot2c_f32_f16 | VOP2 | Compute the dot product of two packed 2-D half-precision float inputs in the single-precision float domain and accumulate with the single-precision… |
| v_dot2c_i32_i16 | v_dot2c_i32_i16 | VOP2 | Compute the dot product of two packed 2-D signed 16-bit integer inputs in the signed 32-bit integer domain and accumulate with the signed 32-bit… |
| v_dot4c_i32_i8 | v_dot4c_i32_i8 | VOP2 | Compute the dot product of two packed 4-D signed 8-bit integer inputs in the signed 32-bit integer domain and accumulate with the signed 32-bit… |
| v_dot8c_i32_i4 | v_dot8c_i32_i4 | VOP2 | Compute the dot product of two packed 8-D signed 4-bit integer inputs in the signed 32-bit integer domain and accumulate with the signed 32-bit… |
| v_exp_bf16 | v_exp_bf16 | VOP1 | AMDGPU VOP1 vector instruction. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_exp_f16 | v_exp_f16 | VOP1 | Calculate 2 raised to the power of the half-precision float input and store the result into a vector register. |
| v_exp_f32 | v_exp_f32 | VOP1 | Calculate 2 raised to the power of the single-precision float input and store the result into a vector register. |
| v_exp_legacy_f32 | v_exp_legacy_f32 | VOP1 | AMDGPU VOP1 vector instruction operating on f32 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_ffbh_i32 | v_ffbh_i32 | VOP1 | Count the number of leading bits that are the same as the sign bit of a vector input and store the result into a vector register. |
| v_ffbh_u32 | v_ffbh_u32 | VOP1 | Count the number of leading "0" bits before the first "1" in a vector input and store the result into a vector register. |
| v_ffbl_b32 | v_ffbl_b32 | VOP1 | Count the number of trailing "0" bits before the first "1" in a vector input and store the result into a vector register. |
| v_floor_f16 | v_floor_f16 | VOP1 | Round the half-precision float input down to previous integer and store the result in floating point format into a vector register. |
| v_floor_f32 | v_floor_f32 | VOP1 | Round the single-precision float input down to previous integer and store the result in floating point format into a vector register. |
| v_floor_f64 | v_floor_f64 | VOP1 | Round the double-precision float input down to previous integer and store the result in floating point format into a vector register. |
| v_fma_dx9_zero_f32 | v_fma_dx9_zero_f32 | VOP3 | Multiply and add single-precision values. Follows DX9 rules where 0.0 times anything produces 0.0. |
| v_fma_f16 | v_fma_f16 | VOP3 | Multiply two half-precision float inputs and add a third input using fused multiply add, and store the result into a vector register. |
| v_fma_f16_gfx9 | v_fma_f16_gfx9 | VOP3 | AMDGPU VOP3 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fma_f32 | v_fma_f32 VDST, S0, S1, S2 | VOP3 | Per-lane single-precision fused multiply-add. |
| v_fma_f64 | v_fma_f64 | VOP3 | Multiply two double-precision float inputs and add a third input using fused multiply add, and store the result into a vector register. |
| v_fma_legacy_f16 | v_fma_legacy_f16 | VOP3 | Fused half precision multiply add. Implements IEEE rules and non-standard rule for OPSEL. |
| v_fma_legacy_f32 | v_fma_legacy_f32 | VOP3 | Multiply and add single-precision values. Follows DX9 rules where 0.0 times anything produces 0.0. |
| v_fmaak_f16 | v_fmaak_f16 | VOP2 | Multiply two half-precision float inputs and add a literal constant using fused multiply add, and store the result into a vector register. |
| v_fmaak_f16_fake16 | v_fmaak_f16_fake16 | VOP2 | AMDGPU VOP2 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fmaak_f16_t16 | v_fmaak_f16_t16 | VOP2 | AMDGPU VOP2 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fmaak_f32 | v_fmaak_f32 | VOP2 | Multiply two single-precision float inputs and add a literal constant using fused multiply add, and store the result into a vector register. |
| v_fmaak_f64 | v_fmaak_f64 | VOP2 | AMDGPU VOP2 vector instruction operating on f64 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fmac_f16 | v_fmac_f16 | VOP2 | Multiply two half-precision float inputs and accumulate the result into the destination register using fused multiply add. |
| v_fmac_f16_fake16 | v_fmac_f16_fake16 | VOP2 | AMDGPU VOP2 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fmac_f16_t16 | v_fmac_f16_t16 | VOP2 | AMDGPU VOP2 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fmac_f32 | v_fmac_f32 | VOP2 | Multiply two floating point inputs and accumulate the result into the destination register using fused multiply add. |
| v_fmac_f64 | v_fmac_f64 | VOP2 | Multiply two floating point inputs and accumulate the result into the destination register using fused multiply add. |
| v_fmac_legacy_f32 | v_fmac_legacy_f32 | VOP2 | Multiply two single-precision values and accumulate the result with the destination. Follows DX9 rules where 0.0 times anything produces 0.0. |
| v_fmamk_f16 | v_fmamk_f16 | VOP2 | Multiply a half-precision float input with a literal constant and add a second half-precision float input using fused multiply add, and store the… |
| v_fmamk_f16_fake16 | v_fmamk_f16_fake16 | VOP2 | AMDGPU VOP2 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fmamk_f16_t16 | v_fmamk_f16_t16 | VOP2 | AMDGPU VOP2 vector instruction operating on f16 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fmamk_f32 | v_fmamk_f32 | VOP2 | Multiply a single-precision float input with a literal constant and add a second single-precision float input using fused multiply add, and store the… |
| v_fmamk_f64 | v_fmamk_f64 | VOP2 | AMDGPU VOP2 vector instruction operating on f64 data. (Format and name extracted from LLVM's AMDGPU backend source - semantics not yet curated.) |
| v_fract_f16 | v_fract_f16 | VOP1 | Compute the fractional portion of a half-precision float input and store the result in floating point format into a vector register. |
| v_fract_f32 | v_fract_f32 | VOP1 | Compute the fractional portion of a single-precision float input and store the result in floating point format into a vector register. |
| v_fract_f64 | v_fract_f64 | VOP1 | Compute the fractional portion of a double-precision float input and store the result in floating point format into a vector register. |
| v_frexp_exp_i16_f16 | v_frexp_exp_i16_f16 | VOP1 | Extract the exponent of a half-precision float input and store the result as a signed 16-bit integer into a vector register. |
| v_frexp_exp_i32_f32 | v_frexp_exp_i32_f32 | VOP1 | Extract the exponent of a single-precision float input and store the result as a signed 32-bit integer into a vector register. |
| v_frexp_exp_i32_f64 | v_frexp_exp_i32_f64 | VOP1 | Extract the exponent of a double-precision float input and store the result as a signed 32-bit integer into a vector register. |
| v_frexp_mant_f16 | v_frexp_mant_f16 | VOP1 | Extract the binary significand, or mantissa, of a half-precision float input and store the result as a half- precision float into a vector register. |
| v_frexp_mant_f32 | v_frexp_mant_f32 | VOP1 | Extract the binary significand, or mantissa, of a single-precision float input and store the result as a single- precision float into a vector… |
| v_frexp_mant_f64 | v_frexp_mant_f64 | VOP1 | Extract the binary significand, or mantissa, of a double-precision float input and store the result as a double- precision float into a vector… |
| v_interp_mov_f32 | v_interp_mov_f32 | VOP3 | Given an attribute specifier and a parameter ID (P0, P10 or P20), load one of the parameter values from the local data share into a vector register. |
| v_interp_p1_f32 | v_interp_p1_f32 | VOP3 | Given the I coordinate in a vector register and an attribute specifier, load parameter data from the local data share, compute the first part of… |