v_cvt_scalef32_sr_pk_fp4_f32 GPU Native ISA AMD Vector

Vector Arithmetic

v_cvt_scalef32_sr_pk_fp4_f32

Scale a packed 2-component single-precision float input using the exponent provided by the third single-precision float input, then convert the…

Also written as v_cvt_scale_sr_pk_fp4_f32. AMD's machine-readable ISA specification lists this name for the same instruction.

Encoding

Binary Layout (AMD CDNA 4)
NEG
63:61
OMOD
60:59
SRC2
58:50
SRC1
49:41
SRC0
40:32
110100
31:26
1000111110
25:16
CLAMP
15
OP_SEL
14:11
ABS
10:8
VDST
7:0
 

The ENC_VOP3 layout from AMD's machine-readable ISA specification (AMD CDNA 4). Opcode 574 in OP rebuilds 0x00000000D23E0000, an identifier AMD lists for this encoding. No other generation in the specification defines this instruction.

Format VOP3
Width 64 bits
Opcode 574
Identifier 0x00000000D23E0000

Operands

  • VDST
    Written. N/A Data format: N/A (OPR_VGPR, FMT_NUM_PK2_FP4)
  • SRC0
    Read. N/A Data format: N/A (OPR_SRC_VGPR, FMT_NUM_PK2_F32)
  • SRC1
    Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_U32)
  • SRC2
    Read. N/A Data format: N/A (OPR_SRC_SIMPLE, FMT_NUM_F32)

In AMD's order, as its machine-readable ISA specification lists them for the ENC_VOP3 encoding (AMD CDNA 4).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in Vector Arithmetic

Reference

AMDGPU / GFX ISA

Description

Scale a packed 2-component single-precision float input using the exponent provided by the third single-precision float input, then convert the values to a packed FP4 float value with stochastic rounding using seed data from the second input. Store the result into 8 bits of a vector register using OPSEL[3:2] to determine which byte of the destination to overwrite.

Sources