global_atomic_max_num_f32 GPU Native ISA AMD Vector

Vector/Global Memory

global_atomic_max_num_f32

Select the IEEE maximumNumber() of two single-precision float inputs, given two values stored in the data register and a location in the global…

Also written as global_atomic_fmax, global_atomic_max_f32. AMD's machine-readable ISA specification lists these names for the same instruction.

Encoding

Binary Layout (AMD CDNA 5)
IOFFSET
95:72
VADDR
71:64
unassigned
63
VSRC
62:55
TH
54:52
SCOPE
51:50
SVE
49
SCALE_OFFSET
48
unassigned
47:40
VDST
39:32
11101110
31:24
unassigned
23:22
01010010
21:14
unassigned
13:8
NV
7
SADDR
6:0
 

The ENC_VGLOBAL layout from AMD's machine-readable ISA specification (AMD CDNA 5). Opcode 82 in OP rebuilds 0x0000000000000000EE148000, an identifier AMD lists for this encoding. Bits marked unassigned have no field in the specification. Not the same in every generation: AMD RDNA 4 (field layout).

Format GLOBAL
Width 96 bits
Opcode 82
Identifier 0x0000000000000000EE148000

Operands

  • VDST
    Written. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: Value can be anything. (OPR_VGPR, FMT_ANY)
  • VADDR
    Read. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: Value can be anything. (OPR_VGPR, FMT_ANY)
  • VSRC
    Read. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: Value can be anything. (OPR_VGPR, FMT_ANY)
  • SADDR
    Read. Any scalar GPR operand including VCC and NULL. Data format: Value can be anything. (OPR_SREG, FMT_ANY)

In AMD's order, as its machine-readable ISA specification lists them for the ENC_VGLOBAL encoding (AMD CDNA 5).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in Vector/Global Memory

Reference

AMDGPU / GFX ISA

Description

Select the IEEE maximumNumber() of two single-precision float inputs, given two values stored in the data register and a location in the global aperture. Update the global aperture with the selected value. Store the original value from global aperture into a vector register iff the temporal hint enables atomic return.

Example

global_atomic_max_num_f32 v0, v2, s[0:1] offset:64

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (gfx12_asm_vflat.s). Not from an AMD document, and not authored here.

Sources