ds_bvh_stack_push4_pop1_rtn_b32 GPU Native ISA AMD Vector

LDS / Data Share

ds_bvh_stack_push4_pop1_rtn_b32

Ray tracing involves traversing a BVH which is a kind of tree where nodes have up to 4 children.

Encoding

Binary Layout (AMD RDNA 4)
VDST
63:56
DATA1
55:48
DATA0
47:40
ADDR
39:32
110110
31:26
11100000
25:18
unassigned
17:16
OFFSET1
15:8
OFFSET0
7:0
 

The ENC_VDS layout from AMD's machine-readable ISA specification (AMD RDNA 4). Opcode 224 in OP rebuilds 0x00000000DB800000, an identifier AMD lists for this encoding. Bits marked unassigned have no field in the specification. No other generation in the specification defines this instruction.

Format DS
Width 64 bits
Opcode 224
Identifier 0x00000000DB800000

Operands

  • VDST
    Written. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: 32 bits of arbitrary data. (OPR_VGPR, FMT_NUM_B32)
  • ADDR
    Written. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: 32 bits of arbitrary data. (OPR_VGPR, FMT_NUM_B32)
  • DATA0
    Read. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: 32 bits of arbitrary data. (OPR_VGPR, FMT_NUM_B32)
  • DATA1
    Read. Operand must be a vector GPR. Uses an 8-bit operand field. Data format: 128 bits of arbitrary data. (OPR_VGPR, FMT_NUM_B128)

In AMD's order, as its machine-readable ISA specification lists them for the ENC_VDS encoding (AMD RDNA 4).

GFX Target Compatibility

Per-target GFX compatibility has not yet been verified for this instruction.

Related

More in LDS / Data Share

Reference

AMDGPU / GFX ISA

Description

Ray tracing involves traversing a BVH which is a kind of tree where nodes have up to 4 children. Each shader thread processes one child at a time, and overflow nodes are stored temporarily in LDS using a stack. This instruction supports pushing/popping the stack to reduce the number of VALU instructions required per traversal and reduce VMEM bandwidth requirements.

Example

ds_bvh_stack_push4_pop1_rtn_b32 v1, v0, v1, v[2:5]

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (gfx12_asm_ds.s). Not from an AMD document, and not authored here.

Sources