ds_swizzle_b32 GPU Native ISA AMD Vector

DS SWIZZLE B32 LDS / Data Share

ds_swizzle_b32

Dword swizzle, no data is written to LDS memory.

Encoding

Verified bit-level encoding data is not yet available for this instruction. The instruction-format classification below (DS) is well-documented and stable; exact per-target opcode/field bit positions have not yet been imported from a verified source.
Format DS
Execution Unit

Operands

Operand details have not yet been curated for this instruction.

GFX Target Compatibility

TargetSupport
gfx1100✅ Supported

Related

More in LDS / Data Share

Reference

AMDGPU / GFX ISA

Description

Dword swizzle, no data is written to LDS memory. Swizzles input thread data based on offset mask and returns; note does not read or write the DS memory banks. Note that reading from an invalid thread results in 0x0. This opcode supports two specific modes, FFT and rotate, plus two basic modes which swizzle in groups of 4 or 32 consecutive threads.

Semantics

The FFT mode (offset >= 0xe000) swizzles the input based on offset[4:0] to support FFT calculation. Example swizzles using input {1, 2, … 20} are: Offset[4:0]: Swizzle 0x00: {1,11,9,19,5,15,d,1d,3,13,b,1b,7,17,f,1f,2,12,a,1a,6,16,e,1e,4,14,c,1c,8,18,10,20} 0x10: {1,9,5,d,3,b,7,f,2,a,6,e,4,c,8,10,11,19,15,1d,13,1b,17,1f,12,1a,16,1e,14,1c,18,20} 0x1f: No swizzle The rotate mode (offset >= 0xc000 and offset < 0xe000) rotates the input either left (offset[10] == 0) or right (offset[10] == 1) a number of threads equal to offset[9:5]. The rotate mode also uses a mask value which can alter the rotate result. For example, mask == 1 swaps the odd threads across every other even thread (rotate left), or even threads across every other odd thread (rotate right). Offset[9:5]: Swizzle 0x01, mask=0, rotate left: {2,3,4,5,6,7,8,9,a,b,c,d,e,f,10,11,12,13,14,15,16,17,18,19,1a,1b,1c,1d,1e,1f,20,1} 0x01, mask=0, rotate right: {20,1,2,3,4,5,6,7,8,9,a,b,c,d,e,f,10,11,12,13,14,15,16,17,18,19,1a,1b,1c,1d,1e,1f} 0x01, mask=1, rotate left: {1,4,3,6,5,8,7,a,9,c,b,e,d,10,f,12,11,14,13,16,15,18,17,1a,19,1c,1b,1e,1d,20,1f,2} 0x01, mask=1, rotate right: {1f,2,1,4,3,6,5,8,7,a,9,c,b,e,d,10,f,12,11,14,13,16,15,18,17,1a,19,1c,1b,1e,1d,20} If offset < 0xc000, one of the basic swizzle modes is used based on offset[15]. If offset[15] == 1, groups of 4 consecutive threads are swizzled together. If offset[15] == 0, all 32 threads are swizzled together. The first basic swizzle mode (when offset[15] == 1) allows full data sharing between a group of 4 consecutive threads. Any thread within the group of 4 can get data from any other thread within the group of 4, specified by the corresponding offset bits --- [1:0] for the first thread, [3:2] for the second thread, [5:4] for the third thread, [7:6] for the fourth thread. Note that the offset bits apply to all groups of 4 within a wavefront; thus if offset[1:0] == 1, then thread0 grabs thread1, thread4 grabs thread5, etc. The second basic swizzle mode (when offset[15] == 0) allows limited data sharing between 32 consecutive threads. In this case, the offset is used to specify a 5-bit xor-mask, 5-bit or-mask, and 5-bit and-mask used to generate a thread mapping. Note that the offset bits apply to each group of 32 within a wavefront. The details of the thread mapping are listed below. Some example usages: SWAPX16 : xor_mask = 0x10, or_mask = 0x00, and_mask = 0x1f SWAPX8 : xor_mask = 0x08, or_mask = 0x00, and_mask = 0x1f SWAPX4 : xor_mask = 0x04, or_mask = 0x00, and_mask = 0x1f SWAPX2 : xor_mask = 0x02, or_mask = 0x00, and_mask = 0x1f SWAPX1 : xor_mask = 0x01, or_mask = 0x00, and_mask = 0x1f REVERSEX32 : xor_mask = 0x1f, or_mask = 0x00, and_mask = 0x1f REVERSEX16 : xor_mask = 0x0f, or_mask = 0x00, and_mask = 0x1f REVERSEX8 : xor_mask = 0x07, or_mask = 0x00, and_mask = 0x1f REVERSEX4 : xor_mask = 0x03, or_mask = 0x00, and_mask = 0x1f REVERSEX2 : xor_mask = 0x01 or_mask = 0x00, and_mask = 0x1f BCASTX32: xor_mask = 0x00, or_mask = thread, and_mask = 0x00 BCASTX16: xor_mask = 0x00, or_mask = thread, and_mask = 0x10 BCASTX8: xor_mask = 0x00, or_mask = thread, and_mask = 0x18 BCASTX4: xor_mask = 0x00, or_mask = thread, and_mask = 0x1c BCASTX2: xor_mask = 0x00, or_mask = thread, and_mask = 0x1e Pseudocode follows: offset = offset1:offset0; if (offset >= 0xe000) { // FFT decomposition mask = offset[4:0]; for (i = 0; i < 64; i++) { j = reverse_bits(i & 0x1f); j = (j >> count_ones(mask)); j |= (i & mask); j |= i & 0x20; thread_out[i] = thread_valid[j] ? thread_in[j] : 0; } } elsif (offset >= 0xc000) { // rotate rotate = offset[9:5]; mask = offset[4:0]; if (offset[10]) { rotate = -rotate; } for (i = 0; i < 64; i++) { j = (i & mask) | ((i + rotate) & ~mask); j |= i & 0x20; thread_out[i] = thread_valid[j] ? thread_in[j] : 0; } } elsif (offset[15]) { // full data sharing within 4 consecutive threads for (i = 0; i < 64; i+=4) { thread_out[i+0] = thread_valid[i+offset[1:0]]?thread_in[i+offset[1:0]]:0; thread_out[i+1] = thread_valid[i+offset[3:2]]?thread_in[i+offset[3:2]]:0; thread_out[i+2] = thread_valid[i+offset[5:4]]?thread_in[i+offset[5:4]]:0; thread_out[i+3] = thread_valid[i+offset[7:6]]?thread_in[i+offset[7:6]]:0; } } else { // offset[15] == 0 // limited data sharing within 32 consecutive threads xor_mask = offset[14:10]; or_mask = offset[9:5]; and_mask = offset[4:0]; for (i = 0; i < 64; i++) { j = (((i & 0x1f) & and_mask) | or_mask) ^ xor_mask; j |= (i & 0x20); // which group of 32 thread_out[i] = thread_valid[j] ? thread_in[j] : 0; } }

Example

ds_swizzle_b32 v8, v2

A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (gfx11). Not from an AMD document, and not authored here.

Sources