s_andn2_wrexec_b64 GPU Native ISA AMD Scalar
S ANDN2 WREXEC B64 Scalar Arithmetic
Calculate bitwise AND on the scalar input and the negation of the EXEC mask, store the calculated result into the EXEC mask and also into the scalar…
Encoding
Operands
Operand details have not yet been curated for this instruction.
GFX Target Compatibility
| Target | Support |
|---|---|
| gfx1100 | ✅ Supported |
In particular, the following sequence of waterfall code is optimized by using a WREXEC instead of two separate scalar ops: // V0 holds the index value per lane // save exec mask for restore at the end s_mov_b64 s2, exec // exec mask of remaining (unprocessed) threads s_mov_b64 s4, exec loop: // get the index value for the first active lane v_readfirstlane_b32 s0, v0 // find all other lanes with same index value v_cmpx_eq s0, v0 <OP> // do the operation using the current EXEC mask. S0 holds the index. // mask out thread that was just executed // s_andn2_b64 s4, s4, exec // s_mov_b64 exec, s4 s_andn2_wrexec_b64 s4, s4 // replaces above 2 ops // repeat until EXEC==0 s_cbranch_scc1 loop s_mov_b64 exec, s2
Related
More in Scalar Arithmetic
Reference
Description
Semantics
Example
s_andn2_wrexec_b64 s[0:1], 0A real instruction accepted by the LLVM assembler, taken verbatim from LLVM's own AMDGPU MC test suite (gfx11). Not from an AMD document, and not authored here.
Sources
- User Guide for AMDGPU Backend ↗ - LLVM Project
-
"AMD Instinct MI300" Instruction Set Architecture: Reference Guide ↗
- Advanced Micro Devices, Inc.
Reference Guide, page 131. - LLVM MC assembler tests for AMDGPU (gfx11) ↗ - LLVM Project