red GPU Virtual ISA NVIDIA
Reduction Parallel Synchronization and Communication Instructions
red.space.op.type [a], b;
Atomically read-modify-write a memory location without returning the prior value.
Encoding
PTX is a virtual instruction set. It has no single, stable native
binary encoding - the compiler lowers this instruction to different native machine
code depending on the selected NVIDIA target architecture (compute capability).
This page intentionally shows no bit-diagram; see the target/version requirements
below for what governs how this instruction compiles.
Syntax Forms
One mnemonic covers many type / state-space / scope / modifier combinations - each row below is an independently valid form.
| Syntax | Data Types | State Space(s) | Modifiers | Min. Target | Description |
|---|---|---|---|---|---|
| red.space.op.type [a], b; | b32, b64, s32, u32, u64, f32, f64 | global, shared | add, min, max, and, or, xor, inc, dec | sm_11 | Same read-modify-write as atom, but discards the prior value - cheaper when the old value isn't needed. |
Operands
-
a
Memory address -
b
Operand value
At a Glance
Related AMDGPU Concepts
Atomic Minimum and Maximum ↗
equivalent with restrictions
global_atomic_smin
(AMDGPU)
global_atomic_umax
(AMDGPU)
ds_min_rtn_i32
(AMDGPU)
Related
More in Parallel Synchronization and Communication Instructions
Reference
NVIDIA PTX ISA
Description
Performs a reduction operation with operand b and the value in location a, and stores the
result of the specified operation at location a, overwriting the original value. Operand a specifies a location in the specified state space. If no state space is given, perform the memory
accesses using Generic Addressing. red with scalar type may
be used only with.global and.shared spaces and with generic addressing, where the address
points to.global or.shared space. (see the official PTX ISA docs for the full description)
Semantics
*a = op(*a, b).
Examples
red.global.add.s32 [a],1;
red.shared::cluster.max.u32 [x+4],0;
@p red.global.and.b32 [p],my_val;
red.global.sys.add.u32 [a], 1;
red.global.acquire.sys.add.u32 [gbl], 1;
red.add.noftz.f16x2 [a], b;
// (truncated - see the official PTX ISA docs for the full example)Reproduced from NVIDIA's official PTX ISA documentation for technical accuracy.
Sources
-
Parallel Thread Execution ISA ↗
- NVIDIA Corporation, Chapter 9 - Instruction Set
Deep-linked directly to this instruction's section.