sub GPU Virtual ISA NVIDIA

Subtract Arithmetic

sub.type d, a, b;

Subtract the second operand from the first, with optional saturation for signed 32-bit integers.

Encoding

PTX is a virtual instruction set. It has no single, stable native binary encoding - the compiler lowers this instruction to different native machine code depending on the selected NVIDIA target architecture (compute capability). This page intentionally shows no bit-diagram; see the target/version requirements below for what governs how this instruction compiles.
PTX ISA Version Introduced PTX ISA 1.0
Minimum Target sm_10

Syntax Forms

One mnemonic covers many type / state-space / scope / modifier combinations - each row below is an independently valid form.

Syntax Data Types State Space(s) Modifiers Min. Target Description
sub.type d, a, b; s16, s32, s64, u16, u32, u64, f32, f64 sm_10 Generic subtract across integer and floating-point types.
sub.sat.s32 d, a, b; s32 sat sm_10 Signed 32-bit subtract with saturation.

Operands

  • d
    Destination register
  • a
    Minuend
  • b
    Subtrahend

At a Glance

Data Types f32, f64, s16, s32, s64, u16, u32, u64
Modifiers sat

Related AMDGPU Concepts

Integer Subtraction ↗
equivalent with restrictions
v_sub_u32 (AMDGPU)
s_sub_u32 (AMDGPU)

Related

More in Arithmetic

Reference

NVIDIA PTX ISA

Description

Performs subtraction and writes the resulting value into a destination register. For.f16x2 and.bf16x2 instruction type, forms input vectors by half word values from source operands. Half-word operands are then subtracted in parallel to produce.f16x2 or.bf16x2 result in destination. For.f16 instruction type, operands d, a and b have.f16 or.b16 type. For.f16x2 instruction type, operands d, a and b have.b32 type. For.bf16 instruction type, operands d, a, b have.b16 type. For.bf16x2 instruction type, operands d, a, b have.b32 type.

Semantics

d = a - b.

Examples

sub.s32 c,a,b;
sub.u8x4 p, q, r;

sub.f32 c,a,b;
sub.rn.ftz.f32  f1,f2,f3;

// scalar f16 subtractions
sub.f16        d0, a0, b0;
sub.rn.f16     d1, a1, b1;
sub.bf16       bd0, ba0, bb0;
sub.rn.bf16    bd1, ba1, bb1;
// (truncated - see the official PTX ISA docs for the full example)

.reg .f32 fc, fd;
.reg .f16 ha;
sub.rz.f32.f16.sat   fd, ha, fc;

Reproduced from NVIDIA's official PTX ISA documentation for technical accuracy.

Sources