vset4 GPU Virtual ISA NVIDIA

vset4 SIMD Video Instructions

// SIMD instruction with secondary SIMD merge operation
vset4.atype.btype.cmp d{.mask}, a{.asel}, b{.bsel}, c;

Four-way SIMD parallel comparison with secondary operation.

Encoding

PTX is a virtual instruction set. It has no single, stable native binary encoding - the compiler lowers this instruction to different native machine code depending on the selected NVIDIA target architecture (compute capability). This page intentionally shows no bit-diagram; see the target/version requirements below for what governs how this instruction compiles.
PTX ISA Version Introduced PTX ISA 3.0
Minimum Target sm_30

Syntax Forms

One mnemonic covers many type / state-space / scope / modifier combinations - each row below is an independently valid form.

Syntax Data Types State Space(s) Modifiers Min. Target Description
// SIMD instruction with secondary SIMD merge operation vset4.atype.btype.cmp d{.mask}, a{.asel}, b{.bsel}, c; sm_30 Four-way SIMD parallel comparison with secondary operation. (see the official PTX ISA docs for the full description)

Operands

At a Glance

Data Types -

Related

More in SIMD Video Instructions

Reference

NVIDIA PTX ISA

Description

Four-way SIMD parallel comparison with secondary operation. Elements of each quad byte source to the operation are selected from any of the eight bytes in the two source operands a and b using the asel and bsel modifiers. The selected bytes are then compared in parallel. The intermediate result of the comparison is always unsigned, and therefore the bytes of destination d and operand c are also unsigned. For instructions with a secondary SIMD merge operation: For byte positions indicated in mask, the selected byte results are copied into destination d. (see the official PTX ISA docs for the full description)

Semantics

// extract quads of bytes and sign- or zero-extend // based on operand type Va = extractAndSignExt_4( a, b, .asel, .atype ); Vb = extractAndSignExt_4( a, b, .bsel, .btype ); Vc = extractAndSignExt_4( c ); for (i=0; i<4; i++) { t[i] = compare( Va[i], Vb[i], cmp ) ? 1 : 0; } // secondary accumulate or SIMD merge mask = extractMaskBits( .mask ); if (.add) { d = c; for (i=0; i<4; i++) { d += mask[i] ? t[i] : 0; } } else { d = 0; for (i=0; i<4; i++) { d |= mask[i] ? t[i] : Vc[i]; } }

Examples

vset4.s32.u32.lt      r1, r2, r3, r0;
vset4.u32.u32.ne.max  r1, r2, r3, r0;

Reproduced from NVIDIA's official PTX ISA documentation for technical accuracy.

Sources