vset2 GPU Virtual ISA NVIDIA

vset2 SIMD Video Instructions

// SIMD instruction with secondary SIMD merge operation
vset2.atype.btype.cmp d{.mask}, a{.asel}, b{.bsel}, c;

Two-way SIMD parallel comparison with secondary operation.

Encoding

PTX is a virtual instruction set. It has no single, stable native binary encoding - the compiler lowers this instruction to different native machine code depending on the selected NVIDIA target architecture (compute capability). This page intentionally shows no bit-diagram; see the target/version requirements below for what governs how this instruction compiles.
PTX ISA Version Introduced PTX ISA 3.0
Minimum Target sm_30

Syntax Forms

One mnemonic covers many type / state-space / scope / modifier combinations - each row below is an independently valid form.

Syntax Data Types State Space(s) Modifiers Min. Target Description
// SIMD instruction with secondary SIMD merge operation vset2.atype.btype.cmp d{.mask}, a{.asel}, b{.bsel}, c; sm_30 Two-way SIMD parallel comparison with secondary operation. (see the official PTX ISA docs for the full description)

Operands

At a Glance

Data Types -

Related

More in SIMD Video Instructions

Reference

NVIDIA PTX ISA

Description

Two-way SIMD parallel comparison with secondary operation. Elements of each dual half-word source to the operation are selected from any of the four half-words in the two source operands a and b using the asel and bsel modifiers. The selected half-words are then compared in parallel. The intermediate result of the comparison is always unsigned, and therefore the half-words of destination d and operand c are also unsigned. For instructions with a secondary SIMD merge operation: For half-word positions indicated in mask, the selected half-word results are copied into destination d. (see the official PTX ISA docs for the full description)

Semantics

// extract pairs of half-words and sign- or zero-extend // based on operand type Va = extractAndSignExt_2( a, b, .asel, .atype ); Vb = extractAndSignExt_2( a, b, .bsel, .btype ); Vc = extractAndSignExt_2( c ); for (i=0; i<2; i++) { t[i] = compare( Va[i], Vb[i], .cmp ) ? 1 : 0; } // secondary accumulate or SIMD merge mask = extractMaskBits( .mask ); if (.add) { d = c; for (i=0; i<2; i++) { d += mask[i] ? t[i] : 0; } } else { d = 0; for (i=0; i<2; i++) { d |= mask[i] ? t[i] : Vc[i]; } }

Examples

vset2.s32.u32.lt      r1, r2, r3, r0;
vset2.u32.u32.ne.add  r1, r2, r3, r0;

Reproduced from NVIDIA's official PTX ISA documentation for technical accuracy.

Sources