mul GPU Virtual ISA NVIDIA
Multiply Arithmetic
mul.mode.stype d, a, b;
Multiply two operands, selecting the low, high, or widened part of an integer product.
Encoding
PTX is a virtual instruction set. It has no single, stable native
binary encoding - the compiler lowers this instruction to different native machine
code depending on the selected NVIDIA target architecture (compute capability).
This page intentionally shows no bit-diagram; see the target/version requirements
below for what governs how this instruction compiles.
Syntax Forms
One mnemonic covers many type / state-space / scope / modifier combinations - each row below is an independently valid form.
| Syntax | Data Types | State Space(s) | Modifiers | Min. Target | Description |
|---|---|---|---|---|---|
| mul.mode.stype d, a, b; | s16, s32, s64, u16, u32, u64 | lo, hi, wide | sm_10 | Integer multiply; mode selects which part of the full product is written to d. | |
| mul.f32 d, a, b; | f32 | sm_10 | Single-precision floating-point multiply. | ||
| mul.rn.f64 d, a, b; | f64 | rn | sm_10 | Double-precision floating-point multiply with explicit round-to-nearest-even. |
Operands
-
d
Destination register -
a
First source operand -
b
Second source operand
Reference
NVIDIA PTX ISA
Description
Performs multiplication and writes the resulting value into a destination register.
For.f16x2 and.bf16x2 instruction type, forms input vectors by half word values from source
operands. Half-word operands are then multiplied in parallel to produce.f16x2 or.bf16x2 result in destination.
For.f16 instruction type, operands d, a and b have.f16 or.b16 type. For.f16x2 instruction type, operands d, a and b have.b32 type. For.bf16 instruction type, operands d, a, b have.b16 type. For.bf16x2 instruction type,
operands d, a, b have.b32 type.
Semantics
d = a * b, truncated to the selected result slice for integer forms.
Examples
mul.wide.s16 fa,fxs,fys; // 16*16 bits yields 32 bits
mul.lo.s16 fa,fxs,fys; // 16*16 bits, save only the low 16 bits
mul.wide.s32 z,x,y; // 32*32 bits, creates 64 bit result
mul.ftz.f32 circumf,radius,pi // a single-precision multiply
// scalar f16 multiplications
mul.f16 d0, a0, b0;
mul.rn.f16 d1, a1, b1;
mul.bf16 bd0, ba0, bb0;
mul.rn.bf16 bd1, ba1, bb1;
// (truncated - see the official PTX ISA docs for the full example)Reproduced from NVIDIA's official PTX ISA documentation for technical accuracy.
Sources
-
Parallel Thread Execution ISA ↗
- NVIDIA Corporation, Chapter 9 - Instruction Set
Deep-linked directly to this instruction's section.