fma GPU Virtual ISA NVIDIA
Fused Multiply-Add Arithmetic
fma.rn.f64 d, a, b, c;
Compute (a * b) + c with a single rounding step for improved precision over mad.
Encoding
PTX is a virtual instruction set. It has no single, stable native
binary encoding - the compiler lowers this instruction to different native machine
code depending on the selected NVIDIA target architecture (compute capability).
This page intentionally shows no bit-diagram; see the target/version requirements
below for what governs how this instruction compiles.
Syntax Forms
One mnemonic covers many type / state-space / scope / modifier combinations - each row below is an independently valid form.
| Syntax | Data Types | State Space(s) | Modifiers | Min. Target | Description |
|---|---|---|---|---|---|
| fma.rn.f64 d, a, b, c; | f64 | rn | sm_13 | Double-precision fused multiply-add. | |
| fma.rn.f32 d, a, b, c; | f32 | rn | sm_20 | Single-precision fused multiply-add. |
Operands
-
d
Destination register -
a
Multiplicand -
b
Multiplier -
c
Addend
Reference
NVIDIA PTX ISA
Description
Performs a fused multiply-add with no loss of precision in the intermediate product and addition.
For.f16x2 and.bf16x2 instruction type, forms input vectors by half word values from source
operands. Half-word operands are then operated in parallel to produce.f16x2 or.bf16x2 result in destination.
For.f16 instruction type, operands d, a, b and c have.f16 or.b16 type. For.f16x2 instruction type, operands d, a, b and c have.b32 type. For.bf16 instruction type, operands d, a, b and c have.b16 type. For.bf16x2 instruction type, operands d, a, b and c have.b32 type.
Semantics
d = round_once(a * b + c), the product is not rounded before the addition.
Examples
fma.rn.ftz.f32 w,x,y,z;
@p fma.rn.f64 d,a,b,c;
fma.rp.ftz.f32x2 p,q,r,s;
// scalar f16 fused multiply-add
fma.rn.f16 d0, a0, b0, c0;
fma.rn.f16 d1, a1, b1, c1;
fma.rn.relu.f16 d1, a1, b1, c1;
fma.rn.oob.f16 d1, a1, b1, c1;
fma.rn.oob.relu.f16 d1, a1, b1, c1;
// (truncated - see the official PTX ISA docs for the full example)
.reg .f32 fc, fd;
.reg .f16 ha, hb;
fma.rz.sat.f32.f16.sat fd, ha, hb, fc;Reproduced from NVIDIA's official PTX ISA documentation for technical accuracy.
Sources
-
Parallel Thread Execution ISA ↗
- NVIDIA Corporation, Chapter 9 - Instruction Set
Deep-linked directly to this instruction's section.