ld GPU Virtual ISA NVIDIA
Load Data Movement and Conversion Instructions
ld.space.type d, [a];
Load a value from the specified state space into a register.
Encoding
PTX is a virtual instruction set. It has no single, stable native
binary encoding - the compiler lowers this instruction to different native machine
code depending on the selected NVIDIA target architecture (compute capability).
This page intentionally shows no bit-diagram; see the target/version requirements
below for what governs how this instruction compiles.
Syntax Forms
One mnemonic covers many type / state-space / scope / modifier combinations - each row below is an independently valid form.
| Syntax | Data Types | State Space(s) | Modifiers | Min. Target | Description |
|---|---|---|---|---|---|
| ld.space.type d, [a]; | b8, b16, b32, b64, s8, s16, s32, s64, u8, u16, u32, u64, f16, f32, f64 | global, local, shared, param, const | sm_10 | Load from an explicit state space. | |
| ld.global.nc.type d, [a]; | global | nc | sm_35 | Load through the read-only (non-coherent) data cache. |
Operands
-
d
Destination register -
a
Source address
At a Glance
Related AMDGPU Concepts
Global-Memory Load ↗
equivalent with restrictions
global_load_dword
(AMDGPU)
flat_load_dword
(AMDGPU)
buffer_load_dword
(AMDGPU)
Private and Scratch Memory Access ↗
equivalent with restrictions
scratch_load_dword
(AMDGPU)
scratch_store_dword
(AMDGPU)
Related
Reference
NVIDIA PTX ISA
Description
Load register variable d from the location specified by the source address operand a in
specified state space. If no state space is given, perform the load using Generic Addressing.
If no sub-qualifier is specified with.shared state space, then::cta is assumed by default.
Supported addressing modes for operand a and alignment requirements are described in Addresses as Operands
If no sub-qualifier is specified with.param state space, then:::func is assumed when access is inside a device function.::entry is assumed when accessing kernel function parameters from entry function. (see the official PTX ISA docs for the full description)
Semantics
d = *a, from the given state space.
Examples
ld.global.f32 d,[a];
ld.shared.v4.b32 Q,[p];
ld.const.s32 d,[p+4];
ld.local.b32 x,[p+-8]; // negative offset
ld.local.b64 x,[240]; // immediate address
// (truncated - see the official PTX ISA docs for the full example)Reproduced from NVIDIA's official PTX ISA documentation for technical accuracy.
Sources
-
Parallel Thread Execution ISA ↗
- NVIDIA Corporation, Chapter 9 - Instruction Set
Deep-linked directly to this instruction's section.