{"api_version":"1","snapshot":"2026-10-08","licence":"CC0-1.0","architecture":"PTX","records":190,"mnemonics":[{"mnemonic":"abs","slug":"abs","records":1,"summary":"Compute the absolute value of a signed or floating-point operand.","page":"https://instructionsets.com/ptx/abs/","api":"https://instructionsets.com/api/v1/ptx/abs.json"},{"mnemonic":"activemask","slug":"activemask","records":1,"summary":"Query the bitmask of currently active (converged) lanes in the executing warp.","page":"https://instructionsets.com/ptx/activemask/","api":"https://instructionsets.com/api/v1/ptx/activemask.json"},{"mnemonic":"add","slug":"add","records":1,"summary":"Add two operands of the same type, with optional saturation for signed 32-bit integers.","page":"https://instructionsets.com/ptx/add/","api":"https://instructionsets.com/api/v1/ptx/add.json"},{"mnemonic":"add.cc","slug":"add_cc","records":1,"summary":"Performs integer addition and writes the carry-out value into the condition code register.","page":"https://instructionsets.com/ptx/add_cc/","api":"https://instructionsets.com/api/v1/ptx/add_cc.json"},{"mnemonic":"addc","slug":"addc","records":1,"summary":"Performs integer addition with carry-in and optionally writes the carry-out value into the condition\ncode register.","page":"https://instructionsets.com/ptx/addc/","api":"https://instructionsets.com/api/v1/ptx/addc.json"},{"mnemonic":"alloca","slug":"alloca","records":1,"summary":"The alloca instruction dynamically allocates memory on the stack frame of the current function and updates the stack pointer accordingly.","page":"https://instructionsets.com/ptx/alloca/","api":"https://instructionsets.com/api/v1/ptx/alloca.json"},{"mnemonic":"and","slug":"and","records":1,"summary":"Bitwise AND of two operands.","page":"https://instructionsets.com/ptx/and/","api":"https://instructionsets.com/api/v1/ptx/and.json"},{"mnemonic":"applypriority","slug":"applypriority","records":1,"summary":"The applypriority instruction applies the cache eviction priority specified by the.level::eviction_priority qualifier to the address range [a..a+size)","page":"https://instructionsets.com/ptx/applypriority/","api":"https://instructionsets.com/api/v1/ptx/applypriority.json"},{"mnemonic":"atom","slug":"atom","records":1,"summary":"Atomically read-modify-write a memory location and return the prior value.","page":"https://instructionsets.com/ptx/atom/","api":"https://instructionsets.com/api/v1/ptx/atom.json"},{"mnemonic":"bar.cta","slug":"bar_cta","records":1,"summary":"Performs barrier synchronization and communication within a CTA.","page":"https://instructionsets.com/ptx/bar_cta/","api":"https://instructionsets.com/api/v1/ptx/bar_cta.json"},{"mnemonic":"bar.warp.sync","slug":"bar_warp_sync","records":1,"summary":"bar.warp.sync will cause executing thread to wait until all threads corresponding to membermask have executed a bar.warp.sync with the same membermask value before resuming execution.","page":"https://instructionsets.com/ptx/bar_warp_sync/","api":"https://instructionsets.com/api/v1/ptx/bar_warp_sync.json"},{"mnemonic":"barrier","slug":"barrier","records":1,"summary":"Block threads in a CTA at a named barrier until the expected number of threads has arrived.","page":"https://instructionsets.com/ptx/barrier/","api":"https://instructionsets.com/api/v1/ptx/barrier.json"},{"mnemonic":"barrier.cluster","slug":"barrier_cluster","records":1,"summary":"Performs barrier synchronization and communication within a cluster.","page":"https://instructionsets.com/ptx/barrier_cluster/","api":"https://instructionsets.com/api/v1/ptx/barrier_cluster.json"},{"mnemonic":"barrier.cta","slug":"barrier_cta","records":1,"summary":"Performs barrier synchronization and communication within a CTA.","page":"https://instructionsets.com/ptx/barrier_cta/","api":"https://instructionsets.com/api/v1/ptx/barrier_cta.json"},{"mnemonic":"bfe","slug":"bfe","records":1,"summary":"Extract bit field from a and place the zero or sign-extended result in d.","page":"https://instructionsets.com/ptx/bfe/","api":"https://instructionsets.com/api/v1/ptx/bfe.json"},{"mnemonic":"bfi","slug":"bfi","records":1,"summary":"Align and insert a bit field from a into b, and place the result in f.","page":"https://instructionsets.com/ptx/bfi/","api":"https://instructionsets.com/api/v1/ptx/bfi.json"},{"mnemonic":"bfind","slug":"bfind","records":1,"summary":"Find the bit position of the most significant non-sign bit in a and place the result in d.","page":"https://instructionsets.com/ptx/bfind/","api":"https://instructionsets.com/api/v1/ptx/bfind.json"},{"mnemonic":"bmsk","slug":"bmsk","records":1,"summary":"Generates a 32-bit mask starting from the bit position specified in operand a, and of the width specified in operand b.","page":"https://instructionsets.com/ptx/bmsk/","api":"https://instructionsets.com/api/v1/ptx/bmsk.json"},{"mnemonic":"bra","slug":"bra","records":1,"summary":"Continue execution at the target.","page":"https://instructionsets.com/ptx/bra/","api":"https://instructionsets.com/api/v1/ptx/bra.json"},{"mnemonic":"brev","slug":"brev","records":1,"summary":"Perform bitwise reversal of input.","page":"https://instructionsets.com/ptx/brev/","api":"https://instructionsets.com/api/v1/ptx/brev.json"},{"mnemonic":"brkpt","slug":"brkpt","records":1,"summary":"Suspends execution.","page":"https://instructionsets.com/ptx/brkpt/","api":"https://instructionsets.com/api/v1/ptx/brkpt.json"},{"mnemonic":"brx.idx","slug":"brx_idx","records":1,"summary":"Index into a list of possible destination labels, and continue execution from the chosen label.","page":"https://instructionsets.com/ptx/brx_idx/","api":"https://instructionsets.com/api/v1/ptx/brx_idx.json"},{"mnemonic":"call","slug":"call","records":1,"summary":"The call instruction stores the address of the next instruction, so execution can resume at that point after executing a ret instruction.","page":"https://instructionsets.com/ptx/call/","api":"https://instructionsets.com/api/v1/ptx/call.json"},{"mnemonic":"clmad","slug":"clmad","records":1,"summary":"Performs a carryless multiplication of a and b, followed by a carryless addition of c, and writes the result into destination register d.","page":"https://instructionsets.com/ptx/clmad/","api":"https://instructionsets.com/api/v1/ptx/clmad.json"},{"mnemonic":"clusterlaunchcontrol.query_cancel","slug":"clusterlaunchcontrol_query_cancel","records":1,"summary":"Instruction clusterlaunchcontrol.query_cancel can be used to decode opaque response written by instruction clusterlaunchcontrol.try_cancel.","page":"https://instructionsets.com/ptx/clusterlaunchcontrol_query_cancel/","api":"https://instructionsets.com/api/v1/ptx/clusterlaunchcontrol_query_cancel.json"},{"mnemonic":"clusterlaunchcontrol.try_cancel","slug":"clusterlaunchcontrol_try_cancel","records":1,"summary":"The clusterlaunchcontrol.try_cancel instruction requests atomically cancelling the launch of a cluster that has not started running yet.","page":"https://instructionsets.com/ptx/clusterlaunchcontrol_try_cancel/","api":"https://instructionsets.com/api/v1/ptx/clusterlaunchcontrol_try_cancel.json"},{"mnemonic":"clz","slug":"clz","records":1,"summary":"Count the number of leading zero bits in an integer operand.","page":"https://instructionsets.com/ptx/clz/","api":"https://instructionsets.com/api/v1/ptx/clz.json"},{"mnemonic":"cnot","slug":"cnot","records":1,"summary":"Compute the logical negation using C/C++ semantics.","page":"https://instructionsets.com/ptx/cnot/","api":"https://instructionsets.com/api/v1/ptx/cnot.json"},{"mnemonic":"copysign","slug":"copysign","records":1,"summary":"Copy sign bit of a into value of b, and return the result as d.","page":"https://instructionsets.com/ptx/copysign/","api":"https://instructionsets.com/api/v1/ptx/copysign.json"},{"mnemonic":"cos","slug":"cos","records":1,"summary":"Find the cosine of the angle a (in radians).","page":"https://instructionsets.com/ptx/cos/","api":"https://instructionsets.com/api/v1/ptx/cos.json"},{"mnemonic":"cp.async","slug":"cp_async","records":1,"summary":"cp.async is a non-blocking instruction which initiates an asynchronous copy operation of data from the location specified by source address operand src to the location specified by destination address operand dst.","page":"https://instructionsets.com/ptx/cp_async/","api":"https://instructionsets.com/api/v1/ptx/cp_async.json"},{"mnemonic":"cp.async.bulk","slug":"cp_async_bulk","records":1,"summary":"cp.async.bulk is a non-blocking instruction which initiates an asynchronous bulk-copy operation from the location specified by source address operand srcMem to the location specified by destination address operand dstMem.","page":"https://instructionsets.com/ptx/cp_async_bulk/","api":"https://instructionsets.com/api/v1/ptx/cp_async_bulk.json"},{"mnemonic":"cp.async.bulk.commit_group","slug":"cp_async_bulk_commit_group","records":1,"summary":"cp.async.bulk.commit_group instruction creates a new per-thread bulk async-group and batches all prior cp{.reduce}.async.bulk{.prefetch}{.tensor} instructions satisfying the following conditions into…","page":"https://instructionsets.com/ptx/cp_async_bulk_commit_group/","api":"https://instructionsets.com/api/v1/ptx/cp_async_bulk_commit_group.json"},{"mnemonic":"cp.async.bulk.prefetch","slug":"cp_async_bulk_prefetch","records":1,"summary":"cp.async.bulk.prefetch is a non-blocking instruction which may initiate an asynchronous prefetch of data from the location specified by source address operand srcMem, in.src statespace, to the L2 cache.","page":"https://instructionsets.com/ptx/cp_async_bulk_prefetch/","api":"https://instructionsets.com/api/v1/ptx/cp_async_bulk_prefetch.json"},{"mnemonic":"cp.async.bulk.prefetch.tensor","slug":"cp_async_bulk_prefetch_tensor","records":1,"summary":"cp.async.bulk.prefetch.tensor is a non-blocking instruction which may initiate an asynchronous prefetch of tensor data from the location in.src statespace to the L2 cache.","page":"https://instructionsets.com/ptx/cp_async_bulk_prefetch_tensor/","api":"https://instructionsets.com/api/v1/ptx/cp_async_bulk_prefetch_tensor.json"},{"mnemonic":"cp.async.bulk.tensor","slug":"cp_async_bulk_tensor","records":1,"summary":"cp.async.bulk.tensor is a non-blocking instruction which initiates an asynchronous copy operation of tensor data from the location in.src state space to the location in the.dst state space.","page":"https://instructionsets.com/ptx/cp_async_bulk_tensor/","api":"https://instructionsets.com/api/v1/ptx/cp_async_bulk_tensor.json"},{"mnemonic":"cp.async.bulk.wait_group","slug":"cp_async_bulk_wait_group","records":1,"summary":"cp.async.bulk.wait_group instruction will cause the executing thread to wait until only N or fewer of the most recent bulk async-groups are pending and all the prior bulk async-groups committed by the executing threads are complete.","page":"https://instructionsets.com/ptx/cp_async_bulk_wait_group/","api":"https://instructionsets.com/api/v1/ptx/cp_async_bulk_wait_group.json"},{"mnemonic":"cp.async.commit_group","slug":"cp_async_commit_group","records":1,"summary":"cp.async.commit_group instruction creates a new cp.async-group per thread and batches all prior cp.async instructions initiated by the executing thread but not committed to any cp.async-group into the new cp.async-group.","page":"https://instructionsets.com/ptx/cp_async_commit_group/","api":"https://instructionsets.com/api/v1/ptx/cp_async_commit_group.json"},{"mnemonic":"cp.async.mbarrier.arrive","slug":"cp_async_mbarrier_arrive","records":1,"summary":"Causes an arrive-on operation to be triggered by the system on the mbarrier object upon the completion of all prior cp.async operations initiated by the executing thread.","page":"https://instructionsets.com/ptx/cp_async_mbarrier_arrive/","api":"https://instructionsets.com/api/v1/ptx/cp_async_mbarrier_arrive.json"},{"mnemonic":"cp.async.wait_all","slug":"cp_async_wait_all","records":1,"summary":"cp.async.wait_all instruction will cause the executing thread to wait until all the prior cp.async operations are complete. It is equivalent to cp.async.commit_group immediately followed by cp.async.wait_group 0.","page":"https://instructionsets.com/ptx/cp_async_wait_all/","api":"https://instructionsets.com/api/v1/ptx/cp_async_wait_all.json"},{"mnemonic":"cp.async.wait_group","slug":"cp_async_wait_group","records":1,"summary":"cp.async.wait_group instruction will cause executing thread to wait till only N or fewer of the most recent cp.async-group s are pending and all the prior cp.async-group s committed by the executing threads are complete.","page":"https://instructionsets.com/ptx/cp_async_wait_group/","api":"https://instructionsets.com/api/v1/ptx/cp_async_wait_group.json"},{"mnemonic":"cp.reduce.async.bulk","slug":"cp_reduce_async_bulk","records":1,"summary":"cp.reduce.async.bulk is a non-blocking instruction which initiates an asynchronous reduction operation on an array of memory locations specified by the destination address operand dstMem with the source array whose location is specified by the source address operand srcMem.","page":"https://instructionsets.com/ptx/cp_reduce_async_bulk/","api":"https://instructionsets.com/api/v1/ptx/cp_reduce_async_bulk.json"},{"mnemonic":"cp.reduce.async.bulk.tensor","slug":"cp_reduce_async_bulk_tensor","records":1,"summary":"cp.reduce.async.bulk.tensor is a non-blocking instruction which initiates an asynchronous reduction operation of tensor data in the.dst state space with tensor data in the.src state space.","page":"https://instructionsets.com/ptx/cp_reduce_async_bulk_tensor/","api":"https://instructionsets.com/api/v1/ptx/cp_reduce_async_bulk_tensor.json"},{"mnemonic":"createpolicy","slug":"createpolicy","records":1,"summary":"The createpolicy instruction creates a cache eviction policy for the specified cache level in an opaque 64-bit register specified by the destination operand cache_policy.","page":"https://instructionsets.com/ptx/createpolicy/","api":"https://instructionsets.com/api/v1/ptx/createpolicy.json"},{"mnemonic":"cvt","slug":"cvt","records":1,"summary":"Convert a value between integer and/or floating-point types with an explicit rounding mode.","page":"https://instructionsets.com/ptx/cvt/","api":"https://instructionsets.com/api/v1/ptx/cvt.json"},{"mnemonic":"cvt.pack","slug":"cvt_pack","records":1,"summary":"Convert two 32-bit integers a and b into specified type and pack the results into d.","page":"https://instructionsets.com/ptx/cvt_pack/","api":"https://instructionsets.com/api/v1/ptx/cvt_pack.json"},{"mnemonic":"cvta","slug":"cvta","records":1,"summary":"Convert a const, Kernel Function Parameters (.param ), global, local, or shared address to a generic address, or vice-versa.","page":"https://instructionsets.com/ptx/cvta/","api":"https://instructionsets.com/api/v1/ptx/cvta.json"},{"mnemonic":"discard","slug":"discard","records":1,"summary":"Semantically, this behaves like a weak write of an unstable indeterminate value: reads of memory locations with unstable indeterminate values may return different bit patterns each time until the memory is overwritten.","page":"https://instructionsets.com/ptx/discard/","api":"https://instructionsets.com/api/v1/ptx/discard.json"},{"mnemonic":"div","slug":"div","records":1,"summary":"Divide the first operand by the second.","page":"https://instructionsets.com/ptx/div/","api":"https://instructionsets.com/api/v1/ptx/div.json"},{"mnemonic":"dp2a","slug":"dp2a","records":1,"summary":"Two-way 16-bit to 8-bit dot product which is accumulated in 32-bit result.","page":"https://instructionsets.com/ptx/dp2a/","api":"https://instructionsets.com/api/v1/ptx/dp2a.json"},{"mnemonic":"dp4a","slug":"dp4a","records":1,"summary":"Four-way byte dot product which is accumulated in 32-bit result.","page":"https://instructionsets.com/ptx/dp4a/","api":"https://instructionsets.com/api/v1/ptx/dp4a.json"},{"mnemonic":"elect.sync","slug":"elect_sync","records":1,"summary":"elect.sync elects one predicated active leader thread from among a set of threads specified by membermask.","page":"https://instructionsets.com/ptx/elect_sync/","api":"https://instructionsets.com/api/v1/ptx/elect_sync.json"},{"mnemonic":"ex2","slug":"ex2","records":1,"summary":"Raise 2 to the power a.","page":"https://instructionsets.com/ptx/ex2/","api":"https://instructionsets.com/api/v1/ptx/ex2.json"},{"mnemonic":"exit","slug":"exit","records":1,"summary":"Ends execution of a thread.\nBarriers exclusively waiting on arrivals from exited threads are always released.","page":"https://instructionsets.com/ptx/exit/","api":"https://instructionsets.com/api/v1/ptx/exit.json"},{"mnemonic":"fabric.submit","slug":"fabric_submit","records":1,"summary":"Submits prior fabric operations issued by the current thread.","page":"https://instructionsets.com/ptx/fabric_submit/","api":"https://instructionsets.com/api/v1/ptx/fabric_submit.json"},{"mnemonic":"fabric.try_get","slug":"fabric_try_get","records":1,"summary":"Asynchronously copies size bytes from fabric handle [srcLeId, srcDataOff] to destination memory [dst], where srcLeId is a 32-bit unsigned value denoting the logical endpoint identifier, and…","page":"https://instructionsets.com/ptx/fabric_try_get/","api":"https://instructionsets.com/api/v1/ptx/fabric_try_get.json"},{"mnemonic":"fabric.try_pullred","slug":"fabric_try_pullred","records":1,"summary":"Initiates asynchronous loads from multiple resources pointed to by multicast fabric handle [srcLeId, srcDataOff], of size bytes, and performs element-","page":"https://instructionsets.com/ptx/fabric_try_pullred/","api":"https://instructionsets.com/api/v1/ptx/fabric_try_pullred.json"},{"mnemonic":"fabric.try_put","slug":"fabric_try_put","records":1,"summary":"Asynchronously copies size bytes from [src] to destination fabric handle [dstLeId, dstDataOff], where dstLeId is a 32-bit unsigned value denoting the logical endpoint identifier, and dstDataOff is a…","page":"https://instructionsets.com/ptx/fabric_try_put/","api":"https://instructionsets.com/api/v1/ptx/fabric_try_put.json"},{"mnemonic":"fabric.try_red","slug":"fabric_try_red","records":1,"summary":"Asynchronously copies size bytes from [src] to destination fabric handle [dstLeId, dstDataOff] with element-wise reduction, where dstLeId is a 32-bit unsigned value denoting the logical endpoint…","page":"https://instructionsets.com/ptx/fabric_try_red/","api":"https://instructionsets.com/api/v1/ptx/fabric_try_red.json"},{"mnemonic":"fabric.wait","slug":"fabric_wait","records":1,"summary":"Fabric-read completion mechanism instruction fabric.wait waits on the local shared memory (.shared::cta ) reads of submitted fabric operations.","page":"https://instructionsets.com/ptx/fabric_wait/","api":"https://instructionsets.com/api/v1/ptx/fabric_wait.json"},{"mnemonic":"fence","slug":"fence","records":1,"summary":"The fence instruction establishes an ordering between memory accesses requested by this thread, as described by the memory consistency model.","page":"https://instructionsets.com/ptx/fence/","api":"https://instructionsets.com/api/v1/ptx/fence.json"},{"mnemonic":"fma","slug":"fma","records":1,"summary":"Compute (a * b) + c with a single rounding step for improved precision over mad.","page":"https://instructionsets.com/ptx/fma/","api":"https://instructionsets.com/api/v1/ptx/fma.json"},{"mnemonic":"fns","slug":"fns","records":1,"summary":"Given a 32-bit value mask and an integer value base (between 0 and 31), find the n-th (given by offset) set bit in mask from the base bit, and store the bit position in d.","page":"https://instructionsets.com/ptx/fns/","api":"https://instructionsets.com/api/v1/ptx/fns.json"},{"mnemonic":"getctarank","slug":"getctarank","records":1,"summary":"Write the destination register d with the rank of the CTA which contains the address specified in operand a.","page":"https://instructionsets.com/ptx/getctarank/","api":"https://instructionsets.com/api/v1/ptx/getctarank.json"},{"mnemonic":"griddepcontrol","slug":"griddepcontrol","records":1,"summary":"The griddepcontrol instruction allows the dependent grids and prerequisite grids as defined by\nthe runtime, to control execution in the following way:","page":"https://instructionsets.com/ptx/griddepcontrol/","api":"https://instructionsets.com/api/v1/ptx/griddepcontrol.json"},{"mnemonic":"isspacep","slug":"isspacep","records":1,"summary":"Write predicate register p with 1 if generic address a falls within the specified state space window and with 0 otherwise.","page":"https://instructionsets.com/ptx/isspacep/","api":"https://instructionsets.com/api/v1/ptx/isspacep.json"},{"mnemonic":"istypep","slug":"istypep","records":1,"summary":"Write predicate register p with 1 if register a points to an opaque variable of the\nspecified type, and with 0 otherwise. Destination p has type.pred;","page":"https://instructionsets.com/ptx/istypep/","api":"https://instructionsets.com/api/v1/ptx/istypep.json"},{"mnemonic":"ld","slug":"ld","records":1,"summary":"Load a value from the specified state space into a register.","page":"https://instructionsets.com/ptx/ld/","api":"https://instructionsets.com/api/v1/ptx/ld.json"},{"mnemonic":"ld.global.nc","slug":"ld_global_nc","records":1,"summary":"Load register variable d from the location specified by the source address operand a in the global state space, and optionally cache in non-coherent read-only cache.","page":"https://instructionsets.com/ptx/ld_global_nc/","api":"https://instructionsets.com/api/v1/ptx/ld_global_nc.json"},{"mnemonic":"ldmatrix","slug":"ldmatrix","records":1,"summary":"Collectively load one or more matrices across all threads in a warp from the location indicated by the address operand p, from.shared state space into destination register r.","page":"https://instructionsets.com/ptx/ldmatrix/","api":"https://instructionsets.com/api/v1/ptx/ldmatrix.json"},{"mnemonic":"ldu","slug":"ldu","records":1,"summary":"Load read-only data into register variable d from the location specified by the source address operand a in the global state space, where the address is guaranteed to be the same across all threads in the warp.","page":"https://instructionsets.com/ptx/ldu/","api":"https://instructionsets.com/api/v1/ptx/ldu.json"},{"mnemonic":"lg2","slug":"lg2","records":1,"summary":"Fast hardware approximation of log2(x).","page":"https://instructionsets.com/ptx/lg2/","api":"https://instructionsets.com/api/v1/ptx/lg2.json"},{"mnemonic":"lop3","slug":"lop3","records":1,"summary":"Compute bitwise logical operation on inputs a, b, c and store the result in destination d.","page":"https://instructionsets.com/ptx/lop3/","api":"https://instructionsets.com/api/v1/ptx/lop3.json"},{"mnemonic":"mad","slug":"mad","records":1,"summary":"Compute (a * b) + c as two rounding steps (unlike fma, which fuses them into one).","page":"https://instructionsets.com/ptx/mad/","api":"https://instructionsets.com/api/v1/ptx/mad.json"},{"mnemonic":"mad24","slug":"mad24","records":1,"summary":"Compute the product of two 24-bit integer values held in 32-bit source registers, and add a third, 32-bit value to either the high or low 32-bits of the 48-bit result.","page":"https://instructionsets.com/ptx/mad24/","api":"https://instructionsets.com/api/v1/ptx/mad24.json"},{"mnemonic":"mad.cc","slug":"mad_cc","records":1,"summary":"Multiplies two values, extracts either the high or low part of the result, and adds a third value.","page":"https://instructionsets.com/ptx/mad_cc/","api":"https://instructionsets.com/api/v1/ptx/mad_cc.json"},{"mnemonic":"madc","slug":"madc","records":1,"summary":"Multiplies two values, extracts either the high or low part of the result, and adds a third value along with carry-in.","page":"https://instructionsets.com/ptx/madc/","api":"https://instructionsets.com/api/v1/ptx/madc.json"},{"mnemonic":"mapa","slug":"mapa","records":1,"summary":"Get address in the CTA specified by operand b which corresponds to the address specified by operand a.","page":"https://instructionsets.com/ptx/mapa/","api":"https://instructionsets.com/api/v1/ptx/mapa.json"},{"mnemonic":"match.sync","slug":"match_sync","records":1,"summary":"match.sync will cause executing thread to wait until all non-exited threads from membermask have executed match.sync with the same qualifiers and same membermask value before resuming execution.","page":"https://instructionsets.com/ptx/match_sync/","api":"https://instructionsets.com/api/v1/ptx/match_sync.json"},{"mnemonic":"max","slug":"max","records":1,"summary":"Select the larger of two operands.","page":"https://instructionsets.com/ptx/max/","api":"https://instructionsets.com/api/v1/ptx/max.json"},{"mnemonic":"mbarrier.arrive","slug":"mbarrier_arrive","records":1,"summary":"A thread executing mbarrier.arrive performs an arrive-on operation\non the mbarrier object at the location specified by the address operand addr. The 3","page":"https://instructionsets.com/ptx/mbarrier_arrive/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_arrive.json"},{"mnemonic":"mbarrier.arrive_drop","slug":"mbarrier_arrive_drop","records":1,"summary":"A thread executing mbarrier.arrive_drop on the mbarrier object at the location specified by the address operand addr performs the following steps: Decrements the expected arrival count of the mbarrier object by the value specified by the 32-bit integer operand count.","page":"https://instructionsets.com/ptx/mbarrier_arrive_drop/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_arrive_drop.json"},{"mnemonic":"mbarrier.check_layout","slug":"mbarrier_check_layout","records":1,"summary":"The layout of the opaque mbarrier object can be queried using mbarrier.check_layout.","page":"https://instructionsets.com/ptx/mbarrier_check_layout/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_check_layout.json"},{"mnemonic":"mbarrier.complete_tx","slug":"mbarrier_complete_tx","records":1,"summary":"A thread executing mbarrier.complete_tx performs a complete-tx operation on the mbarrier object at the location specified by the address operand addr.","page":"https://instructionsets.com/ptx/mbarrier_complete_tx/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_complete_tx.json"},{"mnemonic":"mbarrier.expect_tx","slug":"mbarrier_expect_tx","records":1,"summary":"A thread executing mbarrier.expect_tx performs an expect-tx operation on the mbarrier object at the location specified by the address operand addr.","page":"https://instructionsets.com/ptx/mbarrier_expect_tx/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_expect_tx.json"},{"mnemonic":"mbarrier.init","slug":"mbarrier_init","records":1,"summary":"mbarrier.init initializes the mbarrier object at the location specified by the address operand addr with the unsigned 32-bit integer count.","page":"https://instructionsets.com/ptx/mbarrier_init/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_init.json"},{"mnemonic":"mbarrier.inval","slug":"mbarrier_inval","records":1,"summary":"mbarrier.inval invalidates the mbarrier object at the location specified by the address operand addr.","page":"https://instructionsets.com/ptx/mbarrier_inval/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_inval.json"},{"mnemonic":"mbarrier.pending_count","slug":"mbarrier_pending_count","records":1,"summary":"The pending count can be queried from the opaque mbarrier state using mbarrier.pending_count.","page":"https://instructionsets.com/ptx/mbarrier_pending_count/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_pending_count.json"},{"mnemonic":"mbarrier.test_wait","slug":"mbarrier_test_wait","records":1,"summary":"The test_wait and try_wait operations test for the completion of the current or the immediately preceding phase of an mbarrier object at the location specified by the operand addr.","page":"https://instructionsets.com/ptx/mbarrier_test_wait/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_test_wait.json"},{"mnemonic":"mbarrier.try_wait","slug":"mbarrier_try_wait","records":1,"summary":"The test_wait and try_wait operations test for the completion of the current or the immediately preceding phase of an mbarrier object at the location specified by the operand addr.","page":"https://instructionsets.com/ptx/mbarrier_try_wait/","api":"https://instructionsets.com/api/v1/ptx/mbarrier_try_wait.json"},{"mnemonic":"membar","slug":"membar","records":1,"summary":"Order this thread's prior memory accesses relative to later ones, visible to a given scope.","page":"https://instructionsets.com/ptx/membar/","api":"https://instructionsets.com/api/v1/ptx/membar.json"},{"mnemonic":"min","slug":"min","records":1,"summary":"Select the smaller of two operands.","page":"https://instructionsets.com/ptx/min/","api":"https://instructionsets.com/api/v1/ptx/min.json"},{"mnemonic":"mma","slug":"mma","records":1,"summary":"Cooperative, warp-wide matrix-multiply-accumulate executed on tensor-core hardware.","page":"https://instructionsets.com/ptx/mma/","api":"https://instructionsets.com/api/v1/ptx/mma.json"},{"mnemonic":"mov","slug":"mov","records":1,"summary":"Copy a value into a register, or materialize an address/immediate.","page":"https://instructionsets.com/ptx/mov/","api":"https://instructionsets.com/api/v1/ptx/mov.json"},{"mnemonic":"movmatrix","slug":"movmatrix","records":1,"summary":"Move a row-major matrix across all threads in a warp, reading elements from source a, and writing the transposed elements to destination d.","page":"https://instructionsets.com/ptx/movmatrix/","api":"https://instructionsets.com/api/v1/ptx/movmatrix.json"},{"mnemonic":"mul","slug":"mul","records":1,"summary":"Multiply two operands, selecting the low, high, or widened part of an integer product.","page":"https://instructionsets.com/ptx/mul/","api":"https://instructionsets.com/api/v1/ptx/mul.json"},{"mnemonic":"mul24","slug":"mul24","records":1,"summary":"Compute the product of two 24-bit integer values held in 32-bit source registers, and return either\nthe high or low 32-bits of the 48-bit result.","page":"https://instructionsets.com/ptx/mul24/","api":"https://instructionsets.com/api/v1/ptx/mul24.json"},{"mnemonic":"multimem.cp.async.bulk","slug":"multimem_cp_async_bulk","records":1,"summary":"Instruction multimem.cp.async.bulk initiates an asynchronous bulk-copy operation from source address range [srcMem, srcMem + size) to memory locations residing on each GPU’s memory referred to by the destination multimem address range [dstMem, dstMem + size).","page":"https://instructionsets.com/ptx/multimem_cp_async_bulk/","api":"https://instructionsets.com/api/v1/ptx/multimem_cp_async_bulk.json"},{"mnemonic":"multimem.cp.reduce.async.bulk","slug":"multimem_cp_reduce_async_bulk","records":1,"summary":"Instruction multimem.cp.reduce.async.bulk initiates an element-wise asynchronous reduction operation with elements from source memory address range [srcMem, srcMem + size) to memory locations residing on each GPU’s memory referred to by the multimem destination address range [dstMem, dstMem + size).","page":"https://instructionsets.com/ptx/multimem_cp_reduce_async_bulk/","api":"https://instructionsets.com/api/v1/ptx/multimem_cp_reduce_async_bulk.json"},{"mnemonic":"multimem.ld_reduce","slug":"multimem_ld_reduce","records":1,"summary":"The multimem.* operations operate on multimem addresses and accesses all of the multiple memory\nlocations which the multimem address points to.\nMultimem addresses can be accessed only by multimem.* operations. Accessing a multimem address\nwith ld, st or any other memory operations results in undefined behavior.\nRefer to CUDA programming guide for creation and management of the multimem addresses.","page":"https://instructionsets.com/ptx/multimem_ld_reduce/","api":"https://instructionsets.com/api/v1/ptx/multimem_ld_reduce.json"},{"mnemonic":"multimem.red","slug":"multimem_red","records":1,"summary":"The multimem.* operations operate on multimem addresses and accesses all of the multiple memory\nlocations which the multimem address points to.\nMultimem addresses can be accessed only by multimem.* operations. Accessing a multimem address\nwith ld, st or any other memory operations results in undefined behavior.\nRefer to CUDA programming guide for creation and management of the multimem addresses.","page":"https://instructionsets.com/ptx/multimem_red/","api":"https://instructionsets.com/api/v1/ptx/multimem_red.json"},{"mnemonic":"multimem.red.async","slug":"multimem_red_async","records":1,"summary":"multimem.red.async is a non-blocking instruction which initiates an asynchronous reduction operation specified by.op, with operand b and the value at memory locations residing on each GPU’s memory referred to by the destination multimem address operand a.","page":"https://instructionsets.com/ptx/multimem_red_async/","api":"https://instructionsets.com/api/v1/ptx/multimem_red_async.json"},{"mnemonic":"multimem.st","slug":"multimem_st","records":1,"summary":"The multimem.* operations operate on multimem addresses and accesses all of the multiple memory\nlocations which the multimem address points to.\nMultimem addresses can be accessed only by multimem.* operations. Accessing a multimem address\nwith ld, st or any other memory operations results in undefined behavior.\nRefer to CUDA programming guide for creation and management of the multimem addresses.","page":"https://instructionsets.com/ptx/multimem_st/","api":"https://instructionsets.com/api/v1/ptx/multimem_st.json"},{"mnemonic":"multimem.st.async","slug":"multimem_st_async","records":1,"summary":"multimem.st.async is a non-blocking instruction which initiates an asynchronous store operation that stores the value specified by source operand b to the memory locations residing on each GPU’s memory referred to by the destination multimem address operand a.","page":"https://instructionsets.com/ptx/multimem_st_async/","api":"https://instructionsets.com/api/v1/ptx/multimem_st_async.json"},{"mnemonic":"nanosleep","slug":"nanosleep","records":1,"summary":"Suspends the thread for a sleep duration approximately close to the delay t, specified in nanoseconds.","page":"https://instructionsets.com/ptx/nanosleep/","api":"https://instructionsets.com/api/v1/ptx/nanosleep.json"},{"mnemonic":"neg","slug":"neg","records":1,"summary":"Negate a signed or floating-point operand.","page":"https://instructionsets.com/ptx/neg/","api":"https://instructionsets.com/api/v1/ptx/neg.json"},{"mnemonic":"not","slug":"not","records":1,"summary":"Bitwise complement of an operand.","page":"https://instructionsets.com/ptx/not/","api":"https://instructionsets.com/api/v1/ptx/not.json"},{"mnemonic":"or","slug":"or","records":1,"summary":"Bitwise OR of two operands.","page":"https://instructionsets.com/ptx/or/","api":"https://instructionsets.com/api/v1/ptx/or.json"},{"mnemonic":"pmevent","slug":"pmevent","records":1,"summary":"Triggers one or more of a fixed number of performance monitor events, with event index or mask specified by immediate operand a.","page":"https://instructionsets.com/ptx/pmevent/","api":"https://instructionsets.com/api/v1/ptx/pmevent.json"},{"mnemonic":"popc","slug":"popc","records":1,"summary":"Count the number of set bits in an integer operand.","page":"https://instructionsets.com/ptx/popc/","api":"https://instructionsets.com/api/v1/ptx/popc.json"},{"mnemonic":"prefetch","slug":"prefetch","records":1,"summary":"The prefetch instruction brings the cache line containing the specified address in global or\nlocal memory state space into the specified cache level.","page":"https://instructionsets.com/ptx/prefetch/","api":"https://instructionsets.com/api/v1/ptx/prefetch.json"},{"mnemonic":"prefetchu","slug":"prefetchu","records":1,"summary":"The prefetchu instruction brings the cache line containing the specified generic address into the specified uniform cache level. A prefetch to a shared memory location performs no operation.","page":"https://instructionsets.com/ptx/prefetchu/","api":"https://instructionsets.com/api/v1/ptx/prefetchu.json"},{"mnemonic":"prmt","slug":"prmt","records":1,"summary":"Pick four arbitrary bytes from two 32-bit registers, and reassemble them into a 32-bit destination register.","page":"https://instructionsets.com/ptx/prmt/","api":"https://instructionsets.com/api/v1/ptx/prmt.json"},{"mnemonic":"rcp","slug":"rcp","records":1,"summary":"Compute 1/a, store result in d.","page":"https://instructionsets.com/ptx/rcp/","api":"https://instructionsets.com/api/v1/ptx/rcp.json"},{"mnemonic":"rcp.approx.ftz.f64","slug":"rcp_approx_ftz_f64","records":1,"summary":"Compute a fast, gross approximation to the reciprocal as follows: extract the most-significant 32 bits of.f64 operand a in 1.11.20 IEEE floating-point format (i.e., ignore the least-significant 32…","page":"https://instructionsets.com/ptx/rcp_approx_ftz_f64/","api":"https://instructionsets.com/api/v1/ptx/rcp_approx_ftz_f64.json"},{"mnemonic":"red","slug":"red","records":1,"summary":"Atomically read-modify-write a memory location without returning the prior value.","page":"https://instructionsets.com/ptx/red/","api":"https://instructionsets.com/api/v1/ptx/red.json"},{"mnemonic":"red.async","slug":"red_async","records":1,"summary":"red.async is a non-blocking instruction which initiates an asynchronous reduction operation specified by.op, with the operand b and the value at destination shared memory location specified by operand a.","page":"https://instructionsets.com/ptx/red_async/","api":"https://instructionsets.com/api/v1/ptx/red_async.json"},{"mnemonic":"redux.sync","slug":"redux_sync","records":1,"summary":"redux.sync will cause the executing thread to wait until all non-exited threads corresponding to membermask have executed redux.sync with the same qualifiers and same membermask value before resuming execution.","page":"https://instructionsets.com/ptx/redux_sync/","api":"https://instructionsets.com/api/v1/ptx/redux_sync.json"},{"mnemonic":"rem","slug":"rem","records":1,"summary":"Compute the integer remainder of division.","page":"https://instructionsets.com/ptx/rem/","api":"https://instructionsets.com/api/v1/ptx/rem.json"},{"mnemonic":"ret","slug":"ret","records":1,"summary":"Return execution to caller’s environment.","page":"https://instructionsets.com/ptx/ret/","api":"https://instructionsets.com/api/v1/ptx/ret.json"},{"mnemonic":"rsqrt","slug":"rsqrt","records":1,"summary":"Fast hardware approximation of 1/sqrt(x).","page":"https://instructionsets.com/ptx/rsqrt/","api":"https://instructionsets.com/api/v1/ptx/rsqrt.json"},{"mnemonic":"rsqrt.approx.ftz.f64","slug":"rsqrt_approx_ftz_f64","records":1,"summary":"Compute a double-precision (.f64 ) approximation of the square root reciprocal of a value. The\nleast significant 32 bits of the double-precision (.f64","page":"https://instructionsets.com/ptx/rsqrt_approx_ftz_f64/","api":"https://instructionsets.com/api/v1/ptx/rsqrt_approx_ftz_f64.json"},{"mnemonic":"sad","slug":"sad","records":1,"summary":"Adds the absolute value of a-b to c and writes the resulting value into d.","page":"https://instructionsets.com/ptx/sad/","api":"https://instructionsets.com/api/v1/ptx/sad.json"},{"mnemonic":"selp","slug":"selp","records":1,"summary":"Select between two operands based on a predicate, without branching.","page":"https://instructionsets.com/ptx/selp/","api":"https://instructionsets.com/api/v1/ptx/selp.json"},{"mnemonic":"set","slug":"set","records":1,"summary":"Compare two operands and write a numeric (not predicate) 0/1 or all-ones/all-zeros result.","page":"https://instructionsets.com/ptx/set/","api":"https://instructionsets.com/api/v1/ptx/set.json"},{"mnemonic":"setmaxnreg","slug":"setmaxnreg","records":1,"summary":"setmaxnreg provides a hint to the system to update the maximum number of per-thread registers owned by the executing warp to the value specified by the imm-reg-count operand.","page":"https://instructionsets.com/ptx/setmaxnreg/","api":"https://instructionsets.com/api/v1/ptx/setmaxnreg.json"},{"mnemonic":"setp","slug":"setp","records":1,"summary":"Compare two operands and write the boolean result to a predicate register.","page":"https://instructionsets.com/ptx/setp/","api":"https://instructionsets.com/api/v1/ptx/setp.json"},{"mnemonic":"shf","slug":"shf","records":1,"summary":"Shift the 64-bit value formed by concatenating operands a and b left or right by the amount specified by the unsigned 32-bit value in c.","page":"https://instructionsets.com/ptx/shf/","api":"https://instructionsets.com/api/v1/ptx/shf.json"},{"mnemonic":"shfl","slug":"shfl","records":1,"summary":"Exchange a value directly between lanes of the same warp.","page":"https://instructionsets.com/ptx/shfl/","api":"https://instructionsets.com/api/v1/ptx/shfl.json"},{"mnemonic":"shl","slug":"shl","records":1,"summary":"Shift bits left, filling with zero.","page":"https://instructionsets.com/ptx/shl/","api":"https://instructionsets.com/api/v1/ptx/shl.json"},{"mnemonic":"shr","slug":"shr","records":1,"summary":"Shift bits right, arithmetic or logical depending on the operand's signedness.","page":"https://instructionsets.com/ptx/shr/","api":"https://instructionsets.com/api/v1/ptx/shr.json"},{"mnemonic":"sin","slug":"sin","records":1,"summary":"Fast hardware approximation of sin(x).","page":"https://instructionsets.com/ptx/sin/","api":"https://instructionsets.com/api/v1/ptx/sin.json"},{"mnemonic":"slct","slug":"slct","records":1,"summary":"Conditional selection.","page":"https://instructionsets.com/ptx/slct/","api":"https://instructionsets.com/api/v1/ptx/slct.json"},{"mnemonic":"sqrt","slug":"sqrt","records":1,"summary":"Compute sqrt( a ) and store the result in d.","page":"https://instructionsets.com/ptx/sqrt/","api":"https://instructionsets.com/api/v1/ptx/sqrt.json"},{"mnemonic":"st","slug":"st","records":1,"summary":"Store a register value into the specified state space.","page":"https://instructionsets.com/ptx/st/","api":"https://instructionsets.com/api/v1/ptx/st.json"},{"mnemonic":"st.async","slug":"st_async","records":1,"summary":"st.async is a non-blocking instruction which initiates an asynchronous store operation that stores the value specified by source operand b to the destination memory location specified by operand a.","page":"https://instructionsets.com/ptx/st_async/","api":"https://instructionsets.com/api/v1/ptx/st_async.json"},{"mnemonic":"st.bulk","slug":"st_bulk","records":1,"summary":"st.bulk instruction initializes a region of shared memory starting from the location specified by destination address operand a.","page":"https://instructionsets.com/ptx/st_bulk/","api":"https://instructionsets.com/api/v1/ptx/st_bulk.json"},{"mnemonic":"stackrestore","slug":"stackrestore","records":1,"summary":"Sets the current stack pointer to source register a.","page":"https://instructionsets.com/ptx/stackrestore/","api":"https://instructionsets.com/api/v1/ptx/stackrestore.json"},{"mnemonic":"stacksave","slug":"stacksave","records":1,"summary":"Copies the current value of stack pointer into the destination register d.","page":"https://instructionsets.com/ptx/stacksave/","api":"https://instructionsets.com/api/v1/ptx/stacksave.json"},{"mnemonic":"stmatrix","slug":"stmatrix","records":1,"summary":"Collectively store one or more matrices across all threads in a warp to the location indicated by the address operand p, in.shared state space.","page":"https://instructionsets.com/ptx/stmatrix/","api":"https://instructionsets.com/api/v1/ptx/stmatrix.json"},{"mnemonic":"sub","slug":"sub","records":1,"summary":"Subtract the second operand from the first, with optional saturation for signed 32-bit integers.","page":"https://instructionsets.com/ptx/sub/","api":"https://instructionsets.com/api/v1/ptx/sub.json"},{"mnemonic":"sub.cc","slug":"sub_cc","records":1,"summary":"Performs integer subtraction and writes the borrow-out value into the condition code register.","page":"https://instructionsets.com/ptx/sub_cc/","api":"https://instructionsets.com/api/v1/ptx/sub_cc.json"},{"mnemonic":"subc","slug":"subc","records":1,"summary":"Performs integer subtraction with borrow-in and optionally writes the borrow-out value into the\ncondition code register.","page":"https://instructionsets.com/ptx/subc/","api":"https://instructionsets.com/api/v1/ptx/subc.json"},{"mnemonic":"suld","slug":"suld","records":1,"summary":"suld.b.{1d,2d,3d} Load from surface memory using a surface coordinate vector.","page":"https://instructionsets.com/ptx/suld/","api":"https://instructionsets.com/api/v1/ptx/suld.json"},{"mnemonic":"suq","slug":"suq","records":1,"summary":"Query an attribute of a surface.","page":"https://instructionsets.com/ptx/suq/","api":"https://instructionsets.com/api/v1/ptx/suq.json"},{"mnemonic":"sured","slug":"sured","records":1,"summary":"Reduction to surface memory using a surface coordinate vector.","page":"https://instructionsets.com/ptx/sured/","api":"https://instructionsets.com/api/v1/ptx/sured.json"},{"mnemonic":"sust","slug":"sust","records":1,"summary":"sust.{1d,2d,3d} Store to surface memory using a surface coordinate vector.","page":"https://instructionsets.com/ptx/sust/","api":"https://instructionsets.com/api/v1/ptx/sust.json"},{"mnemonic":"szext","slug":"szext","records":1,"summary":"Sign-extends or zero-extends an N-bit value from operand a where N is specified in operand b.","page":"https://instructionsets.com/ptx/szext/","api":"https://instructionsets.com/api/v1/ptx/szext.json"},{"mnemonic":"tanh","slug":"tanh","records":1,"summary":"Take hyperbolic tangent value of a.","page":"https://instructionsets.com/ptx/tanh/","api":"https://instructionsets.com/api/v1/ptx/tanh.json"},{"mnemonic":"tcgen05.alloc","slug":"tcgen05_alloc","records":1,"summary":"tcgen05.alloc is a blocking instruction which dynamically allocates the specified number of columns in the Tensor Memory and writes the address of the allocated Tensor Memory into shared memory at the location specified by address operand dst.","page":"https://instructionsets.com/ptx/tcgen05_alloc/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_alloc.json"},{"mnemonic":"tcgen05.commit","slug":"tcgen05_commit","records":1,"summary":"The instruction tcgen05.commit is an asynchronous instruction which makes the mbarrier object, specified by the address operand mbar, track the completion of all the prior asynchronous tcgen05 operations, as listed in mbarrier based completion mechanism, initiated by the executing thread.","page":"https://instructionsets.com/ptx/tcgen05_commit/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_commit.json"},{"mnemonic":"tcgen05.cp","slug":"tcgen05_cp","records":1,"summary":"Instruction tcgen05.cp initiates an asynchronous copy operation from shared memory to the location specified by the address operand taddr in the Tensor Memory.","page":"https://instructionsets.com/ptx/tcgen05_cp/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_cp.json"},{"mnemonic":"tcgen05.dealloc","slug":"tcgen05_dealloc","records":1,"summary":"tcgen05.dealloc is a blocking instruction which de-allocates the Tensor Memory specified by the Tensor Memory address taddr. The operand nCols specifies the number of columns to be de-allocated.","page":"https://instructionsets.com/ptx/tcgen05_dealloc/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_dealloc.json"},{"mnemonic":"tcgen05.fence","slug":"tcgen05_fence","records":1,"summary":"The instruction tcgen05.fence::before_thread_sync orders all the prior asynchronous tcgen05 operations with respect to the subsequent tcgen05 and the execution ordering operations.","page":"https://instructionsets.com/ptx/tcgen05_fence/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_fence.json"},{"mnemonic":"tcgen05.ld","slug":"tcgen05_ld","records":1,"summary":"Instruction tcgen05.ld asynchronously loads data from the Tensor Memory at the location specified by the 32-bit address operand taddr into the destination register r, collectively across all threads of the warps.","page":"https://instructionsets.com/ptx/tcgen05_ld/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_ld.json"},{"mnemonic":"tcgen05.mma","slug":"tcgen05_mma","records":1,"summary":"Instruction tcgen05.mma is an asynchronous instruction which initiates an MxNxK matrix multiply and accumulate operation, D = A*B+D where the A matrix is MxK, the B matrix is KxN, and the D matrix is MxN.","page":"https://instructionsets.com/ptx/tcgen05_mma/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_mma.json"},{"mnemonic":"tcgen05.mma.sp","slug":"tcgen05_mma_sp","records":1,"summary":"Instruction tcgen05.mma.sp is an asynchronous instruction which initiates an MxNxK matrix multiply and accumulate operation of the form D = A*B+D where the A matrix is Mx(K/2), the B matrix is KxN, and the D matrix is MxN.","page":"https://instructionsets.com/ptx/tcgen05_mma_sp/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_mma_sp.json"},{"mnemonic":"tcgen05.mma.ws","slug":"tcgen05_mma_ws","records":1,"summary":"Instruction tcgen05.mma.ws is an asynchronous instruction which initiates an MxNxK matrix multiply and accumulate operation, D = A*B+D where the A matrix is MxK, the B matrix is KxN, and the D matrix is MxN.","page":"https://instructionsets.com/ptx/tcgen05_mma_ws/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_mma_ws.json"},{"mnemonic":"tcgen05.mma.ws.sp","slug":"tcgen05_mma_ws_sp","records":1,"summary":"Instruction tcgen05.mma.ws.sp is an asynchronous instruction which initiates\nan MxNxK matrix multiply and accumulate operation, D = A*B+D where the A","page":"https://instructionsets.com/ptx/tcgen05_mma_ws_sp/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_mma_ws_sp.json"},{"mnemonic":"tcgen05.relinquish_alloc_permit","slug":"tcgen05_relinquish_alloc_permit","records":1,"summary":"tcgen05.relinquish_alloc_permit specifies that the CTA of the executing thread is relinquishing the right to allocate Tensor Memory.","page":"https://instructionsets.com/ptx/tcgen05_relinquish_alloc_permit/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_relinquish_alloc_permit.json"},{"mnemonic":"tcgen05.shift","slug":"tcgen05_shift","records":1,"summary":"Instruction tcgen05.shift is an asynchronous instruction which initiates the shifting of 32-byte elements downwards across all the rows, except the last, by one row.","page":"https://instructionsets.com/ptx/tcgen05_shift/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_shift.json"},{"mnemonic":"tcgen05.st","slug":"tcgen05_st","records":1,"summary":"Instruction tcgen05.st asynchronously stores data from the source register r into the Tensor Memory at the location specified by the 32-bit address operand taddr, collectively across all threads of the warps.","page":"https://instructionsets.com/ptx/tcgen05_st/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_st.json"},{"mnemonic":"tcgen05.wait","slug":"tcgen05_wait","records":1,"summary":"Instruction tcgen05.wait::st causes the executing thread to block until all prior tcgen05.st operations issued by the executing thread have completed.","page":"https://instructionsets.com/ptx/tcgen05_wait/","api":"https://instructionsets.com/api/v1/ptx/tcgen05_wait.json"},{"mnemonic":"tensormap.cp_fenceproxy","slug":"tensormap_cp_fenceproxy","records":1,"summary":"The tensormap.cp_fenceproxy instructions perform the following operations in order: Copies data of size specified by the size argument, in bytes, from the location specified by the address operand…","page":"https://instructionsets.com/ptx/tensormap_cp_fenceproxy/","api":"https://instructionsets.com/api/v1/ptx/tensormap_cp_fenceproxy.json"},{"mnemonic":"tensormap.replace","slug":"tensormap_replace","records":1,"summary":"The tensormap.replace instruction replaces the field, specified by.field qualifier, of the tensor-map object at the location specified by the address operand addr with a new value.","page":"https://instructionsets.com/ptx/tensormap_replace/","api":"https://instructionsets.com/api/v1/ptx/tensormap_replace.json"},{"mnemonic":"testp","slug":"testp","records":1,"summary":"testp tests common properties of floating-point numbers and returns a predicate value of 1 if True and 0 if False.","page":"https://instructionsets.com/ptx/testp/","api":"https://instructionsets.com/api/v1/ptx/testp.json"},{"mnemonic":"tex","slug":"tex","records":1,"summary":"tex.{1d,2d,3d} Texture lookup using a texture coordinate vector.","page":"https://instructionsets.com/ptx/tex/","api":"https://instructionsets.com/api/v1/ptx/tex.json"},{"mnemonic":"tld4","slug":"tld4","records":1,"summary":"Texture fetch of the 4-texel bilerp footprint using a texture coordinate vector.","page":"https://instructionsets.com/ptx/tld4/","api":"https://instructionsets.com/api/v1/ptx/tld4.json"},{"mnemonic":"trap","slug":"trap","records":1,"summary":"Abort execution and generate an interrupt to the host CPU.","page":"https://instructionsets.com/ptx/trap/","api":"https://instructionsets.com/api/v1/ptx/trap.json"},{"mnemonic":"txq","slug":"txq","records":1,"summary":"Query an attribute of a texture or sampler.","page":"https://instructionsets.com/ptx/txq/","api":"https://instructionsets.com/api/v1/ptx/txq.json"},{"mnemonic":"vadd","slug":"vadd","records":1,"summary":"Perform scalar arithmetic operation with optional saturate, and optional secondary arithmetic operation or subword data merge.","page":"https://instructionsets.com/ptx/vadd/","api":"https://instructionsets.com/api/v1/ptx/vadd.json"},{"mnemonic":"vadd2","slug":"vadd2","records":1,"summary":"Two-way SIMD parallel arithmetic operation with secondary operation.","page":"https://instructionsets.com/ptx/vadd2/","api":"https://instructionsets.com/api/v1/ptx/vadd2.json"},{"mnemonic":"vadd4","slug":"vadd4","records":1,"summary":"Four-way SIMD parallel arithmetic operation with secondary operation.","page":"https://instructionsets.com/ptx/vadd4/","api":"https://instructionsets.com/api/v1/ptx/vadd4.json"},{"mnemonic":"vmad","slug":"vmad","records":1,"summary":"Calculate (a*b) + c, with optional operand negates, plus one mode, and scaling.\nThe source operands support optional negation with some restrictions.","page":"https://instructionsets.com/ptx/vmad/","api":"https://instructionsets.com/api/v1/ptx/vmad.json"},{"mnemonic":"vote","slug":"vote","records":1,"summary":"Combine a per-lane predicate across the warp using any/all/ballot reduction.","page":"https://instructionsets.com/ptx/vote/","api":"https://instructionsets.com/api/v1/ptx/vote.json"},{"mnemonic":"vset","slug":"vset","records":1,"summary":"Compare input values using specified comparison, with optional secondary arithmetic operation or subword data merge.","page":"https://instructionsets.com/ptx/vset/","api":"https://instructionsets.com/api/v1/ptx/vset.json"},{"mnemonic":"vset2","slug":"vset2","records":1,"summary":"Two-way SIMD parallel comparison with secondary operation.","page":"https://instructionsets.com/ptx/vset2/","api":"https://instructionsets.com/api/v1/ptx/vset2.json"},{"mnemonic":"vset4","slug":"vset4","records":1,"summary":"Four-way SIMD parallel comparison with secondary operation.","page":"https://instructionsets.com/ptx/vset4/","api":"https://instructionsets.com/api/v1/ptx/vset4.json"},{"mnemonic":"vshl","slug":"vshl","records":1,"summary":"vshl Shift a left by unsigned amount in b with optional saturate, and optional secondary arithmetic operation or subword data merge.","page":"https://instructionsets.com/ptx/vshl/","api":"https://instructionsets.com/api/v1/ptx/vshl.json"},{"mnemonic":"vshr","slug":"vshr","records":1,"summary":"vshl Shift a left by unsigned amount in b with optional saturate, and optional secondary arithmetic operation or subword data merge.","page":"https://instructionsets.com/ptx/vshr/","api":"https://instructionsets.com/api/v1/ptx/vshr.json"},{"mnemonic":"vsub","slug":"vsub","records":1,"summary":"Perform scalar arithmetic operation with optional saturate, and optional secondary arithmetic operation or subword data merge.","page":"https://instructionsets.com/ptx/vsub/","api":"https://instructionsets.com/api/v1/ptx/vsub.json"},{"mnemonic":"vsub2","slug":"vsub2","records":1,"summary":"Two-way SIMD parallel arithmetic operation with secondary operation.","page":"https://instructionsets.com/ptx/vsub2/","api":"https://instructionsets.com/api/v1/ptx/vsub2.json"},{"mnemonic":"vsub4","slug":"vsub4","records":1,"summary":"Four-way SIMD parallel arithmetic operation with secondary operation.","page":"https://instructionsets.com/ptx/vsub4/","api":"https://instructionsets.com/api/v1/ptx/vsub4.json"},{"mnemonic":"wgmma.commit_group","slug":"wgmma_commit_group","records":1,"summary":"wgmma.commit_group instruction creates a new wgmma-group per warpgroup and batches all prior wgmma.mma_async instructions initiated by the executing warp but not committed to any wgmma-group into the new wgmma-group.","page":"https://instructionsets.com/ptx/wgmma_commit_group/","api":"https://instructionsets.com/api/v1/ptx/wgmma_commit_group.json"},{"mnemonic":"wgmma.fence","slug":"wgmma_fence","records":1,"summary":"wgmma.fence instruction establishes an ordering between prior accesses to any warpgroup registers and subsequent accesses to the same registers by a wgmma.mma_async instruction.","page":"https://instructionsets.com/ptx/wgmma_fence/","api":"https://instructionsets.com/api/v1/ptx/wgmma_fence.json"},{"mnemonic":"wgmma.mma_async","slug":"wgmma_mma_async","records":1,"summary":"Instruction wgmma.mma_async issues a MxNxK matrix multiply and accumulate operation, D = A*B+D, where the A matrix is MxK, the B matrix is KxN, and the D matrix is MxN.","page":"https://instructionsets.com/ptx/wgmma_mma_async/","api":"https://instructionsets.com/api/v1/ptx/wgmma_mma_async.json"},{"mnemonic":"wgmma.mma_async.sp","slug":"wgmma_mma_async_sp","records":1,"summary":"Instruction wgmma.mma_async issues a MxNxK matrix multiply and accumulate operation, D = A*B+D, where the A matrix is MxK, the B matrix is KxN, and the D matrix is MxN.","page":"https://instructionsets.com/ptx/wgmma_mma_async_sp/","api":"https://instructionsets.com/api/v1/ptx/wgmma_mma_async_sp.json"},{"mnemonic":"wgmma.wait_group","slug":"wgmma_wait_group","records":1,"summary":"wgmma.wait_group instruction will cause the executing thread to wait until only N or fewer of the most recent wgmma-groups are pending and all the prior wgmma-groups committed by the executing threads are complete.","page":"https://instructionsets.com/ptx/wgmma_wait_group/","api":"https://instructionsets.com/api/v1/ptx/wgmma_wait_group.json"},{"mnemonic":"wmma","slug":"wmma","records":1,"summary":"Higher-level warp matrix-multiply-accumulate built from explicit load/mma/store steps.","page":"https://instructionsets.com/ptx/wmma/","api":"https://instructionsets.com/api/v1/ptx/wmma.json"},{"mnemonic":"xor","slug":"xor","records":1,"summary":"Bitwise XOR of two operands.","page":"https://instructionsets.com/ptx/xor/","api":"https://instructionsets.com/api/v1/ptx/xor.json"}]}
