What a zero is worth: checking a GPU decoder

· 5 min read

The AMD GPU decoder was checked against 162,054 lines LLVM assembled and named none of them wrongly. How far that zero reaches, why most lines get no answer, and the first version of the check, which would have reported zero while checking nothing.

The decoder takes the bytes of an AMD GPU instruction and names the instruction, with every field broken out. A decoder that names the wrong instruction is worse than one that says nothing, because nobody double-checks a tool. So before it shipped it was checked against a second source, and the check reported no wrong answers. This note is about what that zero is worth, because it very nearly meant nothing.

Where the layouts come from

AMD publishes a machine-readable specification of its GPU instruction sets, one file per generation, from CDNA 1 to CDNA 5 and RDNA 1 to RDNA 4. Each instruction's bit layout comes from there, and a layout is kept only if the instruction's opcode, placed in its field beside the encoding's fixed bits, rebuilds one of the encoding identifiers AMD lists. That makes a layout consistent with itself. It does not make it right.

A second source

LLVM's assembler has a test suite, and many of its files print the bytes the assembler produced beside each instruction:

v_add_f32 v5, v1, v2
// GFX11: v_add_f32_e32 v5, v1, v2   ; encoding: [0x01,0x05,0x0a,0x06]

Those bytes come from LLVM's own backend, not from AMD's files. The check takes every such line, decodes it with the same JavaScript the decoder page runs, within the GPU generation the test file targets, and sorts the answer three ways: the instruction LLVM assembled, a different one, or no answer. Where the decoder shows a register, such as v5, that is compared with the registers the assembly names.

Lines decoded to the instruction LLVM assembled25,923
Lines decoded to a different instruction0
Register readings contradicting the assembly0 of 23,105
Lines declined136,131

A zero from nothing

The first version of the check recognised a test file's GPU only when the target was written -mcpu=gfx1100. LLVM writes it as -triple=amdgpu11.00 in 779 of its 783 test files, and the one other file that prints encodings targets a GPU none of AMD's files here describe. That version would have checked nothing at all, and it would have reported, truthfully, that nothing decoded wrongly.

A zero with no denominator is not a result. So the check now says how much it reached, and it sorts every line it could not decode by cause, since an unexplained "no answer" can hide exactly the fault the check exists to find.

How far the zero reaches

A layout here is one instruction in one generation. The lines LLVM's tests contain reach 1,320 of the 2,965 layouts in the nine generations they cover, and very unevenly:

GenerationLayouts confirmed
RDNA 3368 of 368
RDNA 190 of 90
RDNA 4624 of 834
RDNA 3.554 of 422
CDNA 478 of 384
CDNA 263 of 262
CDNA 339 of 314
RDNA 22 of 90
CDNA 12 of 201
CDNA 5not checked, 1,059 layouts

So the same zero is strong evidence for RDNA 3, where every layout was decoded correctly at least once, and very weak evidence for CDNA 1, where two were. CDNA 5 is not checked at all, because the check maps no LLVM target to it.

Declining is an answer

Most lines, 136,131 of them, got no answer. That is allowed: the decoder only carries layouts it could verify, and says nothing rather than guess. But "declined" is only reassuring if you know why, so the check counts the causes:

  • 40.9% are longer forms of an instruction whose layout for that generation is carried: VOP3, DPP, SDWA, or a trailing literal constant.
  • 39.6% are instructions whose layout is carried only for other generations.
  • 15.3% have no record here under the name LLVM uses.
  • 4.2% have a record but no verified layout in any generation.
  • None have a layout for their own generation, of the right width, that fails to match.

The last line is the one that would mean a wrong layout, and the check now fails if it is ever not zero. The third is a finding about the data rather than the decoder: nearly all of those lines name an instruction that AMD's specification documents under exactly that name and this dataset has no record of, from RDNA 3's dual-issue v_dual_* forms to v_add_nc_u32. That gap is in the data, and it is now measured.

Register numbers take rules, not arithmetic

The obvious reading of a register field is that it holds the register number, and for most fields it does. The decoder shows a register only where its reading was checked this way, and leaves the raw value where the value alone does not say:

  • a 16-bit operand on RDNA 3 and later, where the field selects a half of a register, so v1.h is written 129 and "v129" would be wrong;
  • a vector field in a word that sets an ACC bit, which makes it an accumulation register;
  • scalar register numbers above 101, which are generation-specific names such as flat_scratch rather than general registers;
  • scalar descriptor fields, which hold the register number divided by four.

And a buffer instruction's address register is unused, written off, when neither of its offset and index bits is set. With those rules, none of the 23,105 readings contradicts the assembly.

What this does not say

Agreement between two sources is not proof: a mistake both made would pass unseen. The check reaches only what LLVM's tests contain, which is why coverage is so uneven and CDNA 5 is unchecked. And it says nothing about the descriptions on each instruction's page, only about the bits.

Check it yourself

Paste [0x01,0x05,0x0a,0x06] into the decoder and choose RDNA 3 to see v_add_f32 taken apart. The method, the figures and the script that produces them are in ACCURACY.md, which ships with the public-domain dataset, and the first note, Measuring the reference instead of the data, covers the rest of the corpus. If the decoder names something wrongly, tell us.