What a zero is worth: checking a GPU decoder
The AMD GPU decoder was checked against 162,054 lines LLVM assembled and named none of them wrongly. How far that zero reaches, why most lines get no answer, and the first version of the check, which would have reported zero while checking nothing.
The decoder takes the bytes of an AMD GPU instruction and names the instruction, with every field broken out. A decoder that names the wrong instruction is worse than one that says nothing, because nobody double-checks a tool. So before it shipped it was checked against a second source, and the check reported no wrong answers. This note is about what that zero is worth, because it very nearly meant nothing.
Where the layouts come from
AMD publishes a machine-readable specification of its GPU instruction sets, one file per generation, from CDNA 1 to CDNA 5 and RDNA 1 to RDNA 4. Each instruction's bit layout comes from there, and a layout is kept only if the instruction's opcode, placed in its field beside the encoding's fixed bits, rebuilds one of the encoding identifiers AMD lists. That makes a layout consistent with itself. It does not make it right.
A second source
LLVM's assembler has a test suite, and many of its files print the bytes the assembler produced beside each instruction:
v_add_f32 v5, v1, v2
// GFX11: v_add_f32_e32 v5, v1, v2 ; encoding: [0x01,0x05,0x0a,0x06]
Those bytes come from LLVM's own backend, not from AMD's files. The check takes
every such line, decodes it with the same JavaScript the decoder page runs,
within the GPU generation the test file targets, and sorts the answer three
ways: the instruction LLVM assembled, a different one, or no answer. Where the
decoder shows a register, such as v5, that is compared with the
registers the assembly names.
| Lines decoded to the instruction LLVM assembled | 25,923 |
| Lines decoded to a different instruction | 0 |
| Register readings contradicting the assembly | 0 of 23,105 |
| Lines declined | 136,131 |
A zero from nothing
The first version of the check recognised a test file's GPU only when the target
was written -mcpu=gfx1100. LLVM writes it as
-triple=amdgpu11.00 in 779 of its 783 test files, and the one other
file that prints encodings targets a GPU none of AMD's files here describe. That
version would have checked nothing at all, and it would have reported, truthfully,
that nothing decoded wrongly.
A zero with no denominator is not a result. So the check now says how much it reached, and it sorts every line it could not decode by cause, since an unexplained "no answer" can hide exactly the fault the check exists to find.
How far the zero reaches
A layout here is one instruction in one generation. The lines LLVM's tests contain reach 1,320 of the 2,965 layouts in the nine generations they cover, and very unevenly:
| Generation | Layouts confirmed |
|---|---|
| RDNA 3 | 368 of 368 |
| RDNA 1 | 90 of 90 |
| RDNA 4 | 624 of 834 |
| RDNA 3.5 | 54 of 422 |
| CDNA 4 | 78 of 384 |
| CDNA 2 | 63 of 262 |
| CDNA 3 | 39 of 314 |
| RDNA 2 | 2 of 90 |
| CDNA 1 | 2 of 201 |
| CDNA 5 | not checked, 1,059 layouts |
So the same zero is strong evidence for RDNA 3, where every layout was decoded correctly at least once, and very weak evidence for CDNA 1, where two were. CDNA 5 is not checked at all, because the check maps no LLVM target to it.
Declining is an answer
Most lines, 136,131 of them, got no answer. That is allowed: the decoder only carries layouts it could verify, and says nothing rather than guess. But "declined" is only reassuring if you know why, so the check counts the causes:
- 40.9% are longer forms of an instruction whose layout for that generation is carried: VOP3, DPP, SDWA, or a trailing literal constant.
- 39.6% are instructions whose layout is carried only for other generations.
- 15.3% have no record here under the name LLVM uses.
- 4.2% have a record but no verified layout in any generation.
- None have a layout for their own generation, of the right width, that fails to match.
The last line is the one that would mean a wrong layout, and the check now
fails if it is ever not zero. The third is a finding about the data rather than
the decoder: nearly all of those lines name an instruction that AMD's
specification documents under exactly that name and this dataset has no record
of, from RDNA 3's dual-issue v_dual_* forms to
v_add_nc_u32. That gap is in the data, and it is now measured.
Register numbers take rules, not arithmetic
The obvious reading of a register field is that it holds the register number, and for most fields it does. The decoder shows a register only where its reading was checked this way, and leaves the raw value where the value alone does not say:
- a 16-bit operand on RDNA 3 and later, where the field selects a half of a register, so
v1.his written 129 and "v129" would be wrong; - a vector field in a word that sets an ACC bit, which makes it an accumulation register;
- scalar register numbers above 101, which are generation-specific names such as
flat_scratchrather than general registers; - scalar descriptor fields, which hold the register number divided by four.
And a buffer instruction's address register is unused, written off,
when neither of its offset and index bits is set. With those rules, none of the
23,105 readings contradicts the assembly.
What this does not say
Agreement between two sources is not proof: a mistake both made would pass unseen. The check reaches only what LLVM's tests contain, which is why coverage is so uneven and CDNA 5 is unchecked. And it says nothing about the descriptions on each instruction's page, only about the bits.
Check it yourself
Paste [0x01,0x05,0x0a,0x06] into the
decoder and choose RDNA 3 to see
v_add_f32 taken apart.
The method, the figures and the script that produces them are in
ACCURACY.md, which ships with the
public-domain dataset, and the first note,
Measuring the reference instead of the data,
covers the rest of the corpus. If the decoder names something wrongly,
tell us.