Measuring the reference instead of the data

· 5 min read

How we measured the accuracy of an instruction reference built by a language model, the 18 wrong RISC-V encodings it found, and the four times the audit itself was the problem.

The instruction records on this site were extracted from vendor specifications by a language model and then repaired over many passes. A validator checks that the result hangs together: no duplicates, syntax that agrees with its own mnemonic, cross-references that resolve. None of that is evidence the records are true. For a reference, correctness is the product, so we measured it, with a script anyone can rerun and a method stated in advance.

The method

A seeded random sample of 300 records per architecture (seed 20260901) is checked against the text of the specification each record was extracted from. Four questions, each aimed at a different kind of error:

  • Mnemonic: does this instruction appear in the specification at all? This catches invented instructions.
  • Quoted opcode: does the opcode appear as the manual prints it? This catches wrong x86 encodings.
  • Derived opcode: does the stored instruction word match the record's own bit pattern? This catches wrong computed encodings.
  • Format: is the encoding form a real form in that specification?

Proportions come with Wilson 95% confidence intervals, which stay honest at these sample sizes where the usual normal approximation does not.

The results

ArchitectureMnemonicOpcodeFormat
x8697.9%97.2%98.6%
ARM96.3%not yet checkednot yet checked
RISC-V96.3%100%89.7%
PowerISA94.0%not checkable98.7%

Across those four architectures, 94% to 98% of sampled records are corroborated by their own specification on every claim that can be tested mechanically, and the failures are concentrated rather than spread thinly everywhere. RISC-V's format score first came out at 79.3% for a dull reason: this corpus labels some forms R-type (atomic), the specification's form plus a gloss of our own, which is house style rather than error. Matching the base form gives the 89.7% above.

PTX and AMDGPU score higher, 99.5% and 100%, and those numbers mean less. Their records were imported programmatically from the very sources they are checked against, so the check measures fidelity to source (did the importer drop or invent anything?) rather than independent corroboration. A 100% there is worth less than PowerISA's 94%.

Four times, the audit was the problem

Four of the alarming results this audit produced were flaws in the audit, not the data, and three of them were the same mistake: checking a record against a document that could not contain it.

  • Opcodes no manual prints. Searching the manuals for PowerISA and RISC-V opcodes scored 2.0% and 0.5%. Those architectures store a fully assembled instruction word such as 0x04000057, which no manual ever prints, while x86 stores the opcode the way Intel writes it, FF /0. Computed words are now checked against the record's own bit pattern instead.
  • AMD instructions in Intel's manual. Nine of the first ten apparent x86 failures were XOP, 3DNow! and TBM, which are AMD extensions Intel correctly never documented.
  • AArch32 in an A64 specification. Checking ARM against the A64 profile alone scored 63.7%, and the misses were overwhelmingly AArch32 and T32 instructions. With the AArch32 profile included, it is 96.3%.
  • A heading written in the singular. PTX first scored 96.8% because a few instructions sit under headings such as "Warp-level matrix load instruction: ldmatrix", singular and lowercase, which a case-sensitive match for "Instructions:" missed.

This is the part worth saying plainly. The first-order risk in work like this is measuring the reference instead of the data, and each of these initially looked like a catastrophic data problem. Every correction is written down in the accuracy report, so the numbers can be read knowing how they were reached.

What it found that was real

The derived-opcode check found 18 RISC-V records whose stored opcode contradicted their own bit pattern. They failed in both directions:

  • NOT stored 0x00004013, which encodes xori rd, rs, 0, a register move. not rd, rs is xori rd, rs, -1, so the right word is 0xFFF04013. The opcode was wrong and the bit pattern right.
  • The VAMO* vector atomics carried bit patterns with a spurious leading field, so there the opcode was right and the pattern wrong.
  • Four .vv forms of VMSGT and VMSGE do not exist in the specification at all. It provides the .vx and .vi forms and notes that the assembler builds the others by swapping the operands of vmslt and vmsle.

All 18 were checked by hand against the specification and repaired by a script that refuses to write unless every record it touches agrees with itself afterwards. That rule caught a bug in the repair itself: its first version gave the 32-bit forms the 64-bit width code. The check now reports 406 of 406 RISC-V records in agreement.

Checking against a second source

AMDGPU bit layouts now come from AMD's machine-readable ISA specification, and each of the 1,535 is kept only if the instruction's opcode rebuilds one of AMD's own encoding identifiers. That is still one source checking itself, so the layouts were also read against the bytes LLVM's assembler produces in its own test suite, which come from a different codebase entirely.

That found a real fault. AMD lists a comparison's vcc result as an operand without an encoding field, the import had dropped every such operand, and so comparison pages such as v_cmp_lt_f32 listed their operands without the destination the assembly writes first. LLVM's tests settled it: 11,290 assembled comparison lines write vcc first. 123 records were corrected.

What this does not say

Corroboration is not correctness. Finding a mnemonic in the specification proves the instruction is real, not that its description or pseudocode is right. Judging those takes a person reading both documents, and this audit does not attempt it. Treat every figure here as a lower bound on fidelity, and ARM's encodings are the obvious next thing to measure.

Check it yourself

The method, the numbers and every correction are in ACCURACY.md, which ships with the public-domain dataset. If you find a record that is wrong, tell us. Reports like that are how this gets better.