Non-ranking diagnostics
Semantic pass rate on the original_binary / native oracle remains the only ranking axis (core C PE headline). The tables below are form quality, readability proxies, and language · ISA · format · opt pivots for multi-corpus investigation only.
Policy
Bare-compile uses minimal headers + gcc -c. Readability reports source similarity, AST tree-edit similarity, proxy score, generic naming, goto / nest / temp / flag density. For Fission, semantic rows use NIR; readability proxies prefer HIR when dual layers are present. ELF uses host/qemu recompile ABI (not wine). Study pack: benchmark/readability/.
Re-evaluation provenance
Source-CFG contract: preprocessed-tu-v1. Compiler-matched preprocessed TU rows: 432; authored-source fallback: 0; missing provenance: 0. Official releases publish binaries, decompiler rows, expected cells, and serialized source CFGs in the eval kit. Open latest eval-kit index →
EXT · Recompilation bytematch
Decompiled C is rebuilt for the original PE/ELF ABI with the matching compiler family and optimization level, then compared as ordered, relocation-normalized assembly. Missing peer measurements stay in the shared denominator. This is diagnostic evidence, not a ranking axis.
| Decompiler | Observed / shared | Compilable | Observed mean | Exact / shared |
|---|
| fission | 216 / 216 | 196 | 22.4% | 6 / 216 (2.8%) |
| ghidra | 216 / 216 | 169 | 29.8% | 14 / 216 (6.5%) |
EXT · Bare-compile rate
| Decompiler | Attempted | OK | Fail | OK rate |
|---|
| fission | 213 | 199 | 14 | 93.4% |
| ghidra | 215 | 206 | 9 | 95.8% |
EXT · Type correctness (DWARF)
Recovered variable types checked against the DWARF debug info baked into each corpus binary. Mean accuracy describes observed rows; perfect rate uses the shared subject denominator, so a peer-measurable missing row is not hidden. Diagnostic evidence only, not ranking.
| Decompiler | Observed / shared | Mean accuracy | Perfect |
|---|
| fission | 213 / 216 | 68.3% | 75 / 216 (34.7%) |
| ghidra | 215 / 216 | 79.8% | 121 / 216 (56.0%) |
EXT · Structural correctness (GED)
Graph edit distance between the decompiled function's control-flow graph and the exact compiler-variant preprocessed TU's CFG (both parsed with Joern for structural comparability). Lower is better — 0.0 means the control-flow shape matches exactly. Mean GED describes observed rows; exact-match rate uses the shared subject denominator. Diagnostic evidence only, not a ranking axis.
| Decompiler | Observed / shared | Mean GED | Exact match (GED=0) |
|---|
| fission | 213 / 216 | 11.61 | 26 / 216 (12.0%) |
| ghidra | 215 / 216 | 7.92 | 59 / 216 (27.3%) |
EXT · Readability · source sim · AST
Policy · diagnostics only
Source similarity, AST tree-edit proxies, and readability proxies (goto / temps / generic names / flag soup) are not ranking axes. Semantic pass rate on the original-binary oracle remains the only tool ranking signal. Table order is fixed (Fission, Ghidra, then alphabetical) — not sorted by proxy score. Human study materials: benchmark/readability/ (Phase 3 before any composite).
| Decompiler | Rows | Src sim | AST sim | Proxy | GNR↑ | Goto | Nest | Temp/LOC | Flag/LOC |
|---|
| fission | 426 | 0.164 | 0.579 | 0.534 | 0.209 | 0.09 | 3.38 | 0.355 | 0.008 |
| ghidra | 430 | 0.384 | 0.701 | 0.595 | 0.411 | 0.07 | 3.00 | 0.257 | 0.000 |
EXT · Language · ISA · format · opt
By language
| Language | Rows | Tested | Mean pass | Perfect | Timeouts |
|---|
c | 432 | 432 | 68.4% | 262 | 4 |
By track
| Track | Rows | Tested | Mean pass | Perfect | Timeouts |
|---|
| dev | 432 | 432 | 68.4% | 262 | 4 |
By ISA
| ISA | Rows | Mean pass | Timeouts |
|---|
x86_32 | 144 | 80.2% | 1 |
x86_64 | 288 | 62.5% | 3 |
By format
| Format | Rows | Mean pass |
|---|
pe | 432 | 68.4% |