Adaptation runs

One row per run of the benchadapt closed loop, newest first; each row links to the full export (contract, every agent call, the diff, verification, numbers). Outcome, attempts, wall clock and blocks come from the run record; the description and the equivalence column are maintained by hand.

datedesignmodeoutcomeattemptsmodelswallblocksregimeequivalencewhat it was for
2026-09-14e388bfraygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude5m 43s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vThird without the inherited prompt, the same nine sites and nine blocks again; its editor simulated its own edit against a model of the block it wrote in scratch before check.sh ever ran.
2026-09-149be831raygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude7m 36s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vSecond without the inherited prompt, nine sites and nine blocks; its editor wrote a testbench and a Python driver into scratch to check the edit itself, and its plan reviewer read the top-level consumers before approving.
2026-09-14ef78a2raygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude5m 59s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vFirst of three runs with the inherited assistant prompt removed from every agent: the same nine sites and nine blocks as the three runs before it, so the framing that prompt carried was not what chose them. Its edit kept 58 more flip-flops than the other two and came out 1.3 ns slower.
2026-09-13642729raygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude5m 18s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vThird of the three, establishing that the nine-region choice repeats at this prompt set; its planner reached for ripgrep and fell back, and its reviewer probed with three successive find patterns.
2026-09-13d2386araygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude6m 33s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vSecond of the three, and the first with every agent call inside a bubblewrap sandbox; its acceptance reviewer searched the whole filesystem with find and a recursive grep from /, and both returned only its own working copy.
2026-09-13ed0a74raygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude9m 41s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vFirst of the three runs at the corrected prompt set, and the last before the sandbox existed: the same nine sites as the two that follow, but its acceptance reviewer was refused a cat of a copy of the block's behavioral model kept in the developer's notes outside the repository, read the same file with grep instead, and read an architecture copy the same way. The read the sandbox was built to stop.
2026-09-13c7b4deraygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude7m 51s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vThird repeat at one commit, the same nine sites and the same chain shape again; the plan reviewer approved and the loop re-ran the planner anyway, which is where the verdict-parse defect surfaced.
2026-09-13033296raygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude6m 54s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vSecond repeat at the same commit, same nine sites and same shape; its editor wrote and ran a 20,000-vector check of its own edit, which survives in the run directory as _bl_tb.v.
2026-09-126553b1raygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude6m 54s9currentequivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 124/124; tb_equiv_bilin_rgb_run4.v, comb_mult_add_16.vFirst of three repeats at one commit establishing that the planner takes all nine applicable sites: nine blocks across all three channels, where the 09-06 run had taken three over red alone.
2026-09-129ca8d3raygentopcomb_mult_add_16_mode
needs restructuring
accepted2GPT4m 12s1currentequivalent; sim 598 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 180/180; tb_equiv_bilin_rgb_run4gpt.v, comb_mult_add_16.vGPT against Claude on the same design and mode: one block fusing two of red's three products, with delayed operand copies keeping the terms aligned. The acceptance reviewer rejected the first attempt for summing a current-cycle product with a registered one.
2026-09-07b07463dla_like.smallcomb_sop_2_18_mode
direct swap
accepted1Claude45m 42s72legacy*equivalent; sim 1030 cycles on dsp_block_16_8_false.resulta/chainout, 0 mismatches; LEC 73/73; tb_equiv_dsp_block.v, comb_sop_2_18.vThe Koios result: every one of the 72 sum-of-products sites in dla_like.small converted onto comb_sop_2_18 from a single module edit. Its acceptance reviewer checked 200,000 vectors against a model of the block it wrote itself.
2026-09-0668ab30raygentopcomb_mult_add_16_mode
needs restructuring
accepted1Claude8m 7s3legacy*equivalent; sim 526 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 166/166; tb_equiv_bilin_rgb.v, comb_mult_add_16.vFirst rewrite proven equivalent: three chained comb_mult_add_16 blocks over the red channel of bilinearintrp. Both the editor and the acceptance reviewer simulated it against their own models of the block before the framework checked anything.
2026-09-055363c7raygentopcomb_mult_add_16_mode
needs restructuring
rejected gate_failed4GPT4m 44s1legacy*not equivalent; sim 526 cycles on bilinearintrp.r/g/b, 473 mismatches on g (0 on r, 0 on b); LEC 173/180 (unproven on g); tb_equiv_bilin_rgb.v, comb_mult_add_16.vFour editor attempts against a gate no edit could satisfy: the netlist aliasing labelled a single region a fusion, so the multi-region checks fired on every attempt. The rewrite was also wrong on the green channel.
2026-08-27b498feraygentopsop_4_mode
needs restructuring
accepted1Claude6m 39s1legacy*equivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 166/166; tb_equiv_bilin_rgb_run4.v, int_sop_4_model.v (1-stage)The Claude half of the controlled pair run once the mode contract was complete: the same single int_sop_4 rewrite as its GPT twin up to two wire names, and cycle-exact where the pre-fix edits had been a cycle late.
2026-08-278d3f03raygentopsop_4_mode
needs restructuring
accepted1GPT2m 6s1legacy*equivalent; sim 542 cycles on bilinearintrp.r/g/b, 0 mismatches; LEC 166/166; tb_equiv_bilin_rgb_run4.v, int_sop_4_model.v (1-stage)The GPT half of that controlled pair, and the rewrite the LEC study was built on: one int_sop_4 driving r directly, with the external register the pre-contract runs had added now gone.

equivalence: <verdict>; sim <cycles> on <module>.<outputs>, <mismatches>; LEC <proven>/<points>; <testbench>, <block model>. equivalent = 0 mismatches and every LEC point proven; inconclusive = 0 mismatches, LEC short; not equivalent = any mismatch. Mismatches are counted on the same cycle. Testbenches are in verify/equiv and block models in verify/models of the benchadapt repository.

regime: the VPR settings the run was measured under (legacy = auto device, minimum-width search, seed 1; current = fixed width 300, koios_extra_small, seed 1; a star marks a run recorded before the settings were written to the record). blocks: packed blocks in the target mode. attempts: editor calls against the one approved plan. Generated 2026-09-14 18:11 -0700.