RE-Bench Is a Systems Benchmark: What Its Scorers and Selection Rules Actually Support
LatentAtlas
Abstract
RE-Bench should be read as evidence about versioned agent systems and allocation protocols, not base models in isolation. A pinned-source audit shows why. When a valid score exists, the pinned Kernel package retains a selected single-call timing, while an all-invalid aggregate falls back to numeric 0; the Scaling Law scorer evaluates a submitted JSON against an embedded curve without training the proposed target configuration; and the LLM Foundry validity gate uses parameter distance rather than behavioral equivalence. These findings do not invalidate the paper's historical results. They show that scorer design, source provenance, repeated attempts, and selection rules are part of the capability claim.
This review statically inspected all seven task packages at the pinned public revision. It also ran eight bounded contract tests over four hash-pinned files in Kernel, Scaling Law, and LLM Foundry. Those checks reproduced the named wrapper and source-contract behaviors; they did not execute a participant solution, protected scorer service, benchmark hardware, historical model, or paper-reported result.
Two evidence surfaces remain deliberately separate. Paper-reported claims describe the authors' historical study. Pinned-source-confirmed claims describe the public repository at the commit above. The exact task, runner, dependency, and artifact mapping between that commit and the historical run environments has not been established. No pinned-source finding is therefore transferred backward to a historical result unless the paper independently documents the same property. Not run means no comparable runtime result was produced here; it does not mean zero, failure, or success.
Paper and research resources
- Full paper — PDF, 8 pages
The original published PDF, hosted directly on LatentAtlas.
- Frozen source and contract checks (v0.1)
The original source archive released with this note.
- Inspect the pinned source on GitHub
A fixed commit corresponding to this publication, rather than the changing main branch.
- Archived publication — Zenodo, version 0.1
The existing DOI identifies this work and its deposited version.
Publication status and scope
AI-assisted pinned-source contract audit; not peer reviewed or independently adjudicated. No benchmark-hardware or historical headline-result reproduction is claimed.
The paper is available under CC BY 4.0. Code and dependencies retain the licenses specified in their releases.
Cite this work
Buldurgan, Huseyin. (2026). RE-Bench Is a Systems Benchmark: What Its Scorers and Selection Rules Actually Support (Version 0.1). Zenodo. https://doi.org/10.5281/zenodo.22089195