ProgramBench: Can Language Models Rebuild Programs From Scratch?
A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.
Reference2026Active3 models
Data verified 21 Jul 2026 · Methodology 1.6.0
Benchmark score on ProgramBench: Can Language Models Rebuild Programs From Scratch?
A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
What does ProgramBench: Can Language Models Rebuild Programs From Scratch? measure?
A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.
Which model scores highest on ProgramBench: Can Language Models Rebuild Programs From Scratch??
Kimi K3 by Moonshot AI currently leads with 77.8%.
How many models are evaluated on ProgramBench: Can Language Models Rebuild Programs From Scratch??
3 models in the LuminaBench cohort have a qualifying score on this benchmark.
Does this affect overall Lumina rank?
No. This benchmark is display-only and does not enter the overall Lumina composite.