Skip to the content.

Benchmark Glossary

A standardized task, dataset, procedure, and scoring rule used to compare system performance under defined conditions.

Benchmark

A benchmark is a standardized task, dataset, procedure, and scoring rule used to compare system performance under defined conditions.

A benchmark is not merely a bag of questions. Its result depends on the model, prompt, harness, tool access, time and compute budget, sampling settings, evaluator, and definition of success. Change the conditions and you may change the result.

Benchmarks are useful because anecdotes do not compare cleanly. They are dangerous because a clean number invites a larger claim than the test supports. A coding score is not general intelligence. A vulnerability-discovery score is not autonomous exploitation. Parity on one evaluation is not equivalence of systems.

A benchmark result should therefore be read by asking: What exactly was measured, under which conditions, and how much of the real-world task does that measurement represent?

See also

Benchmark Parity · Capability Gap · ExploitGym · Harness

Return to Dictionary All Entries (A–Z) For Students Other Writing Capstone 2.0