lulupedia
Nordfriisk 版本暂未收录,当前展示 English 内容。

Benchmark

10473 words·9/24/2026·English
0

A benchmark is a standard, reference point, or basis of comparison against which the performance, quality, cost, or other characteristics of an object, process, system, or organization can be measured or evaluated. Benchmarks are used in many fields, including computing, business, finance, surveying, education, and public policy. They may take the form of numerical thresholds, standardized tests, indices, datasets, or physical markers. A benchmark provides a common frame of reference that enables meaningful comparison across time, environments, organizations, or competing alternatives.

Etymology and general concept

The word benchmark originally derives from surveying. A bench mark was a mark cut into a durable surface, such as a stone wall or rock, at a known height above a datum. Surveyors could place a leveling rod on a horizontal ledge or “bench” attached to the mark, allowing accurate elevation measurements to be repeated from the same reference point. Over time, the term broadened to mean any fixed standard or point of reference used for comparison.

In modern usage, a benchmark is not necessarily a physical object. It can be a standard dataset, a performance index, a best practice, a regulatory threshold, or a set of criteria. The common thread is that a benchmark supplies an external reference that helps interpret measurements and guide decisions.

Computing

In computing, a benchmark is a test, program, or set of tasks used to assess the performance of computer hardware, software, algorithms, or systems. Computing benchmarks are commonly used to compare processors, graphics cards, storage devices, databases, compilers, networking equipment, and cloud services.

Types of computing benchmarks

Computing benchmarks range from narrow, low-level tests to broad, application-level workloads.

  • Microbenchmarks measure a single operation or component, such as memory latency, integer arithmetic, floating-point operations, disk input/output, or network throughput. They are useful for isolating specific performance characteristics but may not represent real-world behavior.
  • Macrobenchmarks measure entire systems or applications under realistic workloads, such as web servers, video encoding, database transactions, or office productivity tasks.
  • Synthetic benchmarks are artificial workloads designed to stress particular components in a repeatable way. Examples include graphics benchmarks that render fixed scenes or CPU tests that perform mathematical calculations.
  • Application-based benchmarks use real software with scripted or recorded workloads. They are often considered more representative than synthetic benchmarks but can be harder to standardize.

Common metrics reported by computing benchmarks include execution time, throughput, latency, frames per second, transactions per second, operations per second, and energy consumption per task.

Common benchmark suites and standards

Several organizations and industry groups maintain widely used benchmark suites:

  • SPEC (Standard Performance Evaluation Corporation) produces CPU, graphics, storage, and energy benchmarks, such as SPECint and SPECfp.
  • LINPACK measures floating-point computing performance and is used in the TOP500 ranking of supercomputers.
  • Geekbench and Cinebench are popular cross-platform benchmarks for consumer processors.
  • PCMark and 3DMark assess overall personal computer performance and graphics performance.
  • TPC (Transaction Processing Performance Council) publishes benchmarks for database and transaction processing systems, such as TPC-C and TPC-H.
  • MLPerf measures performance in machine learning training and inference tasks.

Benchmarking in computing practice

Computing benchmarks must be carefully controlled. Results can be affected by hardware configuration, operating system, driver versions, compiler optimizations, cooling, background processes, and test parameters. To improve comparability, benchmark developers often specify exact hardware and software settings, run multiple iterations, and report median or average results with measures of variability.

Benchmark results are frequently used in product marketing, procurement decisions, academic research, and hardware reviews. However, results should be interpreted with caution, because a system optimized for a particular benchmark may not perform equally well on other workloads.

Business and management

In business and management, benchmarking is the systematic process of comparing an organization’s processes, products, services, or performance metrics with those of industry leaders, competitors, or recognized best practices. The goal is to identify gaps, understand superior performance, and implement improvements.

The practice was popularized in the late twentieth century, particularly through the quality management movement. Xerox is often cited as an early corporate adopter of formal competitive benchmarking in the 1980s. Since then, benchmarking has become a standard tool in total quality management, Six Sigma, ISO management standards, and continuous improvement programs.

Types of business benchmarking

Business benchmarking is commonly classified into several types:

  • Internal benchmarking compares processes or performance across units, departments, or locations within the same organization.
  • Competitive benchmarking compares an organization with direct competitors or industry peers.
  • Functional benchmarking compares similar functions or processes across different industries.
  • Generic or process benchmarking examines best-in-class practices for a particular process, regardless of industry.

Typical performance measures used in business benchmarking include cost per unit, cycle time, defect rate, customer satisfaction score, employee productivity, and return on investment.

Benchmarking process

A standard benchmarking process often includes the following steps:

  1. Identify the process or metric to be benchmarked.
  2. Define the comparison group or best-practice source.
  3. Collect data through public reports, surveys, site visits, or third-party research.
  4. Analyze performance gaps and underlying causes.
  5. Develop and implement improvement actions.
  6. Monitor results and repeat the cycle.

Finance and investment

In finance, a benchmark is a standard or index against which the performance of an investment portfolio, fund, asset class, or strategy is measured. Common financial benchmarks include broad market indices such as the S&P 500, FTSE 100, Nikkei 225, MSCI World, and Bloomberg Aggregate Bond Index.

A portfolio’s return relative to its benchmark is often expressed as excess return or alpha. Tracking error measures how closely a portfolio follows its benchmark. Passive investment funds, such as index funds and exchange-traded funds, aim to replicate benchmark performance with minimal tracking error. Active managers attempt to exceed benchmark returns after fees and risk adjustments.

Benchmarks also play a role in asset allocation, performance evaluation, and risk management. They provide a neutral reference for judging whether investment decisions added value. In some cases, investors use customized benchmarks that reflect a specific blend of asset classes or investment constraints.

Benchmark rates

In financial markets, a benchmark rate or reference rate is an interest rate or price used as a standard for contracts, loans, mortgages, derivatives, and other financial instruments. Examples include central bank policy rates, the Secured Overnight Financing Rate (SOFR), the Euro Interbank Offered Rate (EURIBOR), and formerly the London Interbank Offered Rate (LIBOR). Benchmark rates are important because they affect borrowing costs, valuation, and the settlement of financial contracts. Regulators often oversee the administration and reform of benchmark rates to ensure transparency and reliability.

Surveying and geodesy

In surveying and geodesy, a benchmark is a permanent or semi-permanent marker with a known elevation relative to a specified vertical datum. Benchmarks form part of a vertical control network used for mapping, construction, floodplain management, and scientific measurement of land movement.

Traditional benchmarks are often brass disks, bolts, chiseled marks, or concrete monuments set in stable structures such as buildings, bridges, bedrock, or concrete posts. Surveyors use precise leveling instruments to transfer elevations from one benchmark to another. Government agencies, such as the United States National Geodetic Survey, the United Kingdom Ordnance Survey, and similar national mapping organizations, maintain databases of benchmark locations and elevations.

A benchmark in the original sense commonly had a horizontal shelf or “bench” cut into the surface so that a leveling staff could be placed consistently. Modern geodetic control points may also include three-dimensional coordinates determined by satellite positioning, such as GPS or GNSS.

Other applications

Benchmarks are also used in many other domains:

  • Education: Benchmarks are learning standards or performance levels that students are expected to reach at particular grades or stages.
  • Healthcare: Quality benchmarks compare hospitals, clinics, or treatment outcomes against national or international standards.
  • Public policy: Governments and international organizations use benchmarks to track progress on goals such as poverty reduction, environmental quality, and infrastructure development.
  • Manufacturing and engineering: Benchmarks may be reference products, materials, or tolerances used in testing and quality assurance.
  • Energy and environment: Efficiency benchmarks compare buildings, vehicles, appliances, or industrial plants with best-performing equivalents.

Methodological principles

A useful benchmark generally exhibits several characteristics:

  • Relevance: It should measure something important to the decision or objective.
  • Representativeness: It should reflect realistic conditions or workloads.
  • Reproducibility: It should produce consistent results under the same conditions.
  • Comparability: Results should be meaningful across different systems, organizations, or time periods.
  • Transparency: Methods, data, and assumptions should be clearly documented.
  • Stability: The benchmark should not change too frequently or arbitrarily, unless the underlying context requires revision.

When constructing or selecting a benchmark, it is important to define the baseline, normalize inputs and outputs where appropriate, report uncertainty, and avoid hidden biases. Version control matters, especially in computing, where benchmark versions must remain stable to compare results over time.

Criticisms and limitations

Benchmarks are valuable but have limitations. In computing, synthetic benchmarks may reward narrow optimizations that do not improve real-world performance. Compiler and firmware optimizations can be designed specifically to inflate benchmark scores, a practice sometimes called “benchmark gaming” or “benchmarketing.” A benchmark that becomes a target can also distort behavior: according to Goodhart’s law, when a measure becomes a target, it ceases to be a good measure.

In business, benchmarking can encourage imitation of competitors rather than genuine innovation. It may rely on incomplete or non-comparable data, and best practices in one organization may not transfer to another context.

In finance, a benchmark may not match an investor’s actual objectives or risk tolerance. Short-term performance against a benchmark can also encourage excessive risk-taking or herding behavior among fund managers.

In surveying, physical benchmarks can be destroyed, displaced, or affected by ground movement, requiring regular verification and updates.

Despite these limitations, benchmarks remain important tools for measurement, evaluation, and improvement when they are carefully designed, transparently reported, and interpreted in context.

Comments (0)

U

No comments yet. Be the first to comment!

Related Articles