How to read AI power and energy measurements

Pnebbula · · 4 min

Before comparing AI energy claims, identify what was measured: a component rating, system power during a workload, or energy per completed task. These quantities use different boundaries and cannot be substituted for one another.

Match the workload, duration and output-quality condition before interpreting an efficiency difference.

Watts and watt-hours answer different questions

Power describes a rate of energy use. Energy describes an amount used over an interval. A lower power figure does not automatically mean less energy for the completed task, because the task may take longer.

Here is a deliberately simplified calculation, not a benchmark result. A system drawing a constant 200 watts for half an hour uses 100 watt-hours. A system drawing a constant 100 watts for two hours uses 200 watt-hours. The second power value is lower, but the energy for the stated run is higher. Real measurements need the observed power over the relevant interval rather than an assumed constant value.

Watts and watt-hours answer different questions
Quantity Unit in this example Question answered
Power W At what rate is energy being used?
Duration h How long does the measured run last?
Energy Wh How much energy is used over that interval?
Energy per accepted result Wh per result How much measured energy accompanies useful completed work?
Lower power does not always mean less energy
Illustrative calculation with constant power: 200 watts for 0.5 hour uses 100 watt-hours; 100 watts for 2 hours uses 200 watt-hours. These are synthetic values, not measured hardware results.

Do not insert a component's rated power into this calculation and label the output measured system energy. A rating and an observation are different inputs.

Draw the measurement boundary before comparing numbers

The MLCommons Power working group develops power measurement techniques and metrics for machine-learning systems. Its work makes the measurement method part of the result. A number without a boundary is difficult to interpret: does it describe an accelerator, a whole server or a larger installation?

A comparison between accelerator-only power and whole-system power can be misleading even when both measurements are accurate. Storage, host processors and other components may belong to one boundary and not the other. State what is included before calculating an efficiency ratio.

The time boundary matters too. Is the run already warmed up? Are setup and idle periods included? You do not need to impose one boundary on every use case, but you do need to keep the compared boundaries compatible or explain the difference.

Count accepted work as well as generated output

A system that produces a long answer quickly can still fail the task. If the question is operational efficiency, define the quality condition and count the outputs accepted under it. Retries and discarded results remain part of the observed workload.

For example, energy per generated token and energy per accepted document extraction have different denominators. Neither is automatically wrong. The first describes one aspect of generation; the second relates the measurement to a task with an acceptance rule.

The result still does not establish a complete environmental footprint. Manufacturing, electricity sourcing and other lifecycle factors require their own evidence and boundaries. Keep a measured energy claim specific enough that another reader can tell exactly what was compared and what would be needed to extend the conclusion.

Keep useful work constant

The MLPerf Inference documentation defines workloads, scenarios and quality requirements. An efficiency comparison needs an appropriate match across these dimensions.

If one system completes a different task or accepts lower-quality output, the energy difference cannot be attributed only to better hardware efficiency.

For a custom application, define successful work before calculating energy per result. A response that must be discarded or repeated still consumed resources, even if the initial request looked inexpensive.

For broader questions about hardware dependence, see the semiconductor and sovereignty dossier, available in French. Keep that industrial analysis separate from the measurement boundary of an efficiency result.

Sources and reference documents