Skip to content
SpaceXSIDocumentationWebsite
Docs/Foundations

Efficiency & measurement

Useful work per watt is a principle. Reproducible measurement makes it meaningful.

Research reviewed 4 October 2026 Download chapter

Define the measurement boundary

A measurement may describe a chip, server, rack or entire facility. Changing the boundary changes the number. Accelerator telemetry should not be presented as the complete energy cost of a service.

A proposed benchmark should identify hardware, software version, workload, input size, warm-up, measurement period and instrument. It should disclose whether idle power, networking, storage and cooling are included.

Use a portfolio of metrics

MetricDefinitionPurpose
Energy per taskMeasured energy ÷ accepted tasksCompare equivalent useful work.
Useful throughputAccepted results ÷ elapsed timeSeparate valid output from activity.
LatencyRequest-to-result durationCheck service requirements.
UtilizationUsed capacity ÷ available capacityIdentify idle resources.
PUEFacility energy ÷ IT equipment energyDescribe facility overhead.
Carbon estimateEnergy × emissions factorEvaluate electricity-related emissions with stated assumptions.

A facility overhead example

The Green Grid defines PUE as total facility energy divided by IT equipment energy. It does not measure the usefulness of the computation. A low-PUE facility can still run an inefficient workload.

If IT equipment uses 100 kWh, PUE 1.4 implies 140 kWh at the facility boundary. At PUE 1.2, the same IT load implies 120 kWh: a reduction of 20 kWh, or about 14.3% of baseline facility energy. This example assumes equivalent IT work and measurement periods.

PUE = E_facility / E_IT
Saving (%) = (E_baseline − E_candidate) / E_baseline × 100

Preserve output quality

MLCommons develops MLPerf power measurement methods to connect energy use with benchmark performance. The relevant lesson is to measure completed work and its acceptance criteria alongside electricity.

A model that saves power but produces unusable output is not an equivalent improvement. Neither is a batched service that misses the required deadline. Quality, latency and reliability must remain inside the agreed envelope.

Reading the homepage model

The homepage uses a baseline of 100 relative units. An assumed 35% improvement yields 65 units for the same workload. The slider changes an assumption; it does not gather infrastructure measurements.

A future benchmark should replace assumed savings with repeated measurements, documented baselines and a distribution of results rather than one best run.

Candidate energy = Baseline × (1 − assumed saving / 100)
SpaceXSI Documentation Research edition · October 2026