Data centre growth needs a measured energy story. Compute demand, AI workloads, cooling, networking, and storage all affect electricity use. A credible technology plan connects workload growth to power, efficiency, location, grid conditions, and operating decisions.

The U.S. Department of Energy’s Federal Energy Management Program publishes data-centre efficiency guidance that shows why the relationship needs careful measurement. A broad forecast is not a site plan, and a site plan is not a guarantee of lower impact.

On this page

Separate workload from facility demand

Start with the workload: compute type, utilisation, storage, network traffic, training or inference pattern, service level, and growth. Then map it to servers, power, cooling, and facility operation.

A headline about AI or cloud demand hides the engineering variables that operators can actually change.

  • Workload class and utilisation.
  • Compute, storage, and network profile.
  • Growth period and service level.

Use the right efficiency measures

Power usage effectiveness can help compare facility overhead with IT load, but it does not describe the total environmental or business result by itself.

Record measurement boundary, weather, utilisation, redundancy, water, carbon intensity, and workload output where relevant. Avoid comparing numbers built from different boundaries.

  • Boundary and measurement period.
  • IT load and facility overhead.
  • Utilisation, water, and grid context.

Plan for the grid and the site

A facility’s energy story depends on where it is, how it connects, when it draws power, what generation and transmission exist, and how demand competes with other users.

Site selection should include connection time, resilience, local constraints, backup generation, cooling resources, and the evidence behind supply claims.

  • Grid connection and capacity.
  • Peak, average, and time profile.
  • Resilience, cooling, and local constraint.

Measure utilisation before adding capacity

Low utilisation can make both cost and impact worse because infrastructure is powered and cooled without producing proportional useful work. Improve scheduling, right-sizing, workload placement, and decommissioning before assuming more hardware is the only answer.

Capacity planning still needs headroom. The point is to make slack visible and deliberate.

  • Capacity and utilisation.
  • Idle and stranded resource.
  • Scheduling and decommissioning.

Treat cooling as part of the architecture

Cooling design depends on equipment density, climate, water availability, air or liquid systems, redundancy, and maintenance. It can limit where and how a workload is placed.

Test cooling assumptions under peak load and failure conditions. A design that works only at average demand is not a resilient design.

  • Density and thermal profile.
  • Cooling method and resource.
  • Peak and failure test.

Connect efficiency with useful output

A lower facility ratio does not automatically mean a better service. Compare energy with useful output such as transactions, model inferences, storage service, or delivered compute, while stating the measurement boundary.

Consistent definitions matter more than a single impressive ratio.

  • Output definition.
  • Energy per useful unit.
  • Comparison period and assumptions.

Make claims auditable

Keep a ledger of source data, assumptions, calculation method, supplier claims, forecasts, and uncertainty. Distinguish measured results from estimates and scenarios.

This helps technical, finance, sustainability, and procurement teams discuss the same decision without treating a forecast as a fact.

  • Source and calculation.
  • Measured, estimated, or scenario.
  • Owner, date, and uncertainty.

What does not matter as much as a single headline

A large forecast, a low efficiency ratio, or a renewable-energy label cannot answer every site, workload, or resilience question. The useful story is a chain from workload to facility to grid to measured output.

Keep improving the chain as workloads and infrastructure change.

  • Do not use one ratio as the whole story.
  • Do not confuse certificate with physical availability.
  • Do not forecast without workload assumptions.

Turn the design into an operating control

A design becomes an operating control when a named person can perform it, another person can review it, and the organisation can show evidence that it happened. Write the trigger, the action, the expected result, and the exception path in language an operator can use during a busy day.

Keep the control close to the workflow. If staff must leave one system, search an unrelated document, and ask another team before acting, the control will be skipped when pressure rises. Reduce that friction without hiding the decision.

  • Name the trigger and operator.
  • State the expected result.
  • Record the exception and escalation.

Test the failure path

Happy-path demonstrations are useful for learning, but they do not prove resilience or security. Test incomplete data, unavailable dependencies, expired credentials, unexpected volume, delayed input, and a human decision that disagrees with the system output.

A failed test is useful when it produces an owner, a correction, a retest date, and a decision about whether the remaining risk is acceptable. Do not quietly convert a failed test into a passing narrative.

  • Choose realistic failure cases.
  • Record evidence and observed impact.
  • Assign correction and retest dates.

Measure the result without false precision

Choose a small set of measures that show whether the control or workflow is working. Define the denominator, time period, data source, owner, and action that follows a meaningful change.

Use estimates and scenarios honestly. A precise-looking number built on incomplete data is less useful than a range with a clear boundary and a plan to improve measurement.

  • Keep definitions stable.
  • Separate measured, estimated, and projected results.
  • Connect each measure to a decision.

Review change and ownership

Technology environments change through releases, suppliers, data, policies, identities, and user behaviour. A control that was adequate at launch may not remain adequate after a material change.

Set a review trigger as well as a calendar review. When the owner, dependency, data, exposure, or failure mode changes, revisit the design and keep the decision record with the evidence. Keep the next review date visible.

  • Record version and change.
  • Review after material events.
  • Keep owner, date, and decision visible.

Keep the handoff explicit

Most operational failures occur between teams, systems, or stages of work. State what one owner must provide, what the next owner checks, and what happens when the handoff is late, incomplete, or rejected.

This simple contract improves incident response and day-to-day work. It also makes automation safer because the input, output, and exception are visible rather than implied.

  • Name the sender and receiver.
  • Define the input and acceptance check.
  • Record rejection, retry, and escalation.

Operating rule: Name the owner, the evidence, and the action before calling a technology control complete.

Comparison table

Area Practical question Evidence to request
Workload What is the energy serving? Compute, storage, network, output
Facility What overhead is required? Power, cooling, redundancy, utilisation
Grid What supplies the site? Connection, profile, resilience, context
Evidence What can be claimed? Measurement, boundary, source, uncertainty

FAQ

Does AI have one fixed energy cost?

No. Energy depends on model, hardware, utilisation, data movement, cooling, location, workload pattern, and measurement boundary.

Is PUE enough?

No. PUE describes a facility relationship between total facility energy and IT energy. It does not capture every workload, grid, water, carbon, or useful-output question.

Why measure useful output?

It connects energy to the service delivered and helps compare efficiency when workload mix and utilisation change.

What should a data centre plan publish?

Publish the boundary, source data, assumptions, measurement period, workload context, resilience, and uncertainty behind important claims.

How can a team start without rebuilding everything?

Start with one important workflow, define the owner and evidence, test the failure path, and expand only after the operating result is understood.

What should be recorded after a review?

Record the scope, date, evidence, decision, owner, unresolved risk, and next review or correction. A short honest record is more useful than an impressive but untraceable claim.

When should the design change?

Change it when the workflow, data, identity, dependency, supplier, exposure, user group, or failure mode changes materially. A calendar review alone may miss the event that changed the risk.

What is a useful first metric?

Choose a measure close to an operating decision, define its denominator and time period, and state what action follows when it crosses the agreed threshold.

Conclusion

The useful technology decision is the one that can be tested. Define the operating problem, record the evidence, assign ownership, and review the result after launch. Clear scope beats a large claim, and a measured workflow beats a polished demo.

Sources

Previous post Digital Identity Design Is More Than a Login
Next post Supplier Cyber Risk Needs Evidence, Not a Questionnaire