ICANEWS

Hierarchical Semi-Markov Model for AI Data Center Power Dynamics and Job Scheduling

arXiv CS · · 3 min read · Engineering & Technology

Read research and analysis on Hierarchical Semi-Markov Model for AI Data Center Power Dynamics and Job Scheduling published by ICANEWS, a global research journal for emerging researchers.

Key Takeaways

  • Facility-wide power swings and peak demand originate from job arrival and scheduling processes.
  • The HSM-DC model matches mean power (fit score 0.9997), its spread (0.92), and peak-to-average ratio (0.82) for a reference facility.
  • The model accurately predicts the share of queued jobs to within one point at high load.
  • Conventional industrial load models and within-job only models fail to capture true peak-to-average ratios for AI data centers.

Why This Matters

Accurate grid planning for AI data centers requires modeling job arrival and scheduling, not just scaling single-node power curves. This understanding is critical for effectively managing the distinct power demands introduced by these facilities.

Overview

AI data centers introduce a distinct load class with power dynamics fundamentally different from conventional industrial loads. This research addresses the challenge of accurately modeling these dynamics, which exhibit both rapid, within-job fluctuations and slower, facility-wide shifts. The study introduces a hierarchical semi-Markov Data-Center (HSM-DC) load model designed to couple these two distinct timescales and their associated processes: job scheduling and bulk-synchronous-parallel (BSP) power dynamics.

Research Context

Traditional modeling approaches that focus solely on within-job behavior or treat a facility as a fixed set of busy nodes are identified as insufficient. Such approaches tend to smooth out power swings and underestimate the true peak-to-average ratio in AI data centers. The unique characteristics of AI workloads, specifically the bulk-synchronous-parallel algorithm, cause rapid power fluctuations. Within a training job, nodes transition through compute, sync, and checkpoint steps, leading to power swings between full load and near idle conditions within seconds. Concurrently, jobs arrive, occupy blocks of nodes for durations ranging from hours to days, and then complete. This job scheduling process results in daily, weekly, and yearly variations in the number of busy nodes, which in turn drives facility-wide power swings and dictates the peak demand that determines grid link sizing. The explicit coupling of these two layers—the faster within-job dynamics and the slower facility-wide job scheduling—is critical for an accurate representation of AI data center power profiles.

Approach

The developed HSM-DC load model integrates two hierarchical layers:

  • Job-scheduling layer: This layer simulates job creation through a non-homogeneous compound-Poisson process, which is shaped by daily, weekly, and seasonal patterns. Each simulated job is assigned a heavy-tailed node count and duration. Jobs are then placed on a fixed pool of nodes based on a first-come basis.
  • Within-job layer: For each busy node, this layer models the internal power dynamics using a five-state semi-Markov chain. This chain represents the various steps of the BSP algorithm (compute, sync, checkpoint). State-based Ornstein-Uhlenbeck noise is incorporated into this layer.

Facility power is derived from the dynamically changing count of busy nodes and the per-node power consumption. The per-node power is configured to align with measured node data and the facility's straight-line power-versus-load curve. The model was configured to a reference facility, matching its scale for validation.

Findings

When configured to a reference facility of equivalent scale, the HSM-DC load model exhibited strong agreement with observed data:

  • It matched the mean power with a fit score of 0.9997.
  • It matched the spread of power with a fit score of 0.92.
  • It matched the peak-to-average ratio across various load levels with a fit score of 0.82.
  • The model also accurately matched the share of queued jobs, achieving agreement to within one percentage point at high load conditions.

A key finding from this modeling effort is that facility-wide power swings and peak demand are generated by how jobs arrive and are scheduled, rather than solely by the scaling of individual node power curves.

Why This Matters

The research suggests that for grid planning purposes related to AI data centers, it is essential to model the job arrival and scheduling processes. Simply scaling up a single node's power curve without considering these processes would lead to inaccuracies in predicting facility-wide power behavior, particularly peak demand. Understanding these dynamics is crucial for infrastructure planning and management.

Research Information

Institution
arXiv CS
Original Study
View Publication
Source
arXiv CS

About ICANEWS

ICANEWS is a global research journal for emerging researchers, publishing student and emerging researcher work across all fields.