Platform

Every part testable on its own.

HexaGrid is a scheduler, not a suite. Live grid prices come in, a demand forecast says where your load is going, a constraint solver decides what runs when, and a test harness checks the whole thing against cases where the right answer is already known. This page is the detail behind that.

01 — Data in

Live grid and carbon feeds

HexaGrid integrates wholesale price data for all five US ISO regions through the EIA API, cached with a 15-minute TTL, plus carbon intensity for the same regions. In a deployment these are the live inputs every scheduling decision is built on.

Read this before the numbers below. The feed integration covers all five regions, but every figure on this site comes from synthetic price curves built to match the daily shape of real ISO data — not from the live feeds. Validating against real market prices is the next piece of work and it could move these numbers. Scheduling is also single-region: HexaGrid does not move work between ISOs.
02 — Forecast

What your site will draw

A neural network predicts facility demand up to two hours ahead, so deferral decisions are made against where load is heading rather than where it just was. It is measured against the honest baseline — assuming demand stays where it is now — on data the model has never seen.

HorizonHexaGrid errorLast-value baselineBetter by
30 min0.60 kW2.9 kW5×
60 min0.56 kW5.2 kW9×
120 min0.59 kW7.3 kW12×

The point is the shape rather than the ratio. HexaGrid's error stays flat as the horizon grows, while the baseline degrades steadily — which is what makes a two-hour deferral window usable at all.

This forecasts demand, not price. We also built a day-ahead price model and it was beaten by a simple daily-average baseline on every grid we tested, so it is not in the product.
03 — The scheduler

A constraint solver, not a heuristic

Jobs, cluster capacity and deadlines are handed to CP-SAT as a real scheduling model — intervals, cumulative capacity constraints and element lookups for time-varying prices. That formulation is what keeps it tractable across a 300-run sweep.

Every solve returns a certified bound: not just a schedule, but a proof of how far it could possibly be from the best one. Across 300 solves the median certified gap was 0.32% and the worst was 3.50%, with 102 proved exactly optimal. You are told how good the answer is rather than asked to assume it.

Densely loaded clusters take longer to solve and land on looser bounds. That is a computational property, not a physical one — we tested whether dense sites have less to gain and found no evidence that they do.
04 — How it is measured

Three arms, including a ceiling we cannot beat

Every benchmark runs three ways on the same jobs and the same real price data. The middle arm is the product. The outer two exist so you can see how much room there is on either side of it.

On arrival

Every job runs the moment it lands. No scheduling. This is the baseline the 13.8% is measured against.

HexaGrid

Scheduled against the forecast, then billed at the prices that actually occurred. Forecast error costs you real money here, as it would in production.

Perfect foresight

Scheduled with tomorrow's prices known in advance. Impossible to build. On this data it saves 16.0% — the hard ceiling for the whole idea, and we capture 88% of it.

05 — Carbon

A trade you make, not a score we invent

The same schedule can be tuned toward cost or toward emissions. HexaGrid solves for a point on that frontier and reports both outcomes in their own units — dollars and tonnes — and never adds them into a single number, because the exchange rate between them is yours to set, not ours.

We publish no carbon percentage. The carbon signal in our benchmark is synthetic, scaled to an EPA annual average rather than measured, so any figure would reflect our own assumptions rather than your grid. The mechanism is built and the trade is real; the magnitude is unverified until it runs on measured intensity data.
06 — Guardrails

Tests that are allowed to fail

Nineteen regression tests run on every build. Several are null cases where the correct answer is known in advance — given a flat price curve the scheduler must report exactly zero saving, savings must increase with price volatility, and zero deferrable work must yield zero benefit. An earlier version of this code reported 14.8% savings on a flat curve. The current one reports exactly 0.00% across all 50 such runs, and the test that catches it runs on every build.

On the hardware side, NVML telemetry polls every GPU every ten seconds — temperature, power draw, memory use, ECC counts, fan speed — and an anomaly detector watches for combined-metric failure signatures that no single threshold would catch. Nothing gets scheduled onto hardware that is struggling.

Runs beside your cluster

No agents on your nodes, no hardware changes, no vendor lock-in. A single service process, a dashboard, and two API keys. Nothing to evaluate costs you anything.

Bare metal
WSL2
AWS EC2
Docker
Azure
Any NVIDIA GPU, Pascal or later