A Decision Minds product concept · AI & HPC infrastructure
The decision-intelligence layer for AI factories. GPU clouds run five layers of stack on fifty tools and none of them talk. Yantra unifies the telemetry, from silicon to substation, and turns it into decisions: which node fails next, which GPUs are stranded, when to shed power.
“The infrastructure layer is the current beneficiary of spend… it's the people building compute who are doing the spending. It continues to become harder to deploy compute as the demand curve goes vertical.” Nikesh Arora, CEO Palo Alto Networks, Jul 31 2026. When deployment is the bottleneck, the operator who converts megawatts into sellable, reliable GPUs fastest wins the cohort. That conversion (burn-in, health, topology truth, power) is a data problem. Yantra is that data layer.
A week later the framing hardened: “Power is THE binding constraint… not fanciful plans for power, future forecasts of BTM or distributed batteries — but energized power today,” with neoclouds judged “solely measured by energized compute online today.” Chamath Palihapitiya, Aug 7 2026. If financing turns on energized compute online today, then megawatt-to-sellable-GPU stops being an efficiency metric and becomes a financing one. That is the number every layer below moves.
Five days after that, an operator put a price on it. Nebius's Q2 2026 shareholder letter publishes annual contract value per megawatt as a management metric, reports payback falling to 1 year 10 months from a historical two-to-three years, and raises the year-end contracted-power target to 5 GW. Three opinions became one disclosed number. Everything below is a lever on it.
Contracted value on deals signed, much of it against capacity not yet built. It is not realized revenue over active megawatts. The tile above is that, and the two do not belong on one axis.
Power keeps arriving faster than the plan. The conversion from contracted megawatt to sellable GPU is what has to keep up, and it is the part no procurement contract buys for you.
$6.7 trillion of capital expenditure lands in data centers through 2030, and almost none of that clock is measured in days. Site due diligence through permit-ready design runs 10 to 12 months. The power behind the site is a multi-year queue. Yantra owns the last 34 days. That is the point rather than a limitation: those 34 days are the only stretch an operator can still compress this quarter, and at Q2 2026 contract rates a week of a 10 MW hall carries $2.3M to $8.7M.
One number from the design end of the clock makes the deeper argument. Traditional pre-construction explores three or four site concepts, because a human draws each one; Marengo (YC S26) says its tooling explores more than a thousand in parallel and cuts the cycle to five or six months. The bottleneck was never engineering judgment. It was how many options one engineer can hold at once. That failure mode does not stop at the fence line.
Three of these bars belong to somebody else. The orange one is Berkeley Lab's median from interconnection request to commercial operation for generation built in 2025, and it is a proxy: large-load interconnection is a separate queue that study does not measure. Use it for the order of magnitude, not the decimal.
Nothing on the teal bar is cheap because it is short. It is the last gate before a contracted megawatt starts billing, it is the one gate an operator owns outright, and it is the only one where a week of work shows up in this quarter's revenue.
Every module below answers with one recommendation: drain these nodes, migrate this job, cap power here. That is the three-or-four-concepts habit, moved indoors. Crucible runs the candidate plans against the fleet twin and hands back the trade-off frontier instead, so the operator picks the risk appetite rather than accepting ours.
| Decision | Space | Traded against | Today |
|---|---|---|---|
| Maintenance window | 37 nodes × 6 windows | revenue at risk · spares · tech hours | one drain list |
| Job placement | ~10⁴ migrations | stranded HBM · fragmentation · SLA jitter | greedy heuristic |
| Oversubscription | ratio × class × tier | runtime penalty · SLA credits · revenue per GPU | one fixed ratio |
| Model residency | models × chip groups × swap windows | swap cost · traffic forecast · tail latency | pinned at deploy |
| Request routing | requests × engines × admission policy | capacity · fairness · tail latency | least-loaded |
| Power envelope | 24 h × 3 tiers | energy cost · 4CP position · SLA risk | one plan |
| Bring-up order | rack × contract start | days to first revenue · burn-in confidence | first in, first out |
Nebius, on receiving its first Vera Rubin NVL72 systems, describes using them to “validate compute, networking, and orchestration together as a single system” before offering them in production (Q2 2026 letter). That is this layer, described by an operator.
Every burn-in, every validation run, every RMA becomes structured data. Forge closes the loop between what fails in the fleet and what the ODM builds next.
| Signature | 30 d | Trend | Attribution |
|---|---|---|---|
| HBM ECC double-bit | 214 | ▼ 18% | vendor lot 24-C · RMA batch open |
| NVLink CRC storm | 96 | ▲ 41% | switch fw 2.4.1 · rollback queued |
| VRM thermal drift | 61 | ▲ 12% | E1.S custom node · ECN filed |
| PCIe retimer dropout | 23 | ▼ 6% | cable batch replaced wk 28 |
One health score per node, built from DCIM, BMC, and job telemetry. Failure prediction drains nodes before the job dies, not after.
| Node | Signal | Est. TTF | Action |
|---|---|---|---|
| b07-r112-n3 | ECC error slope 4.2×/day | 31 h | drain now |
| b02-r018-n7 | HBM temp drift +6 °C/wk | 58 h | drain at ckpt |
| a11-r201-n2 | NVLink retrain events ×9 | 64 h | drain at ckpt |
| c04-r077-n5 | PSU ripple anomaly | 6 d | watch |
Every incident becomes training data. Sentinel correlates job telemetry, fabric events, and change history into a ranked root cause before the bridge call starts.
A live graph of every GPU: where it is, what it's wired to, who holds it, and whether it can actually be sold. The cheapest capacity on the market is the capacity already inside the fence. Fragmentation is revenue lying on the floor, and it strands memory long before it strands GPUs.
Operators sell GPU-hours, but serving is bound by memory, not by arithmetic. A tenant sizes its KV cache for the longest context it might ever get, runs at a quarter of that, and every dashboard still reads 100% allocated. GPU counters cannot see inside HBM.
Agentic traffic widens the gap. Nebius reports production inference more than tripling in Q2, with a growing share agentic, where “a single task drives many model calls” and consumption scales with the complexity of the work rather than with user count. Context lengths spread further apart, so a reservation sized for the worst case wastes more of the fleet, not less.
That 352 GB bar is one instance of a claim that survives the architecture: what an operator bills for is not what runs out. Change the silicon and the scarce unit changes with it. The strand does not go away, and on the fleets with no HBM at all it gets worse, because there is nothing to page out.
| Fleet | Sold as | What binds | The strand |
|---|---|---|---|
| 8 × H100 | GPU-hours | HBM capacity | KV reserved for a context nobody reaches |
| Groq LPU | tokens | on-die SRAM residency | chips pinned to a model taking no traffic |
| Cerebras CS-3 | system-time · tokens | the wafer and its weight stream | a wafer serving a model too small to fill it |
| SambaNova SN40L | tokens | which tier holds which model | dead models resident in the DDR tier |
Groq is the hard case, not the exception. Weights live on-die, 230 MB of SRAM per chip, so a 70B model is pinned across hundreds of chips whether traffic arrives or not, and a deterministic schedule leaves no room to oversubscribe the gap. Cerebras holds 44 GB of SRAM on the wafer and streams weights from external memory, so its strand is granularity: nobody buys a third of a wafer. SambaNova keeps 64 GB of HBM3 per socket behind 1.5 TB of DDR5, which makes it the three-tier case rather than the HBM-free one.
Power is the scarcest input in the AI buildout. A GB300 NVL72 rack draws 135–150 kW, so a 512-rack hall is a ~65 MW grid asset. Flux treats it as schedulable: price and carbon signals in, checkpoint-aware power envelopes out.
| Window | Envelope | Mechanism | SLA impact |
|---|---|---|---|
| 00–14 h | 62 MW | full rate | none |
| 14–17 h | 48 MW | ckpt + GPU freq cap on preemptible tier | none |
| 17–24 h | 62 MW | catch-up burst, deferred jobs first | none |
Decision Minds has spent 15 years building enterprise data platforms. The AI factory is the largest new data-platform problem in the world: a 50k-GPU fleet emits more telemetry than most banks. Same discipline, new physics.