A Decision Minds product concept · AI & HPC infrastructure

Yantra.

The decision-intelligence layer for AI factories. GPU clouds run five layers of stack on fifty tools and none of them talk. Yantra unifies the telemetry, from silicon to substation, and turns it into decisions: which node fails next, which GPUs are stranded, when to shed power.

GPU cloud operators (tier-2 neoclouds) AI landlords · miner-pivot & colo (Cipher, TeraWulf class) Sovereign AI factories (19 EuroHPC sites) Frontier labs operating leased halls Enterprise HPC & research computing Treasury & lenders underwriting GPU-backed facilities
GPUs under management
48,640
across 6 halls · 3 sites
Revenue per MW · yr
$10.5M
+1 pt sold ≈ $114K/MW
Cluster MFU
41.3%
▲ 2.1 pt since predictive migration
Predicted failures · 72 h
37 nodes
34 already draining
Bring-up · last hall
34 days
energize → sellable · every week saved ≈ $3.5M on 8k GPUs
Energy signal · day-ahead
$47/MWh
ERCOT avg · summer on-peak $110–165
WHY NOWCompute is the moat. Operating it is the margin.

“The infrastructure layer is the current beneficiary of spend… it's the people building compute who are doing the spending. It continues to become harder to deploy compute as the demand curve goes vertical.” Nikesh Arora, CEO Palo Alto Networks, Jul 31 2026. When deployment is the bottleneck, the operator who converts megawatts into sellable, reliable GPUs fastest wins the cohort. That conversion (burn-in, health, topology truth, power) is a data problem. Yantra is that data layer.

A week later the framing hardened: “Power is THE binding constraint… not fanciful plans for power, future forecasts of BTM or distributed batteries — but energized power today,” with neoclouds judged “solely measured by energized compute online today.” Chamath Palihapitiya, Aug 7 2026. If financing turns on energized compute online today, then megawatt-to-sellable-GPU stops being an efficiency metric and becomes a financing one. That is the number every layer below moves.

Five days after that, an operator put a price on it. Nebius's Q2 2026 shareholder letter publishes annual contract value per megawatt as a management metric, reports payback falling to 1 year 10 months from a historical two-to-three years, and raises the year-end contracted-power target to 5 GW. Three opinions became one disclosed number. Everything below is a lever on it.

Contracted ACV per MW · Nebius, disclosed Aug 12 2026

$50M $25M 0 $12M >$20M $40–50M 2026 base Q2 deals signed Q3 short-term

Contracted value on deals signed, much of it against capacity not yet built. It is not realized revenue over active megawatts. The tile above is that, and the two do not belong on one axis.

Contracted power target, same company, five raises

5 GW Aug 25 Nov 25 Feb 26 May 26 Aug 26

Power keeps arriving faster than the plan. The conversion from contracted megawatt to sellable GPU is what has to keep up, and it is the part no procurement contract buys for you.

What a repriced megawatt does to the pitch

×3.7
A week of bring-up, repriced
One week of a 10 MW hall carries ~$2.3M of contract value at the $12M base, ~$3.8M at the Q2 rate and ~$8.7M at the short-term rate. Same week, same work. Yantra's velocity argument repriced with the market and no claim on this page had to change.
1 yr 10 mo
Payback is the CFO's version of this page
Payback is capex, ACV per MW, time-to-revenue and utilization. Yantra moves three of the four. Talk stranded GPUs with the VP of infrastructure; talk months off payback with the CFO.
SOFR+250
Lenders now underwrite working GPUs
The first ~$775M facility is secured on deployed GPU infrastructure and contracted cash flows. Fleet health and delivery-against-contract became collateral quality, and the evidence layer became a treasurer's problem.
4 → 5 yr
Useful life is an evidence argument
Server life was extended a year on "usage patterns and current utilization commitments." Failure rates, thermal history and real duty cycles are what an auditor asks for. Forge and Sentinel already collect them.
BEFORE LAYER 01 The clock starts years before the megawatt Where the 34 days sits · and the search problem underneath it

$6.7 trillion of capital expenditure lands in data centers through 2030, and almost none of that clock is measured in days. Site due diligence through permit-ready design runs 10 to 12 months. The power behind the site is a multi-year queue. Yantra owns the last 34 days. That is the point rather than a limitation: those 34 days are the only stretch an operator can still compress this quarter, and at Q2 2026 contract rates a week of a 10 MW hall carries $2.3M to $8.7M.

One number from the design end of the clock makes the deeper argument. Traditional pre-construction explores three or four site concepts, because a human draws each one; Marengo (YC S26) says its tooling explores more than a thousand in parallel and cuts the cycle to five or six months. The bottleneck was never engineering judgment. It was how many options one engineer can hold at once. That failure mode does not stop at the fence line.

Where the 34 days sits, drawn to scale

Design, traditional Design, accelerated Power behind it Energize → sellable 10–12 months 5–6 months median 5 yr 34 days 0 1 yr 2 3 4 5 yr
grid-side design-side ours

Three of these bars belong to somebody else. The orange one is Berkeley Lab's median from interconnection request to commercial operation for generation built in 2025, and it is a proxy: large-load interconnection is a separate queue that study does not measure. Use it for the order of magnitude, not the decimal.

Nothing on the teal bar is cheap because it is short. It is the last gate before a contracted megawatt starts billing, it is the one gate an operator owns outright, and it is the only one where a week of work shows up in this quarter's revenue.

Counterpoint Halving design time moves nothing if the critical path is the interconnect, and often it is. That objection is the Flux argument stated backwards: when new power is queued behind years, a watt recovered inside a connection you already hold is worth more than a watt you are waiting on.

Crucible · one engine across all five layers

Every module below answers with one recommendation: drain these nodes, migrate this job, cap power here. That is the three-or-four-concepts habit, moved indoors. Crucible runs the candidate plans against the fleet twin and hands back the trade-off frontier instead, so the operator picks the risk appetite rather than accepting ours.

DecisionSpaceTraded againstToday
Maintenance window37 nodes × 6 windowsrevenue at risk · spares · tech hoursone drain list
Job placement~10⁴ migrationsstranded HBM · fragmentation · SLA jittergreedy heuristic
Oversubscriptionratio × class × tierruntime penalty · SLA credits · revenue per GPUone fixed ratio
Model residencymodels × chip groups × swap windowsswap cost · traffic forecast · tail latencypinned at deploy
Request routingrequests × engines × admission policycapacity · fairness · tail latencyleast-loaded
Power envelope24 h × 3 tiersenergy cost · 4CP position · SLA riskone plan
Bring-up orderrack × contract startdays to first revenue · burn-in confidencefirst in, first out
concept
Not a sixth layer
Forge through Flux stay five products in a fixed order. Crucible is the search that runs over all of them, and it only works once the data plane underneath is real, which is what the land-and-expand motion builds first.
LAYER 01 Forge Hardware Engineering: NPI, custom server design, GPU validation, co-design

Nebius, on receiving its first Vera Rubin NVL72 systems, describes using them to “validate compute, networking, and orchestration together as a single system” before offering them in production (Q2 2026 letter). That is this layer, described by an operator.

Every burn-in, every validation run, every RMA becomes structured data. Forge closes the loop between what fails in the fleet and what the ODM builds next.

Burn-in first-pass yield by SKU · last 30 days

80% 90% 100% 98.1 94.5 96.3 91.0 GB300 NVL72 HGX B200 MI355X E1.S custom
flagged: yield below 93% gate, auto-opened ECN w/ ODM

Failure fingerprints · validation + fleet, merged

Signature30 dTrendAttribution
HBM ECC double-bit214▼ 18%vendor lot 24-C · RMA batch open
NVLink CRC storm96▲ 41%switch fw 2.4.1 · rollback queued
VRM thermal drift61▲ 12%E1.S custom node · ECN filed
PCIe retimer dropout23▼ 6%cable batch replaced wk 28
Use case Meta’s Llama 3 run logged a failure every ~3 h on 16,384 H100s, of which ~54% traced to GPU/HBM (arXiv:2407.21783). Forge feeds those fingerprints to the ODM in weeks; burn-in catches the 3–8% infant mortality before it ships.
LAYER 02 Fleet Datacenter Infrastructure & Fleet Management: DCIM, fleet health, node lifecycle

One health score per node, built from DCIM, BMC, and job telemetry. Failure prediction drains nodes before the job dies, not after.

Fleet health · Hall B · 512 racks (each cell = 4 racks)

health low → high predicted failure

Predictive maintenance queue

NodeSignalEst. TTFAction
b07-r112-n3ECC error slope 4.2×/day31 hdrain now
b02-r018-n7HBM temp drift +6 °C/wk58 hdrain at ckpt
a11-r201-n2NVLink retrain events ×964 hdrain at ckpt
c04-r077-n5PSU ripple anomaly6 dwatch
Use case Meta logged 419 unexpected interruptions in 54 days on 16k GPUs. At posted B200 rates (~$5/GPU-hr) that cluster burns ~$82k/hr. Draining predicted failures at checkpoints instead of mid-job is worth ≈$19M/yr.
LAYER 03 Sentinel Production Engineering: reliability, incident command, AI guardrails

Every incident becomes training data. Sentinel correlates job telemetry, fabric events, and change history into a ranked root cause before the bridge call starts.

Job interruptions per 1,000 node-days · 12 weeks

0 5 10 4.6 wk 19 wk 30
interruptions / 1k node-days (Llama 3 baseline ≈ 3.8)

Auto-RCA · last 24 h

confidence 0.91
NVLink flap cascade · Hall B
Correlated w/ switch fw 2.4.1 rollout (wk 29). 14 jobs touched. Rollback queued via change mgmt.
confidence 0.84
Checkpoint stalls · storage pool 3
Metadata server GC pauses align w/ stall windows. Suggested: pin GC off-peak.
guardrail
ADLC audit: remediation agent
312 auto-actions this week · 100% within policy · 2 escalations to human, both approved.
Use case Auto-RCA cut MTTR 55% in reference design. Reliability review packs generate themselves: SEV list, blast radius, corrective actions, trend.
LAYER 04 Atlas Cloud Platforms: capacity control plane, inventory & topology, observability, auto-remediation

A live graph of every GPU: where it is, what it's wired to, who holds it, and whether it can actually be sold. The cheapest capacity on the market is the capacity already inside the fence. Fragmentation is revenue lying on the floor, and it strands memory long before it strands GPUs.

Sellable vs stranded GPUs · 8 weeks

48k 45k 42k sellable 47.4k stranded 1.2k
sellable (topology-clean) stranded (fragmented / misinventoried)

Remediation recommendations

+512 GPUs
Defrag island · Hall C, pods 7–9
3 tenants movable at next checkpoint → contiguous NVL domain restored.
+288 GPUs
Inventory drift · 36 nodes “in repair”, actually healthy
BMC + burn-in evidence attached. One-click return to sellable pool.
auto
Observability rollup
Golden-signal SLOs per pod; 3 pods burning error budget. Throttle recommendation issued.

Stranded HBM · one 8 × H100 node serving Llama 3 70B

WHAT THE TENANT IS BILLED 8 / 8 GPUs · 100% allocated WHAT THE 640 GB OF HBM IS DOING 141 GB weights 107 GB live 352 GB reserved 352 GB idle ≈ 4.4 GPUs of HBM, on a node billed as full
model weights runtime KV cache in use KV reserved, untouched

Why nobody sees it

Operators sell GPU-hours, but serving is bound by memory, not by arithmetic. A tenant sizes its KV cache for the longest context it might ever get, runs at a quarter of that, and every dashboard still reads 100% allocated. GPU counters cannot see inside HBM.

Agentic traffic widens the gap. Nebius reports production inference more than tripling in Q2, with a growing share agentic, where “a single task drives many model calls” and consumption scales with the complexity of the work rather than with user count. Context lengths spread further apart, so a reservation sized for the worst case wastes more of the fleet, not less.

+55% sessions
Paged KV on the four largest inference tenants
Same nodes, same tenants. Reservation follows real context length instead of the worst case.
reprice
Bill GB-hours of HBM, not GPU-hours
Charges for the input that is actually scarce, and pays the tenant to stop over-reserving.

The strand is not an HBM fact

That 352 GB bar is one instance of a claim that survives the architecture: what an operator bills for is not what runs out. Change the silicon and the scarce unit changes with it. The strand does not go away, and on the fleets with no HBM at all it gets worse, because there is nothing to page out.

FleetSold asWhat bindsThe strand
8 × H100GPU-hoursHBM capacityKV reserved for a context nobody reaches
Groq LPUtokenson-die SRAM residencychips pinned to a model taking no traffic
Cerebras CS-3system-time · tokensthe wafer and its weight streama wafer serving a model too small to fill it
SambaNova SN40Ltokenswhich tier holds which modeldead models resident in the DDR tier

Groq is the hard case, not the exception. Weights live on-die, 230 MB of SRAM per chip, so a 70B model is pinned across hundreds of chips whether traffic arrives or not, and a deterministic schedule leaves no room to oversubscribe the gap. Cerebras holds 44 GB of SRAM on the wafer and streams weights from external memory, so its strand is granularity: nobody buys a third of a wafer. SambaNova keeps 64 GB of HBM3 per socket behind 1.5 TB of DDR5, which makes it the three-tier case rather than the HBM-free one.

limit
Shape asserted, figure withheld
The 352 GB bar rests on PagedAttention's measured 60–80% waste band for static reservation. No equivalent published measurement exists for wafer-scale or SRAM-resident fleets, so these rows carry no number until one does.
Outside number Gartner, 17 Aug 2026: inference cost per agentic workflow will rise more than fivefold through 2028, even as model prices fall. The scope is the point. Not inference cost in general, but the cost of one job: routing a task to a reasoning agent instead of a chatbot costs the provider at least five times as much, before task complexity is counted. Gartner names it the Inference Paradox, better unit economics driving total cost up, and prescribes inference tiering, routing and orchestration to protect margin. Price per token falls while cost per job multiplies, and an operator watching only revenue per GPU-hour sees neither blade. That is why TOK sits on the tile above, and why routing and admission control belong in Crucible rather than in a config file.
Outside number Cast AI measured average GPU utilization at 5% across tens of thousands of Kubernetes clusters on AWS, Azure and GCP, against the 50% it calls healthy. That is not the 4% below, and the two do not add: Cast AI counts duty cycle on GPUs already allocated in enterprise clusters, while the reference model counts sellable capacity lost to fragmentation and bad inventory in a fleet already contracted out. They point the same way. Thunder Compute raised a $13M Series A on 19 Aug 2026 to sell software that reclaims the first kind. Crusoe ships MemoryAlloy, a cluster-wide KV cache with a routing gateway, against the second. Two funded companies selling shovels for the capacity this section measures is the best evidence that it is there. Reclaiming it is one job; knowing which reclamations are safe, what they cost in SLA jitter, and who gets billed after is the other.
Use case Reference model: 4% of a 48k fleet stranded at any time. Recovering it ≈ 1,900 GPUs ≈ $29–58M/yr at posted H100→B200 on-demand rates, 70% utilization. Memory strands harder than silicon. On the inference share of the fleet, the untouched KV reservation is worth 4.4 GPUs of HBM per 8-GPU node. HBM is allocated years out, so that capacity cannot be bought at any price.
LAYER 05 Flux Flexible Compute: software-defined power, power as a control input

Power is the scarcest input in the AI buildout. A GB300 NVL72 rack draws 135–150 kW, so a 512-rack hall is a ~65 MW grid asset. Flux treats it as schedulable: price and carbon signals in, checkpoint-aware power envelopes out.

Day-ahead price vs shiftable load · today

$120 $60 $0 $118 peak 00:00 14:00–17:00 24:00
day-ahead $/MWh · shaded = recommended curtailment window

Software-defined power plan · Hall B

WindowEnvelopeMechanismSLA impact
00–14 h62 MWfull ratenone
14–17 h48 MWckpt + GPU freq cap on preemptible tiernone
17–24 h62 MWcatch-up burst, deferred jobs firstnone
today
Projected: $5k energy avoided · 4CP position preserved
ERCOT peak avoidance is worth $40–67k/MW-yr, or $2.5–4M/yr on this hall. Zero SLA-bearing jobs touched.
Use case Emerald AI showed a 25% power cut for 3 h w/o service disruption (Nature Energy, 2026); Google runs 1 GW of DC demand response. ERCOT 4CP avoidance alone ≈ $2.5–4M/yr on a 62 MW hall, before energy arbitrage.

Two ways in

OPERATORS
Full stack · all five modules
You run your own AI cloud. Yantra is the decision layer over your fleet: Forge through Flux, licensed per GPU under management.
LANDLORDS
Facility side · Fleet + Flux
Your tenant brings their orchestration; you still own the power, cooling, and uptime SLAs, and the penalty exposure. Hall-level health scoring and ERCOT-aware load planning, w/o touching the tenant stack. You rent space and power; whoever runs compute on top of it earns the margin on the same megawatt. Fleet + Flux is the on-ramp to closing that gap.

Why Decision Minds

01 · LAND
Fleet data plane engagement
We build the unified telemetry lake (DCIM, BMC, fabric, job scheduler) in 8–12 weeks. Pure consulting revenue; the customer keeps the data model.
02 · EXPAND
Decision modules
Predictive maintenance, fragmentation recovery, power optimization, deployed as Yantra modules on the data plane, licensed per GPU under management.
03 · COMPOUND
Co-design flywheel
Cross-fleet failure fingerprints (anonymized) become the industry benchmark, the data moat neither the ODM nor the hyperscaler has.

Decision Minds has spent 15 years building enterprise data platforms. The AI factory is the largest new data-platform problem in the world: a 50k-GPU fleet emits more telemetry than most banks. Same discipline, new physics.