A Decision Minds product concept · AI & HPC infrastructure

Yantra.

The decision-intelligence layer for AI factories. GPU clouds run five layers of stack on fifty tools and none of them talk. Yantra unifies the telemetry — from silicon to substation — and turns it into decisions: which node fails next, which GPUs are stranded, when to shed power.

Neoclouds (Crusoe, CoreWeave, Nebius class) Sovereign AI factories Enterprise HPC & research computing Colo + energy operators entering AI
GPUs under management
48,640
across 6 halls · 3 sites
Fleet health index
96.2
▲ 0.8 this week
Cluster MFU
41.3%
▲ 2.1 pt since predictive migration
Predicted failures · 72 h
37 nodes
34 already draining
Energy signal · day-ahead
$47/MWh
ERCOT avg · summer on-peak $110–165
LAYER 01 Forge Hardware Engineering — NPI, custom server design, GPU validation, co-design

Every burn-in, every validation run, every RMA becomes structured data. Forge closes the loop between what fails in the fleet and what the ODM builds next.

Burn-in first-pass yield by SKU · last 30 days

80% 90% 100% 98.1 94.5 96.3 91.0 GB300 NVL72 HGX B200 MI355X E1.S custom
flagged: yield below 93% gate — auto-opened ECN w/ ODM

Failure fingerprints · validation + fleet, merged

Signature30 dTrendAttribution
HBM ECC double-bit214▼ 18%vendor lot 24-C · RMA batch open
NVLink CRC storm96▲ 41%switch fw 2.4.1 · rollback queued
VRM thermal drift61▲ 12%E1.S custom node · ECN filed
PCIe retimer dropout23▼ 6%cable batch replaced wk 28
Use case Meta’s Llama 3 run logged a failure every ~3 h on 16,384 H100s — ~54% traced to GPU/HBM (arXiv:2407.21783). Forge feeds those fingerprints to the ODM in weeks; burn-in catches the 3–8% infant mortality before it ships.
LAYER 02 Fleet Datacenter Infrastructure & Fleet Management — DCIM, fleet health, node lifecycle

One health score per node, built from DCIM, BMC, and job telemetry. Failure prediction drains nodes before the job dies, not after.

Fleet health · Hall B · 512 racks (each cell = 4 racks)

health low → high predicted failure

Predictive maintenance queue

NodeSignalEst. TTFAction
b07-r112-n3ECC error slope 4.2×/day31 hdrain now
b02-r018-n7HBM temp drift +6 °C/wk58 hdrain at ckpt
a11-r201-n2NVLink retrain events ×964 hdrain at ckpt
c04-r077-n5PSU ripple anomaly6 dwatch
Use case Meta logged 419 unexpected interruptions in 54 days on 16k GPUs. At posted B200 rates (~$5/GPU-hr) that cluster burns ~$82k/hr — draining predicted failures at checkpoints instead of mid-job is worth ≈$19M/yr.
LAYER 03 Sentinel Production Engineering — reliability, incident command, AI guardrails

Every incident becomes training data. Sentinel correlates job telemetry, fabric events, and change history into a ranked root cause before the bridge call starts.

Job interruptions per 1,000 node-days · 12 weeks

0 5 10 4.6 wk 19 wk 30
interruptions / 1k node-days (Llama 3 baseline ≈ 3.8)

Auto-RCA · last 24 h

confidence 0.91
NVLink flap cascade — Hall B
Correlated w/ switch fw 2.4.1 rollout (wk 29). 14 jobs touched. Rollback queued via change mgmt.
confidence 0.84
Checkpoint stalls — storage pool 3
Metadata server GC pauses align w/ stall windows. Suggested: pin GC off-peak.
guardrail
ADLC audit: remediation agent
312 auto-actions this week · 100% within policy · 2 escalations to human, both approved.
Use case Auto-RCA cut MTTR 55% in reference design. Reliability review packs generate themselves — SEV list, blast radius, corrective actions, trend.
LAYER 04 Atlas Cloud Platforms — capacity control plane, inventory & topology, observability, auto-remediation

A live graph of every GPU: where it is, what it's wired to, who holds it, and whether it can actually be sold. Fragmentation is revenue lying on the floor.

Sellable vs stranded GPUs · 8 weeks

48k 45k 42k sellable 47.4k stranded 1.2k
sellable (topology-clean) stranded (fragmented / misinventoried)

Remediation recommendations

+512 GPUs
Defrag island — Hall C, pods 7–9
3 tenants movable at next checkpoint → contiguous NVL domain restored.
+288 GPUs
Inventory drift — 36 nodes “in repair”, actually healthy
BMC + burn-in evidence attached. One-click return to sellable pool.
auto
Observability rollup
Golden-signal SLOs per pod; 3 pods burning error budget — throttle recommendation issued.
Use case Reference model: 4% of a 48k fleet stranded at any time. Recovering it ≈ 1,900 GPUs ≈ $29–58M/yr at posted H100→B200 on-demand rates, 70% utilization.
LAYER 05 Flux Flexible Compute — software-defined power, power as a control input

Power is the scarcest input in the AI buildout — a GB300 NVL72 rack draws 135–150 kW, so a 512-rack hall is a ~65 MW grid asset. Flux treats it as schedulable: price and carbon signals in, checkpoint-aware power envelopes out.

Day-ahead price vs shiftable load · today

$120 $60 $0 $118 peak 00:00 14:00–17:00 24:00
day-ahead $/MWh · shaded = recommended curtailment window

Software-defined power plan · Hall B

WindowEnvelopeMechanismSLA impact
00–14 h62 MWfull ratenone
14–17 h48 MWckpt + GPU freq cap on preemptible tiernone
17–24 h62 MWcatch-up burst, deferred jobs firstnone
today
Projected: $5k energy avoided · 4CP position preserved
ERCOT peak avoidance is worth $40–67k/MW-yr — $2.5–4M/yr on this hall. Zero SLA-bearing jobs touched.
Use case Emerald AI showed a 25% power cut for 3 h w/o service disruption (Nature Energy, 2026); Google runs 1 GW of DC demand response. ERCOT 4CP avoidance alone ≈ $2.5–4M/yr on a 62 MW hall — before energy arbitrage.

Why Decision Minds

01 · LAND
Fleet data plane engagement
We build the unified telemetry lake — DCIM, BMC, fabric, job scheduler — in 8–12 weeks. Pure consulting revenue; the customer keeps the data model.
02 · EXPAND
Decision modules
Predictive maintenance, fragmentation recovery, power optimization — deployed as Yantra modules on the data plane, licensed per GPU under management.
03 · COMPOUND
Co-design flywheel
Cross-fleet failure fingerprints (anonymized) become the industry benchmark — the data moat neither the ODM nor the hyperscaler has.

Decision Minds has spent 15 years building enterprise data platforms. The AI factory is the largest new data-platform problem in the world: a 50k-GPU fleet emits more telemetry than most banks. Same discipline, new physics.