A Decision Minds product concept · AI & HPC infrastructure
The decision-intelligence layer for AI factories. GPU clouds run five layers of stack on fifty tools and none of them talk. Yantra unifies the telemetry — from silicon to substation — and turns it into decisions: which node fails next, which GPUs are stranded, when to shed power.
Every burn-in, every validation run, every RMA becomes structured data. Forge closes the loop between what fails in the fleet and what the ODM builds next.
| Signature | 30 d | Trend | Attribution |
|---|---|---|---|
| HBM ECC double-bit | 214 | ▼ 18% | vendor lot 24-C · RMA batch open |
| NVLink CRC storm | 96 | ▲ 41% | switch fw 2.4.1 · rollback queued |
| VRM thermal drift | 61 | ▲ 12% | E1.S custom node · ECN filed |
| PCIe retimer dropout | 23 | ▼ 6% | cable batch replaced wk 28 |
One health score per node, built from DCIM, BMC, and job telemetry. Failure prediction drains nodes before the job dies, not after.
| Node | Signal | Est. TTF | Action |
|---|---|---|---|
| b07-r112-n3 | ECC error slope 4.2×/day | 31 h | drain now |
| b02-r018-n7 | HBM temp drift +6 °C/wk | 58 h | drain at ckpt |
| a11-r201-n2 | NVLink retrain events ×9 | 64 h | drain at ckpt |
| c04-r077-n5 | PSU ripple anomaly | 6 d | watch |
Every incident becomes training data. Sentinel correlates job telemetry, fabric events, and change history into a ranked root cause before the bridge call starts.
A live graph of every GPU: where it is, what it's wired to, who holds it, and whether it can actually be sold. Fragmentation is revenue lying on the floor.
Power is the scarcest input in the AI buildout — a GB300 NVL72 rack draws 135–150 kW, so a 512-rack hall is a ~65 MW grid asset. Flux treats it as schedulable: price and carbon signals in, checkpoint-aware power envelopes out.
| Window | Envelope | Mechanism | SLA impact |
|---|---|---|---|
| 00–14 h | 62 MW | full rate | none |
| 14–17 h | 48 MW | ckpt + GPU freq cap on preemptible tier | none |
| 17–24 h | 62 MW | catch-up burst, deferred jobs first | none |
Decision Minds has spent 15 years building enterprise data platforms. The AI factory is the largest new data-platform problem in the world: a 50k-GPU fleet emits more telemetry than most banks. Same discipline, new physics.