NeoX
BlogCapabilitiesRequest Audit
Engineering notes

Operating inference,
written down.

Notes on running open-weight models on your own GPUs — what the metrics actually mean, where the cost hides, and why the layer above the serving stack keeps ending up unowned. Written for the people who get paged.

Industry6 min
28 July 2026

Inference has no owner

Infrastructure owns the GPUs. ML owns the models. Applications own the product. The layer that turns one into the other has no operational owner, and it shows up in the incident review.

Read→
Technical7 min
28 July 2026

Cost per task, and why nobody can calculate it

Cost per task is the one inference metric that platform teams and finance both understand. Attributing it correctly is harder than it looks, and here is where the arithmetic breaks.

Read→
Technical6 min
28 July 2026

What effective GPU utilization actually measures

nvidia-smi showing 90% doesn't mean your GPUs are working. Here's the difference between device utilization, effective utilization, and the number that governs your inference bill.

Read→

Know your numbers.

Most teams cannot say what a task costs them, or how much of their GPU fleet is actually doing work. The Inference Expert Audit measures both on the stack you already run — read-only, no migration.

Schedule your audit
NVIDIA Inception Program member© 2026 NeoX. All rights reserved.
HomeSecurityContact