HACKER Q&A
📣 plainviewinstru

How do you handle uncertain GPU capacity needs months in advance?


If your product has variable or spiky compute requirements, how do you secure capacity when you know you may need a large GPU block in a few weeks or months, but are uncertain about the timing or quantity?

The apparent choices are to reserve capacity in advance and risk underutilizing it, or wait and accept price and availability risk in the on-demand market. I’m curious how inference providers, enterprises running fine-tunes or evals, and teams with batch or seasonal workloads handle this in practice.

In particular: - How far forward do you reserve capacity? - How frequently do you end up underusing reservations? - Can providers resize, defer, or release commitments?

I know that larger cloud providers like AWS and CoreWeave have flexixbility/credits, but if you're largely getting bare metal capacity from Neoclouds, how do you handle this?