AI Infrastructure

Accelerated compute, arranged and governed — with the unit economics attached.

GPU capacity is now one of the largest discretionary lines in the technology budget, and it is being purchased with less rigour than a mid-size SaaS renewal. Aminexus works both sides of it: standing supply relationships across hyperscalers and specialist AI clouds, and the workload analysis that decides how much capacity you should be holding. We place it, stand up the environment around it, then report what it costs per training run and per million tokens.

  • AWS Partner Network
  • Google Cloud Partner
  • Claude Partner Network — Anthropic

The problem

Two very different markets, bought as though they were one.

H200 has become a commodity — tracked across more than twenty-five providers, with real price competition and wide regional supply. Blackwell Ultra is the opposite: scarce, with almost no public on-demand pricing, moving through reserved allocation. A single procurement approach applied to both produces the worst of each: overpaying where you had leverage, and missing out where you needed to move early.

01

Committed too early

A multi-year reservation signed before the workload is characterised, then carried at a fraction of its capacity.

02

Wrong silicon

Frontier-tier GPUs serving inference that a prior generation would handle at materially lower cost per token.

03

No unit economics

One monthly invoice, no cost per run, per model or per team — so no one can say whether it was worth it.

04

Single-source exposure

One provider, one region, one contract, and no alternative quoted when the renewal lands on the desk.

What we do

Four responsibilities, held by one team.

Workload characterisation

Model size, context length, batch behaviour, memory ceiling and interconnect sensitivity. We establish whether you are memory-bound, compute-bound or network-bound before anyone selects a SKU, because that single finding drives every decision after it.

Supply and placement

We hold relationships across AWS, Azure and Google Cloud alongside specialist AI clouds, so a requirement can be placed against capacity already available to us rather than worked up from a public price list. Reserved terms typically span one to thirty-six months; the position against on-demand is negotiated, not published.

Landing zone and operations

Capacity is not a platform. Identity, networking, storage throughput, job scheduling, queueing and observability decide whether the cluster is used or sits idle. We build that layer and hand it over documented.

Cost governance

Utilisation against commitment, cost per training run and per million tokens, chargeback to the team that generated the spend, and a renewal position prepared before the renewal date rather than after it.

The fleet

A mixed fleet beats a single-SKU strategy.

Current-generation capacity for production serving, frontier-tier reserved for the training that genuinely requires it. Specifications below are NVIDIA's published figures.

GPU Memory Bandwidth Best fit Market
NVIDIA H200 141 GB HBM3e 4.8 TB/s Production inference, long-context serving and fine-tuning. Around 1.4× H100 on training and up to 1.8× on inference for memory-bound work. Broad supply
NVIDIA B200 192 GB HBM3e ~8 TB/s Blackwell-generation training and high-throughput inference using FP4 and FP8 paths. Contracted
NVIDIA B300 288 GB HBM3e 8 TB/s Blackwell Ultra. Reasoning-model training and inference where per-GPU memory is the binding constraint. Reserved only
GB300 NVL72 ~20 TB per rack 130 TB/s NVLink 72 Blackwell Ultra GPUs and 36 Grace CPUs in a single NVLink domain, around 1.1 exaFLOPS FP4. Roughly 120 kW per rack — a facilities decision as much as a compute one. Allocation

Market column reflects what can realistically be placed through current supply arrangements rather than a published price list. Supply, region and terms are confirmed against the live position at the point of engagement.

Commercial models

The contract shape decides the economics.

Most organisations need a blend. The judgement is in the proportion: enough committed capacity to earn the discount, enough elastic capacity to absorb peaks without paying for them year-round.

On-demand

Evaluation and burst

  • No commitment, highest unit rate
  • Right for proof-of-concept work and unpredictable spikes
  • Expensive as a steady state

Dedicated cluster

Single-tenant, configured to you

  • Control over topology, interconnect and isolation
  • For sustained frontier-scale training
  • Power, cooling and network enter the decision early

How an engagement runs

Four phases. Stop after any of them.

Each phase produces an artefact you own and could hand to another party.

1. Characterise

Profile the workloads, find the binding constraint. Output: a sizing model with its assumptions written down and open to challenge.

2. Place

Match the requirement against what is currently available to us and normalise the options onto comparable terms. Output: a written position — what can be secured, on what commitment, at what unit cost — that your procurement team can defend.

3. Land

Stand up access, networking, storage, scheduling and observability. Output: a cluster your engineers can submit work to on day one.

4. Govern

Utilisation, unit economics, chargeback and a renewal position. Output: a standing review your finance team can actually read.

Questions

What clients ask before committing.

Do you provide the capacity, or advise on it?

Both, and we tell you which one applies before any work starts. We hold supply relationships on one side and run the workload analysis on the other, so in most engagements the capacity can be arranged through us rather than leaving you to approach the market alone. Where contracting directly is the better outcome for you, we will say so. Either way the commercial structure, including how we are paid, is on the table up front.

We only run inference. Do we need Blackwell Ultra?

Usually not. Frontier-tier silicon earns its premium on training and on reasoning workloads where per-GPU memory is the constraint. A large share of production inference is more economic on H200. Cost per million tokens should decide it, not the generation number.

How quickly can capacity be secured?

Current-generation capacity in a common region can move quickly, and we can usually tell you what is available against your requirement in the first conversation. Blackwell Ultra and NVL72-class allocations are planned in quarters. You will know which situation you are in before you commit to a delivery date.

Can you work with capacity we already hold?

Yes, and it is a common starting point. Where a commitment is signed and under-used, the work is recovering value: raising utilisation, restructuring how it is shared across teams, and building a defensible position for renegotiation.

Do you handle the facilities side?

For hosted capacity the provider does. For dedicated or on-premises deployments it becomes central — an NVL72-class rack draws on the order of 120 kW, with cooling requirements most existing floor space cannot meet. We raise it early, because finding out late is expensive.

Start with the workload, not the purchase order.

A scoping conversation takes thirty minutes and does not require an approved budget. If the answer is that you should not be buying GPU capacity yet, we will say so.

Book a scoping call