- Provider constraints become architecture constraints.
- Sustained workloads can change the cost equation.
Compute infrastructure
Who operates the runtime?
Model infrastructure
Who operates the model?
Model ownership
Who owns the model?
Compute infrastructure
Who operates the runtime?
Low operational burden and rapid scaling, with less control over runtime shape and platform behaviour.
- State, data placement and failure domains become design concerns.
- Platform-engineering capability becomes part of the cost.
Strong scheduling and workload control where scale or complexity justify one of the highest-operational-burden options.
- Complexity can become ossified into machine state.
- IaC and reproducible builds are needed to avoid pets.
- Systems-administration skill remains material.
Flexible and familiar, but the OS, packages, kernels, drivers and machine lifecycle sit inside the operating responsibility.
Model infrastructure
Who operates the model?
Cloud-managed services impose bandwidth, instance and latency constraints, with cost implications that vary by workload.
Cost is workload-specific and requires modelling to optimise.
Managed-service limits can become performance or availability constraints as demand grows.
Model ownership
Who owns the model?
Frontier capability is accessible. Many AI deployments should start here.
- Latency and service quality are externally determined.
- The service is vulnerable to reduced capability, withdrawal or price increase.
Latency, data residency and sovereignty become directly controllable.
- Cost becomes sensitive to utilisation.
- You may not have access to frontier capability.
Adapt or build the model only where prompting, retrieval and model selection cannot meet the required behaviour, performance or economics.
- Training data, evaluation and model supply chain become first-class architecture concerns.
- Cost, specialist capability and governance burden rise sharply.