Deployment
The quiet return of on-premise deployment
Data residency, confidentiality, and cost at volume are pushing a meaningful share of inference back inside the perimeter — and the tooling has finally caught up.
7 min read
For several years the default answer to where a model runs was somebody else's infrastructure. For a large class of work that remains correct, and the operational simplicity is genuine.
But three pressures push the other way. Regulated and sovereign environments often cannot send the data out at all. Confidentiality obligations sometimes make an external API call a contractual problem regardless of the provider's assurances. And at sustained high volume, the arithmetic on hosted inference stops being obviously favourable.
What has changed is that the open-weight model landscape is now good enough that this is a real choice rather than a sacrifice. The engineering work is no longer exotic: hardware sizing, quantization, throughput planning, and the same evaluation discipline you would apply anywhere else.
The decision should be made on the specific workload rather than as a matter of policy. Some things belong inside the perimeter, some do not, and a sensible architecture supports both without treating either as the exception.