Glossary
On-premises LLM
A large language model deployed on infrastructure the organisation controls, so data is processed inside its own security perimeter rather than a vendor cloud.
An on-premises LLM is a large language model that runs on infrastructure operated and controlled by the organisation using it — in its own data centre, in a private rack, or in a dedicated private-cloud environment — rather than being consumed as a hosted API from a third-party provider. Prompts, retrieved documents and generated outputs are processed within the organisation’s own security perimeter, so sensitive data does not have to be sent to an external service to obtain the model’s capability.
The main driver for on-premises deployment is control: over where data is processed and stored, over who can access it, and over the audit evidence a risk or compliance function can produce. It is the practical answer for organisations whose data-handling obligations, contractual confidentiality or regulatory regime make sending regulated material to a public AI service untenable. Open-weight models make this feasible, because the model can be downloaded and run locally rather than only being reachable through a vendor’s cloud.
On-premises deployment is not the same as air-gapping. An on-premises system may still have controlled outbound connectivity for updates or supporting services; an air-gapped system removes that connectivity entirely. On-premises also transfers operational responsibility — capacity planning, hardware, patching and monitoring — to the organisation, which is a genuine cost to weigh against the control it provides.
Discuss a secure AI deployment
We help organisations adopt AI inside their own security and compliance perimeter — vendor-neutral, and designed around the constraints you actually operate under.
Get in touch