Designing an AI Platform for Data That Cannot Leave.md
The architecture for running language models on infrastructure an organization controls, where the data boundary is decided before the model is.
The first question organizations ask when implementing generative AI is which model is safe. That is the wrong question, and answering it wastes months. An open-weight model can run inside a restricted network boundary or on a public cloud. The model weights are identical either way. The environment determines whether sensitive records can exist in the request at all.
Define the data boundary first
Architecture must follow data classification. The initial step is identifying what data exists, determining the sensitivity of each tier, and defining the processing requirements for that sensitivity. I designed a system that supports two hosting postures behind a single interface. Sensitive records stay strictly on-premises where the inference tier has no outbound internet access. Public content routes to a managed cloud for elastic scale. The application talks to one endpoint and does not care which environment answered.
Choosing a model before defining these boundaries forces you to retrofit policy onto an existing tool. That is an expensive mistake.
Enforce a single mandatory gateway
Every model request and response transits a single gateway. There are no exceptions and no alternate routes. This gateway executes the non-negotiable work: request guardrails, sensitive data redaction, audit logging, and providing a unified API surface.
Applications never communicate with the inference tier directly. An advisory control in a regulated environment is not a control. Forcing all traffic through a mandatory hop converts a written policy into a verifiable technical mechanism. Early designs routed generated output directly back to the application. That left output guardrails and audit trails covering only half the interaction. Routing the response back through the gateway corrected that asymmetry.
Execute access control before retrieval
When a system retrieves evidence to answer a prompt, the access filter must execute during assembly, not after generation. Once unauthorized data enters the context window, the security boundary is compromised. You cannot instruct a model to forget what it just read.
The filter must fail closed. If the user lacks clearance or the retrieved evidence scores below a relevance threshold, the system issues a specific refusal and directs the user to an authoritative source. Fabrication is treated as a systemic defect requiring a root cause analysis, not a prompt engineering issue. Permissions continuously resynchronize with the identity system to prevent access drift.
Architect for vendor independence
The inference tier runs on a standard API shape backed by open-weight models. Swapping models is a configuration change. No application code is written against a specific vendor SDK.
This prevents architectural lock-in. An organization that must rewrite applications to change model providers has acquired unquantifiable technical debt. The principle is identical to managing a storage backend: own the interface, rent the implementation, and maintain control over the integration seam. High availability is reserved for the inference tier where failures are expensive. Undifferentiated redundancy across the entire stack turns a prototype into an unapproved budget.
Separate the built from the designed
Honesty about system maturity is an architectural requirement. Request-side guardrails are actively running in the integration layer. Response-side controls, output validation, and citation enforcement are specified for the prototype phase. A system with request-side guardrails only can still generate incorrect outputs. Claiming the architecture is complete misrepresents the current risk posture.
If you are building an AI platform, classify the data and finalize the network boundary before evaluating a single model. That decision dictates whether the system is viable for actual organizational work, and it is the only decision that is difficult to reverse. Every subsequent choice regarding the model, the vector store, or the serving runtime is just a component behind an interface.