If asked: "How many AI systems is all customer data flowing through, and who is keeping the logs for each of those systems?" could you answer within 5 minutes?

For most organizations that have deployed AI on a project-by-project basis a document lookup chatbot last quarter, a profile scoring tool this month, a call summarization assistant next week the answer is usually no. It is not due to a lack of caution. For every project deployed, the right questions were asked: "Is the data stored in Vietnam?" and "Is the provider reputable?" And every time, the answer was yes.

But "every tool is fine" does not mean "the entire system is controlled." Those are two different questions, and the gap between them is where risk silently accumulates: vector data from one tool, query logs from another, backups of a fine-tuned model from a third all scattered, with no one having the full picture, until there is an audit, a major client request, or an incident that forces an exact answer about where the data resides.

This is why data sovereignty in AI operations cannot be assessed at the individual tool level. It is an attribute of the overall architecture how the infrastructure, data, models, and application layers are designed to operate together under a consistent control mechanism, rather than a mere accumulation of isolated "seems fine" decisions.

This article shares an architectural perspective: the layers that make up an AI platform, who is responsible at each layer, and the questions infrastructure managers should ask to get that full picture before having to answer it in a much more urgent situation.

1. Architecture: The Deciding Factor for Data Sovereignty

Returning to the scenario above: if forced to answer right now, "Where is customer data residing within our entire AI system?", many organizations would realize the answer is not as simple as when evaluating each tool individually. This is because an AI system operates across multiple layers—compute infrastructure, data platform, model serving layer, and application layer—and each layer might be operated by different parties, in different jurisdictions, even if the primary infrastructure is located domestically.

An often-overlooked point: even if the data (data plane) is hosted domestically, if the control and monitoring layer management console, operational dashboards, metadata systems is operated by a foreign entity, the organization's actual control capacity remains limited. This is why evaluating data sovereignty requires looking at the entire architecture—every layer, every component rather than just asking a single question about location.

In other words: good architecture creates true sovereignty; infrastructure location is just one variable within that architecture. This is also why the following sections dive deep into each layer comprising an AI platform, giving infrastructure managers a more complete map when evaluating current systems or designing new ones.

2. The Layers Constituting an AI Platform

According to the standard stratification in the AI infrastructure industry, a fully operational AI platform consists of five technical layers stacked from the bottom up, plus a monitoring and governance layer that runs through all five:

  • Infrastructure Layer: Compute (GPU/CPU), physical storage, and network. This is the foundation providing processing and storage capacity for everything above it, determining performance, scalability, and resource location.
  • Data Layer: Where data is collected, stored (databases, data lakes), cleaned, and standardized before being fed into models. For AI systems with retrieval, this is also where vector data is stored.
  • Model Development Layer: Where models are selected, trained, or fine-tuned on the organization's data. This layer dictates who truly controls the model: whether the organization has access to the weights, the training pipeline, and the data used for fine-tuning, or if this entirely belongs to an external model provider.
  • Model Deployment Layer: Where the completed model is packaged and serves actual inference requests, typically via APIs or microservices.
  • Application Layer: Where the model is integrated into actual business systems: chatbots, credit scoring tools, document processing assistants. This is the layer end-users interact with directly.

greennode_ai_architecture_english_website_16x9.png

A crucial point to note: The model development layer and model deployment layer pose two different questions, though they are often lumped together during infrastructure evaluations. An organization might self-host model serving (inference) entirely domestically meeting location requirements at the deployment layer but if the underlying model is a product of a foreign provider where the organization has no access to weights or the fine-tuning process, the core "cognitive" capability of the AI system still depends entirely on external parties at the development layer. This is a form of data sovereignty leakage that is easily missed because it lies not in the question "where is the data?" but rather "who owns the model?".

3. The Singular Governance Layer: Consistent Control Across 5 Technical Layers

Unlike the five technical layers above, there is one layer that does not sit "inside" the architecture but runs across the entire system: the observability & governance layer. From a data sovereignty perspective, the three core components of this layer are:

  • Identity and Access Management (IAM): Who has access to which resources, at what level, and how those permissions are granted or revoked.

  • Encryption Key Management: Who holds the data encryption keys, and whether key management authority is segregated from data access permissions.

  • Logging and Monitoring: Whether every access action, configuration change, or data processing event is fully recorded and retrievable when needed.

When this governance layer is well-designed and uniformly applied across all five technical layers, the organization gains a system that is both easy to operate since all permissions and activities are clear and trackable and ready to answer how data is being processed at any time. This forms the foundation that enables smoother AI scaling, as the operations team doesn't have to rebuild the control mechanism every time a new AI application is deployed.

4. Components Retaining Data the Longest but Often Excluded from Evaluation

When evaluating an AI system, the most attention is usually given to the input data and the core model. However, there are four other components that also retain data sometimes for much longer periods yet are easily left out of the evaluation scope:

  • Logs: Query logs and system logs are often stored by default for months or years for operational purposes, but are rarely reviewed for sensitive content that might appear within them.

  • Backups: Data in backups usually exists longer than the original data and may reside in a different storage location, making it easy to overlook during data location audits.

  • Vector Databases: Data converted into embeddings for retrieval tasks (e.g., internal document lookup chatbots) can still contain information that could be inferred from the original content.

  • Fine-tuned Models: If a model is additionally trained on the organization's proprietary data, the model itself becomes a form of data storage, even if not visually apparent like a standard data file.

Including all four of these components in the infrastructure evaluation scope alongside the original data and core model helps organizations gain a more complete picture of where data truly "lives" across the entire system, thereby enabling the design of appropriate control mechanisms and retention lifecycles for each component.

5. Delineation of Responsibilities Between the Enterprise and the Provider

Under the new legal framework for Personal Data Protection (Law No. 91/2025/QH15 and Decree 356/2025/ND-CP effective January 1, 2026), the boundaries of legal responsibility between the Data Controller and Data Processor are clearly defined according to each technical service model.

No.Service Model UsedGreenNode's Legal RoleEnterprise's Responsibility (Data Controller)GreenNode's Responsibility (Data Processor / Sub-processor)
1Self-managed Cloud Infrastructure (vServer, vStorage, vDB, VKS)Data Processor  Manage OS, application configurations, access permissions, application-layer data encryption.Ensure safety of physical infrastructure, virtualization (hypervisor), network infrastructure, and hardware availability.
2Self-managed Cloud Infrastructure (vServer, vStorage, vDB, VKS)Data Processor  Define processing purposes, manage business data access permissions, approve retention policies.Operate managed platform, apply patches, monitor system security, and retain audit logs.
3AI Platform (Training, Fine-tuning, Inference)Data Processor  Manage input datasets, obtain data subject consent, configure AI pipelines.Protect AI computing environments, isolate tenants, ensure customer data is not arbitrarily used to train shared models.
4Dedicated AI Services (IDP, OCR, Smart Document AI)Data Processor  Control uploaded document content, ensure legal basis for processing sensitive personal data.Process and extract data accurately according to service directives, encrypt processing workflows, and delete temporary data post-processing.
5Service Supply Chain (GreenNode as Subcontractor)Sub-processor  Monitor the entire supply chain, establish DPA agreements with the primary Processor.Strictly process data under valid authorization from the primary Processor, comply with equivalent protection standards.
6Multi-region / Delivery Services (CDN, Multi-region, Backup)Data Processor  Approve data routing policies, conduct Cross-Border Data Transfer Impact Assessments (CTIA/DPIA Form 09) if stored outside VN.Securely store, backup, and transmit data strictly according to the geographic region configurations set by the customer.

Table: Matrix delineating Personal Data Protection responsibilities across 6 GreenNode service models.

Properly identifying which model you are in and how responsibilities are divided within that model helps infrastructure teams avoid a common pitfall: assuming the provider "takes full responsibility" when in reality, a critical portion (e.g., data usage purpose, application-tier access configuration) still rests with the enterprise.

6. How GreenNode Builds This Architecture: Three Deployment Models and Provider Portability

In reality, there is no single "standard" architecture suitable for every enterprise you must choose a deployment model aligned with current contracts, data scale, and readiness for transition. GreenNode currently supports three collaboration models, varying by infrastructure location and the legal entity responsible for operations:

Model 1: Data is stored and processed entirely in Vietnam; no data is transferred abroad.

Model 2: Data is processed by a regional legal entity, accompanied by appropriate cross-border data transfer impact assessments prior to deployment.

Model 3: A multi-layered model where a domestic partner acts as the primary focal point, and GreenNode participates in a technical support role under authorization.

The common thread across all three models is that provider portability (exit & portability) is built-in from the start. Enterprises can export data, migrate workloads, or change collaboration models as needs evolve, without having to rebuild the entire governance mechanism from scratch. This is one of the most practical values of designing a sovereign architecture from the beginning: the organization always retains the right to choose, rather than being locked into a single vendor.

7. The 9 Criteria for Evaluating a Sovereign AI Cloud and Checklist for Infrastructure Managers

To help IT Infrastructure Managers and Chief Information Security Officers (CISOs) establish a basis for assessing the actual capabilities of Cloud/AI providers, below is the Sovereign AI Cloud control criteria set comprising 9 core indicators:

  1. Storage Residency: Source data, backup data, and metadata are accurately stored and processed within the committed geographic scope.
  2. Operator Independence: The enterprise has full autonomy to operate workloads on the platform; the provider has no right to intervene in business data beyond the agreed scope.
  3. Encryption Key Control: Customers hold and manage encryption keys (KMS); data access permissions are completely segregated from key management authority.
  4. Infrastructure Control: Data Center locations are transparent; only properly authorized personnel in Vietnam have the authority to operate the system.
  5. Regulatory Jurisdiction: Data and infrastructure fully comply with the Vietnamese legal system and are not subject to interference by foreign jurisdictions.
  6. Verifiable Auditability: The system maintains detailed trail logs (who accessed, when, from where) ready for generating compliance audit reports.
  7. Exit & Portability: Provides standardized tools to extract data and models, along with secure data wiping mechanisms upon service termination.
  8. Governance Compliance: Meets international information security certifications and industry-specific regulations (ISO 27001, SOC 2, SBV Circular 09/12).
  9. No Foreign Transfer: No automatic transmission or storage of data to infrastructure zones outside Vietnam without formal written approval.

Banner on page (8).png

Checklist: 3 Questions Infrastructure Managers Must Clarify Before Signing a Contract

  • Verify the Control Plane: "Where is the service's control layer (console, metadata, logging) located, and which legal entity directly operates it?"
  • Responsibility Boundaries per Service Model: "To what extent is the provider responsible for securing models and access logs for AIaaS / Managed AI services?"
  • Exit Strategy: "When the contract needs to be terminated, how does the process of extracting all vector embeddings, log data, and model weights take place, and how is the commitment to Data Wiping verified?"

Conclusion

Data sovereignty in the AI era cannot be achieved by accumulating individual tool approval decisions. It requires a standardized architectural blueprintfrom physical infrastructure, data processing layers, and model training/serving layers, to an overarching governance framework.

The GreenNode Sovereign AI Cloud ecosystem is built to give enterprises the ability to accelerate AI deployment while maintaining substantial control over their data assets and ensuring full compliance with Vietnam's legal framework.

kien-truc-sovereign-ai-cloud (5).png