We’ve seen it before in big business. The formula for Coca-Cola is kept in a vault in Atlanta, reportedly accessible to only a handful of people, none of whom are allowed to travel together in case something happens to the plane.
Similarly, KFC keeps its famous eleven herbs and spices split across two manufacturers, so neither one has the full recipe. These stories get told and retold because everyone immediately understands the stakes: in these cases, the formula is the business, and losing control of it is an existential threat.
Frontier AI models can be a vastly larger investment than a soft drink recipe, often hundreds of millions of dollars in compute and years of research, and yet the industry's instinct for protecting them barely resembles how Coke protects its soda formula.
Model weights get loaded into rented GPU memory on infrastructure the model owner doesn't control, where administrators with sufficient privileges can, in principle, read them directly.
There's no vault, only a service agreement.
This setup is increasingly unacceptable. Model builders haven’t historically had a technical way to enforce the same kind of control over their weights that a Coke or KFC take for granted over a recipe, which means the decision to deploy a model on someone else's infrastructure has always involved an act of trust that the rest of the business world would find uncomfortably casual for an asset of this value.
This guide is for those who build models and must decide where those models are allowed to run, ultimately closing that gap with something more durable than trust.
Examining the Threat Model You're Defending Against
Before getting into defenses, you first need to know what you're defending against. Oftentimes, the phrase "model theft" gets used as a catch-all for several distinct threats that require different countermeasures.
- Direct weight extraction: The most severe threat and the one this guide primarily focuses on. It happens when someone with access to the infrastructure your model runs on reads the model's weights directly from memory or storage. It doesn't require any machine learning sophistication on the attacker's part, only system-level access, which is a much lower bar than building a competing model from scratch. A disgruntled employee at a hosting provider, a compromised administrator credential, or an attacker who with the right privileges on a shared system are all capable of executing this kind of extraction if the infrastructure doesn't actively prevent it.
- Query-based extraction: This is also called model distillation or model stealing, and it works differently. An attacker with no infrastructure access submits a large volume of queries to your model's public API and uses the input-output pairs to train a new model that aims to approximate your model's behavior. This doesn't require breaching anything. It only requires patience and API access.
- Side-channel leakage: A more subtle threat where information about the model leaks through indirect signals: timing variations in response generation, power consumption patterns, or memory access patterns that can reveal architectural details even without direct access to the weights.
These three threats require different defenses, and a common mistake model builders make is assuming that solving one solves all three. Spoiler alert: it doesn't.
Hardware-enforced isolation is a strong, often decisive answer to direct weight extraction. It does very little against query-based distillation, which has to be addressed through API governance, rate limiting and usage monitoring.
Meanwhile, side-channel risks call for attention to specific hardware implementation and are an active area of confidential computing research.
A complete IP protection strategy has to account for all three, although this guide concentrates on the first, because it's the threat that hardware-based confidential computing is specifically designed to eliminate.
Conventional Infrastructure Fails the Direct Extraction Test
When your model is deployed for inference, its weights have to be loaded into memory, system RAM for some operations, and GPU memory for actual computations, so that the processor can use them to compute outputs.
For that computation to happen, the weights have to exist in a form the processor can read directly: plaintext. In other words, the moment your model starts doing the thing it was built to do, the protection that kept it safe at rest is gone.
On conventional infrastructure, that plaintext memory space is visible to several parties beyond your model's own process. The host operating system can read it, as can an administrator with root-level access to the physical or virtual machine.
If the deployment is on shared infrastructure, in principle, certain vulnerabilities could expose memory across tenant boundaries (though well-implemented virtualization makes this difficult).
The fundamental problem is that none of these access paths requires “breaking” your model's security in any meaningful sense. They're using the infrastructure exactly as designed. Root access is supposed to allow reading system memory; that's what root access means. The vulnerability is the inability to compute data that even a fully privileged operator cannot read.
Here’s the Architecture That Closes the Gap
Confidential computing addresses this by changing what "memory" means for the duration of the computation. Instead of plaintext memory that any sufficiently privileged process can read, your model runs inside a trusted execution environment (TEE), a hardware-enforced boundary that keeps the model and data confidential while in use.
The TEE is physically rooted in the processor itself. Modern confidential computing CPUs, like those supporting Intel TDX or AMD SEV-SNP, include dedicated hardware for memory encryption that operates below the operating system layer. This means the OS, the hypervisor, and any administrator interacting with the system see only encrypted bytes when they attempt to inspect that memory region.
For AI workloads, this protection has to extend to the GPU, which is where the architecture gets more specific to model deployment than general confidential computing use cases. Your model's actual computation happens on the GPU, not the CPU.
NVIDIA's Hopper and Blackwell architectures are the first GPU generations to support native confidential computing, extending the same memory encryption and isolation principles to the GPU.
Without this, you could have a perfectly secure CPU enclave that hands your model's weights off to a GPU operating in a completely conventional, unprotected memory space, which would ultimately defeat the purpose.
This is where composite attestation comes in for model owners. Attestation is the process by which the hardware cryptographically proves that it’s genuine and unmodified before any sensitive material is released to it.
A composite attestation architecture verifies both the CPU TEE and the GPU TEE, which is the standard needed for truly secure AI inference. If your deployment only attests to the CPU side, your model weights and the actual computation running on them on the GPU are outside the verified boundary, and you have zero assurance about what's happening to your model where it matters most.
Inside the Vault: How the Deployment Actually Works
Translating this architecture into a real-world deployment workflow involves a sequence that should be transparent to a model owner when evaluating a hosting environment, neocloud, or AI token factory partner.
Your model is encrypted before it ever leaves your environment, using keys you control or those managed through a key management system with strict access policies. The encrypted model artifact can then be distributed to the deployment infrastructure without exposing the weights in transit or at rest on the destination system, since the encryption accompanies the model. That’s basic enough.
But before your model is decrypted anywhere, the receiving environment must pass attestation. The hardware generates a cryptographic report describing its identity, firmware state, and the exact code that's about to run.
The report is then verified against the expected values for a legitimate, unmodified deployment environment. Only once that verification succeeds does the process continue.
Encryption keys for your model are released only after successful attestation, and they're released directly into the verified enclave, never into general system memory where an administrator could intercept them.
Attestation-gated key release should be a hard requirement: if the attestation fails for any reason, such as a hardware mismatch, an unexpected firmware version, or unauthorized software, the keys simply won’t be released, and your model won’t be decrypted on that system.
Inside the verified enclave, your model is decrypted and runs inference exactly as it would on any other infrastructure, with the GPU performing the same computations at close to native speed. Since confidential computing is handled largely in hardware, it helps limit performance degradation.
It doesn’t stop there: throughout the deployment operation, the environment is periodically monitored for integrity. If anything changes unexpectedly, like a firmware update that wasn't approved, a software modification, or any other deviation from the attested state, the system can trigger re-attestation or terminate the workload immediately, rather than continuing to run your model in a now-unverified environment.
How to Evaluate Your Deployment Environment
For model builders and AI engineering leads assessing whether a given hosting environment, AI factory, or cloud platform actually protects your IP, there's a specific set of questions that can help you separate genuine protection from marketing language:
- Does the GPU get attested, not just the CPU? This is the most common gap. Plenty of confidential computing deployments protect CPU environments thoroughly and say nothing meaningful about the GPU, where your model's actual computation happens.
- Is attestation a hard gate on key release, or an advisory signal? Some architectures perform attestation and log the result without making it a strict precondition for decryption. If keys are released regardless of the attestation outcome, then attestation is just a monitoring feature, not a security control.
- What happens during disconnected operations? If your deployment scenario involves air-gapped or sovereign environments without reliable connectivity to external attestation services, confirm the platform supports local verification using pre-seeded reference measurements, rather than assuming a live connection that may not exist in your actual deployment.
- Can the protection be independently verified, or do you have to take the operator's word for it? A genuine attestation architecture produces a cryptographic report that you and your security team can verify against known good values. If the assurance you're given is a description of the architecture rather than a verifiable artifact, you're relying on trust rather than proof.
- Finally, what's the performance overhead? Confidential computing can introduce some performance overhead, but the impact depends on the implementation and workload. NVIDIA’s benchmarks on Blackwell Ultra GPUs show Confidential Computing delivering up to 98% of the performance of the same workload without CC enabled, with minimal throughput and time-per-output-token overhead across the tested configurations. Attestation typically occurs once at startup, so it does not add latency to individual inference requests. When evaluating a solution, look for benchmark data on workloads that resemble yours, including model size, concurrency, input/output sequence lengths, and inference framework.
A Standard Worth Building Toward
The deployment question that can stall projects now has a real answer. It's not a contract or a promise about administrator behavior. It's a hardware architecture that makes certain types of access technically impossible, regardless of privilege level, and that produces cryptographic proof of its integrity and doesn't depend on anyone's word.
For model builders, the choice of deployment infrastructure has become a meaningful technical decision. The model you've built, however much investment it represents, is only as protected as the environment you choose to run it in. Evaluating that environment with the same rigor you'd apply to your model's architecture is the standard worth building toward.
It's increasingly the standard that enterprise customers and partners are going to expect you to meet before they deploy your model on infrastructure that touches their sensitive data.
After all, the same architecture that protects your weights (and IP) also protects the data that flows through your model.

