Fortanix Confidential AI Protects Proprietary Model IP and Data for Secure AI Inference in Enterprise AI Factories.

Learn More

AI Inference

What is the SOC 2 requirement for AI inference?

SOC 2 doesn’t have specific technical requirements for AI inference as a distinct category. But it defines trust service criteria, and AI inference workloads should account for those criteria in the same way any other data processing service would. 

When it comes to AI inference, you’ll find the trust service criteria of most importance to be Security (covering access controls and the like), Confidentiality for the protection of sensitive data over its entire lifecycle, and Availability for system uptime and performance. 

Take a token factory undergoing a SOC 2 AI compliance audit. The auditor is there to see if the controls put in place to satisfy those criteria are soundly designed; a Type II report will go further and tell you if they were effective over the course of the audit period. In essence, SOC 2 is about holding an organization to what it says it does with data and verifying that its controls live up to that description. But don’t expect cryptographic evidence from a SOC 2 audit that your data was shielded at the hardware level during processing. For that reason, the best way to handle compliance is to have the token factory operator’s SOC 2 Type II certification backed by confidential computing infrastructure for hardware-level protection, which is something the SOC 2 can point to as a control that has been put in place.

How do I ensure my training data stays private during inference?

This question usually comes from teams that have fine-tuned a model on proprietary or sensitive data and are now trying to figure out whether that investment is safe to deploy. 

There are multiple layers to this risk. First, there's the data itself: if you fine-tuned on customer records, internal documents or licensed datasets you don't have rights to expose, that data is embedded in your model's weights in some form, and a model extraction event would expose it.  

There's also the inference-time risk: even with a model that's already trained, the prompts and context submitted at inference time may contain sensitive information that needs protection in its own right. 

Confidential computing addresses the first concern by running the fine-tuning process within a trusted execution environment, ensuring that proprietary datasets are never exposed outside the protected enclave during training.  

The resulting fine-tuned weights are encrypted and protected in the same way as a base model's weights. The training data informs the model, but deploying the model doesn't require re-exposing that data. 

Confidential computing for LLMs extends this protection to inference. When someone submits a query to your fine-tuned model, that query (along with the model's internal computations and the response it generates) stays inside the hardware-isolated enclave throughout the operation. 

Your investment, along with any proprietary value it represents, is protected throughout the training process, in the resulting weights, and in every inference operation that uses them. 

How is inference output protected?

The output a model generates can be as sensitive as the input that produced it, and sometimes more, since it’s synthesized insight rather than raw data. 

In a conventional inference pipeline, the output is generated in the same exposed memory space as everything else: visible to the infrastructure operator, the host OS, and anyone with sufficient system access at the moment it's produced, before it's ever encrypted for storage or transmission.  

Confidential model serving protects across the full request lifecycle: the input arrives encrypted, is decrypted only inside the verified enclave, the model processes it and generates a response entirely within that protected boundary, and the response is encrypted again before it exits. At no point in that sequence does the input, computation, or output exist in a form that's readable outside the hardware-enforced boundary. 

For model builders offering inference as a service, this means you're not creating a secondary exposure point for your own model's behavior. Outputs can reveal a great deal about how a model was trained and tuned, so protecting them with the same rigor you apply to the model weights themselves is part of complete IP protection.

Fortanix-logo

4.6

star-ratingsgartner-logo

As of January 2026

SOCISOPCI DSS CompliantFIPSGartner Logo

US

Europe

India

Singapore

4500 Great America Parkway, Ste. 270
Santa Clara, CA 95054

+1 408-214 - 4760|info@fortanix.com

High Tech Campus 5,
5656 AE Eindhoven, The Netherlands

+31850608282

UrbanVault 460,First Floor,C S TOWERS,17th Cross Rd, 4th Sector,HSR Layout, Bengaluru,Karnataka 560102

+91 080-41749241

T30 Cecil St. #19-08 Prudential Tower,Singapore 049712