The most useful AI tools are the ones that work with your own data: contracts, customer records, support tickets, financial reports. That's also where the hardest question comes up. Where does that data go when you send it to a model, who can see it and how long is it kept?
There's no single answer, because there are several ways to use a model, and each one handles data differently. Here are the main options, from the simplest to the most controlled.
Consumer apps vs business plans vs APIs
The first distinction is between the chat apps people use personally and the business products and APIs that companies use to build tools.
- Free and personal chat apps may use conversations to improve models, depending on the provider and the user's settings. Staff pasting company data into personal accounts is often the biggest real risk.
- Business and enterprise plans of the major chat products generally don't train on your data by default, and add admin controls, single sign-on and data retention settings.
- APIs from the major providers generally don't use your inputs and outputs for training by default, but they may keep data for a limited period for abuse monitoring.
Policies differ between providers and change over time, so read the current data usage terms of the exact product you use, not a summary from a blog post (including this one).
Zero data retention and no-training terms
For sensitive work, many providers offer stronger terms for business customers. These usually take the form of zero data retention, where inputs and outputs aren't stored after the response is returned, or a contractual commitment not to train on your data.
These options are often available only on request, on certain plans, or for certain features. Some features, like stored files, conversation history or batch processing, may need data to be kept by design. Check which features you plan to use are covered.
Also look for a data processing agreement (DPA), the regions where data is processed, and the provider's list of subprocessors. If you handle personal data under privacy laws like GDPR, you'll need these anyway.
Models hosted in your cloud account
The large cloud platforms (AWS, Google Cloud and Azure) offer access to leading models inside your own cloud account. Requests stay within that cloud provider's infrastructure and are covered by the same agreements, security controls and region choices as the rest of your cloud setup.
This is often the easiest route for companies that already run on one of these clouds and have security reviews built around it. Model availability, features and release timing can differ from using the model provider directly, so confirm the model you need is offered in your region.
Self-hosted open models
Open-weight models can run on servers you control, either your own hardware or rented GPU machines. Data never leaves your environment, which makes this the strongest option for data control.
- You get full control over where data lives, how long it's kept and who can access it.
- You take on the work: setting up GPU servers, scaling, updates, monitoring and security.
- Open models have improved a lot, but the largest closed models are still usually stronger at the hardest tasks. Test on your actual use case.
- Costs shift from per-token pricing to paying for servers, which is cheaper only at steady, high volume.
Redaction: send less in the first place
Whatever option you choose, the safest data is the data you never send. Often the model doesn't need names, emails or account numbers to do its job.
- Identify which fields are sensitive: names, contact details, ID numbers, financial and health information.
- Replace them with placeholders before sending, like [CUSTOMER_1] or [EMAIL_2].
- Keep the mapping on your side and put the real values back into the response afterwards.
- Test the redaction on real samples. Automated detection misses things, especially in free text and other languages.
Original: "Refund the customer at buyer@example.com, phone 555-0142, for order 48213."
Sent: "Refund the customer at [EMAIL_1], phone [PHONE_1], for order [ORDER_1]."Choosing the right option
Match the option to how sensitive the data is, not to the most cautious choice for everything.
- Public or low-risk data: a business API with standard terms is usually fine.
- Customer and internal data: a business API with no-training terms and limited retention, or a model in your own cloud account, plus redaction where practical.
- Highly regulated or contractually restricted data: zero-retention terms, a cloud-hosted model with strict controls, or a self-hosted model. Involve whoever owns compliance early.
Two things apply in every case. Give staff an approved tool, so they don't paste company data into personal accounts. And log what's sent and by whom, so you can answer questions later.
Deeraf builds AI internal tools with the data path chosen to fit your requirements, and a Tech Check can review how an existing AI feature handles sensitive data.