Think about the last question you asked an AI. Maybe you pasted a work email and asked for a summary. Maybe you described a symptom, a money problem, a client's name. The instant you hit send, that text stopped living only on your screen: it traveled across the internet to the server of a company you cannot see, in a country you probably did not choose, and it was stored in some log. The conversation felt private. The path the data took was not.
The short answer, to lead with it: what you write to an AI leaves your device and lands on a server owned by the company that runs it; what happens to that text afterward depends on the product, the plan you use, and the settings you have touched. In some cases your conversation can be used to train future models; in others, it cannot. The difference is no technical footnote: it is the line between a piece of data that evaporates and one that stays. And it almost always sits a couple of clicks away, in a menu nobody opens.
This article is about that line. Not to frighten you (the tool is useful and you are going to keep using it) but so that you use it knowing where the text you write ends up.
01 · the journeyWhat happens to my data when I use ChatGPT and other AI
It helps to picture the prompt as a letter, not a thought. A thought stays inside your head; a letter has a sender, a recipient, and, above all, copies. When you type into a chatbot, your text passes through at least three stations.
First: transit. The message leaves your browser or app encrypted and reaches the provider's server. That encryption protects the data along the way (nobody spies on it mid route) but says nothing about what happens once it arrives. It is the difference between a sealed envelope and what the recipient does with the letter once it is open.
Second: processing. The model reads your text turned into tokens, generates a reply, and hands it back. Nothing odd so far: it is the job you asked for.
Third, and the one that matters: retention. The provider keeps the conversation. It almost always does, for a while at least, for legitimate reasons (showing you your history, detecting abuse, obeying the law) and sometimes also to train future versions of the model with what people write. This is where your letter can turn into study material.
Encryption protects the data's journey. It does not decide what happens once the data arrives.
That third station changes with the product. And the biggest change is not between one brand and another, but between the type of account you use.
02 · the modeThe decisive difference: free account, settings, and business plans
If there is a single idea worth taking away, it is this: not every use of the same AI treats your data the same way. The same ChatGPT, the same assistant, behaves differently depending on how you walk in.
In the consumer version (the free or personal account) the provider usually reserves the right to use your conversations to improve its models, unless you switch it off. OpenAI, for instance, offers a setting in ChatGPT to exclude your content from training; it comes turned on against your interest by default, so you have to go looking for it [1]. The option exists. The trouble is that almost nobody knows it does.
Under Settings → Data controls there is a setting along the lines of "Improve the model for everyone". Turning it off tells the provider not to use your new chats for training. It does not delete what you already sent, but it cuts off the flow going forward. The exact menu names change with each version; the concept (data controls) is stable.
At the other extreme are the business plans and paid APIs. Here the logic flips by design: serious providers commit, by contract, to not train their models on the data that companies send through the API or through Team/Enterprise plans, barring explicit permission. OpenAI states this for its business platform [2]. The reason is commercial before it is ethical: no company would fold an AI into its operations if it suspected that its secrets were feeding the model its competitor will use tomorrow.
The right question is not "is this AI safe?" but "which mode am I using it in?".
| Mode of use | Can it train the model on your text? | Main control |
|---|---|---|
| Free / personal account | Often yes, by default | A "data controls" setting you can switch off |
| "Temporary" chats / no history | Normally not used for training | Turn on temporary mode before you type |
| Team / Enterprise plan | No, by product policy | The provider guarantees it by contract |
| Paid API (integration) | No, barring explicit permission | Data processing agreement (DPA) |
There is a nuance that confuses many people. That a provider does not train on your data does not mean it does not store it. Almost all of them retain conversations for a while (days, weeks) to detect abuse and meet legal requirements, even when they promise not to train on them. Retention and training are two separate taps. You can close one and leave the other open without noticing.
The philosopher Helen Nissenbaum put forward an idea that lights all of this up: privacy is not secrecy, it is contextual integrity [3]. A piece of data is not "private" or "public" in the abstract; what matters is whether it flows according to the norms of the context in which it was shared. Telling a symptom to your doctor is appropriate; that same symptom ending up in the training material of a global model is not. The data did not change. The context changed without anyone warning you. That mismatch, not the lost secret, is what we feel as a violation.
03 · the practiceHow to protect what you write, without giving it up
The good news is that almost all of the risk is governed by simple habits, not by technical know how. Four are enough.
One: separate the data from the task. Most of the time you do not need to hand over the real data to get the help. If you want the AI to polish an email, swap the client's name for "Client A" and the exact figure for an approximate one. The model writes just as well with fictional data. The usefulness does not live in the sensitive data; it lives in the structure of the problem.
Two: adjust the controls once. Go into the data settings of the tool you use every day and decide, consciously, whether you want your chats to feed training. It is one menu, five minutes, and it is done. Many tools also have a "temporary chat" mode that keeps no history: ideal for one off queries you would rather leave no trace.
Three: choose the mode by sensitivity. For the trivial, the personal account is fine. For anything involving client data, health, finances, or your business secrets, use a business plan or an API integration, where the no training clause comes guaranteed by contract. This is not snobbery: it is putting the sensitive data in the mode that protects it by design.
Four: remember that what is sent, is sent. Turning off training today does not rescue what you sent yesterday. Treat each prompt as an irreversible decision at the moment of sending it, not afterward. Prudence goes in front of the button, not behind it.
Before you paste something into an AI, ask yourself: "would it bother me if this text stayed stored on a company's server?". If the answer is yes, either anonymize it, or switch modes, or do not paste it.
None of this asks you to distrust the tool or give it up. It asks you to understand that the comfort of writing to a machine in natural language has a flip side: that natural language is also information, and information travels. Whoever knows the route decides how much of themselves to put in the envelope.
The next time you write to an AI, you will not see an innocent text box. You will see a letter about to leave your home. And you will know, at last, where it is going.
Fuentes
- OpenAI. How your data is used to improve model performance / Data controls in ChatGPT. OpenAI help center: help.openai.com. (Describes the option to exclude user content from model training.)
- OpenAI. Enterprise privacy at OpenAI. openai.com/enterprise-privacy. (States that data sent through the API and the business plans is not used to train the models by default.)
- Nissenbaum, H. (2004). Privacy as Contextual Integrity. Washington Law Review, vol. 79, n.º 1, pp. 101-139. (Framework of contextual integrity as a definition of privacy.)