🔐 Beginner ~11 min Slow · 60d

What You Must Never Paste Into a Chatbot

The exact training toggles and retention windows OpenAI, Anthropic and Google publish, and the data categories to keep out of every assistant.

The AI Dude · Published July 31, 2026 · Verified July 31, 2026

Every piece of advice on this subject repeats the same two sentences: be careful with sensitive data, and check the vendor's policy. Neither tells you what to do on a Tuesday morning with a customer complaint in one hand and a chat window open in the other.

So here are the three vendors' actual published positions, checked on 31 July 2026, with the exact setting names, the retention windows in days and years, and the places their own documents go quiet.

The three assistants ship with three different defaults

ChatGPT on a personal plan trains on your conversations unless you stop it. OpenAI's policy page groups them explicitly, naming ChatGPT, Sora and Operator among the services for individuals where training is on by default. The control is documented in OpenAI's Data Controls FAQ, and the path on the web is: click your profile icon, select Settings, go to Data Controls, and turn off the toggle labelled "Improve the model for everyone". One toggle, applied to the whole account regardless of device.

Claude frames it as permission you grant rather than one you withdraw. Anthropic's privacy article on model training uses conditional language throughout: "If you allow us to use your chats or coding sessions to improve Claude, we may retain your data in a de-identified format for up to 5 years". The control is named Model Improvement and lives at claude.ai/settings/data-privacy-controls. The same page carves out one mode explicitly: "Your Incognito chats are not used to improve Claude, even if you have enabled Model Improvement".

Five years is the longest retention window any of the three publishes for training data, and it is worth sitting with. A conversation you have this week, if you have enabled that setting, can be held in de-identified form until 2031.

Gemini ties training to an activity setting rather than a separate switch. The control is called Keep Activity and sits in Gemini Apps Activity. Google's help page, last updated July 15, 2026, states what happens when you turn it off: "Your future chats won't appear in your Activity, and won't be used to train our AI models, unless you choose to send Google feedback. Future chats are still saved for 72 hours so Gemini can respond to you, process your feedback, and protect Google, its users, and the public."

The default auto-delete period is 18 months, adjustable to 3 months, 36 months, or off entirely so nothing auto-deletes.

AssistantConsumer defaultExact controlPublished retention
ChatGPT (personal plans)Training onSettings, Data Controls, "Improve the model for everyone"Temporary Chats deleted after 30 days; deleted conversations removed within 30 days
Claude (consumer)Framed as something you allowModel Improvement, at claude.ai/settings/data-privacy-controlsUp to 5 years de-identified if enabled; deletions purged from back-end storage within 30 days
Gemini (Gemini Apps)Keep Activity on, 18-month auto-deleteKeep Activity, in Gemini Apps Activity72 hours even when off; up to three years for human-reviewed chats

Turning training off still leaves a retention window

This is the part most guides skip, and it is the part that decides what you can safely paste.

Google is the most explicit about it. The same help page states: "Even if your Keep Activity setting is off or you use temporary chats, Google still uses your chats to respond to you and help protect Google, our users, and the public, including with help from human reviewers." Chats that go to human reviewers are retained for up to three years, disconnected from your account.

OpenAI's developer documentation describes the same shape on the API side, where training has been off by default for years. The doc states that "As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)", and immediately after that: "By default, abuse monitoring logs are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law, or is reasonably necessary to protect our services or any third party from harm." Zero data retention exists as an option, and the same page says those controls are "subject to prior approval by OpenAI and acceptance of additional requirements".

Anthropic publishes longer windows for one specific case. Its article states that where a policy violation is involved, inputs and outputs may be retained for up to 2 years and safety scores for up to 7 years.

So the honest summary of all three: turning off training stops your text feeding a future model. It does not stop the text existing on someone else's servers for somewhere between 72 hours and seven years, depending on vendor and circumstance. Design what you paste around that fact rather than around the toggle.

The categories that stay out of the box

Anthropic publishes the clearest list of the three, in an article about entering personal data. It advises being mindful about sensitive information and names four categories: "Financial information (SSN, credit card numbers, bank account details) Health records or medical information Passwords or private login credentials Confidential business or personal documents".

Read the framing carefully. That page recommends thoughtfulness. It does not prohibit any of it, and none of the three vendors blocks you from pasting any of it. There is no technical guardrail here at all, which is exactly why the rule has to live in your business rather than in the product.

For a small business, four categories cover almost every real incident.

  • Payment data. Card numbers, bank details and anything on an invoice that would let someone take money. There is no upside to it being in a chat log, ever.
  • Credentials. Passwords, API keys, the wifi password for the office, the code to the key safe. Rotate anything you have already pasted rather than hoping.
  • Other people's personal data. Customer addresses, staff records, a client list. You are the data controller for these under most privacy regimes, and pasting them into a consumer chat interface is a processing decision you made on their behalf without telling them.
  • Health information and anything else regulated by name. This one has a specific fix rather than a blanket ban, covered below.

A fifth, less obvious category: material you are contractually forbidden to share. If a client NDA says you will not disclose their documents to third parties, an assistant is a third party. That clause does not have an AI exemption because it was written before the question came up.

What a business tier changes, and the one thing it does not

Upgrading from a personal plan to a business one changes the default, and both OpenAI and Google put that commitment in writing.

OpenAI's policy page states: "By default, we do not train on any inputs or outputs from our products for business users, including ChatGPT Team, ChatGPT Enterprise, and the API." Its enterprise privacy page repeats it as "By default, we do not use your business data for training our models", and adds the compliance apparatus a small business may need to show a client: SOC 2 Type 2 completed for ChatGPT Enterprise, Edu, Healthcare, Business and the API Platform, a Data Processing Addendum available for GDPR purposes, and a HIPAA Business Associate Agreement available for API Platform customers.

That last one is precise in a way that matters. If you handle health information in the United States, the BAA is documented as available for the API Platform, and ChatGPT for Healthcare is described as built with HIPAA compliance support. A personal ChatGPT subscription is not covered by any of it. Pasting patient information into a consumer plan is not a grey area.

Google's Workspace position is the strongest single sentence any of the three publishes. Its admin privacy hub, last updated May 26, 2026, states: "Workspace does not use customer data for training models without customer's prior permission or instruction." It adds that user prompts are considered customer data under the Cloud Data Processing Addendum, and that the commitment sits in the Training Restriction section of the Workspace Service Specific Terms. Retention on that side is administrator-controlled: 90 days to indefinite for Gemini in Workspace, and up to 36 months for the Gemini app.

What a business tier does not change is the retention floor. Abuse-monitoring logs, safety review and legal holds sit underneath every plan. Paying more moves the training default and gives you contracts to point at. It does not make the text disappear.

Three settings and one page of rules, done in twenty minutes

Open each assistant your business uses and set the control by name. In ChatGPT: profile icon, Settings, Data Controls, turn off "Improve the model for everyone". In Claude: claude.ai/settings/data-privacy-controls, check Model Improvement. In Gemini: open Gemini Apps Activity and set Keep Activity, choosing an auto-delete period shorter than the 18-month default if you have no reason to keep the history.

Then write one page and give it to everyone who works for you. Four lines is enough: which assistant is approved, what must never be pasted into it, what to do if someone pastes it anyway, and who to tell. The recovery line matters more than the prohibition, because the incident you want to hear about is the one someone admits to on the same day.

A closing note on verification, since this whole page depends on it. Google dates its policy pages properly, which is how the July 15, 2026 and May 26, 2026 stamps above got here. OpenAI's help articles and Anthropic's privacy articles display relative timestamps instead, reading "Updated: yesterday" and "Over 4 weeks ago" at the time of checking. That is a real gap. It means that when you re-read those pages next quarter, you will be able to see that something changed and not what, or when. Re-check them anyway, and keep your own dated note of what they said.

AI data privacytraining opt-outdata retentionsmall business securitychatbot policy
Changelog (1)
  • July 31, 2026 — First published.