Can We Put Company Secrets into ChatGPT and Claude? The Actual Scope of Generative AI Data Protection
No 'Absolute Confidentiality'... Terms, Settings, Contracts, and Courts Determine the Scope of Generative AI Data Protection
What to Ask Before Asking Whether You Can Trust AI
When considering what can be entered into generative AI, a common question is, "Can AI be trusted?" However, the subject of legal and contractual responsibility is not the chatbot model itself, but the relationship between the service operator and the user or organization. Therefore, what needs to be verified is not whether AI will keep a secret, but how the terms of service, privacy policy, account settings, and company usage policies of the service being used handle that data.
The data protection structure can be broadly divided into three stages. The first is whether the entered information is used for model training or improvement; the second is how long that data remains on the server; and the third is the possibility of human access under the operator's control, or whether it can be provided to third parties according to due process such as lawsuits or investigations. Turning off training utilization in a personal account does not simultaneously close all three of these stages. "Not learning," "not saving," and "never being submitted externally" are distinct concepts.
Turning Off ChatGPT Training Is Different from Erasing Data
In personal services like OpenAI's ChatGPT, user content may be used to improve models. Users can opt out of model training utilization by turning off the "Improve the model for everyone" feature in data controls within settings or through separate privacy procedures. Turning off this setting does not mean existing conversations disappear from chat history, and conversations kept by users remain in the account until deleted.
Temporary Chat is also not accurate if understood as "chat that leaves no records at all." Conversations in a temporary state are not used to model improvement and do not create general chat history or memory unless saved. However, copies may be retained for up to 30 days for safety purposes. As of 2026, users can also save temporary chats separately; in this case, they convert to general chats, after which the account's model improvement and personalization settings apply.
Files also need to be viewed distinctly from chats. For accounts or workspaces where the Library feature is available in ChatGPT, files uploaded or generated in general chats may be stored separately in the library. In this case, deleting the chat does not automatically delete the files in the library. Conversely, files uploaded in Temporary Chat are not saved to the account or Library.
Claude's Retention Structure Varies Depending on Training Permissions
Anthropic announced changes to Claude's consumer terms and privacy policy on August 28, 2025. The targets are Claude Free, Pro, and Max users, along with Claude Code users utilizing those accounts. Users can choose whether their new conversations and resumed conversations are used to improve Claude, and past conversations without separate activity are not newly incorporated into training targets according to this setting.
If model improvement utilization is allowed, those new and resumed conversations can be stored in de-identified form in the training pipeline for up to 5-years. Anthropic explained that the existing 30-day retention framework continues to apply to users who do not choose model improvement, and the current privacy center guides that when users delete conversations, they disappear immediately from chat history and are deleted from backend storage systems within 30 days.
There are exceptions. According to Anthropic's current consumer data retention policy, inputs and outputs of conversations or sessions flagged for terms of service violations by automated trust and safety systems can be stored for up to 2 years, and related trust and safety classification scores can be stored for up to 7 years. A 5-year retention period applies to user-submitted feedback and related data. Longer retention possibilities for legal obligations, dispute resolution, or response to terms of service violations are also separately stipulated.
Personal and Enterprise Conditions Differ Even for the Same AI
Data processing conditions change in enterprise services. OpenAI states that inputs and outputs of ChatGPT Business, Enterprise, Edu, and the API platform are fundamentally not used for model training or improvement. Anthropic similarly states that inputs and outputs of commercial services such as Claude for Work's Team and Enterprise, and APIs, are fundamentally not used for model training.
That does not mean one should interpret that "if it's enterprise, data does not remain at all." Retention periods, administrator settings, data residency, access controls, and external integration structures can vary by product. OpenAI APIs are also fundamentally not used for training, but abuse monitoring logs for some requests can generally be retained for up to 30 days, and depending on API features used, application states are also stored separately.
When complete data non-retention is required, separate conditions must be verified. OpenAI's Zero Data Retention, or ZDR, is a separate data control method provided to eligible and approved API customers. Therefore, one must not understand that "using the API means zero retention across the board." Depending on which endpoints and features are used, the applicability of ZDR can also vary.
Clicking the Delete Button Does Not Make Data Disappear from Servers Instantly
Whether training utilization is applied and data deletion are also separate issues. According to OpenAI's general ChatGPT retention policy, when a user deletes a chat, it is removed immediately from the account screen, and is permanently deleted from systems within 30 days, except when de-identified and separated from the account, or when longer retention for security or legal purposes is required. The archive function is different from deletion; simply archiving a conversation keeps it in the account.
OpenAI previously bore separate legal preservation obligations for some user data during the New York Times-related copyright litigation process in 2025. However, the obligations of previous orders requiring indefinite preservation of new consumer and API data ended after September 26, 2025, and it has stated that it has now returned to general retention policies. Separate legal preservation measures following litigation continued for some past data. Deletion policies of general services and legal preservation orders issued in specific lawsuits need to be viewed separately.
U.S. Courts Do Not View AI Conversations Under a Single Legal Principle
U.S. court judgments surrounding whether conversations had with AI or materials created using AI can receive legal protection also began in earnest entering 2026. However, it is difficult to generalize decisions made so far into a single legal principle that "AI conversations are never confidential." This is because conclusions diverge depending on the nature of the case, the purpose for which AI was used, the involvement of attorneys, and the type of legal protection applied.
A representative case is United States v. Heppner. Bradley Heppner, a former financial firm CEO, used consumer Claude to create documents and defense strategy materials related to his case while under criminal investigation. When about 31 AI-related documents were discovered in seized electronic devices, the defense counsel asserted attorney-client privilege and work-product protection.
On February 10, 2026, Judge Jed Rakoff of the U.S. District Court for the Southern District of New York granted the government's motion following oral argument, and explained the reasons in a written opinion on February 17. The court did not recognize attorney-client privilege and work-product protection based on points including that Claude is not an attorney, Heppner inputted information into a third-party AI platform without an attorney's instruction, and the materials were not created under an attorney's direction.
However, a different judgment emerged on the same February 10 at the U.S. District Court for the Eastern District of Michigan. In the case Warner v. Gilbarco, a plaintiff conducting litigation pro se authored litigation materials using ChatGPT and analyzed emails, and the defendant requested the production of materials related to AI usage. The court viewed those materials as work-product prepared in anticipation of litigation and recognized protection. In particular, it judged with the intent that generative AI programs are "tools, not humans," and waiver of work-product protection becomes an issue when materials are disclosed to an adversary or in a manner likely to be delivered to an adversary.
The two cases yielded different results on the same day, but the applied legal issues and factual relationships were not identical. In Heppner, attorney-client privilege and work-product created without attorney direction were at issue, while in Warner, work-product protection for materials authored by a pro se litigant for litigation was core. Therefore, at the current stage, it is difficult to conclude that legal protection automatically arises or automatically disappears simply by using AI.
What the Order to Produce 20 Million ChatGPT Logs Signifies
Another case is the lawsuit by media companies and copyright holders including The New York Times against OpenAI. On January 5, 2026, Judge Sidney Stein of the U.S. District Court for the Southern District of New York denied the objections raised by OpenAI and maintained the order issued by the magistrate judge to produce a sample of 20 million de-identified consumer ChatGPT logs.
According to court records, plaintiffs originally requested a larger scale of logs, and a 20 million de-identified sample was finalized as the target of production during discussions. OpenAI proposed a plan to apply search terms and filter out only conversations related to the lawsuit, but the court judged that the existing order sufficiently considered user privacy and the relevance of the materials. The sample in question was extracted from consumer ChatGPT conversations between December 2022 and November 2024, and corporate Business, Enterprise, Edu, and API customer data were not included.
This decision does not mean all ChatGPT conversations can be disclosed to litigants at any time. The case is an instance where the relevance and necessity of a specific sample were recognized in the discovery process of a copyright lawsuit, and protection measures such as de-identification and strict access controls were applied together. However, it demonstrates that data actually possessed by service operators has the possibility of becoming the target of valid court orders and discovery procedures.
Interestingly, the January 2026 court order itself specified that OpenAI at the time possessed tens of billions of consumer ChatGPT logs in the ordinary course of business. However, this sentence must not be interpreted to mean OpenAI retains all users' deleted chats indefinitely. General deletion and retention policies, litigation preservation obligations, and past log samples used in discovery are separate matters.
In Practice, It Is Safer to Divide Information into Four Grades
To apply these policies and rulings to practical business, a method of dividing information according to risk level and managing it is useful. This is not an official classification prescribed by laws, but a practical judgment criterion for utilizing generative AI.
Grade 1 consists of already-public information and materials where losses from disclosure are not significant. Public press releases, public statistics, and documents already announced externally fall here. Even if they are materials authored oneself, unpublished core intellectual property rights or business ideas cannot be viewed at the same level. If information is unpublished, one must look at not only training utilization but also retention, access, and contract conditions.
Grade 2 consists of information that can be used restrictively after sufficiently lowering identification risks. When handling consulting interviews, organizational diagnostic materials, and internal client documents, erasing direct identifiers such as names, titles, departments, and company names may not be sufficient. This is because targets can still be identified when specific dates, events, rare job functions, and transaction terms are combined. De-identification is a measure to reduce risk, not a guarantee that completely eliminates re-identification possibilities.
Grade 3 is information for which making it a rule not to input raw text into personal general-purpose AI accounts is appropriate. Unpublished financial figures, personnel evaluations, contract terms, customer data, litigation-related facts, and pre-launch product specifications are representative examples. If business AI processing is unavoidable, enterprise services or API environments should be reviewed, and one should check not merely phrases like "not used for training" but also contracts, data processing agreements (DPAs), retention and deletion periods, access controls, data storage regions, and external integration scopes together.
Grade 4 is information required not to be directly inputted into general-purpose generative AI. Representative examples are information that immediately leads to security incidents the moment they are exposed, such as resident registration numbers, passwords and one-time authentication codes, card numbers, and API keys. The Personal Information Protection Commission also guides generative AI users not to input important information like resident registration numbers, passwords, and authentication codes. Even when handling high-risk personal information such as health information or child and adolescent information for business, separate processing environments equipped with legal grounds and organizational approval should be prioritized over personal general-purpose AI.
Six Checkpoints Presented by the PIPC to Users
Domestic standards are also taking concrete shape. The Personal Information Protection Commission (PIPC) released personal information processing guidance for generative AI development and utilization in August 2025, and revised privacy policy drafting guidelines on April 24, 2026, arranging a separate appendix for generative AI services. The core is requiring operators to clearly state in privacy policies whether they collect prompts or outputs, utilize them for training, how long they store them, and what methods users have to refuse training utilization.
Subsequently, on May 19, 2026, it announced the "Personal Information Protection Guide for Generative AI Service Users." The guide recommends users check whether input data is utilized for AI training and limitation settings, and examine chat history storage and deletion functions and retention policies. The fundamental principle is not to input important information like resident registration numbers, passwords, and authentication codes.
Also, before inputting conversations or documents, users are urged to check whether personal information of third parties, including family and colleagues as well as oneself, is included, and to mind that information appearing fine in a single conversation can lead to personal identification when accumulated. Work materials of companies or institutions must follow internal organizational guidelines, use approved AI services, and comply with material outflow procedures.
When personal information is discovered in AI answers, users are guided to stop transmission or posting, capture screenshots to leave evidence, delete conversations and files, release necessary external integrations, and request action from operators. When connecting generative AI with other services, users are recommended to grant only necessary permissions and re-examine the necessity of integration if excessively broad permissions are requested.
The 'Shadow AI' Problem Originating in Corporations
A representative case where the information security issue of generative AI became known in earnest within corporations emerged from Samsung Electronics in 2023. At the time, it became known that employees in Samsung Electronics' semiconductor division inputted internal source code and meeting-related contents into ChatGPT, and the company subsequently took measures restricting the use of external generative AI on company-owned equipment. The core of the incident lay not in ChatGPT's performance itself, but in the fact that internal information moved to external services outside the company's control.
The risk of so-called 'Shadow AI' also arises here. Even if a company is equipped with enterprise AI contracts and access control systems, if an employee uses their personal account to process customer information or in-house documents, data can move outside the data protection conditions arranged by the company. What organizations must do is go beyond blocking AI usage and pre-determine approved tools, prohibited information, de-identification methods, and external outflow procedures.
Sole Proprietors Must Reduce Materials Themselves Rather Than Relying Only on Settings
For sole proprietors or small consultants who find it difficult to use enterprise contracts, settings and data minimization become the first defense line. When using personal AI, model improvement settings must be checked first, and Temporary Chat or equivalent features can be utilized for sensitive tasks when possible. However, this must be premised on the fact that temporary chats can also be retained for security purposes for a certain period.
It is necessary to make it a basic procedure for customer data to extract only information necessary for the task and remove identifiers before inputting, rather than putting in the entire original text as-is. If there is a confidentiality obligation in a contract, whether external processing using generative AI is permitted at all must also be verified first. If necessary, separately agreeing with the client on the scope of AI usage is safe.
APIs are likewise not an unconditional solution. Commercial APIs from OpenAI and Anthropic fundamentally do not use inputs and outputs for model training, but data retention conditions can vary depending on products, features, and contracts. Especially for materials where non-retention is critical, one must not stop at using general APIs, but go as far as verifying the applicability of separate data retention conditions such as ZDR.
Conclusion: Before Speaking to AI, You Must Ask 'Is This Information Allowed to Go Out?'
How well information is protected in generative AI is dictated by service policies, contracts, technical settings, and legal procedures rather than the model's intelligence. Turning off model training is an important protection measure, but by itself, it does not make data storage or the possibility of legal preservation and submission disappear. Nor does utilizing enterprise services mean all data is automatically non-retained.
U.S. rulings in 2026 also do not provide simple conclusions. In the Heppner case, where legal-related materials were inputted into consumer Claude, legal protection was negated, but in the Warner case, where ChatGPT was utilized as a litigation preparation tool, work-product protection was recognized. Ultimately, results can vary depending on what service and settings were used, what purpose materials were created for, and whether there was direction from an attorney or organization.
The simplest judgment criterion in practice is actually unrelated to technology. It is to first ask whether this information can be handled when it goes out externally. If it can be handled, utilize AI within the necessary scope; if not, go through de-identification and data minimization or use an approved corporate environment. If the risk is still difficult to handle, not inputting it is the surest protection measure.


