Is it safe to share data with AI tools?
Whether it is safe to share data with AI tools depends on two things: what kind of data it is, and what the tool’s terms allow the vendor to do with it. Public, non-personal information can go into almost any approved tool; client files and details about people need the right contract and several further checks.
Hypothetical example: an account manager pastes a client’s contract into a chatbot she signed up for herself. The summary is useful, but the terms she accepted let the vendor use conversations to improve its models and keep them for an unstated period. The client’s confidential terms have left the company’s control.
Sort the data first
Recommendation. A traffic-light rule lets staff decide quickly without asking the legal team every time.
- Green: public and non-personal. Information that is published or meant for publication, that you are allowed to use this way, and that contains no personal or confidential information, such as your own press releases and product descriptions. It can go into any tool the company has approved. Published material can still be protected by copyright; see what to check about AI and copyright.
- Amber: internal. Information not meant for outsiders but not harmful if exposed, such as process notes and draft marketing copy. Use it only in business tools the company has checked.
- Red: confidential and personal. Client documents, trade secrets, source code, financial figures and any information about identifiable people. Red data goes only into approved tools that have passed the checks below.
Where data fits more than one category, the stricter one applies. A staff profile on your website is public, but it is also personal data, so it is red. When in doubt, treat data as red. Removing names lowers the risk, but information that can still be linked to a person remains personal data.
Consumer or business plan: three checks
The same AI tool often comes in a consumer and a business version with very different terms. Before staff share data with AI tools for work, the tool’s owner checks three things.
- Does the model train on your data? Find out whether prompts, files and outputs can be used to improve the vendor’s models, whether this is off by default, and who controls the setting.
- How long is data kept? Find the retention period for prompts, files and logs, and whether the company can delete them.
- Who is the contract with? Check who is party to the contract, which terms apply to your organisation, and whether they include a data processing agreement. The plan’s name alone does not tell you.
A tool that fails any of these checks should receive green data only. Passing all three does not by itself allow red data: the legal and confidentiality checks below still apply. For fuller contract questions, see what to ask an AI vendor.
What the law says
Legal requirement (GDPR, Art. 28 and 35). These GDPR rules apply now. Entering personal data into an AI tool is processing of personal data. Where the vendor processes it on the company’s behalf, it acts as a processor, and GDPR requires a contract with specific terms, commonly called a data processing agreement or DPA (GDPR, Art. 28). A DPA is required for that relationship, but it does not make the processing lawful by itself, and it is not a general contract protecting trade secrets. Check separately:
- a lawful basis and a defined purpose;
- the vendor’s role;
- the contract terms;
- data minimisation and security;
- retention and deletion;
- the conditions for any transfer of data outside the EU;
- a data protection impact assessment (DPIA) where the processing is likely to result in a high risk to people’s rights (Art. 35).
Special categories of personal data, such as health data, may be processed only where one of the conditions in Art. 9 applies; internal sign-off is not enough. The data protection officer (DPO), where there is one, advises and monitors. The decision is taken by the person accountable for the processing, and accountability stays with the organisation.
Voluntary framework. The NIST Generative AI Profile treats leakage of personal and sensitive data as a generative AI risk (NIST AI 600-1). Data can also leak through logs and integrations, as covered in the main AI security risks.
Next step: for one specific use case, complete the three vendor checks and every applicable legal and confidentiality check above. Record the approved data categories, restrictions and accountable owner before telling staff what they may share.
Sources and further reading
- Regulation (EU) 2016/679 (GDPR) — Articles 9, 28 and 35
- NIST AI 600-1, Generative AI Profile
AI Horizon Conference
The AI Horizon Conference returns to Lisbon, once again bringing together entrepreneurs, investors and industry leaders to discuss the future of AI.