AI data safety guide

The risks of uploading sensitive information to AI models

Pasting a contract, résumé, support ticket or spreadsheet into an AI model is rarely just letting a model “take a look.” The content enters a processing path determined by the provider, product, account and organization settings you selected. Who can access it, how long it is retained and whether it is used to improve a service depend on that specific setup.

Risk is not limited to names and phone numbers. Customer lists, pricing, unreleased product plans, source code, medical excerpts and identifying context can all be confidential even when they contain no direct identifier.

Mask locally before translatingConfirmed sensitive entities stay protected locally · review remaining text

Short answer: do not paste unreviewed sensitive content into an AI model

Direct upload broadens the exposure surface. A third-party provider may process the content, and records can exist in operational logs, abuse detection, quality assurance, enterprise admin controls or connected tools. The actual risk depends on terms, account tier, data controls and your organization’s policy; “it is an AI tool” is not an adequate risk assessment.

Common harms go beyond model training

Training is only one concern. More common failures include sharing more than the task requires, using an unapproved provider, adding access through share links, plugins, browser extensions or team workspaces, and failing internal requirements for cross-border transfers, retention, deletion or auditability.

A provider saying that it does not train on inputs does not automatically make the input non-sensitive. Check the applicable product terms, data-processing agreement, retention and access controls, then follow your company policy, client commitments and applicable law.

Information that should not be uploaded directly

First remove direct identifiers and high-risk credentials: names, government IDs, contact details, addresses, bank accounts, login credentials, API keys, access tokens, customer IDs and employee IDs. Never provide passwords, private keys, recovery codes or production secrets to a generative AI service.

Also review indirect sensitive material: contract pricing and terms, unreleased financials, customer complaints, health and HR records, business strategy, source code and context that can re-identify a person, customer or project when combined.

A safer workflow before upload

Minimize the data first: provide only the excerpt needed for the task, and remove attachments, comments, hidden columns and unrelated context. Replace real values with stable placeholders where their position matters, then have a person review for misses and false positives.

Confirm the processing boundary next: use an approved endpoint and account configuration, check retention, access and cross-border handling, and involve security, legal or privacy teams when needed. Masking reduces exposure; it does not make the remaining body text risk-free.

Frequently asked questions

Can I upload content if the provider says it does not train on it?

Not automatically. No-training commitments address one risk only. You still need to assess the processing path, retention, access controls, organization policy, contractual duties and whether the remaining content is confidential.

Is a document safe after names and phone numbers are removed?

Not necessarily. Prices, project names, job titles, time, location and a small amount of context can still identify a person, customer or business. Minimize the content for the task and review it manually.

How does Veil Translate reduce exposure during translation?

It replaces detected and confirmed sensitive entities with placeholders locally in the browser, then sends masked text directly to the translation endpoint you choose. The remaining de-identified body text still leaves as plaintext, so it also needs review.

Related guides