AI at the Gemba · Lesson 2 of 10 · Week of October 19
What's safe to paste
Every question you type into an assistant is a copy of your words sent to someone else's computers. Most of what a Lean practitioner wants help with works fine without names, customers, formulas or passwords, so the skill is learning to take those out before you press enter. This lesson gives you a rule you can say out loud, shows what your account does with a paste, and has you scrub a messy export before an assistant ever sees it.
By the end of this lesson you can:
- sort anything you're about to paste into green, yellow or red;
- say what your plan does with a paste: whether it trains a model, how long it is kept, and who can read it;
- name the add-ons that widen the risk: connectors, extensions, recorders and share links;
- strip the names from an export and still get the analysis;
- say what to do if something red goes in.
Last time: two questions to warm up (not graded)
1. An assistant gives you a neat cause section with no quotes from your notes. What is every claim in it until you check? A guess. Ask which statements come from your text and which are assumptions.
2. Why might it contradict an SOP you pasted sixty messages ago? A long chat crowds the context window, so detail gets dropped or underweighted. Keep the SOP in a Project and start a fresh chat.
This lesson is not legal advice, and it is not your company's policy. Where a rule depends on your employer, your contracts or your license, ask the person who owns that rule.
A paste is a copy that leaves the building
An assistant runs on a vendor's servers. When you paste, your text goes there. What happens to it next is decided by the contract behind your account, not by the brand on the screen. Four words carry most of the idea:
- Training: the vendor learns from saved chats to build future models.
- Retention: how long copies are kept, even after you delete a chat.
- Human review: people at the vendor may read some chats, such as those flagged for safety.
- Admin access: on a company account, your own administrators can often see your chats.
Researchers have pulled verbatim training text, including personal details, out of language models, and NIST lists leakage of personal data as a risk. That is real, but it is rare. The boring risks come first: an account you don't control, a chat kept for years, a shared chat that turns up in a web search, an add-on that reads more than you meant it to. This lesson teaches those first.
Predict first
A supervisor on your company's approved assistant plan pastes a shift log into her own free personal account "because it's quicker". Both accounts are the same brand. Before you read on, write down what is different about the two accounts. The next sections give you the answer.
Three buckets and four questions
Sort anything you're about to paste into one of three buckets. If you can't tell, it is yellow until you've checked. This is the same rule as on the data safety page, which also offers a free one-page version your team can print and adopt.
Four questions before every paste
- Could this identify a person? A name, a face, a badge number, a purchase history.
- Would I be comfortable if a competitor read it? Prices, formulas, volumes, customers.
- Am I allowed? Check your company's policy, your contracts and your license conditions.
- Can I get the same help with the who taken out? Usually, yes.
If any answer makes you pause, strip it or don't paste it.
What your plan does with a paste
As of October 1, 2026. These terms change often. Anthropic changed its consumer defaults in 2025, and vendors rename plans. Read the vendor's own page for the plan you use, listed in the sources, and check again after any notice from them.
The account type sets the defaults. This is a summary of what each vendor's own pages say, not a promise about your contract.
That answers the predict question. The same brand on a personal plan and a company plan are two different contracts. A personal plan follows consumer terms, where training is a setting you have to find and retention can reach years. The company plan is usually not trained on, and the company's administrators may be able to read it.
Training is only one of five exposures
Training, retention, human review, legal holds and your own administrators are all ways a paste can be seen again. On legal holds: OpenAI said it kept data from its Free, Plus, Pro, Team and API accounts under a court order until September 26, 2025, and still holds a slice of April to September 2025.
Three things that look like protection and aren't
- Paid does not mean private. Claude Pro and Max, and ChatGPT Plus and Pro, are consumer plans. On a business plan, your own admins can read staff chats, and the HIPAA agreements some vendors offer leave out specific features: Anthropic's leaves out connectors, its Files API and code execution, Google's leaves out Gemini Notebook and Gemini in Chrome, and Microsoft's leaves out web queries.
- Incognito and temporary chats are not invisible. They change what is saved and trained on. On Claude's Enterprise plan, for example, administrators can still see incognito chats.
- Certificates describe the vendor, not your use. SOC 2 Type II means an auditor tested the vendor's security controls over a period of time. ISO 27001 certifies a managed security system and ISO 42001 a managed AI-governance system. All the big vendors list them. None of them says you may paste supplier prices.
Add-ons reach further than the chat
The plan's promises cover the chat in front of you. Add-ons sit outside them.
- Connectors let the assistant read your drive, email or ERP.
- Browser extensions run inside your browser, right beside your chats.
- Meeting recorders and note-takers record what is said.
- Share links put a chat in front of other people.
A fifth risk, prompt injection, is hidden text in a file or web page that the assistant obeys as if you wrote it. It matters most when one assistant has three things at once: private data, untrusted content to read, and a way to send data out. Lesson 9 returns to it. For now, Anthropic's own advice for custom connectors is to connect only servers you trust, and your company's IT team should decide what gets connected.
What has actually gone wrong
- Samsung, 2023. Trade press reported three incidents in about three weeks after Samsung allowed ChatGPT: staff pasting database code, equipment code and meeting audio. The company capped prompt size, then banned staff use. Samsung did not comment. I found no evidence the pasted code ever surfaced for anyone else. The harm was losing control of where it went.
- A vendor bug, March 20, 2023. OpenAI reported that a bug showed some users the titles of other users' chats, and exposed billing details for 1.2% of active ChatGPT Plus subscribers.
- Share links, August 2025. OpenAI removed an option that let shared chats appear in web search, after chats showed up there.
- An extension, December 2025. Koi Security alleged that a VPN browser extension with over 6 million Chrome users captured prompts and replies from several assistants. The publisher had not replied when it was reported.
None of these is a reason to ban assistants. Some analysts argue a ban pushes use out of sight, which would be worse. They are reasons to decide in advance which account, which add-ons and which data.
Strip the who
An assistant needs the steps, the times, the counts and the problem. It does not need to know who anyone is. Five moves do most of the work:
- Swap names for roles or random codes, not initials. "Trimmer A" and "Packaging lead" tell the assistant there are two people. A role with only one person in it still points to them.
- Swap customers and brands for codes. "Customer C1", "Brand B". Keep the key that says which is which in your own file, never in the chat, and never build a code from the real name.
- Delete the columns you don't need. If the analysis doesn't use a column, don't paste it.
- Start with ten rows. Read what you pasted before you paste the full export.
- Put the names back yourself. Read what the assistant gives you, and add the real names and numbers back on your own side.
Stripping names is not the same as making something anonymous
The US Department of Health and Human Services, writing about health data, says initials, dates finer than a year, device and plate numbers, and any unique code can all defeat de-identification, and that a code must not be derived from the person's own data. Models can also work out personal details by stitching sources together, which NIST lists as a risk. So the safest move for anything about people is the one in the red bucket: don't paste it. For everything else, the more you strip, the better.
Habits that make it routine
- Use only the approved account. Write the plan name down where the team can see it.
- Read the paste before you send it.
- No new extension, connector or meeting bot without IT's sign-off, and switch off share links.
- Review the settings every quarter and after any vendor notice, and write down the date.
- Treat what the assistant reads from a file or a page as untrusted.
- Record a conversation only with everyone's consent. In Massachusetts, secretly recording one is a crime under General Laws chapter 272, section 99. AI note-takers record.
The rules that sit behind it
I am not a lawyer, and what follows is a map, not advice. Ask your compliance lead or counsel what applies to you.
- Trade secrets. Federal law and Massachusetts law both protect a trade secret only if its owner took reasonable steps to keep it secret. Massachusetts says efforts that are "reasonable under the circumstances". Whether a paste counts is a question for your NDA, your policy and counsel.
- Personal information. Massachusetts regulation 201 CMR 17.00 requires a written security program, training and oversight of service providers for residents' personal information, such as a name with a Social Security, license or account number, which is how Massachusetts General Laws chapter 93H, section 1 defines it.
- Health information. HIPAA binds only covered entities and their business associates. A cloud vendor that holds your patient data, even encrypted, is a business associate.
If you work in cannabis in Massachusetts
Under the version of 935 CMR 500 that I could read, current to November 2024, retailers may not record a customer's personal information beyond what a sale needs without written permission (500.140(2)(c)). Delivery details such as name, date of birth, address, phone and email are for delivery only and must be kept confidential (500.140(2)(e) and (f)). Licensees need a written plan for confidential information (500.105(1)(l)), and written cash-handling measures count as security planning documents (500.110(7)(c)). Check the current text with your compliance lead, since the version I read may have been amended since.
If red data goes in
It will happen to someone. The worst response is to hide it.
- Stop. Don't keep going in the same chat.
- Write down the tool, the account, the time and roughly what went in.
- Tell your lead.
- Delete the chat following the vendor's steps for your plan. Deleting is not the same as it never happened, so keep the note.
- Ask your compliance lead or counsel whether 201 CMR 17.00, the Massachusetts breach law (chapter 93H) or a contract requires a notice.
Your turn: find out, then scrub
About 25 minutes. It has two parts and a swap with a colleague, and you paste nothing of your own in any of them.
Your work, your data. The export below is invented: every name, company and number is made up, and any match with a real person or company is a coincidence. The first part pastes nothing at all. In the second, you scrub the invented export, and only then paste it. Use only this export, not anything from work, until you have done the scrub and checked it.
Part 1: find out what your account does (8 minutes)
For the assistant you use, find and write down: the plan name; whether the training setting is on or off, and where it is; how long chats are kept; whether a temporary or incognito mode exists; and who at your company, if anyone, can see your chats. Paste nothing. If you can't find an answer, that is the answer to take to your IT or compliance lead.
Part 2: scrub an invented export (12 minutes)
Here are ten stops from a made-up packaging line. It has operator names, customer names, lot numbers and dates, and two of the reason notes carry a first name.
- Copy the table into a spreadsheet or a text file of your own.
- Delete the columns the analysis doesn't need. A Pareto of stops by reason needs shift, customer, minutes and reason, not the operator or the lot.
- Swap the three customers for C1, C2 and C3. Keep the key in a separate file.
- Replace each date with Day 1 to Day 5.
- Read the reason column for names, and replace each with a role.
- Paste the ten scrubbed rows into your assistant and ask for a Pareto of stops by reason and two questions to ask at the gemba.
- Check the answer the way Lesson 1 taught you. Add the minutes by reason yourself, then compare.
What a good scrub looks like, and the numbers to check
Day, Shift, Customer, Minutes, Reason
Day 1, Day, C1, 12, film splice, operator re-threaded
Day 1, Swing, C2, 25, label jam
Day 2, Day, C1, 8, film splice
Day 2, Swing, C3, 31, waiting on forklift, asked lead
Day 3, Day, C2, 15, label jam
Day 3, Night, C3, 40, changeover
Day 4, Swing, C2, 22, label jam
Day 4, Night, C1, 18, film splice
Day 5, Day, C3, 35, changeover, operator waiting on kit
Day 5, Swing, C2, 20, label jam
The stops add up to 226 minutes. By reason: label jam 82 minutes (36%), changeover 75 (33%), film splice 38 (17%) and waiting on the forklift 31 (14%). If your assistant's Pareto differs, find out why before you believe it. Your own scrub may differ from this one, and that is fine as long as nobody can be named from it.
Last, hand your scrubbed version to a colleague, or look at it again tomorrow, and ask: can anyone guess a person or a customer from this? If so, strip more. Then write your team's three-line rule for what goes in each bucket.
Knowledge check
Five questions, graded for you. Sign in to take the check, save your progress and count this lesson toward your certificate. Sign-in opens on Monday, October 12.
Sign in to take the check
Your task: write the rule, then use it once
- Open the printable worksheet. Fill in what you found about your account, your scrub, and your team's three-line rule.
- If you want a ready-made version to adapt, the data safety page has a free one-page team rule you can print.
- Use the four questions on one real thing this week, before you paste it. Write down what you took out.
Open the printable worksheet
Sources
Where the facts in this lesson come from. Facts last checked October 1, 2026. Plans, settings and rules change quickly, so check the vendor's own page for the plan you use before you rely on a row of the table. If you find a fact that is out of date or wrong, tell me.
- Anthropic: Updates to Consumer Terms, Is my data used for model training?, How long do you store my data?, organization data, Commercial Terms, Incognito chats, BAA coverage, certifications, custom connectors
- OpenAI: How your data is used to improve model performance, Data controls in ChatGPT, Enterprise privacy, Business data privacy, response to data demands, the March 20 outage
- Google: Gemini Apps Privacy Hub and Generative AI in Workspace Privacy Hub
- Microsoft: Data, privacy and security for Microsoft Copilot, Enterprise data protection, Privacy FAQ for personal Copilot
- Law: 18 U.S.C. 1839, M.G.L. c. 93 s. 42, 201 CMR 17.00, M.G.L. c. 93H s. 1, M.G.L. c. 272 s. 99, 935 CMR 500.140 (text current to November 2024)
- HHS: De-identification guidance and HIPAA and cloud computing
- NIST AI 600-1: Generative AI Profile; Extracting Training Data from Large Language Models, Carlini et al.; Scalable Extraction of Training Data from Production Language Models, Nasr et al. (preprint)
- AICPA: SOC 2; ISO standards 27001 and 42001; OWASP Top 10 for LLM Applications; Simon Willison on the lethal trifecta
- Incidents, as reported: Dark Reading on Samsung and CSO Online on the Samsung ban; Malwarebytes on discoverable chats; Infosecurity Magazine on Urban VPN (an allegation)