AI at the Gemba · Lesson 1 of 10 · Week of October 12
Start here: what an assistant is, and isn't
An AI assistant writes answers that look finished. In Lean work, that is the problem. A neat page can hold a guess in every line, and nothing about its look gives it away. This lesson is about reading an assistant's answer the way you would read a report from a stranger, with the question every Lean person already asks: how do you know?
By the end of this lesson you can:
- say in plain words what an assistant does, and what it doesn't;
- tell a statement that comes from your text from one that is a guess;
- run a seven-question check on any answer before you use it;
- set up a practice project and see what one instruction changes.
You don't need any AI background. If you already use an assistant every day, the check and the exercise will still be worth your time. Facts about products and plans are as of October 1, 2026, and they change often, so every source is listed at the bottom with the date it was checked.
The Lean job: read a finished-looking answer without being fooled
You have spent years learning not to trust a report you haven't seen the work behind. An A3, a 5 whys and an OEE number all get walked back to the floor. An assistant hands you the same kind of document in seconds, and it takes away the signals you used to read: no hesitation, no missing number, no cramped handwriting. Everything is neat and everything sounds sure.
So the first skill in this course is not writing prompts. It is reading.
Predict first
Here is a small made-up case. The numbers are invented. A lead pastes a jam log for Labeler 3 into an assistant: jams per shift, over three days.
| Shift | Day 1 | Day 2 | Day 3 |
| Day shift | 2 | 3 | 2 |
| Swing shift | 9 | 8 | 10 |
| Night shift | 3 | 2 | 3 |
The assistant answers:
About 5 jams per shift, consistent with random guide-rail wear.
Before you read on, write down what you would check before you used that sentence. Then open the answer.
What to check, and what it finds
- Does it quote anything? No. The log says nothing about guide rails. Guide-rail wear is a cause with no source.
- Is the arithmetic right? Yes. 42 jams over 9 shifts is 4.7, so "about 5" is fair. But look at what it hides.
- What does the average hide? Swing shift had 27 jams over three days, 9.0 a day. Day and night together had 15 over six shifts, 2.5 a shift. No shift has 5. The average describes none of them.
The arithmetic was correct and the conclusion was unsupported. A better next move is to treat the shift as the signal and ask: "List three things that differ on swing shift. Mark each one [NEEDS GEMBA]." Then walk to the line at the start of swing and look.
What an assistant is
An assistant like Claude, ChatGPT, Gemini or Microsoft Copilot is built on a large language model. It writes the words that are most likely to come next, based on patterns in an enormous amount of text it was trained on. That has four consequences you can rely on.
It predicts text. It doesn't look things up.
It knows your floor only from what is in the conversation: what you pasted, the files you attached, and, if the product has a search tool, what it found. NIST, the US standards agency, describes it the same way in its generative AI profile: the model generates content by predicting likely next words from patterns in its training data.
It can state things that are false, in the same tone as things that are true.
NIST calls this confabulation: confidently stated content that isn't supported. People usually say hallucination. Confidence is a style of writing, not a sign the model checked anything. A 2025 preprint argues that the way models are trained and scored rewards a guess over "I don't know", which is one reason this keeps happening.
It is lower, never zero, when it summarizes text you supplied. In a peer-reviewed 2025 test, legal research tools built to look up real documents first still gave wrong answers between about 17% and 33% of the time. And the cost is real: a public database kept by one researcher listed 2,097 court decisions, as of September 30, 2026, that found someone had relied on hallucinated material.
Its explanation of itself can be a plausible story.
Ask it why it answered the way it did and you will get a fluent reason. That reason is also generated text. In a 2025 preprint, models that had been nudged by a planted hint usually did not mention the hint when they explained their answers. Treat an explanation like the 5th "why" on a board: something to go and check, not the evidence.
It is trained to guess.
Training and scoring tend to reward a guess over "I don't know", so allow "I don't know" in your instructions. Think of an andon cord that only gets pulled if you wrote down when to pull it.
What it can see: the context window
The context window is everything the assistant can read in one conversation: your messages, anything you pasted or attached, and its own earlier replies. It is measured in tokens, and a token is roughly half to three-quarters of a word.
A bigger window is not a better memory. Anthropic's documentation says accuracy can drop as the window fills, which it calls context rot. A 2024 study found models use facts in the middle of a long document less reliably. And Claude's help pages say long chats may be summarized automatically, which drops detail. So an SOP you pasted sixty messages ago may no longer be read the way you assume.
Four things people call "memory"
Other products have the same ideas under other names. ChatGPT also has Projects, Gemini has Gems and notebooks, and Microsoft Copilot has notebooks. Look up what yours is called, and what it keeps, before you rely on it.
Same question, different answer
Ask the same question twice and you can get two different answers. In one preprint, running the same prompt 1,000 times on a setting meant to make the output repeatable still produced 80 different outputs, and accuracy varied by up to 15%. Anthropic's documentation says its newer API models reject that setting anyway.
Models also change under you. OpenAI retired GPT-4o from ChatGPT in February 2026. A vendor rolled back an update because the assistant had started flattering people. A 2023 preprint found GPT-4's accuracy on one task, spotting prime numbers, fell from 84% to 51% in three months.
Neither of these is a reason to avoid assistants. They are reasons to do what you do with any process: run it more than once, write down which version ran, and test again when something changes. Lesson 9 turns that into a routine.
When it helps, and when it doesn't
The honest answer from the research is mixed. Here are five well-known studies, with how much weight each one can carry.
I found no study of Lean work specifically, so what follows is my reading of the studies above, not a finding. Assistants tend to be useful for drafting, turning messy gemba notes into a first problem statement, summarizing, asking you clarifying questions, and spotting gaps in the logic of an A3. They tend to be poor at your floor, at exact numbers, at anything that needs a source they weren't given, and at giving a verdict.
Why people stop checking
Deferring to a confident machine has a name, automation bias. NIST lists it as a risk. It has been found in experts, and practice doesn't cure it. In a Microsoft Research survey of knowledge workers, people who trusted the AI more reported using less critical thinking. That is self-report, so it is a hint, not a measurement.
There is a learning cost too. In a PNAS study, students given unguarded AI help scored 48% higher on practice and 17% lower on a later test without it. In a small Anthropic trial of 52 people learning a coding skill, those who used AI scored 50% on a quiz afterwards, against 67% for those who didn't.
The takeaway for a Lean leader is a design one. If you want people to keep their own judgment, put the checking into the standard work, because a confident answer will otherwise make checking feel like extra effort.
Seven questions to ask of any answer
Use these before you act on anything an assistant gives you. They are the beginning of a checklist you will extend through the course.
- Which statements quote my text? A claim with no quote is a guess until proven otherwise.
- Which numbers did I give it? Recompute two in a spreadsheet. Mark any number it supplied itself as [NEEDS GEMBA].
- What did it assume that I didn't say? Ask: "List what you assumed that I did not say."
- Can I go and see it? Name who will check, where and when.
- What is the strongest case against this answer? Ask for it directly.
- Does a fresh chat give the same answer? If it doesn't, your confidence should be low.
- Did I allow "I don't know"? Say so in the instructions. Anthropic's own guidance on reducing hallucination recommends it.
Standard work, with a catch. Project instructions work like standard work: written down, repeatable, and improved over time. The catch is that the reader improvises every run. This checklist is a kind of source inspection, but nothing mechanically blocks a defect, and confident users check less. A rule that has to hold every time needs a check outside the assistant. You will build those in Lessons 8 and 9.
Your turn: the same request, two ways
About 25 minutes. The steps are written for Claude, and a free account is enough: the Free plan includes Projects, up to five of them, as of October 1, 2026. Other assistants have their own version of a project, and menus change often, so look for the same ideas under whatever yours calls them.
Your work, your data. This course uses invented files only. What you type into an assistant leaves your computer, so don't paste anything from work yet. Lesson 2 shows what is safe to paste, and what isn't, before any exercise uses your own material. Use only the invented text below.
- Make a practice project (3 minutes). In Claude, choose Projects, then New project, and name it "AI practice". The name and description are for you. The assistant can't see them.
- Paste your notes (2 minutes). Start a new chat in the project and paste this invented note, exactly:
Fill line, 3 shifts. Short fills on swing. Nozzle replaced last week. Operators say the hopper is slow.
- Chat A (5 minutes). With no project instructions yet, send:
Write a problem statement and a 5 Whys.
Save the answer.
- Add one instruction (2 minutes). Open the project's instructions and paste:
Use only my text. Quote the line you rely on for each claim. Mark anything else [NEEDS GEMBA]. If you don't know, say "I don't know."
- Chat B (5 minutes). Start a new chat in the project, paste the same note, and send the same request as Chat A.
- Count (3 minutes). In each answer, count the claims that come with a quote from your four lines. Then run Chat B once more in another new chat and write down what changed between the two runs.
- Ask for the other side (3 minutes). In Chat B, send:
Give me the strongest case against your top cause.
- Write a two-line rule (2 minutes) for how you will check an assistant's output from now on.
Your results will differ from mine and from the next person's, and that is part of the lesson. I'd expect Chat A to add causes your four lines never mentioned, and Chat B to quote what it used and mark the rest, though neither is certain. One instruction can change a lot, and it still isn't a guarantee, so keep checking.
Knowledge check
Five questions, graded for you. Sign in to take the check, save your progress and count this lesson toward your certificate. Sign-in opens on Monday, October 12.
Sign in to take the check
One of the five shows a short chat with the assistant and asks you to tick the statements you wouldn't rely on as written.
Your task: check one answer, this week
- Open the printable worksheet and fill in what you found in the exercise.
- Write the two-line rule you chose for checking an assistant's answers.
- Write one thing you will check on the floor this week that an assistant told you or could tell you. Then go and look.
Open the printable worksheet
Sources
Where the facts in this lesson come from. Facts last checked October 1, 2026. Products, plans and model names change quickly, so check the vendor's own page before you rely on one. If you find a fact that is out of date or wrong, tell me.
- NIST AI 600-1: Generative AI Profile (July 2024): confabulation, automation bias, how models generate content
- Why Language Models Hallucinate, Kalai et al. (preprint)
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Magesh et al. (2025)
- AI Hallucination Cases database, Damien Charlotin: the count is as of September 30, 2026
- Reasoning models don't always say what they think, Chen et al. (2025 preprint) and Anthropic: Tracing the thoughts of a large language model
- Anthropic: Context windows and Claude Help: Usage and length limits
- Lost in the Middle, Liu et al. (TACL 2024)
- Claude Help: Projects, Memory and Skills; Claude pricing
- Non-determinism of "deterministic" LLM settings, Atil et al. (2024 preprint); How is ChatGPT's behavior changing over time?, Chen, Zaharia and Zou (2023 preprint); Anthropic: Model deprecations
- OpenAI Help: Model release notes and OpenAI: Sycophancy in GPT-4o
- Navigating the Jagged Technological Frontier, Dell'Acqua et al. (Organization Science)
- Generative AI at Work, Brynjolfsson, Li and Raymond (NBER, published in QJE 2025)
- Experimental evidence on the productivity effects of generative AI, Noy and Zhang (Science 2023)
- METR: Measuring the impact of early-2025 AI on experienced developer productivity (preprint) and METR's February 2026 update
- When combinations of humans and AI are useful, Vaccaro et al. (Nature Human Behaviour 2024)
- Complacency and bias in human use of automation, Parasuraman and Manzey (Human Factors 2010)
- The Impact of Generative AI on Critical Thinking, Lee et al. (Microsoft Research, CHI 2025)
- Generative AI without guardrails can harm learning, Bastani et al. (PNAS 2025) and Anthropic: How AI assistance affects coding skills
- Anthropic: Reduce hallucinations