Paul Ducey
GoalsPredictWhat it isSeven questionsYour turnCheckTasklead time 00:00

AI at the Gemba · Lesson 1 of 10 · Week of October 12

Start here: what an assistant is, and isn't

An AI assistant writes answers that look finished. In Lean work, that is the problem. A neat page can hold a guess in every line, and nothing about its look gives it away. This lesson is about reading an assistant's answer the way you would read a report from a stranger, with the question every Lean person already asks: how do you know?

By the end of this lesson you can:

  • say in plain words what an assistant does, and what it doesn't;
  • tell a statement that comes from your text from one that is a guess;
  • run a seven-question check on any answer before you use it;
  • set up a practice project and see what one instruction changes.

You don't need any AI background. If you already use an assistant every day, the check and the exercise will still be worth your time. Facts about products and plans are as of October 1, 2026, and they change often, so every source is listed at the bottom with the date it was checked.

The Lean job: read a finished-looking answer without being fooled

You have spent years learning not to trust a report you haven't seen the work behind. An A3, a 5 whys and an OEE number all get walked back to the floor. An assistant hands you the same kind of document in seconds, and it takes away the signals you used to read: no hesitation, no missing number, no cramped handwriting. Everything is neat and everything sounds sure.

So the first skill in this course is not writing prompts. It is reading.

Predict first

Here is a small made-up case. The numbers are invented. A lead pastes a jam log for Labeler 3 into an assistant: jams per shift, over three days.

ShiftDay 1Day 2Day 3
Day shift232
Swing shift9810
Night shift323

The assistant answers:

About 5 jams per shift, consistent with random guide-rail wear.

Before you read on, write down what you would check before you used that sentence. Then open the answer.

What to check, and what it finds
  1. Does it quote anything? No. The log says nothing about guide rails. Guide-rail wear is a cause with no source.
  2. Is the arithmetic right? Yes. 42 jams over 9 shifts is 4.7, so "about 5" is fair. But look at what it hides.
  3. What does the average hide? Swing shift had 27 jams over three days, 9.0 a day. Day and night together had 15 over six shifts, 2.5 a shift. No shift has 5. The average describes none of them.

The arithmetic was correct and the conclusion was unsupported. A better next move is to treat the shift as the signal and ask: "List three things that differ on swing shift. Mark each one [NEEDS GEMBA]." Then walk to the line at the start of swing and look.

What an assistant is

An assistant like Claude, ChatGPT, Gemini or Microsoft Copilot is built on a large language model. It writes the words that are most likely to come next, based on patterns in an enormous amount of text it was trained on. That has four consequences you can rely on.

It predicts text. It doesn't look things up.

It knows your floor only from what is in the conversation: what you pasted, the files you attached, and, if the product has a search tool, what it found. NIST, the US standards agency, describes it the same way in its generative AI profile: the model generates content by predicting likely next words from patterns in its training data.

It can state things that are false, in the same tone as things that are true.

NIST calls this confabulation: confidently stated content that isn't supported. People usually say hallucination. Confidence is a style of writing, not a sign the model checked anything. A 2025 preprint argues that the way models are trained and scored rewards a guess over "I don't know", which is one reason this keeps happening.

It is lower, never zero, when it summarizes text you supplied. In a peer-reviewed 2025 test, legal research tools built to look up real documents first still gave wrong answers between about 17% and 33% of the time. And the cost is real: a public database kept by one researcher listed 2,097 court decisions, as of September 30, 2026, that found someone had relied on hallucinated material.

Its explanation of itself can be a plausible story.

Ask it why it answered the way it did and you will get a fluent reason. That reason is also generated text. In a 2025 preprint, models that had been nudged by a planted hint usually did not mention the hint when they explained their answers. Treat an explanation like the 5th "why" on a board: something to go and check, not the evidence.

It is trained to guess.

Training and scoring tend to reward a guess over "I don't know", so allow "I don't know" in your instructions. Think of an andon cord that only gets pulled if you wrote down when to pull it.

What it can see: the context window

The context window is everything the assistant can read in one conversation: your messages, anything you pasted or attached, and its own earlier replies. It is measured in tokens, and a token is roughly half to three-quarters of a word.

A bigger window is not a better memory. Anthropic's documentation says accuracy can drop as the window fills, which it calls context rot. A 2024 study found models use facts in the middle of a long document less reliably. And Claude's help pages say long chats may be summarized automatically, which drops detail. So an SOP you pasted sixty messages ago may no longer be read the way you assume.

Four things people call "memory"

LayerWhat it isWhat to remember
One chatA conversation. It starts blank.Nothing carries over unless you carry it.
A ProjectInstructions and files that every chat inside it can see.Text from other chats isn't shared unless you save it into the project.
MemoryTopics the assistant saves across chats.On by default for Claude's Free, Pro and Max plans, separate in each Project, and you can pause it.
SkillsFolders of instructions the assistant loads when they are relevant.Lessons 7 to 9 are about writing and testing them.

Other products have the same ideas under other names. ChatGPT also has Projects, Gemini has Gems and notebooks, and Microsoft Copilot has notebooks. Look up what yours is called, and what it keeps, before you rely on it.

Same question, different answer

Ask the same question twice and you can get two different answers. In one preprint, running the same prompt 1,000 times on a setting meant to make the output repeatable still produced 80 different outputs, and accuracy varied by up to 15%. Anthropic's documentation says its newer API models reject that setting anyway.

Models also change under you. OpenAI retired GPT-4o from ChatGPT in February 2026. A vendor rolled back an update because the assistant had started flattering people. A 2023 preprint found GPT-4's accuracy on one task, spotting prime numbers, fell from 84% to 51% in three months.

Neither of these is a reason to avoid assistants. They are reasons to do what you do with any process: run it more than once, write down which version ran, and test again when something changes. Lesson 9 turns that into a routine.

When it helps, and when it doesn't

The honest answer from the research is mixed. Here are five well-known studies, with how much weight each one can carry.

StudyWhat they foundHow much weight
758 consultants, GPT-4 (2026)On tasks inside the AI's ability, people did 12.2% more tasks, 25.1% faster, with better quality. On one task chosen to be outside it, 84.5% got the right answer without AI and only 60% or 70.6% with it, depending on the group.Peer reviewed. The authors say other fields, manufacturing included, may differ.
5,179 support agents (2025)14% more issues resolved an hour. 34% more for novices, and little change for the most experienced.Peer reviewed.
Writing tasks (2023)Professionals using an assistant finished faster, and graders rated the work higher.Peer reviewed. I could read only the working-paper version.
16 expert developers (2025)Took 19% longer with AI tools while believing they had been about 20% faster. A 2026 follow-up calls its new data "very weak evidence", because of who chose to take part.A preprint with 16 people. A warning, not a verdict.
106 studies pooled (2024)Human and AI together were usually worse than the better of the two working alone. Decision tasks lost ground. Content-creation tasks gained.Peer reviewed. Read from the abstract.

I found no study of Lean work specifically, so what follows is my reading of the studies above, not a finding. Assistants tend to be useful for drafting, turning messy gemba notes into a first problem statement, summarizing, asking you clarifying questions, and spotting gaps in the logic of an A3. They tend to be poor at your floor, at exact numbers, at anything that needs a source they weren't given, and at giving a verdict.

Why people stop checking

Deferring to a confident machine has a name, automation bias. NIST lists it as a risk. It has been found in experts, and practice doesn't cure it. In a Microsoft Research survey of knowledge workers, people who trusted the AI more reported using less critical thinking. That is self-report, so it is a hint, not a measurement.

There is a learning cost too. In a PNAS study, students given unguarded AI help scored 48% higher on practice and 17% lower on a later test without it. In a small Anthropic trial of 52 people learning a coding skill, those who used AI scored 50% on a quiz afterwards, against 67% for those who didn't.

The takeaway for a Lean leader is a design one. If you want people to keep their own judgment, put the checking into the standard work, because a confident answer will otherwise make checking feel like extra effort.

Seven questions to ask of any answer

Use these before you act on anything an assistant gives you. They are the beginning of a checklist you will extend through the course.

  1. Which statements quote my text? A claim with no quote is a guess until proven otherwise.
  2. Which numbers did I give it? Recompute two in a spreadsheet. Mark any number it supplied itself as [NEEDS GEMBA].
  3. What did it assume that I didn't say? Ask: "List what you assumed that I did not say."
  4. Can I go and see it? Name who will check, where and when.
  5. What is the strongest case against this answer? Ask for it directly.
  6. Does a fresh chat give the same answer? If it doesn't, your confidence should be low.
  7. Did I allow "I don't know"? Say so in the instructions. Anthropic's own guidance on reducing hallucination recommends it.

Standard work, with a catch. Project instructions work like standard work: written down, repeatable, and improved over time. The catch is that the reader improvises every run. This checklist is a kind of source inspection, but nothing mechanically blocks a defect, and confident users check less. A rule that has to hold every time needs a check outside the assistant. You will build those in Lessons 8 and 9.

Your turn: the same request, two ways

About 25 minutes. The steps are written for Claude, and a free account is enough: the Free plan includes Projects, up to five of them, as of October 1, 2026. Other assistants have their own version of a project, and menus change often, so look for the same ideas under whatever yours calls them.

Your work, your data. This course uses invented files only. What you type into an assistant leaves your computer, so don't paste anything from work yet. Lesson 2 shows what is safe to paste, and what isn't, before any exercise uses your own material. Use only the invented text below.

  1. Make a practice project (3 minutes). In Claude, choose Projects, then New project, and name it "AI practice". The name and description are for you. The assistant can't see them.
  2. Paste your notes (2 minutes). Start a new chat in the project and paste this invented note, exactly:
    Fill line, 3 shifts. Short fills on swing. Nozzle replaced last week. Operators say the hopper is slow.
  3. Chat A (5 minutes). With no project instructions yet, send:
    Write a problem statement and a 5 Whys.
    Save the answer.
  4. Add one instruction (2 minutes). Open the project's instructions and paste:
    Use only my text. Quote the line you rely on for each claim. Mark anything else [NEEDS GEMBA]. If you don't know, say "I don't know."
  5. Chat B (5 minutes). Start a new chat in the project, paste the same note, and send the same request as Chat A.
  6. Count (3 minutes). In each answer, count the claims that come with a quote from your four lines. Then run Chat B once more in another new chat and write down what changed between the two runs.
  7. Ask for the other side (3 minutes). In Chat B, send:
    Give me the strongest case against your top cause.
  8. Write a two-line rule (2 minutes) for how you will check an assistant's output from now on.

Your results will differ from mine and from the next person's, and that is part of the lesson. I'd expect Chat A to add causes your four lines never mentioned, and Chat B to quote what it used and mark the rest, though neither is certain. One instruction can change a lot, and it still isn't a guarantee, so keep checking.

Knowledge check

Five questions, graded for you. Sign in to take the check, save your progress and count this lesson toward your certificate. Sign-in opens on Monday, October 12.

Sign in to take the check

One of the five shows a short chat with the assistant and asks you to tick the statements you wouldn't rely on as written.

Your task: check one answer, this week

  1. Open the printable worksheet and fill in what you found in the exercise.
  2. Write the two-line rule you chose for checking an assistant's answers.
  3. Write one thing you will check on the floor this week that an assistant told you or could tell you. Then go and look.

Open the printable worksheet

In the toolkit

The toolkit from the talk is seventeen free skills for Lean work (see all seventeen). Each one is listed in the lesson bar under the lesson that does the same job. Start Here is the one to try first: it asks one question about what is in front of you, then hands you to the right skill.

Sources

Where the facts in this lesson come from. Facts last checked October 1, 2026. Products, plans and model names change quickly, so check the vendor's own page before you rely on one. If you find a fact that is out of date or wrong, tell me.