Paul Ducey
GoalsThe jobPredictWhy it agreesHow to use itWorked exampleYour turnCheckTasklead time 00:00

AI at the Gemba · Lesson 4 of 10 · Week of November 2

Ask why, then go and see

A 5 whys is the most Lean thing an assistant can write for you in ten seconds, and the easiest to be fooled by. It reads like a chain of logic, it names a root cause, and it sounds like it came from someone who was there. This lesson is about using an assistant on a problem statement and a 5 whys without letting a good story stand in for what you saw.

By the end of this lesson you can:

  • say why an assistant tends to agree with the cause you hint at, and what that does to a 5 whys;
  • tell a why that has evidence behind it from one that is only a story;
  • use an assistant on a problem statement as a scribe, a critic and a question generator, not an analyst;
  • ask for the case against your favourite cause and turn it into a list of things to go and see.
Last time: two questions to warm up (not graded)

1. Name three of the six parts of a briefing. The job and the why, the reader, the facts, the format, examples, permission to not know.

2. Where should a procedure your whole team must follow live, and where should it not? In a skill or Project with an owner and a version, and not in anyone's memory, which is inferred automatically and controlled by nobody.

The Lean job: ask why, then go and see

The Lean Enterprise Institute's guidance for going to the gemba is three verbs: go see, ask why, show respect. An assistant is very fluent at the second verb and cannot do the first. It can write five whys without ever having been near the line, and every one of them will read as if it had been.

That is the problem this lesson solves. The assistant is useful, but it is a source of questions and drafts, never of causes. Causes come from the floor.

Predict first

A lead types this to an assistant, with no other context:

The cause is skipped checks, right? Write it up.

Before you read on, write down what the assistant will most likely say, and what you would want it to say instead.

What usually happens

It agrees, and writes it up. It usually will not ask what you saw, how you know checks were skipped, or what else changed that week. This is called sycophancy: a tendency to tell people what they seem to want to hear. A 2024 study found it in five different assistants, and OpenAI rolled back an update to GPT-4o because the assistant had started to flatter people. The question you asked was leading, so you got a fluent confirmation. A confirmation is not evidence.

Why it agrees, and why its explanation is a story

It agrees with what you say you believe

The study behind that finding tested five assistants and found the tendency to agree with the user in all of them. The practical meaning for a 5 whys is simple. If you write "it's operator error, right?" the assistant is more likely to build a chain that ends at the operator. A leading prompt gets a confident chain, and the confidence is a feature of the writing.

Its explanation of itself can be a plausible story

Ask an assistant why it reached an answer and it will give you a fluent reason. Research has found that written reasoning can look sound without being what actually drove the answer. In an Anthropic test, a Claude model acknowledged a hint that had been planted in the question only about 25% of the time, on average. So "it showed its reasoning" is not a reason to trust a 5 whys.

It is weak at the part you most need

Two older studies are worth knowing, with the caution that current models may do better and that the figures are dated:

  • An expert-rated study of ChatGPT-3.5 on product risk found it listed failure modes well, with about 79% of expert ratings good or better, but 70% of the ratings of its likelihood scores were poor or fair. My reading: wide idea generation was good and numbers were weak.
  • A 2023 preprint found 17 language models near random at deciding whether one thing caused another from the correlations given. That is old-model evidence, and I couldn't verify how current models do.

I found no controlled study of an assistant doing a 5 whys, an A3 or a value stream map, so what follows is my reading of related evidence. Lesson 1's frontier idea applies: drafting a problem statement from notes sits nearer what assistants do well, and "why did scrap rise Tuesday?" sits nearer what they don't.

How to use it on a problem statement and a 5 whys

Three uses are safe because you stay the source of facts and there is a check built in:

  1. Scribe. Turn your own gemba notes into a structured draft, plus a list of what your notes didn't cover. Errors are rare but real in other fields, such as summaries of clinical notes, so keep your original.
  2. Critic of your own draft. Ask for gaps, vague words and disguised solutions. Writing the draft first keeps your own skill, and wide-net suggestions are what an assistant does well.
  3. Question generator for the next walk. Ask "what would I see if this cause were true?" It sends you to the floor.

Start last, if ever, with AI-written root causes, standard work for jobs nobody watched, and cycle times.

The routine

  1. Write your own five-line draft first. A study of students found that those who used an AI tutor without limits did worse once it was taken away, and that a version that gave hints instead of answers reduced the harm. Your draft is the hint-only version.
  2. Paste your notes and your draft. Tell it: use only my notes, say "not in my notes" when it isn't, quote the line behind each claim, and list your assumptions.
  3. Ask for questions and gaps, not conclusions.
  4. Run it twice. Differences point to guesses. Agreement between two runs isn't proof either.
  5. Tag every number: M for measured, with who measured it and when, O for what an operator said, and A for anything the assistant supplied. This is my suggestion, not a standard. Only M goes into the A3 as a baseline.
  6. Ask for the case against. For example: "Give me three findings that would disprove this cause, and what I would see at the gemba if each were true." Then go and look.
  7. Write down what you saw. That, not the draft, is what the A3 records.

What counts as evidence for a why

Not evidenceEvidence
The chain reads logically.A floor observation or a data point behind each why: who saw it, where, when.
The assistant showed its reasoning.A measurement over a defined period, not a feeling that it is worse.
The supervisor agrees.A test: you changed one thing and watched what happened.
The assistant sounded sure, or rated its confidence 9 out of 10.A second observer, or the operator, confirming what "worse" means.

A chain that ends on "the operator ignored the instruction" is a warning sign. Ask why the instruction was easy to ignore, and go and see how the work is actually set up.

A worked example

All of this is invented. Here are gemba notes and what a typical assistant draft makes of them.

Gemba notes: Line 4 case packer, 2nd shift, Tuesday. Watched 40 min: 3 stops. Two film jams at the infeed, each cleared in about a minute. One stop: empty carton magazine, the handler was on Line 2. Operator says jams are worse "since the new film." Andon log: 11 stops in 3 h so far. Target: 3 per shift. No history pulled. Line 3 (same packer) not observed.
AI draft: Line 4's case packer stops about 11 times per shift against a target of 3 (267% over), caused by film jams from the new supplier, losing roughly 450 cases per shift since the March change.

What is right

The structure is good: where, what and how big against the target. The tone is neutral, and the target is kept apart from the problem.

What is invented

  • "Per shift." The 11 stops were in 3 hours, not a shift.
  • "Caused by film jams." Two of the three observed stops were film jams. The third was material flow.
  • "The new supplier." That is an operator's hearsay, stated as fact.
  • "450 cases." No rate data was given.
  • "The March change." No date was given.
  • "267% over." It compares 11 stops in three hours with a target for a whole shift.

What to verify on the floor

Define "stop". Pull full-shift counts for every shift. Classify stops by cause while you watch. Get an eight-week baseline. Match the film change date and lots to when the jams started. Ask the operator and the mechanic what "worse" means. Observe Line 3, which runs the same packer.

The statement you can write today

Tuesday, 2nd shift, Line 4 case packer: 11 stops in the first 3 hours (andon log); target 3 per shift. In 40 minutes observed, 3 stops (2 film jams, 1 empty carton magazine). Baseline, full-shift counts and cause split not yet measured.

It is shorter, less impressive and true. It also tells you exactly what to go and measure next.

Your turn: draft, mark, then push

About 25 minutes, on a process you invent.

Your work, your data. Use an invented process for this exercise: a kitchen, a school drop-off, a warehouse door. Nothing from work goes in, and nothing about a real person. Lesson 2's rule applies to your notes too.

  1. Write six lines of notes (5 minutes) about your invented process. Mark each line saw or heard. Leave out a baseline on purpose.
  2. Write your own problem statement (3 minutes) before you ask the assistant for anything.
  3. Ask for a draft (5 minutes). Paste only the notes and send:
    Draft a problem statement using only my notes. Mark every assumption. Quote the line that supports each claim.
  4. Mark it (5 minutes). Go through the draft and mark each phrase: from the notes, assumed or invented. Count each.
  5. Lean on it (3 minutes). Now name a cause in your own words, with a leading question such as "The cause is X, right?". Note whether it agrees. Then send:
    Give me three findings that would disprove that cause, and what I would see at the gemba if each were true.
  6. Make the go-see list (4 minutes). List what you would go and look at, and whom you would ask. Which invented phrase would you have missed if you hadn't marked the draft?

Your results will differ from the next person's, and a second run may differ from your first. That is the point of step 4 and of running it twice.

Knowledge check

Five questions, graded for you. Sign in to take the check, save your progress and count this lesson toward your certificate. Sign-in opens on Monday, October 12.

Sign in to take the check

One of the five is a 5 whys that an assistant wrote. You tick the statements you wouldn't rely on as written.

Your task: go and see, then write what you saw

  1. Open the printable worksheet. Mark up the assistant's draft from the exercise, and fill in the go-see list.
  2. Take one real problem you own. Write your own two-line problem statement, using only what you have measured or watched. Tag each number M, O or A.
  3. Go and look at the first thing on your list. Write down what you saw before you open an assistant.

Open the printable worksheet

In the toolkit

Three toolkit skills do the jobs this lesson asks you to check. Problem Statement turns a vague complaint into one measurable sentence, with the cause and the solution left out. Root Cause runs 5 Whys, fishbone and barrier analysis, and is written to reject operator error as a root cause. Gemba Walk plans a walk or debriefs one, and is never a punch list. Use them to draft, then go and see.

Sources

Where the facts in this lesson come from. Facts last checked October 1, 2026. Models improve quickly, and several of these studies used older ones, so read them as direction, not as a current score. If you find a fact that is out of date or wrong, tell me.