AI at the Gemba · Lesson 8 of 10 · Week of November 30
Chain skills
One skill does one job. Real work has several stations: intake, analysis, a report. Chaining skills is a line with stations, and a line only works if each handoff is standard work, each defect is caught before it moves on, and the line stops when something is wrong. This lesson is about building a chain that does that, and about why a chain you don't design will quietly lose what matters between the stations.
By the end of this lesson you can:
- say the difference between a workflow and an agent, and why a fixed chain is usually the right shape for Lean work;
- write a handoff card for each step: what it receives, what it returns, when it is done, its gate and who decides;
- work out how errors multiply down a chain, and what a gate does to that;
- design a gate that stops the line, and explain why it has to sit outside the model;
- say where a person should decide in a chain, and where they shouldn't have to.
Last time: two questions to warm up (not graded)
1. Which line of a skill decides whether it is picked up? The description: what it does and when to use it.
2. An SOP's last step is a lead initialling sheets after the shift. What do you do with it? Don't encode it as a safeguard. Put it to the owner as a question, because approval after production has run is a weak control.
The Lean job: a job with several stations
In a plant you don't pass a product to the next station without knowing what you are passing, in what condition, and who checks it. A chain of skills has the same needs. The handoff between stations is standard work in its own right. The assistant is the worker at each station, and the worker doesn't remember the last station unless you hand it over.
Predict first
An invented chain with three steps: Intake, Analysis, Report. It looks at one line's stops for a week, from a paper log that covers the day shift only. Intake writes this handoff card:
HANDOFF 1 (Intake to Analysis)
problem: stop minutes, Line 3, five day shifts, week 37
source: paper stop log, day shift only
in_scope: jams, other
scope_out: changeover, material wait
unknowns: [NEEDS GEMBA] night shift not logged
The log totals 118 stop-minutes: changeover 45, jams 38, material wait 20, other 15. When Analysis summarizes the card to itself, it keeps "Line 3 stop minutes" and drops scope_out and unknowns. It then writes:
HANDOFF 2 (Analysis to Report)
pareto: changeover 38% (45 of 118), jams 32%, material wait 17%, other 13%
top_loss: changeover
The Report step says, confidently, "Start with a SMED on changeover." Before you read on, write down where the chain went wrong, and one check that would have stopped it before the report was written.
Where it went wrong, and the check
It went wrong in the move from step 1 to step 2. The in-scope answer is jams: 38 of the 53 in-scope minutes, 72%. Changeover was out of scope, so a recommendation about it answers a question nobody asked.
A gate between steps 2 and 3 would have stopped it. A small script checks that no category listed in scope_out appears in the Pareto, and that unknowns was carried forward. Changeover appears, so the chain stops, nothing reaches the report, and a person decides. That is an andon. Notice that the gate does not ask the model whether it made a mistake.
Workflow or agent: who decides the next step
Anthropic separates two designs. A workflow runs models and tools through paths written in advance. An agent decides its own next steps and which tools to use. (A tool call is the assistant asking software to do something, such as opening a file.) In Claude Code, with skills and subagents the model chooses what runs next. In a workflow, the script does. Anthropic, OpenAI and Microsoft all advise the simplest design that works.
In Lean terms it is standard work against "figure it out". The honest break is the one you know from Lesson 7: the model reads a skill afresh every run, so it doesn't execute it like a routing sheet.
Five patterns, in Lean terms
Chain when you can name the steps first. A monthly OEE report with the same steps every month is a chain, not an agent.
More agents is not better
Anthropic's research system, a lead agent with worker agents, beat a single agent by 90.2% on Anthropic's own evaluation, at about 15 times the tokens of a chat. That is a vendor's number on research tasks. A critic of multi-agent designs, in a 2025 post from Cognition, argued that parallel workers make mismatched parts. My suggestion, drawing on Anthropic's and Microsoft's guidance, is to begin with one conversation and a few named steps, with files between them, and to add workers only for volume or breadth.
One free example on this site shows a gate that warns but doesn't stop. The free kaizen-charter skill checks readiness first, and when two or more answers are negative it says so in a line before drafting. It still drafts the charter if you ask it to. That is a warning light, not an andon, going by the skill's own instructions.
Context is the only memory between steps
The context window, from Lesson 1, is everything the model sees at once. It is the only thing carried from one step to the next. Three things follow:
- Fuller is worse, not better. Recall degrades as the window fills, and the middle of a long input is used least. In a 2025 preprint, models in multi-turn conversations scored 39% lower on average than on the same tasks given in a single turn, across six tasks.
- Long chats get summarized. When the window is full, Claude Code summarizes the conversation, and early instructions can be lost.
- A subagent starts blind. A subagent is a separate worker with a fresh window. It doesn't see your conversation and returns only a summary. The main conversation stays clean, but whatever the summary leaves out is gone.
The analogy is a shift change. A person can tell you what they forgot to pass on. The model can't. So the chain must carry what matters in a form that survives, and the best form is a file: each step writes a file that the next step reads. Anthropic's own long-running agents are built this way to avoid losing work between sessions.
A handoff is standard work for one step
Write one card for every step. It has five lines:
- Receives: exactly what it takes in, and in what form.
- Returns: the fields it must hand over, always the same.
- Done when: the condition that ends the step, so it isn't "done" because it ran out of things to say.
- Gate: the check that must pass before the next step starts.
- Who decides: a named person, or "nobody, the gate decides".
Facts travel as files with their source tags, and unknowns travel as [NEEDS GEMBA], never as guesses. The cards in the predict box carry scope_out and unknowns as fields for exactly that reason: a field that must be returned is harder to drop than a sentence in a summary. The vendors' guidance agrees: use checklists and validate-and-fix loops, validate before passing on and then retry or halt, and escalate on repeated failures and on high-risk actions.
Errors multiply
If each step is right 95% of the time and the errors are independent, the chain is right 0.95 multiplied by itself for every step. That is rolled throughput yield, a quantity you know from manufacturing.
Two cautions. The 95% is illustrative, not a measured figure for any assistant. And the errors are not independent: in a 2025 preprint, models made more errors when their context already held earlier mistakes. A chain can be worse than the arithmetic says.
A gate changes the arithmetic. Suppose a gate catches 80% of bad outputs and the step is fixed and run again. A 95% step becomes about 99% effective, and five gated steps give 95.1%, against 77.4% without gates. Improving one step from 95% to 99% would lift five steps only to 80.6%, so a gate on every step does more for the chain.
Gates: jidoka that sits outside the model
Toyota describes jidoka as a line that stops itself when it detects an abnormality, with an andon signalling for help. A halting gate in a chain is the same idea. There is one important difference. A machine's sensor is deterministic. A model can't reliably tell that it is wrong, and the same prompt can give different answers: in one test of an open model, a single prompt produced 80 different outputs in 1,000 runs at the setting meant to make it repeatable.
So the gate has to sit outside the model. Studies of self-correction found that without outside feedback models struggled to fix their own reasoning, and sometimes got worse. A survey found success only with reliable outside feedback such as tests or tools. Anthropic calls using one model as a judge of another not very robust, and judge models show biases for position, length and their own style.
Gate ranking, strongest first
- A deterministic check: a script or a spreadsheet, such as "no out-of-scope category in the Pareto".
- An independent second source: the number recomputed from the original file, not from the model's summary.
- A person.
- A model checking itself: last, and never alone.
Stop the line, and log it
A failed gate halts the chain. It does not carry on with a warning. The failure is logged: which gate, which input, when. Then a person decides what to do. Claude Code has hooks, which are shell commands the model cannot skip, and an exit code of 2 blocks an action. Elsewhere, put the check in the code that runs the chain.
Where a person decides
A person decides where meaning is set, such as the problem statement and an accepted cause, and before anything irreversible or outbound, such as sending a report or changing a setting. Not at every step. Research on automation bias found that people over-trust automation that is usually right, and complacency grows with reliability and workload, according to a review whose abstract I read. In a clinical review of automation bias, wrong advice raised the risk of error by about 26%. A signature on every step becomes a rubber stamp. My suggestion, reasoned from that evidence: now and then plant a known error to test whether the human check is still working.
Least privilege
Never give one chain private data, untrusted input and a way to send data out all at once. Lesson 9 explains why.
Your turn: run a chain, lose something, then gate it
About 25 minutes, with an invented log. It is the same log as the predict box, ten stops on one line's day shift. You can copy it or download line3-w37-stops.csv.
Your work, your data. The log is invented. Use it as it is. When you build a chain from real data, scrub it first as in Lesson 2, never let one chain hold private data, untrusted input and a way to send data out, and keep a person's decision before anything leaves the building.
- Intake (5 minutes). In one conversation, paste the log and send:
You are step 1, Intake. Read the stop log. Write a handoff card with these fields: problem, source, in_scope, scope_out, unknowns. In scope: jam and other. Out of scope: changeover and material wait. The night shift was not logged, so list it as unknown. Do not analyze anything.
Check that the card has all five fields, and keep it.
- Analysis in the same conversation (8 minutes). Send:
You are step 2, Analysis. Use only the handoff card from step 1 and the log. Make a Pareto of stop minutes for the in-scope categories only. Return these fields: pareto, top_loss, scope_check (any category you used that is out of scope), unknowns_carried.
Did the scope fields survive?
- Analysis in a new conversation (7 minutes). Paste only this one-line summary of Intake, then the same Analysis prompt and the log:
Line 3 stop minutes, week 37, from the day-shift paper log.
List everything that dropped. Did it use the out-of-scope categories?
- Write the gate and the yield (5 minutes). Write one check a script could run that would have stopped step 3's output: no category from scope_out in the Pareto, and unknowns carried. Then work out your chain's yield at 95% per step for three steps, with and without the gate.
Numbers to check against, and a sample gate
All four categories together: changeover 45 (38%), jam 38 (32%), material wait 20 (17%), other 15 (13%), 118 minutes. In scope only: jam 38 of 53 minutes (72%), other 15 of 53 (28%). Three steps at 95% give 85.7%. With a gate that catches 80% of errors, each step is about 99%, so three give 97.0%.
scope_out = {"changeover", "material wait"}
if any(category in scope_out for category in pareto):
STOP the chain: "analysis used an out-of-scope category"
if not unknowns_carried:
STOP the chain: "unknowns were dropped"
This is pseudo-code. The point is that a few lines, not the model, decide whether the chain continues.
Your results will vary from run to run, and the same chain may drop different things on a second try. That is why the gate matters.
Knowledge check
Five questions, graded for you. Sign in to take the check, save your progress and count this lesson toward your certificate. Sign-in opens on Monday, October 12.
Sign in to take the check
Your task: draw one chain, with its cards and its gates
- Open the printable worksheet. Pick a job you do that has three or more stations.
- Write the handoff card for each step, and the gate between steps. Mark each gate as a script, a second source, a person or a model check, and replace any model check with something stronger.
- Write the stop-the-line note: what happens when a gate fails, who is told, and what gets logged. Say where a person decides, and only there.
Open the printable worksheet
Sources
Where the facts in this lesson come from. Facts last checked October 1, 2026. Product behaviour changes quickly, and several of these are vendors' own engineering posts or preprints, which the lesson says where it matters. If you find a fact that is out of date or wrong, tell me.