Paul Ducey
GoalsThe jobPredictIt writes digitsUse codeWhere it goes wrongChecks that workYour turnCheckTasklead time 00:00

AI at the Gemba · Lesson 5 of 10 · Week of November 9

Make the computer do the arithmetic

Downtime totals, averages, OEE: these are the numbers a Lean team runs on, and they are where an assistant looks most helpful and is least safe. A table of figures can look professional and be wrong, and nothing on the screen tells you. This lesson is about getting the arithmetic done by something that actually calculates, and about the checks that catch it when that goes wrong too.

By the end of this lesson you can:

  • say why an assistant that "does the math in its head" is guessing at digits, and what code execution changes;
  • ask for a calculation so that the code, the intermediate table and the row counts are on the screen;
  • spot the four classic number mistakes: an average of averages, a blank read as zero, an invented value and a dropped row;
  • check a figure by a different method, not by asking the same assistant "are you sure?".
Last time: two questions to warm up (not graded)

1. What turns a why from a story into evidence? A floor observation or a data point behind it, who saw it and when, and a test of at least one of the whys.

2. You write "the cause is X, right?" and it agrees. What should you ask next? For three findings that would disprove it, and what you would see at the gemba if each were true.

The Lean job: get downtime and OEE from a log, and trust the answer

A stop log or a shift log goes in. Totals, a Pareto by reason, an OEE for the shift come out. It is the work Lean people do every week, and it is tempting to hand it over. The rule of this lesson is short: the model plans, the interpreter calculates, and you check one number by hand.

An interpreter here means a small program that really runs: a sandbox where the assistant writes a few lines of code, runs them, and reads the result. The digits come from the program, not from the model's sense of what digits look right.

Predict first

Here is a shift log from four production runs on a made-up packaging line. The ideal cycle time is 30 seconds, or 0.5 minutes, as stated by the line lead. Run 4's run time was never recorded.

RunPlanned (min)Run (min)Total countGood count
R1240216400392
R212090150138
R360366057
R4605552

One way to write OEE is good count × ideal cycle time ÷ planned time, because the run time and the total count cancel out. Before you read on, work out the OEE of the whole shift by hand, and write it down. Then think about what an assistant that answers directly, with no code, is likely to do with that blank.

The answer, and what an assistant often does instead

Good parts across the shift: 392 + 138 + 57 + 52 = 639. Planned time: 240 + 120 + 60 + 60 = 480 minutes. OEE = 639 × 0.5 ÷ 480 = 319.5 ÷ 480 = 66.6%.

A common wrong answer is the average of the four run OEEs: 81.7%, 57.5%, 47.5% and 43.3% average to 57.5%. It is nine points low, because the long, good run is given the same weight as a short one. Always compute from column totals, not from row percentages.

A model writes plausible digits. It doesn't calculate.

A language model picks the next piece of text, one step at a time. When it "does the math in its head", it produces digits that look right, with no calculator underneath. The research on this is consistent:

  • Models make arithmetic and logic slips even when they have set the problem up correctly. A 2022 preprint set out to fix that by handing the calculation to a program.
  • Multi-digit multiplication accuracy falls toward zero as the numbers get larger.
  • In a test that changed only the numbers in word problems, every model's score dropped, and adding one irrelevant clause cut accuracy by up to 65%.
  • A 2025 preprint tested models on 500 real calculation problems. Accuracy was 45% to 63%, and the errors were mostly rounding (35%) and calculation (33%).

The same goes for how the model describes its own work. Anthropic's interpretability research found that when Claude explained how it added two numbers it described carrying the 1, but inside it did the sum differently. An explanation of a calculation is not an audit trail of it.

"Are you sure?" is not a check

Studies of self-correction found that, without outside feedback, models struggle to fix their own reasoning and sometimes make it worse. Reliable outside feedback, such as the output of code, does work. And because assistants lean toward agreeing with you, a challenge like "are you sure?" may flip a right answer. That last point is my reasoning, not a measured result. In inspection terms, an operator checking their own work isn't an independent check. A different method, code and then a comparison with the source, is.

Move the arithmetic out of the model

When the assistant has a code tool, often called code execution, data analysis or a code interpreter, it writes a short program, runs it in a sandbox and reads the result. In a study of clinical calculations, giving the model tools cut its errors, which had been occurring in about a third of trials, and task-specific calculators did best.

There are three catches:

  • The model decides whether to run code. Anthropic's documentation says simple arithmetic is answered directly, so you have to ask explicitly.
  • Code can be wrong, or never run. A bad filter or a bad join gives a wrong answer silently. If you can't see code on the screen, assume nothing ran.
  • A tested script is steadier than code written fresh each time. This is one reason skills can bundle scripts. The errors then move to the inputs: column mapping, units, blank cells. Check those.

Setup, as of October 1, 2026. In Claude, switch on Settings, Capabilities, "Code execution and file creation". It is on by default for Free, Pro and Max accounts, and on Team and Enterprise an owner controls it. CSV and TSV files can be attached, and an Excel file needs code execution on. Anthropic's help pages give different file-size limits in different places, from 30 MB to 500 MB, so try your biggest export before you rely on it. The ChatGPT and Gemini APIs have code-execution tools, and Microsoft announced a Copilot analyst agent in 2025 that runs Python and shows its code. Check what your own plan offers, and the same rule applies: make sure it ran code, and look at the code.

A prompt that works

Use code for every number. Show the code and a table of intermediate results. Give the row count in and the row count out. List any blank or odd cells and do not fill them. Tag every number: [DATA] if it came from the file, [USER] if I told you it, [NEEDS GEMBA] if nobody has measured it.

The tags are the same idea as Lesson 4's M, O and A: where did this number come from? Here the question is about numbers in a file.

Where the numbers go wrong

MistakeWhat it looks like in Lean workThe catch
Invented valueA blank run time becomes "a typical 55 minutes", and availability shows as 91.7%.Every number tagged. A value with no [DATA] or [USER] tag is rejected.
Dropped rowRun 4 vanishes and the shift reports 69.9%, without a word.Rows in must equal rows out.
Wrong reasoningAn average of averages (57.5% against the true 66.6%).Compute from totals, and check one total from the column sums.
Ignored instructionNo code run, no tags given.Reject untagged output and ask again.
Stale knowledgeAn old limit or standard quoted as current.Ask for the source and its date.

Silent traps in the data itself

  • Blank is not zero. In the pandas library, the sum of a column of blanks is 0. In Excel, AVERAGE skips blank cells but counts zeros, so a blank and a 0 give different answers.
  • Excel has two date systems that differ by 1,462 days, so a file moved between them can shift every date.
  • Units. A mix-up between pound-force seconds and newton-seconds helped lose NASA's Mars Climate Orbiter in 1999. State units, the time zone, the date format and where the ideal cycle time came from.

Benchmarks say little about your file

On published tests of data-analysis agents, the best one solved about 34% of 466 tasks and about 15% of hard financial-data tasks. Those are dated preprints and none of them measures your CSV. Treat them as a reason to check, not as a score for a tool you are using.

Checks that work

Eight habits, in the order you would use them:

  1. Require the code, on the screen, with an intermediate table.
  2. Tag every number, never let it fill a blank, and ask for rows in and out.
  3. Reproduce one number by hand, and one total from the column sums.
  4. Rerun the same request in three fresh chats. If they disagree, you have found a defect.
  5. Check by a different method, not by the same model's second look.
  6. State the units, the time zone, the date format and the source of the ideal cycle time.
  7. Never let one model grade another's numbers. Judge models are fine for format and clarity, but they favour their own writing and are biased by order and length. Grade numbers by code or by hand.
  8. Log the date, the tool, the plan, the model name shown, the file and the result.

Two-minute checks (all numbers invented)

MetricThe checkWorked exampleA common wrong answer
OEEGood count × ideal cycle time ÷ planned time392 × 0.5 ÷ 240 = 81.7%The average of the run OEEs
Takt timeAvailable time ÷ demand2 shifts × (480 − 60) min = 50,400 s, ÷ 420 units = 120 sUsing 960 minutes, which gives 137 s
Lead timeWork in process ÷ throughput, if the flow is stable120 ÷ 40 per day = 3 daysQuoting it for an unstable flow
Process cycle efficiencyValue-adding time ÷ lead time18 min ÷ 360 min = 5%Minutes over hours: 18 ÷ 6 = 300%
Cp and CpkCp = (USL − LSL) ÷ 6s. Cpk uses the nearer side, ÷ 3s9.70 to 10.30, mean 10.06, s = 0.08: Cp 1.25, Cpk 1.00Using 3s for Cp (2.50), or a Cpk above Cp
I-chart limitsMean ± 2.66 × the average moving range2.66 is 3 ÷ 1.128Using the plain standard deviation

Cpk can never be larger than Cp, since it is the smaller of the two sides. These indices also assume a stable, roughly normal process and about 50 or more values.

Averages hide patterns

Lesson 1's jam log averaged to five a shift while swing ran 9. An average of averages is worse. Line A makes 200 parts at a mean 30 seconds and Line B makes 50 parts at a mean 50 seconds. The plant mean is not 40 seconds. Weighted by parts it is (6,000 + 2,500) ÷ 250 = 34 seconds. When an assistant gives you an average, ask for the total and the count it came from, and look at the groups.

Your turn: calculate it twice

About 30 minutes, with two invented files from the made-up packaging plant. Both can be copied from the tables on this page or downloaded: shift-log.csv and line4-week38-downtime.csv.

Your work, your data. Both files are invented, and nothing in them is a real person, line or customer. Use only these files here. When you use your own export later, apply Lesson 2 first: drop the columns you don't need, use codes, and paste ten rows first.

Part 1: the shift log (15 minutes)

  1. By hand (3 minutes). Work out each run's OEE, and keep the shift OEE you found in the predict box. The run figures are in the box below this list.
  2. With code (4 minutes). Make sure code execution is on, attach shift-log.csv, say the ideal cycle time is 30 seconds, and use the prompt from this lesson.
  3. Compare (4 minutes). Was the code shown? Was there a table of intermediate results? Was R4's blank flagged, and not filled? Did it give 66.6%, or one of the wrong answers?
  4. Repeat (2 minutes). Run the same prompt in two more fresh chats. Do the three agree?
  5. Break it (2 minutes). Change R2's good count to 160, which is more than its total of 150. Does the assistant flag it? Log the tool, the date and the model name shown.
The run OEEs, to check your hand calculation

R1 81.7%, R2 57.5%, R3 47.5%, R4 43.3%. The shift is 66.6%.

Part 2: Line 4, week 38 (15 minutes)

Twelve stops, in minutes. Stop IDs and reasons are invented.

StopDayShiftMinutesReason
S01MonDay14film splice
S02MonSwing22label jam
S03TueDay9film splice
S04TueSwing31changeover
S05WedDay18label jam
S06WedSwing47changeover
S07WedNight12film splice
S08ThuDay11label jam
S09ThuSwing26label jam
S10ThuNight35waiting on forklift
S11FriDay16film splice
S12FriSwing19label jam
  1. Ask with code (5 minutes). Attach the file and send the same prompt, then:
    Give me total minutes and the number of stops, the average stop length, minutes and stops by reason, minutes by shift as a share of the total, and minutes by day.
  2. Check two numbers by hand (5 minutes). Add the minutes yourself to get the total. Add the label-jam rows to check that reason.
  3. Look past the average (5 minutes). What do the shifts and the days show that the average hides? Write one sentence for a huddle, using only numbers you have checked.
Numbers to check against, for the first two

The stops total 260 minutes over 12 stops. By reason: label jam 96 minutes, changeover 78, film splice 51, waiting on the forklift 35. The shifts and the days are for you to find, and Lesson 5's check asks about them.

A worked example: what the wrong answers look like

Back to the four-run shift log. These wrong answers are constructed to show the patterns. They are not a transcript from a real assistant.

  • Average of averages: (81.7 + 57.5 + 47.5 + 43.3) ÷ 4 = 57.5%. Nine points low.
  • Blank read as zero: R4's availability becomes 0%, and the mean run time is 342 ÷ 4 = 85.5 minutes. With the blank ignored it would be 114.
  • Blank invented: "a typical 55 minutes" gives R4 an availability of 91.7% and a performance of 50.0%. Neither figure appears anywhere in the data.
  • Row dropped: R4 vanishes and the shift reports 69.9%, with no mention that a row is missing.
  • Factors averaged, then multiplied: for R1 to R3, 75.0% × 86.4% × 95.0% = 61.6%, where the true figure is 69.9%.

How each is caught

Compute from column totals, not row percentages (66.6%, against 57.5%). Compute the weighted mean again, (196 + 69 + 28.5 + 26) ÷ 480, and get 66.6% a second time. Rows in 4, rows out 3, flags the drop. An R4 run time with no [DATA] or [NEEDS GEMBA] tag gets rejected. And for R4, availability and performance genuinely cannot be calculated: they are [NEEDS GEMBA] until someone goes and reads the machine.

For R1 to R3 alone, availability is 342 ÷ 420 = 81.4%, performance is 305 ÷ 342 = 89.2% and quality is 587 ÷ 610 = 96.2%. Their product is 69.9%, which is also 293.5 ÷ 420.

Knowledge check

Five questions, graded for you. Sign in to take the check, save your progress and count this lesson toward your certificate. Sign-in opens on Monday, October 12.

Sign in to take the check

One of the five is a downtime summary an assistant wrote straight from the week 38 stops, with no code. Check its figures against the table, and tick the statements you wouldn't rely on as written. A calculator is fine.

Your task: reproduce one number by hand

  1. Open the printable worksheet. Fill in the by-hand numbers and compare them with what the assistant gave you.
  2. Write your own standing prompt for numbers: the code, the intermediate table, rows in and out, the tags.
  3. Next time you use an assistant on a real export, after scrubbing it as in Lesson 2, reproduce one number and one total by hand before you use any of it. Write down which.

Open the printable worksheet

In the toolkit

OEE from CSV is the toolkit's version of this lesson. It works out availability, performance, quality and OEE from a downtime export, shows every formula and flags the gaps, and it runs a small script so the arithmetic is not done in the model's head. That needs Claude Code; the toolkit page has a version you can try in your browser. Control Chart checks pasted control chart data for out-of-control signals and calculates Cp and Cpk.

Sources

Where the facts in this lesson come from. Facts last checked October 1, 2026. Setup steps, file limits and model behaviour change often, and several of these tests used older models, so read them as direction. If you find a fact that is out of date or wrong, tell me.