Chapter 3 of 522% of exam

Analyze — Verifying Root Causes

Analyze is where the team moves from 'what is happening' to 'why it is happening,' and — crucially — verifies the answer rather than guessing it. Green Belts brainstorm potential causes with graphical and structured tools, then use basic statistics to test which suspected inputs actually drive the output. The phase discipline is that a cause is not a root cause until data confirms it; opinion narrows the list, evidence closes it.

Structuring causes: fishbone and the 5 Whys

The Ishikawa (fishbone / cause-and-effect) diagram organizes the team's hypotheses about causes into major branches so nothing is overlooked. In manufacturing the classic 6M categories are Man, Machine, Method, Material, Measurement, and Mother Nature (Environment); service processes often use the 4Ps (People, Process, Policies, Plant/Technology). The diagram generates candidate causes but does not prove any of them. The 5 Whys drills from a symptom toward its origin by repeatedly asking 'why' — 'errors occur → why? operators mis-key → why? no validation on the form → why? the form was never updated → root cause.' Progressing from a surface symptom through successive whys to the underlying systemic cause is called root cause analysis; the number five is a guideline, not a rule.

Pareto analysis: focusing on the vital few

A Pareto chart is a bar chart of defect categories sorted from most to least frequent, with a cumulative-percentage line, expressing the 80/20 principle: a small number of categories usually account for most of the problem. Its Analyze role is to direct limited effort at the vital few causes that will return the most, rather than spreading resources across the trivial many. Worked reading from the bank: of 500 defects, 3 of 12 defect types account for 410 defects — the correct Analyze conclusion is to concentrate the investigation on those three types, because they represent 82% of the total. The chart prioritizes; it does not by itself explain the cause of those top categories, which is the next step.

Hypothesis testing: the logic of significance

A hypothesis test decides whether an observed difference is real or just sampling noise. The null hypothesis (H₀) states there is no effect or no difference — the status quo — and the alternative (H₁) states there is one. You compute a p-value, the probability of seeing data at least this extreme if H₀ were true, and compare it to a preset significance level α (commonly 0.05). If p < α you reject H₀ and call the effect statistically significant; if p ≥ α you fail to reject H₀. A p-value of 0.02 against α = 0.05 means reject the null — the effect is significant at the 5% level. Two error types define the risks: a Type I error (α) is rejecting a true null — a false alarm, concluding an effect exists when it does not; a Type II error (β) is failing to reject a false null — a miss, overlooking a real effect. Power = 1 − β is the ability to detect a true effect, and it grows with sample size.

Choosing the right test: t-tests, ANOVA, chi-square

Matching the test to the data types is heavily tested. A two-sample t-test compares the means of exactly two groups on continuous data (e.g., cycle time on line A vs. line B). One-way ANOVA compares the means of three or more groups at once; it is used instead of many separate t-tests because running multiple pairwise tests inflates the overall Type I error rate, while ANOVA holds α across the whole comparison. ANOVA's F-statistic is a ratio of between-group variance to within-group variance — large when the group means differ more than individuals within groups do. Degrees of freedom for one-way ANOVA with k groups and N observations: between = k − 1, within = N − k; for 4 groups and 40 observations that is 3 and 36. The chi-square test of independence works on counts in a contingency table to test whether two categorical variables are associated; its degrees of freedom = (rows − 1) × (columns − 1), so a 3×4 table has (3−1)(4−1) = 6. Rule of thumb: two group means → t-test; 3+ means → ANOVA; categorical vs. categorical counts → chi-square.

Correlation and regression: association vs. causation

Correlation quantifies how strongly two continuous variables move together. The correlation coefficient r ranges from −1 to +1: the sign gives direction and the magnitude gives strength, so r = 0.85 indicates a strong positive linear relationship. Regression goes further and models the relationship as an equation — simple linear regression fits Y = b₀ + b₁X, letting you predict the output from an input and estimate how much Y changes per unit of X. R² (the coefficient of determination) reports the fraction of variation in Y explained by the model. The recurring exam warning: correlation shows association, not proof of causation. Two variables can correlate because a third lurking variable drives both, or by coincidence. A causal claim must be confirmed by a controlled experiment (DOE) in the Improve phase, not asserted from a scatterplot.

FMEA: prioritizing failure risk with RPN

Failure Mode and Effects Analysis (FMEA) is a structured way to rank where a process is most likely to fail and hurt the customer, so prevention effort goes where risk is highest. For each potential failure mode the team rates three factors on a 1-to-10 scale: Severity (how bad the effect is), Occurrence (how likely it is to happen), and Detection (how likely current controls are to catch it before it reaches the customer — note the scale is inverted, so 10 means the failure is almost impossible to detect and 1 means it is almost certain to be caught). The Risk Priority Number is their product: RPN = Severity × Occurrence × Detection. Worked example: S = 8, O = 4, D = 5 gives RPN = 8 × 4 × 5 = 160. The team attacks the highest RPNs first, and a very high Severity alone (safety) can warrant action regardless of RPN. FMEA is used both to analyze existing risk and, in later phases, to confirm that mitigations lowered Occurrence or improved Detection.

Keep going: the full Lean Six Sigma Green Belt (IASSC ICGB) guide covers every section of the exam. Lean Six Sigma Green Belt — Complete Study Guide (2026) — PDF + EPUB, $19.99 · 14-day refund →

Studying in order?

Practice stays free. The full Lean Six Sigma Green Belt (IASSC ICGB) study guide is the material itself, taught start to finish — a downloadable PDF + EPUB you keep.

Get the book — $19.99
Report