Prepare for the Certified Six Sigma Green Belt by treating every DMAIC phase as a decision point, not a tool list. When you read a scenario, ask what choice the belt faces: trust this data or fix the measurement system first, compare these groups with a paired or independent test, or declare this process stable, capable, or neither. Study the concepts that force those distinctions — control versus capability, gage studies versus process studies, paired versus two-sample tests — and drill with written cases until you can justify the choice and state what evidence would change it. For administrative details such as eligibility, format, and reference policies, rely on the issuer's official page rather than secondary summaries.
Map each DMAIC phase to the decision it forces, not to its tool list
Treat Define through Control as a chain of commitments: what problem matters, whether the data can be trusted, what drives the output, which fix to pilot, and how to keep the gain. Name the decision first, then pick the tool that serves it.
In Define and Measure, the decisions are about scope and evidence quality. Define forces you to convert a vague complaint into a problem statement with a measurable output, a bounded process, and a customer who feels the defect. Measure then asks a question people skip: is the measurement system good enough to detect the change you hope to make? A baseline number from an unverified gage is not evidence; it is a hypothesis about your instruments. Practicing this means writing one sentence per phase stating what you would commit to before naming any tool.
Analyze, Improve, and Control each carry a different burden of proof. Analyze must demonstrate that a suspected input actually moves the output, using stratification, regression, or designed comparisons rather than anecdote. Improve must show the fix works under realistic conditions, ideally confirmed with a small designed experiment or pilot rather than a full rollout. Control must prove stability over time, which requires control charts, reaction plans, and an owner, not a celebratory before-and-after snapshot. When you rehearse scenarios, state which burden of proof applies before you choose between tools.
- Define: convert a complaint into a measurable problem statement with scope and customer impact.
- Measure: verify the measurement system before trusting any baseline.
- Analyze: demonstrate a cause-effect relationship, not just a correlation in a fishbone.
- Improve: confirm the fix with a pilot or designed trial under realistic conditions.
- Control: institutionalize the gain with charts, reaction plans, and a named owner.
Statistical control is not capability: separating two verdicts the exam contrasts
A control chart answers whether the process is stable and predictable over time; a capability index answers whether that stable output meets specifications. A process can be in control and out of spec, or vice versa, and the correct response differs in each case.
Scenario 1: A project team charts daily defect rates for a filling process. The chart shows all points within limits, no trends, no rule violations — the process is in statistical control. The team also computes Cpk at roughly 0.8 against the specification. A tempting conclusion: the chart looks healthy, so the process is fine and no action is needed. That conclusion confuses the two verdicts. Stability means the process behaves predictably around its current center and spread; it says nothing about whether that predictable output satisfies customers.
The better decision is to read the pair of results together. An in-control process with poor capability is stable at the wrong target or with too much inherent variation, so improvement requires systemic change — recentering the process, reducing common-cause variation, or redesigning the specification — not hunting for assignable causes that the chart says are absent. Capable but unstable is a different problem: performance is acceptable today and unpredictable tomorrow, so the priority is finding and controlling the special causes. Practicing this pairing, and writing one sentence about which verdict each tool produces, is the exercise that makes the distinction automatic.
Choosing the hypothesis test the scenario actually calls for
Test selection follows three questions: how many groups are compared, whether the observations are paired or independent, and whether the data are continuous or categorical. The common mistakes are pairing ignored, independence assumed, and chi-square forced onto small expected counts.
Build test selection as a decision habit. First ask what you are comparing: one sample against a target, two groups, several groups, or counts in categories. Then ask whether the same units appear in both conditions — before and after on the same parts, same operators, same stations — because that pairing changes the test. Finally check the data type: continuous measurements lead to t-tests and ANOVA, while pass-fail outcomes and defect classifications lead to proportion tests and chi-square. Write the three questions at the top of your practice sheet until the routing is reflexive.
Two cautions keep the routing honest. Independence matters as much as pairing: if one production lot feeds both compared conditions, the observations are not independent and any two-sample conclusion is fragile. And a chi-square test with very small expected counts in some cells gives unreliable results, so the better move is to combine sparse categories sensibly or gather more data rather than to report a shaky p-value. The table below contrasts the everyday tests so you can trace any scenario to one row and justify the match.
| Test | Question it answers | Data condition | Common mismatch to avoid |
|---|---|---|---|
| One-sample t-test | Does a process mean differ from a target value? | Continuous data, one group | Using it on attribute counts or on non-independent repeated readings |
| Two-sample t-test | Do two independent groups differ in mean? | Continuous data, two separate groups | Applying it to before/after data collected on the same units |
| Paired t-test | Does the same set of units change between two conditions? | Continuous data, matched pairs | Discarding the pairing and losing the precision it provides |
| One-way ANOVA | Do three or more group means differ? | Continuous data, independent groups | Running multiple two-sample tests instead and inflating error risk |
| Chi-square test | Is the distribution across categories different between groups? | Categorical counts, adequate expected counts | Forcing it onto small counts in sparse cells |
Scenario walkthrough: a before-and-after comparison read badly, then correctly
Scenario 2 involves cycle times recorded on the same five workstations before and after a layout change. The tempting route is a two-sample t-test on the pooled data; the defensible route is a paired test, and the p-value must be read as evidence strength, not proof probability.
In the scenario, baseline cycle times were recorded on five workstations, the layout was changed, and cycle times were recorded again on the same five stations. A plausible mistake is to stack all ten readings and run a two-sample t-test. This treats station-to-station differences as noise to be averaged, when they are actually the dominant variation, and with so few readings the unpaired test has little power to detect the change the pairing would reveal. The conclusion drawn — that the improvement is not statistically significant — is an artifact of the wrong test choice.
The better decision is a paired t-test: compute each station's difference from its own baseline, then test whether the mean difference differs from zero. Pairing removes the station effect from the comparison, which is exactly why matched designs exist. Read the output with discipline: a small p-value indicates the observed difference would be unlikely if no real change existed; it does not state a probability that the improvement is real, and it says nothing about whether the change is large enough to matter operationally. Reporting both the size of the mean difference and its statistical evidence is the complete answer.
Measurement system analysis: proving the ruler before measuring the process
Gage repeatability and reproducibility studies separate variation from the part, the gage, and the operators; attribute agreement analysis does the same for pass-fail inspections. If measurement error consumes most of the observed variation, process conclusions from that data are not supportable.
For continuous measurements, a Gage R&R study has each of several operators measure several parts multiple times, ideally with parts spanning the actual process range. Repeatability captures variation when one operator remeasures the same part; reproducibility captures variation between operators or instruments. The interpretation decision is a ratio judgment: measurement variation should be small relative to the part-to-part variation and to the tolerance. Practicing the setup matters as much as the math — randomize the measurement order so operators cannot remember earlier readings, and never let parts be identified to the measurer.
When the characteristic is a judgment such as pass-fail visual inspection, attribute agreement analysis replaces Gage R&R. The study presents the same items to the same inspectors more than once and compares inspectors against each other and, where possible, against a known standard. The decision it supports is whether the inspection itself is a trustworthy measurement: if inspectors disagree with themselves or with a standard on a meaningful share of items, defect rates built on those inspections describe the inspectors as much as the process. In scenario practice, noticing that the data pass through human judgment is the observation that triggers this analysis.
Control plans and honest documentation: what makes an improvement stick
A control plan specifies the output and input characteristics to monitor, the method and frequency of checks, the reaction plan when a check fails, and the responsible owner. Documentation also carries an ethical duty: report data and results accurately, including inconvenient findings.
Treat the Control phase as a handoff document, not a summary. A usable control plan names each characteristic being monitored, distinguishes outputs from the key inputs identified in Analyze, states how each is measured and how often, and defines the reaction plan — who does what when a signal appears. It also records the validated settings, procedures, and training changes that produced the improvement, so the process survives personnel turnover. When you build or review practice cases, check for the missing element: a monitoring scheme with no reaction plan, or a reaction plan with no owner, is an incomplete control strategy, and identifying those gaps is a skill worth drilling on paper before any real handoff.
The same documentation carries professional obligations. Data integrity is non-negotiable: measurements are recorded as taken, exclusions are documented and justified, and results are reported whether or not they favor the project. Confidentiality of process and customer information applies to charts, reports, and any shared data. If a green belt observes a safety-relevant condition during data collection, the finding goes to the responsible parties promptly rather than waiting for project reporting cycles. Practicing this means checking each written case for what the belt is obligated to disclose and to whom.
A scenario-drill sequence with a self-check rubric
Run a repeating drill: study one concept cluster, then solve written cases under time pressure, then score yourself on a four-point rubric — decision named, test chosen, assumptions checked, conclusion qualified. Adapt the cycle length to your available weeks.
A practical exercise: take any DMAIC case you can construct from your own work or a public example and interrogate it in writing. For each phase, answer four questions — what decision does this phase force, which single tool best serves it, what assumption does that tool require, and what finding would overturn the conclusion. Then run the paired-scenario drill from this guide: take one before-and-after dataset and analyze it both as paired and as unpaired, observing how much the pairing changes the result and the conclusion. Expected observations: the paired analysis shows a smaller spread of differences and stronger evidence; the unpaired analysis is diluted by unit-to-unit variation.
Score each drill with this rubric, treating the totals as learning milestones rather than predictions of any exam result: 1 — you can name the DMAIC decision and a candidate tool; 2 — you can also state the tool's key assumption; 3 — you can also identify a plausible wrong choice and say why it fails; 4 — you can also state what evidence would change your answer and what belongs in the control plan. A repeatable cycle: one concept cluster with its readings, two or three written cases the same day, rubric scoring, and a re-drill of missed distinctions a few days later. Expand the cycle across clusters in this order: control versus capability, test selection, measurement systems, then control planning.
- Rubric level 1: name the phase decision and a candidate tool.
- Rubric level 2: state the tool's key assumption.
- Rubric level 3: identify a tempting wrong choice and explain its failure.
- Rubric level 4: state what evidence would change your conclusion and what the control plan needs.
References and further reading
Use these references to explore the concepts and check the latest information from the relevant organizations.
