Select Page

A Lean Six Sigma Black Belt project at the U.S. Army Command and General Staff College’s (CGSC) Department of Distance Education used DMAIC and a full factorial DesiDesign of Experiments (DOE) to fix a chronic grading problem in the H400 military history essay course.

The team, led by project champion LTC Scott Grimsey, found that two factors, structured peer review and tutor lab support, were the statistically significant drivers of essay quality, while a factor the team initially suspected (student age) was not.

After standardizing peer review, tutor labs, and a detailed grading rubric, faculty man-hours lost to remedial regrading dropped from roughly 60 hours in the project sample to zero, moving the program toward its Academic Year 2027 target median grade of 87%.

Quick-Reference Table: Case Study at a Glance

Bar chart showing before-and-after results from a Lean Six Sigma Black Belt project
Bar chart showing before-and-after results from a Lean Six Sigma Black Belt project
MetricBaselineAfter ImprovementResult
Median H400 essay grade (6-month baseline)85.50%Target: 87.00% by AY2027Gap narrowed toward target
Faculty man-hours spent monthly on regrading (Cost of Poor Quality)15–20 hrs/monthGoal: reduce by at least 75%11–15 faculty man-hours/month targeted for recovery
Man-hours lost to grading failures (project sample)60.25 hours0 hours100% eliminated
Measurement system variation (Gauge R&R audit, 45 essays)23% of total process variationData judged reliable for decision-makingConfirmed instructors were calibrated to the same standard
Full factorial DOE model fit (R²)—87.9%Most of the variation in grades explained by course design factors
Peer Review effect (t-value; significance threshold 2.03)—6.82Statistically significant driver
Tutor Labs effect (t-value)—5.69Statistically significant driver
FMEA risk score, reading non-completionRPN 300RPN 6080% risk reduction
FMEA risk score, Blackboard/reading-access issueRPN 180RPN 6067% risk reduction

Key Takeaways

  • This project used quantitative Design of Experiments, not just qualitative process mapping, to identify which variables actually predicted essay grades among six candidates: student age, education level, military component, rubric type, peer review, and tutor labs.
  • An initial, more limited analysis suggested student age might matter (R² = 0.78 in a simple Pareto/regression view), but the full factorial model showed age was not statistically significant (t = 0.23) once course design factors were properly isolated.
  • Peer Review and Tutor Labs were confirmed as the two statistically significant drivers of grade outcomes, with t-values of 6.82 and 5.69 against a significance threshold of 2.03.
  • A Detailed Rubric didn’t move grades much as a standalone factor, but interaction analysis found it acted as a safety net, recovering roughly 2.5 points of the grade drop specifically in cases where peer review was unavailable.
  • Two failure modes identified through FMEA (students not completing assigned readings, and Blackboard access issues to reading materials) were addressed with targeted fixes, cutting their risk priority numbers by 80% and 67% respectively.
  • The control plan was built into the existing gradebook system, with an automatic flag if the average grade falls below 89%, so the fix is monitored continuously rather than checked manually.

Background: A Familiar Problem in Professional Military Education

The Department of Distance Education (DDE) at the Command and General Staff College (CGSC) runs the H400 military history course as part of its curriculum. Students complete a series of readings and submit a graded essay as the capstone assignment.

Over a recent six-month period, the median grade on H400 essays was 85.50%, short of the college’s Academic Year 2027 target of 87%. That performance gap wasn’t just an academic concern: it was estimated to cost 15 to 20 faculty man-hours per month in regrading essays that failed to meet the minimum 80% passing threshold, a classic Lean Six Sigma Cost of Poor Quality (CoPQ).

An earlier, smaller baseline sample used for the project’s measurement system analysis showed a median grade of 82.09% across 45 essays. Both figures describe the same underlying process at different points and sample sizes; the 85.50% figure is the official six-month program baseline used in the project charter, while the 82.09% figure comes from the 45-essay sample used to validate the grading measurement system before drawing further conclusions.

Kevin Clay

Public, Onsite, Virtual, and Online Six Sigma Certification Training!

  • We are accredited by the IASSC.
  • Live Public Training at 52 Sites.
  • Live Virtual Training.
  • Onsite Training (at your organization).
  • Interactive Online (self-paced) training,

Define: Setting Clear Scope and a Measurable Target

The project charter defined scope tightly: AOC Teams 3, 4, and 5 within the Department of Distance Education, starting at the H400 readings and ending at paper upload. All other departments were explicitly out of scope, a common and important discipline in Black Belt projects, since scope creep is one of the most common reasons DMAIC projects stall.

The stated goal: increase the median H400 essay grade from 85.50% to 87.00% by the end of Academic Year 2027, while reducing the Cost of Poor Quality by at least 75%, recovering an estimated 11 to 15 faculty man-hours per month currently lost to remedial regrading.

Measure: Mapping the Process and Testing Whether the Data Could Be Trusted

Current-State Value Stream Mapping

The team mapped the full H400 essay workflow across three swim lanes: Educate (orientation and readings, ranging from 2 to 8 hours per module), Draft/Outline Review (essay outline, instructor review, and drafting), and Paper Review (final instructor review). Documenting this before touching anything is standard Lean practice: you can’t remove waste from a process you haven’t actually mapped.

SIPOC, Process Mapping, and Cause-and-Effect Analysis

Working from Subject Matter Expert (SME) input, the team built a SIPOC (Suppliers, Inputs, Process, Outputs, Customers) chart, a detailed process map, and a cause-and-effect matrix. This pointed to three Key Process Input Variables (KPIVs):

  1. Instructors — specifically, grading accuracy and consistency
  2. Students — specifically, failure to complete assigned readings
  3. The Blackboard learning management system — specifically, availability of the book of readings

Measurement Systems Analysis: Can You Trust the Grading Data?

Before analyzing why grades varied, the team had to confirm the grades themselves were measured consistently. This is a step many process improvement efforts skip, and it’s often the reason later “fixes” don’t hold up.

The team ran a Gauge Repeatability and Reproducibility (Gauge R&R) study, auditing 45 essays graded across multiple instructors. The result: instructor grading boxes overlapped significantly, indicating consistent calibration to the same standard, with measurement system variation accounting for 23% of total process variation. That was within an acceptable range, meaning the team could trust the dataset for the analysis that followed.

Also Read: How Six Sigma Principles Align with BAE Systems’ Quality Approach

Analyze: Why the “Obvious” Suspects Weren’t the Real Drivers

This is where the project moves beyond typical process-mapping exercises into genuine statistical inference, and where it produces its most useful lesson for anyone applying Six Sigma outside manufacturing.

First Pass: Age Looked Like It Might Matter

An initial Pareto and regression analysis on essay grades produced an R² of 0.78 tied to student age, suggesting age was a meaningful predictor. But the relationship wasn’t a simple “older students do better” or “younger students do better” pattern; it was cohort-based, with specific age groups performing differently from others in ways a straight-line trend couldn’t capture.

Second Pass: Categorical Cohort Modeling

Shifting from a linear numerical view of age to a categorical “bucket” approach (grouping by age bracket and military component) raised the model’s explanatory power to R² = 0.89, suggesting nearly 90% of grade variance could be explained by combined demographic-style factors. At this stage, it looked like student background mattered a great deal.

Third Pass: Full Factorial Design of Experiments

Chart showing which factors were statistically significant drivers of essay grades:
Chart showing which factors were statistically significant drivers of essay grades:

The team then ran a full factorial regression incorporating course design factors alongside demographic factors: Peer Review (yes/no), Tutor Labs (self-guided vs. tutor-led), Rubric Type (detailed vs. standard), Education Level, Military Component, and Student Age. This model fit at R² = 0.833, and a follow-up full factorial analysis pushed explanatory power to R² = 87.9%.

Once course design factors were included and properly isolated, the picture changed substantially:

Factort-valueStatistically Significant? (threshold = 2.03)
Peer Review: No vs. Yes6.82Yes
Labs: Self-Guided vs. Tutor5.69Yes
Education: Masters vs. Bachelors3.02–2.31 (varied by model)Yes
Military: Reserves vs. Active2.84–2.27 (varied by model)Yes
Rubric: Detailed vs. Standard1.96–1.03 (varied by model)Borderline / not consistently significant
Military: Guard vs. Active1.62–1.35 (varied by model)Not significant
Student Age0.99–0.23 (varied by model)Not significant

The headline finding: student age, which looked like a strong predictor in the first two passes of analysis, was not statistically significant once Peer Review and Tutor Labs were properly modeled. This is a textbook example of why Six Sigma insists on structured experimentation rather than stopping at the first correlation that appears meaningful.

The Interaction Effect: What the Detailed Rubric Actually Does

The full factorial analysis surfaced one more subtle finding. On its own, a Detailed Rubric didn’t move grades much when peer review was already in place. But in an interaction analysis, the Detailed Rubric acted as a safety net specifically when peer review was missing, recovering roughly 2.5 points of the grade drop that would otherwise occur. In other words, the rubric’s value only becomes visible when you test how it interacts with other variables, not when you test it in isolation.

FMEA: Where the Process Was Most Likely to Fail

Failure Mode and Effects Analysis (FMEA) was used to examine how the KPIVs identified earlier could fail in practice. Two failure modes stood out:

  • Students failing to complete assigned readings (Severity 6, Occurrence 5, Detection 10 → Risk Priority Number of 300)
  • Blackboard connectivity issues preventing access to the book of readings (Severity 6, Occurrence 3, Detection 10 → Risk Priority Number of 180)

Improve: The Fixes That Actually Moved the Needle

Based on the DOE and FMEA findings, the team implemented five targeted changes:

  1. Structured Peer Review — identifies issues in a draft before it reaches final grading, addressing the single largest statistically confirmed driver of grade outcomes.
  2. Tutor Lab integration — provides additional writing support resources, the second-largest confirmed driver.
  3. A standardized Detailed Rubric — gives students a consistent grading metric to write toward, and functions as a safety net when peer review isn’t available.
  4. A shared drive of course readings — eliminates the Blackboard-availability failure mode by removing a single point of technical failure.
  5. Reading verification quizzes — directly targets the “failure to complete readings” failure mode by confirming students engaged with the material before submission.

After implementation, man-hours lost to grading failures dropped from 60.25 hours to 0 hours in the post-improvement sample, directly attacking the Cost of Poor Quality identified in the Define phase.

Applying the FMEA-derived corrective actions also reduced both target risk scores substantially: the reading-completion failure mode dropped from a Risk Priority Number of 300 to 60 (an 80% reduction), and the Blackboard/connectivity failure mode dropped from 180 to 60 (a 67% reduction).

Control: Making the Fix Stick

A DMAIC project that stops at Improve tends to drift back to baseline within a year. This project built the fix into two control mechanisms:

  • A live gradebook dashboard tracking Peer Review pass/fail status alongside paper grades, coded to flag organizational leadership automatically if the rolling average grade falls below 89%, triggering a review of procedural accuracy.
  • A recurring instructional cycle: H400 Paper Submission → H400 Grading (using the Detailed Rubric) → H400 After Action Review → Initial Faculty Development (covering rubric training, the peer review process, and course-bank review) → back to Paper Submission.

This closes the loop: the same faculty development cycle that trained instructors on the fix also feeds lessons from each grading cycle back into the next one.

Also Read: Pursue an Exciting Career in Lean Six Sigma

What This Case Study Teaches About Applying Six Sigma Outside Manufacturing

Three lessons generalize well beyond a single essay-grading process:

  • Correlation from a simple analysis can mislead you. Student age looked like a real driver of outcomes in two separate early analyses (R² = 0.78 and R² = 0.89) before full factorial DOE showed it wasn’t statistically significant once course design factors were properly isolated. Teams that stop at the first plausible-looking correlation risk fixing the wrong problem entirely.
  • Measurement Systems Analysis has to come before root-cause analysis. Before analyzing why grades varied, the team confirmed instructors were grading consistently (23% measurement variation, within acceptable range). Skipping this step is one of the most common reasons process improvement conclusions don’t hold up under scrutiny.
  • Interaction effects matter as much as main effects. The Detailed Rubric looked unremarkable as a standalone factor, but its real value only appeared in interaction analysis, as a safety net specifically when peer review was unavailable. A simpler analysis that only tested factors independently would have missed this entirely.

FAQs on Fixing Grading in Military Education

Q: What organization was this Lean Six Sigma Black Belt project based on?

A: The Department of Distance Education (DDE) at the U.S. Army Command and General Staff College (CGSC), specifically the H400 military history essay course, with LTC Scott Grimsey serving as project champion.

Q: What statistical methods did this project use beyond standard DMAIC tools?

A: The project used SIPOC, value stream mapping, cause-and-effect analysis, Measurement Systems Analysis (Gauge R&R), FMEA, Pareto and ANOVA analysis, multiple regression, and a full factorial Design of Experiments (DOE).

Q: What was the main finding of the Design of Experiments?

A: Peer Review and Tutor Labs were confirmed as statistically significant drivers of essay grade outcomes (t-values of 6.82 and 5.69 against a significance threshold of 2.03), while student age was not statistically significant once these factors were properly modeled, despite appearing to matter in earlier, simpler analyses.

Q: What results did the project produce?

A: Man-hours lost to grading failures dropped from 60.25 hours to 0 hours in the project sample, and two FMEA-identified failure modes had their risk priority numbers reduced by 80% and 67% respectively, moving the program toward its Academic Year 2027 target median grade of 87%.

Q: Can Lean Six Sigma and Design of Experiments be applied to education and training programs, not just manufacturing?

A: Yes. This project applied full factorial DOE, normally associated with manufacturing and product design, to a subjective, human-graded academic process, and used it to distinguish real drivers of quality from variables that only appeared significant in simpler analysis.

Final Words

This project is a useful reminder that the tools associated with manufacturing quality control, Design of Experiments, Gauge R&R, FMEA, apply just as directly to subjective, human-judgment-driven processes like essay grading.

The most valuable finding wasn’t the fix itself; it was what the full factorial analysis ruled out. Student age looked like a real driver of grade outcomes in two separate early analyses before disciplined experimentation showed the actual drivers were Peer Review and Tutor Labs.

That distinction, between what correlates and what actually causes an outcome, is the difference between a process fix that holds up and one that doesn’t.

Bring This Kind of Rigor to Your Organization’s Black Belt Projects

SSDSI is an IASSC accredited Lean Six Sigma training organization, rated 5 stars on Google Reviews, with 5,322+ professionals certified across 600+ organizations in 52 cities. Black Belt projects like this one, using full factorial DOE to separate real drivers from plausible-looking correlations, are exactly the kind of rigorous, defensible project work SSDSI’s Black Belt curriculum is built to produce.

If your organization has a process where the “obvious” explanation might not be the right one, whether in education, operations, or research, a structured Black Belt project can tell you which factors actually matter before you spend resources fixing the wrong ones.

About Six Sigma Development Solutions, Inc.

Six Sigma Development Solutions, Inc. offers onsite, public, and virtual Lean Six Sigma certification training. We are an Accredited Training Organization by the IASSC (International Association of Six Sigma Certification). We offer Lean Six Sigma Green Belt, Black Belt, and Yellow Belt, as well as LEAN certifications.

Book a Call and Let us know how we can help meet your training needs.

Secret Link