RSC Cluster: Scrap, Rework and Cost of Poor Quality Reduction

The Scrap, Rework and Cost of Poor Quality Cluster connects quality losses to financial impact and operational root causes. It reframes scrap and rework as symptoms of upstream process and training failures rather than isolated mistakes. The content walks through the full feedback loop from work instructions to nonconformance to corrective action and prevention. This cluster helps operations and finance leaders align improvement work with measurable cost reduction.

  • Sensitivity analysis

    Sensitivity analysis is a method for evaluating how much an output changes when one or more inputs, assumptions, or parameters are varied. It is commonly used in manufacturing, quality, planning, and analytics to understand which factors have the greatest influence on a result.

    In practice, the term usually refers to testing a model, calculation, forecast, or process metric by changing selected variables and observing the effect. Examples include changing demand assumptions in a planning model, adjusting process parameters in a yield model, or examining how scrap rate, cycle time, or supplier lead time affects cost or schedule performance.

    Sensitivity analysis does not by itself prove cause and effect, validate a model, or determine the single correct operating setting. It is an analytical technique for understanding responsiveness, uncertainty, and relative influence.

    Where it applies

    In industrial and regulated operations, sensitivity analysis may appear in:

    • production planning and capacity scenarios

    • cost, margin, and inventory modeling

    • quality and process-improvement studies

    • risk assessments and contingency planning

    • engineering and process development work, including design of experiments and simulation

    It can be performed with simple spreadsheet models or with more formal simulation, statistical, or optimization tools.

    Common confusion

    Sensitivity analysis is often confused with scenario analysis and design of experiments.

    • Scenario analysis usually compares a defined set of combined conditions, such as best case, expected case, and worst case.

    • Sensitivity analysis focuses on how the output responds when specific inputs are changed, often one at a time or across a defined range.

    • Design of experiments (DoE) is a structured experimental method used to study factor effects and interactions in real or simulated processes.

    It is also not the same as measurement sensitivity, which refers to how responsive an instrument or detection method is.

    Manufacturing example

    A planner may test how a 5 percent change in forecast demand, supplier lead time, or machine uptime affects required inventory and promised ship dates. A quality engineer may assess how variation in temperature, dwell time, or torque changes predicted defect rates. In both cases, the goal is to identify which inputs matter most and where tighter control or better data may be needed.

  • Design of experiments (DoE)

    Design of experiments (DoE) is a structured statistical method for planning, running, and analyzing tests so teams can determine how one or more input factors affect a measured output. In manufacturing and quality contexts, it is commonly used to study process settings, material variables, equipment parameters, and environmental conditions in a controlled way.

    DoE is more than trial-and-error testing. It is designed to separate the effect of individual factors and, in many cases, the interaction between factors. This helps explain why a process outcome changes, not just whether it changed.

    What it includes

    A DoE typically defines:

    • the response or output being measured, such as yield, strength, cycle time, or defect rate
    • the factors being varied, such as temperature, pressure, speed, dwell time, or material lot
    • the levels or settings for each factor
    • the test structure, such as randomized runs, replicated runs, or factorial designs
    • the analysis used to identify statistically meaningful effects

    Depending on the objective, DoE may be used for screening important variables, optimizing process settings, characterizing a process window, or supporting root cause analysis.

    What it is not

    DoE is not the same as changing one variable at a time without a formal design. It is also not limited to product design work. In industrial operations, it is often applied to manufacturing processes, inspection methods, formulation work, and validation-related studies.

    DoE does not by itself prove long-term process control or regulatory acceptability. It is a method for generating evidence about relationships between inputs and outputs under the conditions studied.

    How it appears in operations

    In plant and quality workflows, DoE may be used during process development, scale-up, transfer to production, deviation investigation, or continuous improvement. Results are often documented in engineering, quality, or validation records and may inform work instructions, control limits, recipes, or parameter ranges in MES, historians, or related systems.

    Example: a manufacturer may run a DoE to study how cure temperature, hold time, and humidity affect bond strength and scrap rate.

    Common confusion

    DoE vs. A/B testing: A/B testing usually compares one change against another. DoE can evaluate multiple factors at the same time and can reveal interactions between them.

    DoE vs. one-factor-at-a-time testing: One-factor-at-a-time testing is simpler but may miss combined effects. DoE is specifically structured to estimate those effects more reliably.

    DoE vs. statistical process control: Statistical process control monitors an ongoing process for stability and variation. DoE is used to learn how deliberate changes in inputs influence outputs.

  • What are realistic first steps to pilot an enterprise scrap visibility solution?

    The realistic first step is not an enterprise rollout. It is a tightly scoped pilot that proves whether your organization can capture scrap data consistently enough to trust it, connect that data to existing systems without disrupting production, and use the results to drive action.

    In most regulated manufacturing environments, scrap visibility fails first as a data and process problem, not a dashboard problem. Plants often already record scrap somewhere, but the definitions, timing, units of measure, reason codes, operator workflows, and links to orders, lots, parts, and nonconformance records are inconsistent. A pilot should expose those gaps early.

    Recommended pilot scope

    • Pick one site or area with meaningful scrap cost, engaged local leadership, and manageable system complexity.
    • Limit the process scope to one value stream, product family, workcenter group, or operation sequence.
    • Define a small scrap taxonomy with a controlled list of reason codes and clear ownership for changes.
    • Connect scrap to business context such as part number, work order, operation, shift, machine or cell, operator role, and disposition path where available.
    • Decide what counts as scrap for the pilot and what does not. Include rework, yield loss, and MRB-related loss only if you can identify them reliably.

    What to implement first

    1. Baseline the current state. Document where scrap is recorded today across MES, ERP, spreadsheets, quality systems, machine logs, and manual whiteboards. Expect conflicts.
    2. Standardize the minimum data set. Define the fields required to answer basic questions: what was lost, where, when, against which order or lot, why, and who confirmed it.
    3. Map source systems and handoffs. In brownfield environments, scrap events may start in MES, inventory adjustments may post in ERP, and root-cause or disposition details may live in QMS or NCR workflows. Your pilot should make those boundaries explicit rather than hiding them.
    4. Establish a governed reason-code model. Keep it small at first. If every area uses different codes, enterprise comparison will be misleading.
    5. Build one traceable workflow. For example: scrap event captured at operation completion, reviewed by supervisor or quality, then reconciled to inventory and linked to any NCR if required by local process.
    6. Create a simple review cadence. Daily operational review and weekly cross-functional review are usually more valuable than advanced analytics in the pilot phase.
    7. Measure adoption and data quality. Track missing reason codes, late entries, overrides, duplicate events, reconciliation gaps, and exceptions between systems.

    Integration strategy

    Use the lightest integration approach that still preserves traceability. In many plants, that means starting with batch or event-level interfaces to existing MES and ERP rather than trying to replace them. Full replacement is often unrealistic in regulated, long-lifecycle operations because of qualification burden, validation cost, downtime risk, integration complexity, and the need to maintain historical traceability across legacy processes.

    If your current systems are heavily customized, the pilot may need a coexistence model:

    • MES remains the system of execution for order and operation context.
    • ERP remains the financial and inventory system of record.
    • QMS or NCR workflow remains the governed quality record where required.
    • The pilot layer consolidates scrap events, reason codes, and visibility metrics.

    That approach is less elegant than a greenfield design, but it is usually more achievable and lower risk.

    What success should look like

    A good pilot does not prove that every scrap dollar across the enterprise is now visible. It proves narrower things:

    • Operators and supervisors can record scrap with acceptable burden.
    • Reason codes are used consistently enough to compare shifts, products, or cells.
    • Scrap events can be reconciled to inventory and production context.
    • Quality and operations can investigate the same event without competing data sets.
    • Leadership can identify a few repeatable loss patterns worth corrective action.

    If the pilot cannot achieve those basics, scaling it enterprise-wide will amplify confusion.

    Common failure modes

    • Trying to start enterprise-wide. This usually creates definition disputes before value appears.
    • Overloading the first workflow. Too many mandatory fields will drive bypass behavior or poor data quality.
    • Ignoring system-of-record boundaries. If ERP, MES, and QMS disagree, the pilot must show how conflicts are resolved.
    • Confusing scrap visibility with root cause elimination. Visibility helps prioritize losses; it does not replace process engineering, CAPA, or RCCA.
    • Skipping governance. Uncontrolled edits to reason codes, mappings, or formulas can invalidate trend comparisons.
    • Underestimating validation and change control. In regulated settings, even reporting changes may require review depending on how the output is used operationally or for quality evidence.

    Practical 60 to 90 day pilot plan

    • Weeks 1 to 2: select pilot area, define objectives, map current-state data sources, agree minimum data model.
    • Weeks 3 to 4: rationalize reason codes, define workflow ownership, identify integration points and reconciliation rules.
    • Weeks 5 to 8: configure capture and reporting, test with historical and live transactions, train a small user group.
    • Weeks 9 to 12: run in production, review exceptions daily, measure adoption, refine controls, and document scale-up requirements.

    Before expanding, ask whether the pilot produced trusted data with manageable effort. If not, the next step is usually process and data cleanup, not broader software deployment.

  • What metrics best capture the cost of poor quality in MRO environments?

    No single metric is sufficient. In MRO, cost of poor quality is best captured as a small set of linked measures that show both direct cost and operational impact.

    The most useful core metrics are:

    • Rework cost per work order or event: labor hours, replacement parts, consumables, tooling time, and repeat inspection caused by a defect, escape, or incorrect execution.
    • Scrap and beyond-economical-repair value: material or component value written off because the unit cannot be repaired within approved limits or required economics.
    • Nonconformance rate and recurrence rate: count and percentage of jobs, parts, or routings generating NCRs, plus repeat occurrence by defect code, asset family, station, supplier, or technician group.
    • Turnaround time impact from quality events: added elapsed time from hold points, re-inspection, teardown/reassembly, engineering review, quarantine, or waiting for disposition.
    • Inspection yield and first-pass yield: the share of work that clears required checks without rework, re-opened tasks, or document correction. In MRO, this often matters more than a generic plant-level scrap rate.
    • MRB or disposition cycle time: elapsed time from defect identification to approved disposition and release back into the work stream. This exposes hidden queue cost, not just visible repair cost.
    • Deferred revenue or slot utilization loss: the capacity and commercial impact when quality issues consume bay time, test stand time, rotable availability, or planned induction slots.
    • Supplier quality cost: incoming defect cost, outside processing returns, expedited freight, additional inspection, schedule disruption, and administrative recovery effort tied to suppliers.
    • Warranty, return-to-service failure, or repeat visit rate: downstream quality cost after release. This is especially important because some MRO quality losses are shifted into field events rather than captured at the original work center.
    • Documentation correction burden: labor spent fixing records, traceability gaps, task signoff errors, parts history issues, or missing approvals. In regulated environments, this is often material and routinely undercounted.

    If you need a practical starting point, track COPQ in four buckets:

    1. Internal failure: rework, scrap, re-inspection, troubleshooting, and hold time.
    2. External failure: warranty, repeat removals, customer claims, service credits, and field support.
    3. Appraisal overload: extra inspection and verification driven by unstable processes or poor incoming quality.
    4. Administrative recovery: NCR processing, investigations, record correction, engineering review, disposition routing, and change-controlled documentation updates.

    What usually gets missed in MRO

    MRO environments have several COPQ blind spots. The largest are schedule disruption, asset availability loss, and traceability recovery work. A part that requires another inspection loop may create only a modest direct labor charge, but it can delay release, displace other jobs, consume scarce certifying resources, and create knock-on shortages. If you only track scrap and rework, you will understate the real cost.

    Another common miss is mixed causality. A single quality event may involve planning error, outdated work instructions, supplier issue, incomplete maintenance lineage, incorrect material substitution, and technician execution. If coding is weak, the cost gets booked as generic labor variance instead of poor quality.

    How to make the metrics usable

    The best metrics are the ones you can trace back to a work order, serialized asset, task, part, station, supplier, and disposition path. Without that lineage, COPQ becomes a finance estimate instead of an operational control tool.

    In practice, reliable MRO COPQ reporting usually depends on:

    • consistent defect and cause codes in the QMS or NCR workflow
    • accurate labor charging in ERP or MRO systems
    • clear links between work orders, material issues, inspections, and nonconformance records
    • disposition timestamps for queue and approval delays
    • repeatable rules for valuing downtime, slot loss, and outside processing impact

    If those controls are weak, start with trendable measures before trying to monetize everything. A stable recurrence rate, first-pass yield by task family, and MRB cycle time by disposition category are often more actionable than a precise but fragile total-dollar estimate.

    Brownfield reality

    Most organizations will not get these metrics from one system. MRO, ERP, QMS, EAM, supplier portals, and inspection records often hold different parts of the picture, with inconsistent IDs and coding. That means COPQ accuracy depends on integration quality, master data discipline, and change control.

    Full replacement is usually not the right first move in regulated, long-lifecycle environments. Replacing core systems to get cleaner COPQ reporting often fails because qualification burden, validation effort, downtime risk, historical traceability needs, and integration complexity are high. In most brownfield settings, a phased approach works better: standardize codes, improve cross-system linkage, validate calculations, then automate reporting incrementally.

    The tradeoff is speed versus confidence. You can estimate COPQ quickly with finance-side allocations, but operational credibility may be weak. Or you can build traceable, work-order-level metrics, which are more defensible but slower and harder to implement.

    A practical scorecard

    For most MRO operations, a balanced COPQ scorecard includes:

    • rework labor hours as a percentage of direct labor
    • scrap and BER value as a percentage of material value processed
    • NCR rate per 100 work orders or per 1,000 task cards
    • repeat defect rate within 30, 60, or 90 days
    • first-pass yield at final inspection and key in-process gates
    • MRB or disposition cycle time
    • quality-driven turnaround time delay
    • supplier-driven quality cost
    • post-release failure or repeat visit rate
    • documentation correction hours

    That combination usually captures the real cost better than any single KPI. It also makes failure modes visible enough to support corrective action, provided the underlying data is governed and traceable.

  • How can we quantify the impact of a design change on scrap and yield?

    You can quantify it, but only if you can separate the design change from everything else that changed around it.

    The practical method is to compare scrap and yield at the part, operation, and revision level before and after the design change, then adjust for confounders such as supplier lot, material condition, tooling state, routing changes, inspection changes, machine changes, operator mix, and production volume. If you do not control for those factors, the number may be directionally useful but not defensible.

    What to measure

    At minimum, measure the same product family across a defined baseline period and post-change period using consistent definitions.

    • First-pass yield by operation and by finished assembly

    • Scrap rate by count, unit, weight, or cost, depending on how your plant records loss

    • Rework and repair incidence, because some design changes reduce formal scrap but increase hidden recovery effort

    • Nonconformance rate and defect codes tied to the changed characteristics

    • Cost of poor quality, if finance and quality data can be linked reliably

    • Cycle time impact, because yield improvements can come with longer processing or inspection time

    Recommended approach

    1. Define the exact design change window using the approved revision, effectivity, and disposition dates.

    2. Identify the affected part numbers, configurations, routings, work centers, and suppliers.

    3. Build a pre-change baseline and a post-change observation window with enough volume to be meaningful.

    4. Segment results by operation, defect mode, machine family, supplier, and lot where possible.

    5. Exclude or flag units built during transition conditions such as mixed inventory, temporary deviations, pilot builds, training periods, or parallel routings.

    6. Compare the changed design against either its own historical baseline or a matched control population if other plant conditions were moving at the same time.

    7. Validate whether the observed shift is statistically credible, not just operationally interesting.

    How to calculate it

    A simple starting point is:

    • Scrap impact = post-change scrap rate minus pre-change scrap rate

    • Yield impact = post-change first-pass yield minus pre-change first-pass yield

    • Financial impact = change in scrap quantity or rework quantity multiplied by material, labor, outside processing, and delay cost where available

    That is only the first layer. In most brownfield plants, a better estimate comes from stratified analysis or regression that controls for major drivers such as supplier, lot, machine, shift, and inspection method. If the design change altered tolerances, materials, joining method, or inspection characteristics, those factors need explicit treatment.

    What usually makes the analysis fail

    • No reliable link between engineering revision and actual as-built unit genealogy

    • Scrap recorded only at the work order level, not by operation or defect mode

    • Mixed old and new revision inventory consumed in the same period

    • Simultaneous process changes, supplier changes, or tooling changes

    • Inspection sensitivity changed, making defects appear to rise when detection simply improved

    • Too little production volume after the change to distinguish signal from noise

    • Rework moved off the books into concession, deviation, or informal recovery activity

    In those cases, the honest answer is that you can estimate impact, but not attribute it cleanly.

    Brownfield system reality

    In many regulated environments, the needed data lives across PLM, ERP, MES, QMS, SPC, and sometimes spreadsheets. Quantification depends on whether those systems share consistent part, revision, lot, operation, and defect identifiers. If they do not, the work becomes a data reconciliation exercise before it becomes an engineering analysis.

    This is why full replacement is often the wrong first move. Replacing PLM, MES, ERP, or QMS just to answer this question usually creates more qualification, validation, integration, and downtime risk than value in the near term. A narrower approach is usually safer: improve revision effectivity tracking, strengthen genealogy, standardize defect coding, and create a governed cross-system view of scrap, yield, and design revision.

    What a credible answer looks like

    A credible result usually includes:

    • The exact design revision or effectivity studied

    • The units, lots, and time period included and excluded

    • The baseline and post-change sample sizes

    • The scrap and yield deltas by operation and overall

    • The main controlled variables and known uncontrolled variables

    • Any transition effects, temporary work instructions, or training effects

    • Confidence level or at least a clear statement of statistical and practical uncertainty

    That level of traceability matters. Without it, the analysis may still help internal decision-making, but it will not stand up well to skeptical review.

    So the short answer is: quantify the impact by linking approved design revision changes to as-built production records, then compare revision-level scrap and yield with controls for process and supply variation. If your data model, change control, or genealogy is weak, say so explicitly and treat the result as an estimate rather than a clean attribution.

  • Multivariate analysis

    Multivariate analysis commonly refers to a group of statistical methods used to examine more than one variable at the same time. In manufacturing and regulated operations, it is used to understand how process inputs, material attributes, equipment conditions, and quality results vary together rather than looking at each factor in isolation.

    The term includes techniques that model relationships, patterns, correlation structures, and sources of variation across many measurements. Depending on the method, the goal may be to explain variation, detect abnormal conditions, classify outcomes, reduce data dimensionality, or predict a result. It does not simply mean making several separate single-variable analyses in parallel.

    How it is used in operations

    In industrial settings, multivariate analysis often appears in process monitoring, quality investigations, yield analysis, and continuous improvement work. Examples include evaluating how temperature, pressure, mix time, and raw material properties relate to assay, defect rate, or cycle time, or using many sensor signals together to identify a developing process drift.

    It may be applied within MES, historian, LIMS, QMS, or analytics platforms, or performed offline using statistical software. The output can support investigation and decision-making, but the term itself refers to the analytical approach, not to a specific software product, dashboard, or compliance record.

    What it includes and excludes

    • Includes analysis of multiple variables simultaneously.

    • Includes methods such as principal component analysis, cluster analysis, discriminant analysis, factor analysis, and multivariate regression, depending on context.

    • May use historical batch, process, laboratory, or equipment data.

    • Does not automatically imply machine learning, though some machine learning methods are multivariate.

    • Does not mean any dataset with many columns unless a method is actually evaluating relationships among those variables.

    Common confusion

    Multivariate analysis is often confused with multiple regression and multivariable analysis. Multiple regression is one specific statistical technique. Multivariable analysis often means one outcome modeled using several predictors, while multivariate analysis more strictly means analyzing multiple variables or outcomes together. In practice, different disciplines sometimes use these terms less precisely, so the intended method should be checked.