You probably have one of these. I’d be surprised if you only had one.
A defect signature everyone in the fab recognizes on sight. It flares up after a PM cycle, when a new chemistry batch is used, or when Tool X “acts weird” on night shift. You contain it, scrub the lot history, tweak a recipe, and hold your breath… and the KPI line calms down.
For a while.
Then it’s back. Same shape, different lot. Same defect, different chamber. Same “root cause” label… somehow a different corrective action plan.
The real issue is the one nobody wants to admit:
If a problem hasn’t been solved within a reasonable time, it means it will never be solved.
In any fab, “chronic” usually doesn’t mean “hard.” It means you’ve built a repeatable way to contain the pain, but you don’t have a repeatable way to eliminate the mechanism.
Excursions punish that gap. They quietly chew through yield—especially when they go undiscovered long enough to process more material under the wrong conditions.
This isn’t a motivational speech about trying harder. It’s a practical point about how you define “solved.” You’ll see a simple model for moving from recurring symptoms and comforting labels to a testable mechanism with boundaries, evidence, and prevention controls that block recurrence by default.
Because if it keeps coming back, it wasn’t solved. And if you’ve been busy with this problem for a while, it will never be solved — ever — unless you change.

To keep this practical, define “chronic” using the same operational language you already use for yield and excursions.
An excursion is when a process or a piece of equipment moves outside accepted specifications. These events can be a significant contributor to yield loss, particularly when the drift or shift goes undiscovered.
Now, the two separate patterns often get mixed together.
A sporadic problem is a deviation from the current baseline. Your control loop is built for this: detect the abnormal change, diagnose what changed, and restore performance back to the prior level.
A chronic problem is different. Chronic loss is the “built-in” waste level that persists over time until you change the underlying system. Control keeps it from getting worse; improvement is what drives it down.
In fab terms, a yield or defect issue becomes chronic when the pattern is recognizable and recurring (across lots and weeks, sometimes across tools or chambers), yet the work repeatedly ends at stabilization. You regain the baseline, operations resume, and the same thing returns when conditions line up again.
This definition is intentionally no-blame. Chronic problems persist even with strong teams because incentives and information layout tend to push work toward fast containment: data is distributed across functions, evidence arrives at different times, and closure pressure favors a credible explanation and an implemented action over a mechanism that’s proven across boundary conditions.
Operationally, you can use a simple test:
If you can repeatedly contain the impact but you cannot yet reliably prevent recurrence across normal variability and change, you’re dealing with a chronic problem.
When someone says, “no one knows how to solve it,” the practical meaning is usually this:
You can’t reliably reproduce the failure on purpose.
You can observe it, you can contain it, and you can build a plausible narrative around it. What you don’t have yet is a controlled way to make the defect signature appear (and disappear) by changing a specific condition.
That matters because a “root cause” only earns the title after verification. In disciplined problem-solving methods, causes are proven, not selected through confident storytelling. Practically, proof looks like this: you can make the failure come and go by toggling the suspected factor within defined boundaries, and the signature follows.
In fab terms, this is the difference between:
“We think it’s contamination/handling/tool noise,” and
“When Factor X crosses Boundary Y, the signature appears; when we remove it, the signature disappears; and the same relationship holds when we repeat the test.”
Chronic problems get labeled “mysterious” because building that on-demand reproduction is genuinely hard. Evidence often arrives late, after multiple steps and variables have already changed. Under those constraints, teams drift toward correlation and containment because those are fast, while proving a reproducible mechanism takes deliberate experimentation and patience.
So “no one knows how to solve it” often translates to: the mechanism isn’t stable enough yet to test. Until you can reproduce the failure under defined conditions, any “solution” remains vulnerable to the next normal change in tools, materials, recipes, or operating windows.
The same “recurring problem” pattern exists in many industries. In semiconductors, it carries a different price tag because time, evidence, and value are delayed and compounded.
Even in normal conditions, the time constant is long. Wafer fabrication cycle time is measured in weeks, and advanced processes run longer. End-to-end lead time to a finished chip can stretch to months.
That baseline matters because the typical cost pattern for chronic problems is not a single dramatic failure that stays invisible until the very end (though that can happen). The more common pattern is slower and more expensive:
Because the mechanism is not eliminated, the organization keeps paying for repeated containment loops.
Containment stretches cycle time. Lots are held. Extra metrology becomes routine. Split experiments run. Temporary recipe offsets get added. Tool matching and re-qualification work expands. Teams add “just to be safe” checks that quietly become semi-permanent. Each loop adds days or weeks, and the baseline process keeps running underneath it. Meanwhile, WIP continues moving through expensive steps, so the value at risk grows with every additional operation.
This is why chronic problems in semiconductors often feel “unsinkable.” The learning loop is slow, and the cost of each additional loop rises as material moves downstream.
In a fab, “solved” has a higher bar than “we stopped the bleeding.”
A problem is solved when you can explain it as a testable mechanism, prove that mechanism under controlled conditions, and then put controls in place so the mechanism stays blocked during normal operation and normal change.
Start with the mechanism. You need a causal statement with boundaries and predictions, not a label. Proof looks like reproducibility: you can make the failure come and go by toggling the suspected factor, and the defect signature follows.
That forces clarity on three things teams often leave fuzzy:
Then comes prevention by default. “Solved” means the line doesn’t rely on people remembering the story. It relies on engineered controls: limits, interlocks, monitoring with clear triggers, control plan updates, and change discipline that prevents a “fixed” mechanism from being quietly reintroduced during the next adjustment.
The practical definition is simple:
You have a verified mechanism you can reproduce, and you have controls that keep it from recurring under normal variability and normal change.
You don’t need a reorg to stop one chronic issue. You need a workflow that forces mechanism clarity, reproduction, and prevention controls—then preserves the learning.
A pilot-sized loop looks like this:
PRIZ supports this workflow at a high level by keeping the signature, causal model, tests, decisions, and prevention actions in one project space, so the outcome is a mechanism + evidence + prevention package, not a closed ticket that has to be re-learned next quarter.
A chronic problem isn’t hard. It’s made chronic.
It becomes chronic when containment counts as closure, when “back to baseline” is treated as solved, and the mechanism never gets proved and engineered out.
“Solved” means you can reproduce the failure on purpose, prove the mechanism, and block it with prevention-by-default solutions, so normal variability and normal change don’t bring it back.
If you want to break one chronic loop this quarter, run a single-excursion pilot with PRIZ: one signature, one mechanism, one documented prevention package, so it becomes reusable.