On this page
Equipment failure's share of unplanned downtime (Siemens, True Cost of Downtime, vendor-reported).
Pilot on one line or cell, starting in shadow mode at zero risk to production.
Shadow run showing what it would have done, with nothing permitted to act.
Challenge
One bad automated call on a running line can cost a shift of output or a whole batch. The risk is identical whether the call comes from a person or a tool: a fast decision, made under pressure, with nobody signing off and no record of why.
- Equipment failure is the largest cause of unplanned downtime. Roughly 42% of incidents (Siemens, True Cost of Downtime, vendor-reported).
- Alarm fatigue dulls the one that matters. A supervisor can spend much of a shift on alerts that turn out to be nothing.
- Night shift is where it bites. Thin staffing, fewer senior people, and no second opinion on a fast call.
Two calls on the same line
Routine
Consequential
An alarm fires
An alarm fires
One of many that shift
One that stops or holds output
The call is made
The call is made
Proceeds, logged with its reasoning
Waits for a supervisor's sign-off
The line keeps moving
The line waits, deliberately
No supervisor time spent
Attention spent where it changes the outcome
Who decided
Who decided
Recorded, reviewable later
A named person, at the time
Solution
Not every automated call needs a person. The ones that stop a running line do, and the layer decides which is which before anything moves.
- A sign-off desk between the call and the line. Routine, low-risk flags keep moving. High-consequence moves wait.
- The signal is checked before anything acts. Currency, range, cross-source agreement and actual machine state, so a stale reading cannot trigger a move.
- Hard limits on critical equipment. Nothing touches it without a person, whatever any tool decides.
- Plain-language record of every call. What the signal was, what was done, how it turned out, ready for shift handover.
Who Decides What, Written Down Once
This is not a question about models. It is a question about who is allowed to do what on a running line, answered once and enforced on every shift rather than re-argued at 3am by whoever is nearest.
- The routine tier. Low-consequence flags act without waiting for anyone. If a supervisor has to confirm everything, the desk becomes the bottleneck it was put there to remove.
- The sign-off tier. Slowing or stopping a line, holding a batch, changing a setpoint on critical equipment. These wait, and they arrive as one approve or reject with the evidence attached.
- The never tier. Equipment where no automatic action is permitted at any confidence. The limit is set by asset before go-live, and the layer enforces it rather than the tool respecting it.
- The escalation path. Who is pulled in when the supervisor is on the floor and not at a screen, and how long a hold waits before it escalates. Agreed in advance, because the worst moment to design this is during the incident.
Why the Signal Is the Weak Link
The data these calls run on comes off equipment that predates the tools reading it. A reading can be stale, out of range, or simply wrong, and none of those look wrong in the output.
- Current. A reading older than the window set for that asset is treated as unknown rather than as unchanged. Age is the most common defect and the cheapest to check.
- In range. A value outside the physically possible band is a sensor fault. It is recorded as one instead of being dispatched as a process event.
- Agreeing with something else. Where a second source exists, the two are compared. A single source disagreeing with the rest of the cell is a reason to hold, not a reason to act faster.
- Matching the machine's actual state. The asset may be in changeover, already down, or handed over for work. A correct diagnosis against the wrong state is still a wrong move.
Night shift is where this bites.
Thin staffing, fewer senior people, and a fast call with no second opinion. The desk is the second opinion.
Outcome Derived
This is a 60 to 90 day pilot on a single line, cell, category or product family. The figures below are what the pilot measures against a baseline captured in its first two weeks. They are targets and instrumentation, not results already delivered.
- Wrong calls prevented before the floor. Each hold logged with the check that caused it.
- False alerts not chased. Measured as supervisor hours returned against the weeks 1-2 baseline.
- Handover and root cause from the record. Designed so an incident review reads a log rather than reconstructing memory.
Highlights
- A sign-off desk sits between any automated call and the line. Routine flags keep moving; the moves that stop a line, hold a batch or change a setpoint on critical equipment wait.
- Signal currency, range, cross-source agreement and actual machine state are checked before anything acts, so a stale reading cannot trigger a move.
- Hard limits keep critical equipment out of automatic action regardless of what any tool concludes.
- Works whether the line runs on manual calls today or is already testing automated ones, because the risk is the same either way.
- Every call is recorded in plain language, which is what shift handover and the incident review actually read.
Frequently asked questions
Do we need to be running AI on the line for this to help?
No. The risk is the same whether the call comes from a person or a tool: a fast decision under pressure, with nobody signing off and no record of why. The sign-off desk and the record apply to both, and they are what the line needs in place before more of those calls become automatic.
Will this add more alerts for supervisors to chase?
It is built to reduce them. Routine, low-risk flags keep moving without a person. Only high-consequence moves reach a supervisor, and they arrive with the evidence for a single approve or reject rather than as another alarm that has to be investigated from scratch.
What stops a stale sensor reading from stopping a line?
The signal is checked before anything acts: currency, range, agreement with other sources, and whether the machine is actually in the state the tool assumes. A reading that fails any of those holds the action instead of triggering it, and the failure is recorded as a sensor fault.
What does the plant get after an incident?
A plain-language record of what the signal was, what was done, who approved it and how it turned out. That is the artifact handover reads at shift change and the artifact a root-cause review starts from, instead of reconstructing the night from three people's memory.








