The only honest way to design a monitoring system for a hazardous site is to assume that part of it is already broken. Not "might break someday." Broken now, quietly, while the dashboard glows green. Every sensor drifts, every cable corrodes, every radio link drops packets, and every controller eventually hangs. The engineering question was never whether components fail.
The question is whether the rest of the system notices, says so loudly, and keeps watching while you fix it.
A system that cannot answer that question is barely a monitoring system at all; it is closer to a decoration with a power supply. The stakes are not abstract. Over nearly three decades, U.S. chemical safety investigators have examined close to 180 major chemical incidents that together produced more than 200 fatalities and over 1,300 serious injuries, and inoperative gas detection and shutoff systems appear in those case files with dispiriting regularity.
Kill the Single Point of Failure
Start with the failure that hides best: the single point everything else depends on. Centralised architectures feel tidy on a whiteboard, one master node collecting readings from a field of obedient sensors, but that master node is a promise you cannot keep. Research on distributed leak detection for crude oil pipelines has shown how pushing detection logic out to multiple cooperating nodes eliminates exactly this weakness, with several nodes independently able to detect and localise the same event. If one goes dark, four others still raise the alarm. The lesson appears to generalise well beyond pipelines: any reading that matters should reach a decision point by at least two paths, and no single box should be able to silence the site by dying. For practitioners, the audit is straightforward. Draw the data path from every life-safety sensor to the alarm output and circle each element that appears only once. Each circle is a bet you are making with someone else's safety, and most of those bets can be bought out for the price of a second gateway.
People Are Part of the Architecture
Hardware is only half the redundancy story, and arguably the less interesting half. People are a layer of the system too, and they fail in their own predictable ways when nobody prepares them for instrument failure. The 40-hour curricula built around training for hazardous-site personnel drill a habit that engineers should steal wholesale: treat any single reading, and any single silent instrument, as unverified until a second source agrees. A workforce that expects instruments to lie occasionally will cross-check; a workforce that worships the display tends to walk into a vapour cloud because the alarm never sounded. Design the humans in, or accept that you have designed them out.
Silence Is Not Safety
Here is the distinction most builds get wrong: a dead sensor and a lying sensor are different emergencies, and the second one is usually worse. A dead sensor at least produces an absence you can hunt for. A lying sensor produces plausible numbers. Silence and safety look identical on a chart, which is why "no alarm" must never be allowed to mean "all clear." The fix is structural, not aspirational. Every channel needs a heartbeat, a periodic signal that says the sensor is alive, powered, and within its plausible range. Industrial control references treat heartbeat monitoring as the standard defence against stale data, precisely because a network failure can leave a control room acting on readings that stopped being true minutes ago. Miss two heartbeats and the system should escalate on its own, without waiting for a human to wonder why the graph went flat.
Plausibility checking belongs in the same layer. A hydrogen sulphide reading of exactly zero for six straight hours is not reassurance, it is a symptom, because real ambient readings twitch. A value frozen to the decimal point is a sensor talking in its sleep. In practice, three checks catch most of these cases: flag any channel whose variance drops to zero over a rolling window, flag any value that sits outside the sensor's physical range, and flag any reading that disagrees wildly with its nearest redundant neighbour. None of the three requires machine learning. All three require someone deciding, in advance, who gets woken up when the flag trips.
Watchdogs That Watch the Right Thing
The same logic has to run inside the controllers themselves, because firmware hangs are not exotic. Memory corruption, a wedged peripheral, two tasks deadlocked over a mutex grabbed in the wrong order: any of these can freeze a node while its network stack keeps politely answering pings. The embedded community has spent years refining watchdog strategies for exactly this situation, pairing a hardware timer that forces a reset with a software layer that checks whether every critical task is still making forward progress, not merely whether the CPU is awake. A watchdog that only confirms the processor is running will happily feed itself while the gas-reading task sits deadlocked forever. Monitor progress, not existence.
Two practical rules follow. First, size the timeout from worst-case normal timing plus a safety margin, never from the average, or you will trade real protection for nuisance resets. Second, log every watchdog reset somewhere permanent, because a node that reboots itself twice a week is a failure announcement you are choosing not to hear.
Sensors Lie Slowly
Then there are the sensors, which fail slowly and speak in half-truths long before they die. Anyone who has worked with metal-oxide gas sensors knows the ritual: the burn-in period, the drift, the way clean-air baselines wander between individual units of the same model. The documented behaviour of the MQ-2 gas sensor makes the point plainly, since its readings depend on heater warm-up time, storage history, and a resistance baseline that must be calibrated per device rather than copied from a datasheet. The numbers involved should give any project manager pause: an MQ-series element that has sat unused for a month can need 24 to 48 hours of powered warm-up before its output stabilises, against a few minutes for a recently used unit. Deploy such a sensor with a hard-coded threshold and you have not built a detector so much as a random number generator with a siren attached. Calibration is not a commissioning task; it is a maintenance rhythm, and the monitoring system should track its own calibration age the way an aircraft tracks flight hours.
There is a cheap and unglamorous way to internalise all of this, and it does not require a refinery. Build the toy version first. Wiring up a basic Arduino smoke detector on a bench teaches lessons that no architecture diagram will: thresholds that seemed sensible at noon false-alarm at midnight, analog values wobble with supply voltage, and a loose jumper wire reproduces every symptom of a dying sensor. The bench version is where you learn humility about components. Scale that humility, not the circuit.
Fail Toward Safe
A short digression, because it matters more than it seems: fail-safe defaults are a moral choice disguised as an electrical one. Relays that hold a valve open only while energised. Alarms that trigger on loss of signal rather than presence of signal. Normally-closed contacts for anything that protects a life. Each of these decisions costs pennies at design time and determines what the site does in the thirty seconds after a power supply dies. Choose defaults so that the failure state is the safe state, and many of your failure modes stop being emergencies at all. Choose the opposite and you have built a system that protects people only when nothing is wrong, which is to say, only when it is not needed.
Supervise the Supervisor
Supervision needs supervising too, and this is where the architecture earns its keep. A monitored microcontroller with its own internal watchdog can still wedge in ways that fool itself, which is why designers of unattended equipment add external supervisory watchdog circuits, sometimes a dedicated chip, sometimes a second small microcontroller that watches heartbeats and communication lines independently of the thing it guards. Field deployments in remote pipeline monitoring take the idea further and layer the protection, with hardware watchdogs set around 30 seconds and network watchdogs around 120, so that a frozen process, a crashed operating system, and a dropped communication link each meet a defence tuned to its own timescale.
Yes, the supervisor can fail as well. That objection is not clever; it is the whole point. Layers do not promise immortality. They promise that no single failure, anywhere, is enough. Two independent faults must line up before the site goes blind, and you make that alignment rare through diverse hardware, separate power, and separate failure physics.
Break It on Purpose
Test the failures before the site does. Not the happy path, which will pass, but the ugly ones: pull the antenna mid-transmission, brown-out the supply, unplug a sensor and confirm the escalation fires within the promised window. Do it on a schedule, unannounced, the way fire drills work, because a recovery procedure that only ran once during commissioning has quietly expired and nobody filed the paperwork.
Simulation earns its place here, and it costs almost nothing. Modelling an LPG leak detector in a circuit simulator before touching hardware lets you inject faults that would be dangerous, expensive, or simply impossible to stage on a live site, and it turns "we think the fallback works" into "we watched it work, twelve times, under twelve different failures." A fallback that has never been exercised is not a fallback. It is a rumour.
The Sixty-Second Question
Little of this is expensive by the standards of what it protects. A second communication path, a supervisory timer, a calibration schedule, an afternoon of deliberate sabotage in the lab: these tend to cost less than a single day of downtime and immeasurably less than a single injury. What they actually cost is pride, because designing for failure means admitting, in the bill of materials and in front of the client, that your beautiful system will break. Some teams cannot bring themselves to write that down. Their dashboards are the greenest of all.
So here is the assignment, and it fits in one working week. Audit the system you have and ask it one question: what happens in the first sixty seconds after your most trusted component dies? Trace the answer on paper, then prove it with a controlled test. If the answer involves a person noticing something odd on a screen, the design is not finished; add the heartbeat, the plausibility check, or the second path that closes the gap, and put the next failure drill on the calendar before the meeting ends. Because somewhere on a site right now, a sensor has been reading zero for three weeks, the chart is flat, the log is quiet, and everyone agrees the air has never been cleaner.