This article discusses a crucial design consideration for monitoring systems: knowing when to stay quiet, especially for newly configured monitors. It proposes a state-based approach rather than time-based delays, emphasizing that a monitor should "learn" about its target before confidently alerting. This prevents alert fatigue and provides a more honest representation of system state to users.
Read original on Dev.to #systemdesignTraditional monitoring engines often fall into a common pitfall: alerting prematurely on new services or endpoints. This occurs when a monitor triggers an alert based on initial failures during a service's startup phase, DNS propagation, or certificate issuance, before the service has had a chance to stabilize or even become fully operational. This leads to false positives and alert fatigue, eroding trust in the monitoring system.
The article introduces the concept of an AWAKENING state for new monitors. Instead of relying on a fixed time delay (like `new_group_delay` in Datadog), it advocates for a count-based approach. A monitor remains in AWAKENING until it has successfully processed a configured number of pulses (e.g., three pulses). During this state, even if all pulses are failures, no alerts are fired, acknowledging that the system is still learning about the monitored target.
Design Decision: Count Over Time
Using a pulse count (`N` pulses) instead of a fixed time delay (e.g., 60 seconds) for the warm-up period offers greater flexibility and robustness across varying polling frequencies. A time-based delay can be too short for slow-polling monitors or excessively long for fast-polling ones, leading to either premature alerts or prolonged silence. A count ties the warm-up directly to the amount of information gathered, making the monitoring system more adaptable.
private ReliabilityState evaluateReliability(
double healthIndex,
long totalCount,
long pulsesSinceActivation) {
if (pulsesSinceActivation < monitoringProperties.getAwakeningPulseThreshold()) {
return ReliabilityState.AWAKENING;
}
if (healthIndex >= monitoringProperties.getHealthyThreshold()) {
return ReliabilityState.HEALTHY;
}
return ReliabilityState.DEGRADED;
}The monitoring system incorporates a state machine with states like AWAKENING, HEALTHY, and DEGRADED. New monitors begin in AWAKENING. Once the `awakeningPulseThreshold` is met, the monitor transitions to HEALTHY or DEGRADED based on its accumulated health index. Crucially, the alert pipeline is guarded to ignore monitors in the AWAKENING state, preventing unnecessary notifications.
Transparency and User Experience
Beyond the technical implementation, making the warm-up state visible to the user (e.g., displaying an "AWAKENING" badge on the dashboard) is a key product philosophy. It fosters transparency, helps users understand why alerts aren't firing immediately, and builds trust by honestly communicating what the system knows or doesn't yet know about a new monitor.