graph_watchdog design
Role
Detects silent faults in the ROS 2 graph. A GatewayPlugin
hosting a fleet of detectors; each detector observes the ROS 2 graph and raises
faults via a ReportFault service client on the gateway node. Faults reach the
SOVD /faults API via FaultManager - the plugin injects no entities. Its only
HTTP surface is a read-only reliability status route (see Reliability core
below).
Structure
Plugin shell (
GraphWatchdogPlugin): loads via the gateway plugin ABI (v7). Inset_contextit casts the context withas_ros_plugin_context, creates onerclcpp::Client<ReportFault>on the gateway node, and starts a dedicated tick thread. The tick is deliberately NOT a gateway wall timer: detectors do blocking parameter/service reads, and the gateway’s small executor is only safe because blocking never runs on an executor thread. No child node.Every ROS entity the plugin owns - the fault client here and the
LifecycleWatcher’s~/transition_eventsubscriptions below - is created in a callback group of its own withautomatically_add_to_executor_with_node = false, and added only to a privateSingleThreadedExecutorthat the tick thread pumps between ticks. Neither executor is ever spun on a thread of its own. The reason is the same in both cases and is a correctness requirement, not tuning:rclcpp::AnyExecutableholds a strong reference to the subscription or client it dispatches and is a local of the executor thread, so an entity left in the node’s default group can have its last reference released - and its destructor run, mutating the node’s rcl entity registry - on a gateway executor thread, concurrently with this plugin creating an entity on the same node. No lock the plugin holds can reach a gateway executor thread, so the fix is to keep the entities out of every gateway executor; that puts creation, destruction and callback execution on the tick thread alone.For the fault client the drain is load-bearing for a second reason:
DetectorContext::raise_faultis fire-and-forget (it guards onservice_is_ready()and discards the future), and rclcpp keeps per-request state until an executor processes the response, so an unpumped client would leak one pending entry per raised fault for the life of the process.Detector pattern:
Detector+DetectorContext; each detector is one self-registering file undersrc/detectors/(REGISTER_DETECTOR), globbed by CMake so parallel authors never edit a shared file. A detector’s source needs no CMake or registry edit; adding a per-detector unit test is the one shared touch - it adds anament_add_gtestentry to the test block.Fault-code contract: the frozen
GRAPH_*namespace.mode seam:
raise/advisory/offper detector.Reliability core: see the dedicated section below - central warmup and lifecycle gating enforced inside
raise_fault, detector-consulted clock validity, and thex-medkit-watchdogstatus route.
Reliability core
ReliabilityGate composes two independent trackers and is the single entry point both the tick loop and the HTTP route go through:
WarmupTracker(pure, no ROS): an entity arms once continuously present forwarmup_cyclesticks. A disappearance longer than a short forget grace followed by a reappearance (a real mid-run restart) re-warms it from scratch; a transient one-tick discovery gap is absorbed and does not re-warm, so recurring DDS churn cannot re-arm (and thus permanently suppress) a still-running entity. Unknownsource_ids (e.g. a topic, not a tracked app) fall back to a global bringup grace window keyed off the first tick the graph was seen non-empty, which re-arms whenever the graph empties out so a full-stack restart gets the grace again.LifecycleWatcher(event-driven): discovers managedrclcpp_lifecyclenodes from the introspection snapshot, seeds their state via aGetStateservice call, then keeps it fresh by subscribing to each node’s~/transition_eventtopic (reliable + volatile, matching thercl_lifecyclepublisher). A node still cached non-active shortly after discovery is briefly re-seeded viaGetStateso anactivetransition lost during the subscription’s DDS endpoint-matching window self-heals rather than suppressing the node for the process lifetime; the blocking seeds are bounded per tick so a batch bringup cannot stall the tick loop. An entry’s identity is its BINDING - the pair (fqn,GetStatepath) - not itsApp::id, which can survive a graph sweep while pointing at a different node. When either half moves, the entry is dropped (the old binding recorded as departed under its own fqn) and re-seeded from scratch, so a moved binding can never keep enforcing the old node’s label. The subscription callback holds only aweak_ptrto the watcher’s shared state, so a callback in flight during teardown bails instead of touching freed memory. Non-managed nodes are never gated -node_okreturns true for them unconditionally, and so is a managed node whoseGetStatehas never answered: its label stays empty, andnode_okreads an unread label as permission on purpose, since gating on it would silence every detector for a node whose lifecycle service is broken. Only a KNOWN non-active label suppresses.presence_ownership()asks the STRICTER question the presence class needs, without changing that permissive answer, and answers it with a GROUND rather than a boolean: EARNED for a state read asactiveor for a node with no lifecycle at all, PROVISIONAL for one whose state was asked for as often as it ever will be and never came (measurement_pending(), which reads the per-node GetState re-seed budget charged only for reads that actually ran), and neither while the asking is still going on. Only the first is permanent - see thenode_deathsection.Those subscriptions follow the private-callback-group rule described in the plugin shell bullet above: created on the shared gateway node, but in the watcher’s own group and pumped only by the watcher’s own executor, so they are created, run and destroyed on the tick thread alone. This is also why the tick thread drains in short slices rather than once per tick - a transition would otherwise be delayed by up to a whole tick interval, and a bringup burst could overrun the subscription’s queue and lose the intermediate transitions the departure classification reads.
Central enforcement.
DetectorContext::raise_faultruns every raise throughreliability_allows(gate, source_id)before the fault client sends anything; a detector raising about a still-warming-up entity or a lifecycle-inactive node is silently suppressed, transparent to the detector.clear_faultis never gated - a fault can always be cleared. A null gate (not yet wired) always allows, so unit tests that construct a bareDetectorContextare unaffected.Clock validity is detector-consulted, not central.
WatchdogClock::time_is_valid()flags a paused or absent/clockunderuse_sim_timeby comparing wall-clock advancement against sim-time advancement on eachmark_tick(). Unlike warmup/lifecycle, this is not enforced insideraise_fault- a time-based detector must check it itself and skip its own age/grace-period math for the tick when it is false.tf_static QoS helper.
tf_static_qos()returns a depth-1transient_localrclcpp::QoS, available to detectors alongsidetime_is_valid()./tf_staticpublishers latch with transient-local durability, so a late subscriber must match it; the actual/tf_staticsubscription belongs to a later detector.Status endpoint.
GraphWatchdogPlugin::get_routes()registersGET x-medkit-watchdog(mounted at/api/v1/x-medkit-watchdogby the gateway), returningReliabilityGate::status_json(): schema version,warmup_cycles, aglobal_state(armed/warming_up), and one entry per known entity (id,first_seen_tick,armed,state,lifecycle). Returns 503 withERR_SERVICE_UNAVAILABLEif the gate has not been constructed yet or has already been torn down byshutdown().Beside the gate’s own state the payload carries a
detectorsobject, one block per detector whoseDetector::status_json()returns something, omitted entirely when none does. It exists for a condition that belongs to a SINGLE detector and would otherwise live only in the gateway log, which no HTTP client and no e2e assertion can read -lifecycle_expectationreportstracking_saturated(it has refused to track a required node) beside the livetracked_nodescount and thetracked_node_capin force. The handler runs on an HTTP thread whiletick()runs on the tick thread and holds no lock a detector takes, so astatus_json()implementation must build its payload from atomics and must not block.Build note.
LifecycleWatcherreuses the gateway’s own lifecycle-state helpers (lifecycle_status_helpers.cpp,ros2_lifecycle_state_reader.cpp) compiled in viaGATEWAY_SRC_DIR- the same non-header-only reuse pattern other gateway plugins use - rather than reimplementing lifecycle-state parsing.
Detectors
qos_mismatch, orphan, param_drift, lifecycle_expectation and
node_death are the detectors this package ships so far. Two silent-fault classes
remain undelivered, GRAPH_TF_STALE and GRAPH_LATENCY_BUDGET; they land in
follow-up changes, each against its own issue.
qos_mismatch raises GRAPH_QOS_MISMATCH. It
watches every topic’s publisher/subscriber QoS pairs rather than parameter values: each
tick it enumerates every topic via get_topic_names_and_types() and, for each one,
every currently-connected publisher and subscriber (get_publishers_info_by_topic /
get_subscriptions_info_by_topic, which report the RESOLVED live profile - never a
hand-built one), checking each pub x sub pair for incompatibility.
Two levels behind one fault code. Per-pair incompatibility is not reported uniformly.
A subscriber incompatible with EVERY publisher on its topic receives nothing and raises at
SEVERITY_ERROR. A subscriber incompatible with some publisher but not all raises at
SEVERITY_WARN: DDS refuses that one pair, so that producer’s data is silently
discarded while the topic keeps looking alive - a real silent fault, since an
RxO-incompatible pair is a match DDS has already computed, not a heuristic. One fault code
carries one severity, so the emitted fault reflects the worst finding in the sweep. Before
this split the partial case was indistinguishable from healthy, which meant e.g. a second
/tf broadcaster or a hand-rolled BEST_EFFORT publisher on /diagnostics went
unreported forever.
RxO rule. qos_policy.hpp’s qos_incompatibility(pub, sub) implements RxO
(Request <= Offered) compatibility for the QoS policies that silently starve a
subscriber with no error surfaced anywhere in the graph: reliability, durability,
liveliness kind, deadline, and liveliness lease duration. Publisher = offered,
subscriber = requested; a pair is incompatible only in the one strict direction per
policy - a BEST_EFFORT publisher against a RELIABLE subscriber, a VOLATILE
publisher against a TRANSIENT_LOCAL subscriber, an AUTOMATIC-liveliness
publisher against a MANUAL_BY_TOPIC subscriber, or an offered deadline / lease
duration greater than the requested one. The reverse direction (e.g. a RELIABLE
publisher feeding a BEST_EFFORT subscriber) is compatible by design - the subscriber
asked for no more than it is offered - even though the QoS profiles differ, which is
exactly the discriminator the integration and e2e tests each pin with a positive control
(see “Test tiers” below). For deadline and lease, an unspecified/infinite offer is
unbounded (fails any finite request), while an unspecified subscriber value does not
constrain the policy (always compatible); history and depth are not RxO dimensions and
are not checked. The checks match only concrete incompatible enum pairs, so
SYSTEM_DEFAULT/UNKNOWN (never
reported by a live endpoint, which always carries the resolved profile) never raise.
Aggregation via the shared helper. qos_mismatch was the first detector to use the
AggregatedFault helper (aggregated_fault.hpp); orphan and param_drift go
through the same one, so no detector reimplements the level-triggered raise/clear
pattern. The rationale: the fault_manager identifies a fault by
fault_code alone, so one GRAPH_QOS_MISMATCH per mismatched topic would collide
into a single record under the shared code. The detector keeps one AggregatedFault
instance for the whole graph and, each tick, hands it every currently-mismatched
topic’s description (keyed by topic name so a repeat mismatch on the same topic
overwrites rather than duplicates); an empty map on a clean tick clears
(EVENT_PASSED), level-triggered semantics (see
“Closing the loop” in the README for the healing_enabled requirement to actually
reach HEALED). graph_source_id(), also in aggregated_fault.hpp, returns the
constant kGraphWatchdogEntityId ("graph_watchdog") and nothing else, so every
detector’s aggregated faults land under the same source_id. It deliberately does NOT
fall back to the host Component id: a runtime host Component built by HostInfoProvider
never sets external, collect_component_app_fqns only puts a Component’s bare id in
the fault scope set when it does, and a fault raised under it is therefore listed by no
entity endpoint at all. The ctx argument is unused and kept only because every call
site already has it in hand.
Coverage is exhaustive, not budgeted. Unlike a budgeted
round-robin sweep (bounded by a parameter-service round-trip budget),
qos_mismatch reads no external service - get_topic_names_and_types() and the
two get_*_info_by_topic() calls are local graph-cache queries, so every topic is
checked every tick with no coverage-latency trade-off to configure or budget knob to
tune. /rosout and /parameter_events are skipped - ROS 2 system topics with
their own well-known QoS conventions, not useful signal for this detector. The sweep
polls ctx.cancelled between topics so a shutdown mid-sweep bails promptly, the same
shutdown-responsiveness contract every detector honors.
Test tiers. Three tiers each prove a different layer, deliberately not overlapping:
Unit (
test_qos_policy.cpp): pureqos_incompatibility()logic against hand-builtrmw_qos_profile_tvalues - the RxO trio (reliability, durability, liveliness), the reverse-direction compatible case, identical-profile compatibility, and the never-raise guarantee forSYSTEM_DEFAULT/UNKNOWNenums a live endpoint never actually reports.Integration (
test_qos_mismatch_integration.cpp): a realrclcpp::Nodepublisher/subscriber pair over real DDS, read through the actualget_publishers_info_by_topic/get_subscriptions_info_by_topicAPI - proving the detector reads the RESOLVED live profile correctly, not just that the pure comparison function is correct - against a fakeReportFaultservice. A second, concurrent RxO-compatible-but-different-QoS topic pair proves the detector does not over-fire on “QoS differs somewhere in the graph”.E2e (
test/e2e/test_qos_e2e.test.py): the Acceptance gate - a real gateway process with the plugin.soloaded, a real fault_manager, and the operator-visibleGET /api/v1/faultssurface, proving the whole raise/clear story reaches a SOVD fault through the real tick timer, not just the detector’s owntick()called directly.
orphan raises GRAPH_ORPHAN. It looks for the mistake no ROS 2 tool
reports: a topic that exists but has no counterpart, because one side was spelled
slightly differently. Each tick it counts publishers and subscribers per topic, keeps
the one-sided ones, and pairs a publisher-only topic with a subscriber-only topic when
their names are within max_edit_distance Levenshtein edits of each other. Only a
pair is reported, never a lone one-sided topic: a publisher with no subscriber is
normal on almost every robot, so it carries no signal on its own.
Three guards keep this from crying wolf.
Names must sit in the same parent namespace.
/a/scanand/b/scandiffer by one character but belong to two different robots, and pairing them would be nonsense. The leaf and the namespace carry separate edit budgets and the namespace one defaults to zero, so this is an exact match unless an operator raisesnamespace_edit_distance. Raising it is what makes a misspelled namespace (/robott/scanagainst/robot/scan) visible at all, at the price of reporting every robot of a fleet whose namespaces are themselves near-misses, which is why it is opt-in rather than tuned.A pair that matches once every run of digits collapses to one placeholder is an enumeration, not a typo.
/lidar_1next to/lidar_2, each one-sided, is the ordinary state of a multi-sensor machine, and reporting it would fire on every such machine. Collapsing runs rather than comparing character positions also covers/lidar_9next to/lidar_10, where the index changes length. The price is a real typo in a digit going unreported, which is a deliberate trade.graceconsecutive sweeps must agree before anything is raised. During bringup one side of a healthy pair is routinely missing for a moment, which looks exactly like a typo until the other side appears.
An operator who still gets a false alarm puts the topic in allowlist. What this
detector cannot see at all is a namespace added or dropped by accident: /scan
against /robot/scan is far past any sensible edit distance.
Test tiers. test_orphan_policy.cpp pins the pure matching rules, including each
guard and the cases it deliberately lets through. test_orphan_integration.cpp drives
the detector over a real rclcpp graph and a fake ReportFault service, covering
the grace counter, the clear on repair, and the allowlist. test/e2e/test_orphan_e2e.test.py
is the acceptance gate: a real gateway with the plugin loaded, a real fault_manager, and
the fault read back from GET /api/v1/faults. It carries three pairs at once - one
reported, one allowlisted, one visible only under a raised namespace_edit_distance -
so it also proves a string array and an integer survive the parameter path into
configure(), which no C++ tier can show.
param_drift raises GRAPH_PARAM_DRIFT. It reads node parameters over the
parameter service and reports a value that no longer matches its reference. The
reference is either self-captured, the value the parameter had when the node armed, or
pinned in config through expect. A node absent for longer than the window the
reliability gate absorbs loses its captured baseline, so a restart - the documented way to apply
a new configuration - re-captures instead of reporting the change as drift. A shorter gap keeps
the baseline on purpose: DDS discovery drops a node from a single poll routinely, and re-capturing
on that would take the node’s current, possibly already drifted, value as the new reference and
heal a real fault that could then never fire again. Only the pinned form can catch a value that was
already wrong at startup; self-capture by construction treats whatever it first sees as
correct.
A budgeted sweep, unlike every other detector here. qos_mismatch and orphan
read the local graph cache and are therefore exhaustive every tick. This one makes real
service calls to other nodes, so it is bounded by max_reads_per_tick and sweeps the
graph round-robin. The budget is spent in service ROUND TRIPS at the watched node rather
than in transport calls, because one call is several requests there: a self-capture read
is a list plus a batched get, and an expect pin the node declares is a list, a get and a
descriptor read. A pin the node does NOT declare stops at the list, and is charged two rather
than the one it costs, because a second and rarer NOT_FOUND path does cost two and the error
code cannot tell them apart - a load bound may overstate, never understate. That buys a bound
on the load the sweep puts on watched nodes and pays with
coverage latency: a drift is invisible until its node’s turn comes, and so is a repair.
The reads run on a thread the plugin owns rather than the gateway executor, so a node
that stops answering stalls this sweep only.
Three horizons for a stale entry. A node missing from one sweep is not forgotten
immediately. Its captured BASELINE is dropped on the third consecutive sweep without it - the
two the warmup tracker’s forget grace absorbs, plus one - so that a one-tick discovery gap does
not silently reset the baseline and hide a real drift behind a fresh capture. Its FINDING
survives prune_grace consecutive absences and is dropped on the next one, so that the same
gap does not clear the aggregated fault either. Only the finding horizon is configurable:
prune_grace is a per-detector config key, while the baseline horizon is the compile-time
kBaselineForgetGrace. A prune_grace below 2 makes the finding the shorter of the two and
the baseline then goes on the same sweep, because the absence counter it rides on is dropped
with the finding.
The third horizon is not about absence at all. A node that stays PRESENT but leaves the
reliability gate’s active state is never re-evaluated (it is not armed) and never pruned (it
is still there), so its finding would otherwise be re-raised for ever - an operator who fixes
the parameter and leaves the node deactivated would watch GRAPH_PARAM_DRIFT stay CONFIRMED
naming a value that is no longer true. kFrozenHoldTicks (60 consecutive gate-denied ticks,
one minute at the default tick period) bounds that: past it the finding is released and the app
also stops holding back a clear, and both re-derive from a fresh read if it ever arms again.
Because it is a compile-time constant and counted in ticks rather than in rotations, neither
max_reads_per_tick nor tick_interval_ms shortens it.
Severity is WARN, deliberately. The detector knows a value moved; it does not know whether that breaks the robot. Grading the consequence needs knowledge of the machine that the graph does not carry.
Nothing is cleared that was not measured. A clear asserts that nothing is drifting anywhere,
which is only true if every app was actually looked at, so the aggregate emits nothing at all while
any app still has no usable read. That covers three states which are all “nothing is known about
this app”: the round-robin has not reached it, its reads keep failing, or the reliability gate has
not armed it - and the last of those is the ordinary state of every app for the first
warmup_cycles ticks of a run, so without it a restart clears before issuing a single round trip.
A cached read counts as evidence only until a read of that app FAILS, after which the app is
unmeasured again. Each hold is bounded from the other side: an app is given up on after both sixty
failed attempts and sixty ticks unread, and a gate-denied one after the frozen hold. The bound on
attempts alone would not do, because an attempt is not a duration - the reader paces itself off
max_reads_per_tick, so at the accepted ceiling sixty-one attempts are spent in milliseconds and
a node still being discovered would be written off. Finally, a clear withheld for a minute of ticks
is logged once with the count of apps behind it, because from outside the process a correctly
withheld clear and a detector that has nothing to report look exactly the same.
Test tiers. test_param_drift_policy.cpp pins the pure comparison, rendering and
config rules. test_param_drift_integration.cpp drives real nodes with real parameters over
the real parameter service against a fake ReportFault service, including a node that
never answers, which must not hang the tick, and a node that answers past the per-read bound,
which must not be read at all.
test/e2e/test_param_drift_e2e.test.py is the acceptance gate: one gateway launch carrying
the whole raise-and-clear story for the detector’s default self-capturing mode, with a real
fault_manager and both ends read back from GET /api/v1/faults.
test/e2e/test_param_roundtrip_e2e.test.py measures what one parameter read really costs the
node that serves it, against a node that answers the parameter services by hand and counts every
request - so a round trip added to or removed from the TRANSPORT is caught, instead of quietly
invalidating the round-trip figures the budget is built on. Nothing else in the package can
notice that: every other test proves the arithmetic GIVEN those figures, and the stand-in
transport they use counts calls. It does not, however, read the detector’s own charge constants
at all - it runs with param_drift switched off - so what holds those to the measurement is
the timing suite in test_param_drift_integration.cpp, which times the reader against a
configured budget and therefore moves when a charge does. Cost and charge are pinned separately
and on purpose.
test/e2e/test_config_plumbing_e2e.test.py drives three further gateway launches covering the
nested expect sub-object reader, mode: "off", and the bare mode: off that the
ROS parser types as a YAML 1.1 boolean; those three prove configuration delivery and
suppression only, and assert no clear.
lifecycle_expectation raises three independent faults, GRAPH_NODE_INACTIVE,
GRAPH_NODE_UNREADABLE and GRAPH_NODE_NOT_MANAGED. It watches an operator-declared
set of require_active node names rather than topics, parameters, or presence. Each
tick it matches every configured entry against the live apps by App::id OR the
stable effective_fqn() OR that fqn’s bare leaf name - App::id alone is
recomputed each sweep and gets namespaced on a same-bare-name collision, so an id-only
match would silently stop covering a multi-robot graph; a bare name therefore matches
that node in every namespace (use a full FQN to pin one). For each matched present node
it reads the node’s raw lifecycle label via ReliabilityGate::lifecycle_state_of()
and hands the matches - a std::vector<LifecycleMatch> of (entry, node fqn, state) -
to lifecycle_expectation_tracker.hpp’s LifecycleExpectationTracker, the
detector’s pure, ROS-free core, independently unit-tested before the detector ever
touches a live gate.
The model: one observed state per node per tick, two clocks - not three fault codes
each running their own bookkeeping. All three faults are outputs of a SINGLE per-node
state machine inside the tracker. Every tick, a matched node is classified into exactly
one of ACTIVE, INACTIVE (non-empty label, not active), UNREADABLE (a
managed node whose label has never been read - optional("")), or NOT_MANAGED
(nullopt - no tracked lifecycle at all); a node not matched at all this tick is a
fifth case, ABSENT, handled by a wholly separate path. Two clocks track history:
The violation streak advances on
INACTIVE, resets ONLY onACTIVE. Pastgracethe node isGRAPH_NODE_INACTIVE’s content.The unmeasured clock advances on
UNREADABLEORNOT_MANAGED, resets ONLY on a real measurement (ACTIVEorINACTIVE). It is deliberately BLIND to which of the two causes it is seeing on any given tick - a node alternating between “matched but never read” and “not a tracked lifecycle node at all” keeps this ONE clock climbing instead of each cause resetting a separate counter of its own. That blindness closes an entire class of alternation, not one more special case for it: three review rounds each found a pair of “I cannot measure this” causes whose separate counters could erase each other’s progress, leaving a node invisible to every code that existed. Pastunmeasured_hold_ticks(a fixed 60 ticks, not configurable), the clock MATURES and takes ownership of the node.
Ownership is exclusive. A node is content of at most one of the three faults at a
time. The moment the unmeasured clock matures, the violation streak is RELEASED -
actually reset to 0, not merely excluded from GRAPH_NODE_INACTIVE’s content this
tick - so that fault’s clear is free to flow immediately; a node returning to
INACTIVE afterwards re-earns grace from zero. Which of the two unmeasured codes
a matured node reports under is STICKY: it keeps reporting under whichever code the
clock matured under even if the LIVE cause later flips, and only a real measurement
(resetting the whole clock) changes it. The rejected alternative, current-cause-wins,
is more literal but makes a node that flaps between the two causes flap the fault
surface too - a raise/clear/raise churn across two codes for a node whose actual
situation never changed.
Absence continues the last CORROBORATED observation, it never erases. A node not matched
at all this tick is ABSENT, a fact separate from any observed state (there is no label to
classify). For up to absence_grace (a fixed 3 ticks) consecutive absent ticks a node’s
whole state is held unchanged - the blink tolerance. Past absence_grace absence resets
nothing, but what it ADVANCES depends on whether the node’s SETTLED observation is already
CONTENT, and - for a below-grace violation streak only - on whether the presence detector is
CURRENTLY tracking the node: a settled-INACTIVE node ALREADY reported under
GRAPH_NODE_INACTIVE keeps its streak exactly as it was, so the fault stays raised,
regardless of ownership. One not yet past grace splits on that fact, read every tick from
the key set node_death publishes and never remembered: a node in that set is HELD - neither
advanced (no fault born from evidence gathered while nobody could observe the node, since
node_death is tracking this exact departure) nor erased (it
resumes rather than restarts on return) - while a node outside it keeps climbing on absence
exactly as it would on a present tick, because nothing else
in the plugin will ever report its departure either. A settled-UNREADABLE/NOT_MANAGED
node keeps climbing its unmeasured clock under that same cause regardless of maturity or
ownership - that clock’s own absence behaviour is unrelated to this distinction. A
settled-ACTIVE node advances nothing at all - anything it had started but not corroborated
is released, so the entry becomes idle and is reclaimed silently. So a departure never heals a
fault it has already earned, and never starts one out of evidence gathered while the node
could not be observed AND could be reported some other way; what it changes, for an
already-raised fault, is the detail phrase, which then says the node has since left the graph.
What “settled” means, and why an unmeasured reading needs corroborating. A real
measurement (ACTIVE or a non-active label) settles at once - a label is a fact about the
node, and no discovery artifact invents one. An UNMEASURED reading settles at once too on a
node nothing has ever measured, since it is then the only thing known about it (this is what
keeps a crash-looping node accumulating toward its report rather than being released on every
absence). Only when an unmeasured reading would OVERRIDE a real measurement must it first
hold for kDefaultObservationSettleTicks (6) consecutive matched ticks. The transient it
guards against is real and cheap to hit: LifecycleWatcher::update() drops a tracked id
whose get_state path is absent from the current sweep, and discover_apps() can yield
an app with no services when a sweep races service enumeration - so a present, healthy,
managed node reads NOT_MANAGED for a tick. If that tick is the last one before a clean
shutdown, continuing “the last observation” literally would mature a healthy departure into a
permanent GRAPH_NODE_NOT_MANAGED. Six is the same bar the typo warning already sets
against the same transient, and a tenth of the 60-tick unmeasured hold, so R30 is untouched: a
node that is genuinely unmeasurable when it leaves corroborated that long before the hold that
reports it.
Where “a departure never heals a fault” ends. It holds within one gateway lifetime. Across
a RESTART it does not: the restarted tracker has no measurements, the departed node is not in
the graph, and its require_active entry matches nothing - which this detector cannot tell
apart from a misspelt entry, because the only component that knows the difference is the fault
store and it is not read at startup. The never-matched hold is deliberately bounded (a typo
must not block healing forever), so once it lapses the level-triggered clear flows and the
record heals with nothing having been measured. Re-seeding the tracker from the fault store at
startup would change that and is not implemented; the boundary is pinned by the
restart_departed e2e scenario rather than left to be discovered.
Why the erasure went away rather than moving, for the two UNMEASURED causes: every erasure
horizon is an evasion for a node that touches it periodically. A node in a restart loop -
start, crash, respawn delay, start - touches absence by construction, and discarding its
evidence would let it alternate (UNREADABLE, ABSENT x N) or (NOT_MANAGED, ABSENT x N)
forever without ever accumulating enough of anything to be reported.
The VIOLATION streak needed the identical rule for the identical reason once - a node
alternating (INACTIVE, ABSENT x N) would otherwise never accumulate grace + 1 either.
Where the presence class CAN also report the same departure, it no longer does:
GRAPH_NODE_DISAPPEARED (this package’s own node_death detector) independently
reports it, whether or not the node was ever measured not-active first. But node_death
only ever tracks an App once the reliability gate says it OWNS that App’s departure, on one of
one of the two grounds ReliabilityGate::presence_ownership() accepts - so a require_active
node that never reaches active is never tracked by node_death, and neither is a managed
node whose GetState has not answered while the watcher is still asking. The gate would
permit that second one to RAISE - LifecycleWatcher::node_ok() treats an unread label as
permission, deliberately, or a broken lifecycle service would silence qos_mismatch,
orphan and param_drift too - but permission is not knowledge. Once the asking stops the
node IS tracked, provisionally, and a label arriving afterwards over ~/transition_event and
reading non-active takes it straight back out again. A node that is merely RE-WARMING is not
that: warmup says nothing about who a node belongs to, so the gate distinguishes kUnclaimed
(nothing known yet) from kDisowned (measured as another detector’s) and only the second
withdraws a key already held.
That withdrawal happens inside node_death’s own tick, and only on a tick that still sees
the app present, so it leaves a window of about one entity-cache refresh in which a node that
dies right after its label arrives is reported by the presence code anyway (measured: 210 ms
between the label reaching the status route and the release). A withdrawal that HAS happened
is final, which needed its own guard: a dying managed node loses its lifecycle services from
the snapshot before it loses its App entry, and a node with no managed record answers
kEarned - so dying erased the very measurement that disowned the node and re-admitted the
key on the tick before it departed. A released key is therefore readmitted on a label
reading active, never on the mere absence of a label - but the refusal runs for a bounded
window rather than the node’s whole life: after miss_grace + 1 consecutive record-less
PRESENT ticks (this detector’s own “absent this long means dead” span, already floored to
kMinNodeDeathWindowMs of wall clock) the key is taken back, because a node still present
past that window is not mid-death, it has simply stopped being managed - and holding the
refusal open would leave its eventual death reported by nobody at all. It is a mis-attribution of a TRUE
report, not a false positive, and it cannot be closed at report time: the departed record keeps
only the last label, which reads inactive both for a node earned and then deactivated - a
death this detector must report - and for one only ever held provisionally. Separating them
after the fact needs per-key ownership history surviving the departure, which is the unbounded
state the tracked-key prune exists to prevent. So
the split is on whether the presence detector is tracking the node RIGHT NOW - read from the
key set that detector publishes each sweep, not from the reliability gate and not guessed from
the observed label: while it holds the key, a below-grace streak that goes absent is simply
HELD rather than matured on the strength of the absence
alone - two codes standing at once for the same node is not a problem to fix, it is
GRAPH_NODE_INACTIVE and GRAPH_NODE_DISAPPEARED saying different, both-true things.
Reading membership rather than latching it is what makes the hold end when the claim behind it
does: a key handed back on a measured disown, one reclaimed by prune() after a durable
suppression, and one collapsed under the cap all leave the set, and the streak resumes climbing
for each. A node outside the set gets no backstop - GRAPH_NODE_DISAPPEARED
will not report it - so absence keeps advancing its streak exactly as a present tick
would. Reporting a HEALTHY node that left remains GRAPH_NODE_DISAPPEARED’s job as well, for
either kind of node - a healthy departure never starts a violation regardless of ownership.
Content follows the clocks, not the snapshot. Whether a node was in this tick’s matches
decides nothing about what it reports: a node past grace stays in
GRAPH_NODE_INACTIVE’s content while it blinks, and a matured node stays in its own
code’s content. A fault’s content is what has been MEASURED about the node, and a missing
snapshot entry measures nothing either way.
Presence is not enough here - and absence is someone else’s fault. A
present-but-non-active managed node (inactive, unconfigured, finalized) is
exactly the case GRAPH_NODE_INACTIVE exists to catch: on a Nav2 stack, a
controller_server stuck inactive means the robot silently will not act, with
no crash and no signal on /diagnostics or /rosout. A node that vanishes having
only ever been measured HEALTHY is the presence fault class’s business, and this detector
raises nothing for it; a node that vanishes while already reported under any of the three
faults KEEPS that fault, because leaving the graph answers nothing the operator asked -
the integration test’s AbsenceAfterRaiseKeepsTheFaultAndSaysTheNodeIsGone and
HealthyNodeThatVanishesRaisesNothingAtAll pin the two halves of that split.
Safe default: off. require_active is
empty by default, so the detector emits nothing at all - not even clears - until an
operator opts specific nodes in (zero false positives, the same config-scoped posture
param_drift’s expect uses).
Keyed by node, not by config entry. The bare-name form of require_active is
deliberately fleet-wide (“all controller_servers must be active”), so one entry
legitimately covers N nodes. The tracker therefore keys every fact by the node’s
stable fqn: two namesakes are two report entries (one cannot silently replace the
other, and a healthy one cannot mask a broken one), and the fault description names
each offending node’s FQN with the entry that demanded it as context - the
description is the only carrier, since every GRAPH_* fault shares one
source_id. Two entries naming the SAME node advance its clocks once per tick, not
once per matching entry.
Why grace, not the reliability gate (the design decision this detector turns on).
Every detector’s raise is subject to the central reliability_allows(gate,
source_id) gate inside ctx.raise_fault(), which ANDs the entity’s warmup state
with LifecycleWatcher::node_ok() - true only when a managed node’s lifecycle is
active (or it is not tracked at all). Gating THIS detector’s raise the same way on
the required node’s OWN lifecycle would be self-defeating: node_ok() is exactly
FALSE for the inactive node this detector exists to report, so the central gate would
suppress the signal forever. Instead the tracker counts consecutive not-active ticks
itself, per node, entirely independent of reliability_allows. This does not bypass
bringup-quiesce entirely: the aggregated fault’s own outer ctx.raise_fault call is
still subject to the central gate for the aggregate’s own source_id
(graph_watchdog) - only the per-node lifecycle check bypasses it, because that check
IS what this detector reports on.
Bounded by evidence, not by age. The map is NOT simply bounded by the
operator-declared require_active set: keys are node fqns, and one bare-name entry
matches a node of that name in every namespace, so identity churn (nodes reappearing under
ever-new fqns) would grow it without a bound. Since absence no longer erases anything, an
entry carrying a live clock is not reclaimed by age either - those are the same defect seen
twice, and moving a horizon rather than removing it leaves the evasion in place at a
different N. So prune_ticks (the operator’s prune_grace, used as written - there is
no grace + 1 clamp any more, because there is nothing left for one to protect) reclaims
IDLE entries only: both clocks at zero and no matured ownership, i.e. nothing to lose, and
still atomically - ONE map entry per node, gone in the same tick, never partially. A
non-idle entry is never pruned by age. An entry whose UNMEASURED clock is still climbing
cannot grow without bound in time either: past the absence grace it advances every tick, so
it matures within at most 60 + absence_grace + 1 ticks and is reported - the longest
GRAPH_NODE_UNREADABLE or GRAPH_NODE_NOT_MANAGED’s clear can be withheld by one
departed node. A below-grace VIOLATION streak on a node the presence detector is not tracking
shares that same bound, for the same reason it advances at all while absent: it matures within
at most grace + absence_grace + 1 ticks, because nothing else will ever report that node’s
departure either. Only once a node HAS been owned does its below-grace streak lose the
bound: absence then holds it rather than advancing it, so it neither matures nor becomes
idle for as long as the node is away. That is not a new way to withhold
GRAPH_NODE_INACTIVE’s clear - the clear was already gated on every required node’s
status being settled, so one node this indecisive already blocked it; what changes is only
that the fault never NAMES an OWNED node’s departure, since that node’s evidence belongs to
GRAPH_NODE_DISAPPEARED instead. grace is still capped at 300 instead of being
accepted up to INT_MAX - 1, because it independently
bounds how long a PRESENT node may go unreported and how long a returning node takes to
re-mature: at the old maximum that PRESENT-side bound was roughly 24 days at the shipped
cadence, with no warning - silence indistinguishable from a working detector finding
nothing.
What bounds the map is tracked_node_cap (default 512, kDefaultTrackedNodeCap,
accepted range 1..16384): at the cap idle entries are reclaimed first, then entries for
DEPARTED nodes are collapsed into per-code COUNTS - lexicographically last first, keeping at
most kMaxNamedDepartedEntries (3) named, since three maximally-long details are all one
480-character description holds - and only if every tracked node is PRESENT and carrying
evidence is the NEWCOMER refused, never a live violation evicted.
Collapsing follows from the fault being keyed by CODE rather than by node: five hundred
entries for dead identities keep the same one fault raised that a single entry would, so
holding them buys nothing while the slots they occupy can cost total blindness. Under
identity churn a cap full of the dead would refuse a genuinely broken PRESENT node - which
then never reaches affected or pending, is not covered by the never-matched hold
either (its entry HAS matched), and so GRAPH_NODE_INACTIVE would emit a level-triggered
CLEAR every tick while that node sat there not-active. The count is content, so collapsing an
entry heals nothing; it is ordered ahead of the individually named entries but behind
anything crossing on this tick, so a node that just broke is never displaced by it. It only
grows within one tracker lifetime: a collapsed entry’s fqn is no longer known, so a node
returning under it is tracked and measured afresh.
A refused node is a required node going unchecked, so refusing one is never silent: it
withholds GRAPH_NODE_INACTIVE’s clear for as long as it lasts, is logged once per
saturation EPISODE (the latch re-arms when the episode ends, so a later, real saturation is
not silenced by an earlier one), and is reported on GET /x-medkit-watchdog under
detectors.lifecycle_expectation. That withhold is deliberately UNBOUNDED, unlike the
never-matched hold beside it: the never-matched hold is bounded because a typo must not block
healing forever, while saturation - once departed entries can no longer crowd out present
ones - means genuinely more required PRESENT nodes than the cap allows, a capacity condition
the operator resolves rather than a transient that resolves itself.
The clear is withheld until the required set has actually been measured - and this is
entirely about GRAPH_NODE_INACTIVE. The other two faults have no withheld-clear
guard of their own; see “Three independent faults, not one shared record” below for why
they don’t need one. A clear asserts that every required node is free of a CONFIRMED
violation. That is the restart-heal hazard: a gateway restart brings every detector
counter back to zero while the GRAPH_NODE_INACTIVE raised before it is still in the
fault_manager’s store (a separate process), so a clear emitted on the strength of an
empty affected map heals a fault that is still real. TWO things produce that empty map
without the assertion being true, both the ordinary state of bookkeeping that just
started over (a restarted gateway, a reconfigure - configure() rebuilds the tracker
- or a node that respawned stuck):
Not matched yet. Before the entity snapshot catches up with the graph a
require_activeentry matches nothing at all, so the clear is about a node the detector has never once looked at. The plugin ticks as soon as it is loaded, so every restart passes through this window. Bounded the same way node-keyed state is (60 ticks), so a misspelt entry cannot block healing for the process lifetime.A node’s status is UNSETTLED (the tracker’s own
pendingset). A violation streak that has not yet passedgrace, an unmeasured clock still climbing under EITHER cause, or a streak HELD while the node is inside an unmeasured spell. Absence never puts a node here on its own - content follows the clocks, so a node already pastgracestays in the fault’s content through a blink rather than dropping into a withheld limbo. A node whose unmeasured clock has MATURED is deliberately NOT in this set - ownership passed to its own fault code and the violation streak was released, so it stops counting toward this withhold the exact tick it stops being uncertain, rather than continuing to poisonGRAPH_NODE_INACTIVE’s withhold decision the way a shared record used to.
Either reason withholds the emission entirely, neither raise nor clear. A raise is never
withheld: a violation read from the nodes that DID answer is real regardless of the
unmeasured ones. The never-matched leg and both unmeasured causes are bounded: they release
after 60 consecutive ticks (a minute at the shipped cadence, mirroring param_drift’s
frozen hold), whether the node is present or gone. The pending leg releases the same
way - as soon as the node reads active, or a PRESENT tick pushes its streak past grace
and the fault is raised again - but while the node stays absent that release has no timeout
of its own: absence holds a below-grace streak rather than advancing it, so the hold lasts
for exactly as long as the node does not return. Every hold releases by SETTLING the node’s
status, never by giving up on it. Because a correctly
withheld clear and a detector with nothing to report look identical from outside, a hold
that lives past 10 consecutive ticks is explained in the log once per episode, naming
every reason in force and, for the node-keyed ones, the count behind each and one node
by name - the not-managed and unreadable reasons are still named separately even though
both now release the same way (into their own fault code), since an operator reading the
log wants to know WHICH of the two is happening.
Three independent faults, not one shared record. GRAPH_NODE_INACTIVE,
GRAPH_NODE_UNREADABLE and GRAPH_NODE_NOT_MANAGED are each raised through the
shared AggregatedFault helper every GRAPH_* detector uses (one graph-level
record per code, since the fault_manager identifies a fault by fault_code alone),
but as three SEPARATE, fixed-severity class members - GRAPH_NODE_INACTIVE always
SEVERITY_ERROR, the other two always SEVERITY_WARN - the same shape
orphan_detector and param_drift_detector use, rather than one record whose
severity is chosen from that tick’s mixed content the way
qos_mismatch_detector’s any_starved ? kStarvedSeverity : kPartialSeverity does
for its own single code. Raises are fully independent: each fault’s content comes from
its own measurement, and one raising, healing, or changing severity never forces,
blocks, or reflects onto another. GRAPH_NODE_INACTIVE’s own clear is not simply
“nothing CONFIRMED non-active this tick” though - it is withheld exactly as described
above; that is the ORIGINAL withheld-clear guarantee this detector always gave, scoped
to GRAPH_NODE_INACTIVE alone. The other two have no such guard: each one’s own
clear needs nothing beyond its own content going empty, because once the unmeasured
clock has matured “still cannot be measured” is a settled fact, not a pending one. A
node is content of at most one of the three at a time, and healing one never forces,
blocks, or changes the severity of another. Content under either unmeasured code
survives for as long as a node stays that way, with no further bound past the initial
hold.
A re-bind is a fresh binding. LifecycleWatcher keys its tracked map by
App::id, and an id can survive a graph sweep while pointing at a DIFFERENT node
(id assignment shifts under bare-name collisions) - an entry kept across such a move
would keep enforcing the old node’s label, and its old ~/transition_event
subscription, against the new binding. An entry’s binding identity is its fqn plus
its GetState service path, both captured at first sighting, and update()
re-checks that identity every tick. A moved binding is two events at once: the OLD
binding departed (recorded under ITS OWN fqn in recently_departed_, same record
and retention as a vanish), and the id is new again (erased and re-seeded through the
ordinary new-node path: fresh GetState, fresh subscription, fresh self-heal
budget). For this detector the consequence is that a re-bind is never enforced with the
departed node’s label: the new binding starts unknown (benign) until its own label is
read.
No straggler from the old subscription has to be filtered out of the new entry. Erasing
the entry destroys that subscription, on the tick thread: the private executor that
runs these callbacks is pumped by the same thread that runs update() (see the
plugin-shell bullet above), so nothing else holds a reference to it and no message it
had queued is delivered afterwards. The same property removes the other ordering
question this seam used to carry - a re-seed’s blocking GetState cannot be
overtaken by a ~/transition_event, because while update() blocks nothing is
pumping the events at all.
New violations are named first, not alphabetically - in EACH fault’s own description
independently. The description used to list affected nodes in fqn order
(AggregatedFault::emit’s default), which is fine when every entry is equally
interesting but not once the cap is full: a fleet sharing
require_active: ["controller_server"] across a dozen robots fills the 480-char cap
from the alphabetically-earliest ones, and a THIRTEENTH robot going inactive afterward
would be silently invisible forever - one shared fault_code, one record, no way to
tell the operator which of thirteen actually broke. The tracker reports which fqns
entered EACH fault’s content on THIS tick - newly_affected, newly_unreadable,
newly_not_managed - and each list orders only its OWN fault’s
AggregatedFault::emit_ordered call, since the three faults never share a
description: a fresh entry is named FIRST in whichever fault it belongs to; every other
affected node in that same fault still appears, in the same fqn order as before, once
the fresh ones are placed. When every node crosses on the same tick (grace: 0
during a bringup burst, for instance), there is nothing to distinguish them by and the
order degrades to fqn order. Two budgets protect each description, applied before it
ever reaches the 480-char cap: the lifecycle label - which arrives verbatim off a
remote ~/transition_event and is therefore untrusted and unbounded, and only ever
appears in GRAPH_NODE_INACTIVE’s own detail, never the two unmeasured faults’ -
is trimmed to 32 characters before it is interpolated (more than double the longest
label a conforming implementation produces, errorprocessing at 15 characters), and
the whole per-node detail is then capped at 150 characters as a backstop against a
pathological fqn or a long require_active “required by” list, sized so at least
three worst-case details still fit inside the 480-char cap
(3 * 150 + 2 * 2 = 454 <= 480).
Test tiers. Three tiers each prove a different layer, deliberately not overlapping - plus the shared-watcher seam the re-bind behaviour lives in:
Unit (
test_lifecycle_expectation_tracker.cpp): pureLifecycleExpectationTrackerlogic over hand-built matches. Covers the violation streak (stuck-inactive past grace raising keyed by the node with the entry as context, active never raising, reaching active within grace resetting the streak, both namesakes reported separately, two entries not halving the grace, a violating read winning a duplicate match’s tie-break); the unmeasured clock shared by UNREADABLE and NOT_MANAGED (neither ever confirms a violation while climbing; each matures into its OWN fault at the exact hold boundary; the clock resets ONLY on a real measurement, in both directions; maturity RELEASES the violation streak entirely, so a node returning to inactive re-earns grace from zero; the sticky-cause decision, pinned directly against the rejected current-cause-wins alternative); the alternation this redesign closes - a node alternating between unreadable and not-managed, indefinitely and on every single tick, still matures the shared clock; theINACTIVE/NOT_MANAGEDalternation across absence gaps longer than the absence grace, now counting only the MEASURED not-active legs; the violation streak surviving non-maturing unmeasured ticks and RESUMING rather than restarting, counted exactly and read throughpending_violation; absence CONTINUING an already-matured fault exactly as it was on its last present tick - a blink holds either clock unchanged; past the blink tolerance an unmeasured clock still climbing keeps climbing and matures on absence alone, while a below-grace violation streak is instead HELD and RESUMES rather than restarts once the node returns, proven both from one long absence and from many short ones interleaved with matched reads; a matured violation fault survives a departure, a node measured ACTIVE that vanishes raises nothing and is reclaimed, and the two(UNREADABLE/NOT_MANAGED, ABSENT x N)restart-loop shapes are swept at the absence grace and past it (theirINACTIVEsibling pinned to the opposite claim at the same two N: it never crosses grace from absence alone, for a node the presence detector is tracking - a node outside its key set gets the opposite result instead, maturing from absence alone withingrace + absence_grace + 1ticks exactly like the unmeasured clock does, and resuming rather than restarting on return the same way an owned node’s held streak does); thependingset and its per-reason breakdown; new-first ordering for all three fault-shaped maps; the remote-supplied label’s own trim budget ahead of the whole-detail backstop, reapplied to a matured unreadable node’s detail (which carries no label, only a fqn and a “required by” list); the age horizon reclaiming IDLE entries only, atomically, and never one that carries evidence even when it undercuts the absence grace; the SETTLING rule that decides what absence may continue - an uncorroborated unmeasured run before a healthy departure raising nothing, swept across every run length below the bound and with the entry released rather than left holding the clear hostage, the same run one tick longer still being reported, and a single MEASURED not-active read before a departure still confirming; and the tracked-node cap - idle entries reclaimed first, departed entries collapsed into a count so a present broken node is always admitted and reported, at most three left named, the count surviving into the description as content, a node returning after its entry was collapsed measured afresh, the newcomer refused only when every entry is PRESENT and carrying evidence, and saturation reported as a LEVEL on every refused tick with its edge re-arming when an episode ends - swept at a shrunk cap and again at the real shipped 512. The re-bind seam is shared infrastructure and is pinned separately intest_lifecycle_watcher.cpp.Integration (
test_lifecycle_expectation_integration.cpp): the detector driven against a fakeReportFaultservice, with a REALReliabilityGatearming the global bringup grace and feeding the detector its labels throughlifecycle_state_of()- injected viaset_lifecycle_state_for_test()strictly AFTER the lastgate.update()call for most cases, or, for the alternation and not-managed cases (which need a genuinenulloptthe injection seam cannot produce - it can only ever SET a tracked value, never remove tracking), by toggling whether the matched app carries lifecycle services and driving the gate’s real discovery path instead. The fixture cases cover theGRAPH_NODE_INACTIVEraise/clear round trip naming the stuck node; the active-from-arming positive control; the fault SURVIVING the required node’s departure while a healthy node’s departure raises nothing at all; bare-name and full-FQN matching against a namespaced app; the zero-config default measured as zero fault_manager requests of any kind; the unmanaged-entry and no-match warnings; the withheld-clear guard’s releases and its once-per-episode log line; the blink-plus-unread-re-seed sequence; a filler batch sized (from the real detail-building code) to exceed the 480-char cap aggregating into one fault with the fresh crossing named first; a PRESENT node crossing grace fresh while a batch of already-departed, already-matured ones sits absent-and-content, named ahead of them rather than truncated away by them; a required node appearing mid-run; a re-bind under the sameApp::id; and the reconfigure/config-validation edge cases. The unmeasured clock’s own split is pinned directly, for BOTH codes symmetrically: a managed node whoseGetStategenuinely never answers (throughset_managed_appand a real, failing seed - proven by assertinglifecycle_state_of()actually returnsoptional("")first, not assumed) is reported underGRAPH_NODE_UNREADABLE; a node with no tracked lifecycle at all is reported underGRAPH_NODE_NOT_MANAGED(no longer released into silence, the deliberate behaviour change this redesign makes); either hold releases into its report on the exact tick past its bound; a node already reported under one of the two clears once genuinely read, returning to ordinaryGRAPH_NODE_INACTIVEtracking; a confirmed node healing while an unreadable OR not-managed sibling stays present clearsGRAPH_NODE_INACTIVEpromptly without touching either unmeasured fault; content under either code survives with no window and no expiry; 25 nodes of either cause aggregate into one capped fault atSEVERITY_WARN; a node whose hold expires opens its fault’s own description over an already-reported filler batch with new-first ordering; an already-reported node of either cause KEEPING its own record once it vanishes, with its description switching to say the node has left the graph; a node returning from that absence staying reported with no clear/re-raise churn; each of the two UNMEASURED restart-loop shapes raising its own code from absence alone (theirINACTIVEsibling instead pinned to counting only real reads, never crossing from any of the interleaved absence gaps), and theinactive/not-managedalternation across absence gaps now counting only the MEASURED not-active legs; a below-grace streak surviving the tightestprune_gracewithout being confirmed while the node stays absent, and resuming (not restarting) once it returns; a withheldGRAPH_NODE_INACTIVEclear releasing when the absent node’s clock MATURES into its sibling’s content rather than when the node is given up on; the two wire strings pinned as hand-typed literals; and the independence claim in both directions.The
configure()-level cases pin the config contract: the unknown-key warning forrequire_activ(the worst-case typo - it also leaves the detector unconfigured, so the warning must precede the zero-config early return), a fully-valid config producing zero warnings (includinggrace: 0andprune_grace: 0, the documented low endpoints), negative, non-integer, past-the-int-range and exact-kMaxGrace-boundarygrace(both sides),prune_graceout of 0..3600 rejected on the WIDE integer (never truncated into a hair-trigger prune) plus both range endpoints accepted, non-arrayrequire_activeand empty-string and non-string entries each warning, the unclamped prune horizon reaching idle bookkeeping at exactly the configuredprune_grace(0, 1 and 4 - the smallest positive value included, since neither documented endpoint sweeps it) while a node carrying evidence survives it - including thegrace: 0, prune_grace: 0corner (an UNMEASURED clock, the only kind that keeps climbing while absent) and a widegracebeside the tightestprune_gracefor a below-grace VIOLATION streak, whose instrument is that it is neither confirmed nor pruned while the node stays absent, and resumes rather than restarts once it returns -tracked_node_capvalidated at both range endpoints and one value past each with the key proven IN FORCE at both ends, boundedness under identity churn at the real 512-node cap by collapsing the departed rather than refusing the live node, and saturation reported once per EPISODE with a second episode reported again after the first ends.E2e (
test/e2e/test_lifecycle_expectation_e2e.test.py): the acceptance gate - one source file, TWELVE CTest targets (the config-plumbing pattern: the plugin reads its config once atset_context(), so different configs need different gateway launches), each bringing up a real gateway with the plugin.soand a real fault_manager, asserting on the operator-visibleGET /api/v1/faultssurface. Six scenarios drive themanaged_lifecycledemo node (a realrclcpp_lifecycle::LifecycleNode, which always answersGetState); two,unreadableanddeparture_keeps, drive a fixture built to never answer it (see below); the last two,not_managedandrestart_loop, drivecalibration- a PLAIN demo service node with no lifecycle interface at all, so no purpose-built fixture was needed forGRAPH_NODE_NOT_MANAGED, unlikeGRAPH_NODE_UNREADABLE. The main scenario proves raise-survives-CONFIGURE-survives-RESTART-heals-on-ACTIVATE through reallifecycle_msgs/srv/ChangeStatetransitions, including the entity-scoped/apps/graph_watchdog/faultssurface. The restart leg is the withheld-clear guard at the only tier that can reach it: the gateway is SIGTERMed, its port is waited down, the relaunched process is gated on being armed again, and the fault about the still-inactive node must come back withlast_passedunset. The default-config scenario launches with nolifecycle_expectationconfig at all against the same inactive node and holds a sustained silence window, and the negative control does the same with the self-activating variant of the same executable; the discriminating variable is the node’s actual lifecycle state. Both silence scenarios gate on three facts before asserting absence: the plugin is armed;GET /faultsanswers 200 in THIS launch; and the target’s label was actually READ, pinned through the plugin’s ownGET /x-medkit-watchdogroute.The fourth scenario,
healing_threshold, is the withheld-clear guard’s pending leg at the only tier that runs the REAL fault_manager debounce state machine at all. The required node, already reported stuck, is SIGTERM’d and respawned twice under the same name (a real snapshot blink, not a lifecycle transition), each one comfortably inside the tracker’s fixed absence grace, against ahealing_thresholdof 1. This same run is also this package’s only e2e coverage for a node oscillating readable/unreadable producing noGRAPH_NODE_INACTIVEevent churn, read raw offGET /faults/stream: nofault_clearedframe forGRAPH_NODE_INACTIVEat any point, and everyfault_confirmed/fault_updatedframe carryingseverity_label: "ERROR"- its only possible value now that severity is fixed per code.The fifth scenario,
unreadable, is where the 60-consecutive-tick failure that maturesGRAPH_NODE_UNREADABLEis actually reached. It launchesunreadable_lifecycle_node.cpp: a plainrclcpp::Nodethat advertisesget_state/change_statewith the service TYPESfind_lifecycle_get_state_pathactually checks for, whoseget_statehandler stores every request and never answers until its ownstart_answeringparameter is set true. Against that fixture the scenario proves the node is present and matched with its lifecycle label read as"";GRAPH_NODE_INACTIVEstays silent for it throughout;GRAPH_NODE_UNREADABLEraises once the hold expires, names the node, and carriesSEVERITY_WARN- a node is content of at most one of the three, never more than one; and once the fixture starts answering"active"the watcher reads it through a realGetStateround trip andGRAPH_NODE_UNREADABLEclears.The sixth scenario,
departure_keeps, proves what a DEPARTURE does to an already-reported unmeasured fault: nothing. The same fixture is SIGTERM’d permanently onceGRAPH_NODE_UNREADABLEhas raised, confirmed gone fromGET /apps, and past every horizon that could have discarded its evidence the fault is still active, has never once been reported PASSED (last_passed, which catches even a transient clear), and its description now says the node has left the graph. The seventh,not_managed, is the sibling proof for the OTHER cause:calibrationis present and matched from launch,GRAPH_NODE_NOT_MANAGEDraises once its hold expires naming the node atSEVERITY_WARN, and neitherGRAPH_NODE_INACTIVEnorGRAPH_NODE_UNREADABLEever raises for it - the one live proof that the NOT-MANAGED cause specifically never bleeds into the UNREADABLE code it shares a clock with. Its departure leg mirrorsdeparture_keeps’s own shape.The eighth scenario,
restart_loop, is the acceptance gate for the whole evidence-retention model. The same plaincalibrationnode, but respawning, is SIGTERM’d over and over on a cadence that never lets it accumulate the 60 consecutive PRESENT ticks the unmeasured hold would otherwise need. Every cycle’s absence is proven rather than assumed -GET /appsis polled until the node is gone and again until it is back, and the measured gap must EXCEED the 3-tick absence grace, on top of launch’s own enforcedrespawn_delayfloor - so every cycle genuinely crosses the horizon past which evidence used to be discarded.GRAPH_NODE_NOT_MANAGEDmust raise anyway and survive the restarts that follow. A required node in a crash loop is the case this detector most exists to catch, and the one a design that discarded evidence on absence made permanently silent.The last four scenarios are about the bounds themselves.
cap_pressurerunstracked_node_cap: 1against TWO required nodes, so one is refused on every tick - reachable only because the cap is a config key; against a compile-time 512 it would need 513 real lifecycle nodes. It proves the refusal is visible onGET /x-medkit-watchdog, that it WITHHOLDSGRAPH_NODE_INACTIVE’s clear (measured throughlast_passed, which catches even a transient clear, at the moment the tracked node’s unmeasured clock matures and that fault’s content goes empty), and that the entry for a node that then LEAVES is collapsed so the present, still-broken node is admitted and named.unsettled_departureruns two healthy managed nodes built from this package’s owndroppable_lifecycle_node.cpp- which looks managed, answers a chosen label, and stops advertising its lifecycle services on command - and drops the services of each: one is killed immediately (a single missed sweep before a clean shutdown, about which nothing may ever be reported) and the other holds the dropped state past the settling budget before being killed (genuinely not managed when it left, so it must still be reported). The first leg’s window is MEASURED against the settling budget rather than assumed, so it cannot silently turn into the second.wide_graceconfigures the value that used to be the acceptedgracemaximum and proves it is refused and the documented default applied - under it the detector could neither raise nor heal for days.restart_departedrecords where “a departure never heals a fault” ends: the required node is killed while its fault is outstanding, the gateway is restarted, and the record then heals - because the restarted detector cannot tell an entry for a departed node from a misspelt one.
node_death raises GRAPH_NODE_DISAPPEARED. It watches every ARMED App in the graph
for one that goes offline - zero-config, unlike lifecycle_expectation’s
operator-declared require_active list, because every armed App is a candidate.
Liveness is App::is_online, never mere membership in the entity snapshot: in
runtime-only discovery the two happen to coincide (a dead node’s App leaves the snapshot
entirely), but a manifest keeps a bound App present with only is_online cleared once
its ROS binding disappears, and hybrid discovery inherits that shape - counting snapshot
membership alone would make a manifest-declared node immortal. A managed lifecycle node
that merely deactivates keeps is_online: true (its process is still running, only its
ROS 2 lifecycle state changed), so a deactivation is never mistaken for a death either -
that is lifecycle_expectation’s concern, not this one’s. Tracking is keyed on the
STABLE fqn (App::effective_fqn()), never App::id: an id is recomputed every sweep
and only gains a namespace prefix once a same-bare-name collision currently exists
anywhere in the graph, so a live node’s id can change out from under a key built from it.
What is never tracked. A peer-aggregated app (app.source starting peer:)
carries no ROS binding of its own, so effective_fqn() is empty for every one of them;
tracking them would collapse an entire peer fleet onto one "" key. A
_ros2cli_<pid> node - rcl’s own hidden-node naming convention for every ros2 CLI
invocation, matched structurally rather than guessed at - is excluded too, or every CLI
invocation would accumulate one permanently “dead” entry for the life of the gateway. An
App that never comes online is skipped by the is_online filter above and so is never
admitted to tracking in the first place - the identical protection a node still inside
warmup_cycles gets. A MANAGED App whose lifecycle state has not been read YET is excluded
too, and for a different reason: the gate’s raise permission for it is deliberately PERMISSIVE
(an unread label must not silence qos_mismatch, orphan or param_drift), but this
detector reports an ABSENCE and cannot attribute the departure of a node whose state nothing
has read - so tracking asks presence_ownership() instead.
That exclusion is BOUNDED, and the readmission that follows it is PROVISIONAL. The watcher
charges a GetState re-seed budget per node and charges it only for a read that actually ran, so
an empty label with attempts left is “we have not finished asking” and an empty label with none
left is “we asked and failed”. Past that point the node is admitted here after all - nothing
will ASK again, and lifecycle_expectation returns before it looks at anything unless an
operator named the node in require_active, which defaults to empty, so refusing forever
would mean a whole class of deaths reported by nobody. But asking is not the only channel: the
~/transition_event subscription is created independently of the seed budget and is never
torn down when it runs out, so a label can still arrive unasked. One that reads non-active says
the node was lifecycle_expectation’s all along, and this detector releases the key while the
node is still alive rather than reporting a death it was never entitled to. Ownership EARNED
from a measurement is not released that way: a node that ran and then deactivated is still a
death when it stops running. Stickiness is earned by knowledge, never by ignorance.
What that does NOT close: an App counts as managed only once the snapshot shows it advertising
a GetState-typed service, and until then it is indistinguishable from a plain node, so a
managed node whose services have not yet been discovered can still be taken for one. That
window is bounded by the arming warmup, which already needs several consecutive ticks of
presence, and node and services come from the same introspection sweep. Nor does anything here
close the case of a node that appears, lives fewer than warmup_cycles ticks and vanishes:
the gate has always required completed warmup before it arms anything, so such a node is owned
by nobody and reported by nobody. That is the bringup-quiesce trade-off working as designed,
not a gap this boundary introduces. Once
tracked, though, a node’s continued life-or-death judgement rests on PRESENCE alone, not
on staying armed: a tracked node that later goes lifecycle-inactive is not thereby
mistaken for dead. Composable/component nodes hosted in one process die together - if the
container process dies, every node it hosted leaves the snapshot on the same tick, and all
of them are named in the one aggregated fault rather than raising one fault each.
The wall-clock floor on miss_grace. The entity snapshot a tick reads is not
rebuilt every tick: runtime discovery rebuilds it off a graph event debounced to about one
refresh per second, so several consecutive ticks between two refreshes see the IDENTICAL
snapshot. A miss_grace counted purely in ticks does not count independent samples of
the graph - shorten tick_interval_ms enough and one stale cache generation gets
re-counted as several misses. The floor raises miss_grace, for the configured
tick_interval_ms, to the smallest value whose (miss_grace + 1) * tick_interval_ms
window is still at least 3000 ms - one graph-cache refresh cycle plus margin. At the
shipped 1000 ms tick the floor is exactly the shipped default (miss_grace: 2), so the
default configuration is unaffected; only a faster-than-default tick ever raises the
effective grace, with a warning naming the window it actually spans.
Suppression: opt-in, and only by name. Nothing is suppressed unless the operator
names the mechanism in suppress - a configured allowlist that suppress does
not name has NO effect, and a startup warning says so. Two mechanisms exist, activated
independently of each other:
"allowlist"builds anAllowlistSuppressorover the configuredallowlistset - the same three-way match (fqn, bare leaf,App::id)lifecycle_expectation’srequire_activeuses, so an operator who has one working can expect the other to accept the same shapes. Exact match only, never a prefix or a shared suffix."lifecycle"builds a (stateless)LifecycleShutdownSuppressor: it suppresses a departure the reliability gate classifies as a clean managed-lifecycle shutdown - see “Clean-shutdown suppression” below for what counts as clean.
Every field is re-validated from scratch on every configure() call, so a malformed
allowlist entry, an unrecognized suppress name, or a wrongly-typed entry each
produce their own named warning rather than being silently dropped. A candidate is
suppressed the moment ANY active mechanism votes yes, order-independent - which one runs
first never changes the result. Suppression is independent of mode: a suppressed key
never becomes content in the first place, while mode separately decides whether
non-suppressed content is actually sent.
Durability, and why only a durable veto may reclaim bookkeeping. Whether a suppressor
is durable answers whether its “yes” is a standing fact about the key rather than a
condition that can later stop holding. Both mechanisms here are durable: an
operator-declared allowlist entry does not start matching and then stop on its own, and a
departure’s observed shutdown shape does not change after the fact. Durability is what
makes reclaiming a suppressed key’s tracker bookkeeping (prune_grace consecutive
suppressed ticks) sound - a NON-durable veto could lift later even after its key’s
bookkeeping was discarded, leaving a live, unsuppressed death with nothing left to report
it: a false heal that silently outlives the very condition that produced it. A durable
veto never lifts for a given key once it has fired, so reclaiming under it loses nothing.
Suppressor::durable() defaults to false, the safe assumption for a suppressor nobody
has reasoned about yet; both suppressors here override it explicitly to true. Reclaiming
is re-checked against the durable suppressors specifically, never merely “no longer in
this tick’s report”, because a key can be dropped from a tick’s report by ANY suppressor
in the chain while a durable one also happens to cover it - whether a key may be reclaimed
depends on WHICH suppressor vetoed it, not on whether it survived the filter. The generic
chain is reusable (Suppressor plus the free function apply_suppressors(),
independently unit-tested), but this detector’s own filtering is hand-inlined rather than
a call to that helper, because it has to interleave the id-form check above - answerable
only while the node is still present - with the ordinary fqn/leaf dispatch the helper’s
plain string-keyed signature does not have room for.
Clean-shutdown suppression - what counts as clean. LifecycleShutdownSuppressor
reads the lifecycle watcher’s cached label for the departed node’s LAST observed
transition. shuttingdown alone suppresses - it is only reachable via a deliberate
SHUTDOWN transition, so the label alone is enough. finalized suppresses only WITH an
observed transition and never through the error branch; without an observed transition,
or reached through the error branch, it does not, because finalized is also where a
node’s on_error override lands after ON_ERROR_FAILURE/ON_ERROR_ERROR out of
errorprocessing - the standard way a driver reports a hardware fault it cannot
recover from, which is precisely the death an operator needs reported, not silenced.
unconfigured is deliberately excluded even though it looks quiet: it is also the
resting state of a node whose configure() failed or that never activated at all, and
suppressing on it would hide exactly that startup failure. Any other label, or no
departure on record at all, abstains. The suppressor itself is stateless by construction - a
detector’s configure() runs before any per-tick context exists, so it stores nothing
at construction and reads the gate fresh on every call.
The verdict is latched; the evidence is not. node_death remembers that a durable
suppressor vetoed a given departed key and stops re-asking while that key stays tracked and
absent. The label the suppressor reads is retained for a bounded number of ticks, while the
veto must hold to prune()’s reclaim - which needs it UNBROKEN for prune_ticks
consecutive ticks. Re-asking every tick made those windows race, and the retention clock is
anchored on the wrong event to win it: it starts when the node’s lifecycle SERVICES leave the
graph, while the reclaim tick counts from when its App does, one or more sweeps later. Once
the label expired mid-veto the suppressor abstained, the streak reset, and with the label gone
for good the key could never be suppressed again - a cleanly shut-down node named for the life
of the process. Latching is sound because that is what durable() means: the verdict never
lifts for a given key, so remembering it cannot diverge from asking again. Only positive
verdicts from durable suppressors are latched, and an entry is dropped the moment its key is
present again, so one departure’s verdict never carries into the next.
The retention window is still not this detector’s own to size: the plugin predicts the same
prune_ticks_ (max(prune_grace, miss_grace + 1)) and sizes the watcher’s departed-node
retention from it. With the verdict latched, that window only has to outlive the FIRST tick
suppression is evaluated on; it stays sized to the reclaim tick as slack against the
services-before-App lead (GraphWatchdogPlugin::compute_departed_retention_ticks()).
Bounded by evidence, not by age - and what the cap actually bounds. This detector is
zero-config over every armed App in the graph - a strictly larger scope than
lifecycle_expectation’s named require_active set - so identity churn would grow the
tracker’s map without a bound if nothing capped it. An unsuppressed, still-dead entry is
never reclaimed by age (prune_grace only ever reclaims a DURABLY suppressed key), so
tracked_node_cap (default 512, accepted range 1..16384) is what bounds memory -
specifically the DEPARTED subset of the map. A PRESENT/armed key never counts against the
cap and is never evicted to make room for anything, at any map size
(NodeLivenessTrackerCap.PresentEntriesExceedingTheCapAreAllStillIndividuallyReportableOnDeath):
it is bounded by the live graph rather than by churn, so a graph far larger than the cap is
not capacity pressure at all.
Under pressure the cap collapses departed entries into one synthetic count, keeping at most
three individually named
(NodeLivenessTrackerCap.MaturedDepartedEntriesAreCollapsedIntoACountUnderCapPressure) -
but only entries that have actually crossed miss_grace. An entry still mid-grace is
never a collapse candidate, because collapsing erases the identity and would report a death
the node has not earned, permanently
(NodeLivenessTrackerCap.ImmatureDepartedEntriesAreNeverCollapsedOrReportedUnderCapPressure).
Collapsing is one-way, and the cost is worth stating plainly: the identity is DISCARDED, the
collapsed count only grows within a tracker’s life, and nothing decrements it when one of those
nodes returns. So the description stops naming them, and - since the aggregated fault clears
only on an EMPTY dead set - the synthetic collapsed entry keeps that set non-empty, which means
GRAPH_NODE_DISAPPEARED cannot clear again for the life of the process once anything has been
collapsed. It ends by operator acknowledge or by a restart, never by the graph recovering.
Neither half is repairable here: decrementing needs the identity the collapse threw away, and
keeping identities is the unbounded state the cap exists to prevent, while decaying the count by
time would heal a fault by waiting - which this package refuses everywhere else. What bounds the
damage is what collapsing may touch: only a MATURED entry, whose death has already been reported
by name, so nothing becomes UNREPORTED by being collapsed. The attribution and the self-healing
are lost, not the warning.
So saturation IS reachable here: when immature departures alone exceed the cap,
tracking_saturated goes true and the departed set stays oversized for a few ticks -
bounded by how fast departures arrive times miss_grace, not by process uptime, which is
the growth the cap exists to prevent. What differs from LifecycleExpectationTracker is
not whether the flag can fire but what it costs: there a refused newcomer withholds
GRAPH_NODE_INACTIVE’s clear, whereas here nothing is refused and no report is lost, the
map is merely temporarily larger than its target.
The boundary with lifecycle_expectation. GRAPH_NODE_INACTIVE and
GRAPH_NODE_DISAPPEARED can both be raised for the same node at the same time, and that
is CORRECT, not a defect this detector removes: GRAPH_NODE_INACTIVE says a required
node is present but not active, GRAPH_NODE_DISAPPEARED says a node is gone, and an
operator facing each has a different repair - reactivate it, or find out why it left. A
node CONFIRMED GRAPH_NODE_INACTIVE that then departs keeps that fault (its description
switches to say the node has since left the graph) AND now also raises
GRAPH_NODE_DISAPPEARED, since this detector tracks every App it owns independently of
whatever lifecycle_expectation thinks of it.
What changed is narrower than that, and it only applies to a node this detector could ever
have tracked. Before this detector existed, lifecycle_expectation’s own violation streak
let a SUSTAINED absence mature an unconfirmed (below-grace) streak into a confirmed
GRAPH_NODE_INACTIVE - the only way a node stuck in a restart loop, never observed long
enough to cross grace while present, would ever be reported at all. Where this detector
can independently report the same departure - which needs the presence detector to have OWNED
the node at least once, since it only ever tracks an App the reliability gate has admitted for
ownership, and ownership needs a MANAGED node’s lifecycle state known to read active or
asked for as often as it ever will be, plus the node being online - that absence-alone maturity
is gone:
absence CONTINUES a violation that has already matured past grace (the fault stays
raised, still naming the node), but no longer CREATES one that has not. A node that was
briefly non-active and then died is reported as gone (GRAPH_NODE_DISAPPEARED) rather
than also acquiring an inactive fault born from ticks gathered while nobody could observe
it. A require_active node that never reaches active is never admitted, is never tracked
here, and its departure can never raise GRAPH_NODE_DISAPPEARED; the same holds for an app
that is not online. For those, lifecycle_expectation keeps the older behaviour: absence
still matures a below-grace streak, because it is structurally the only detector that will
ever get to report them. See
lifecycle_expectation’s own “Absence continues the last CORROBORATED observation, it never
erases” section above for the mechanism this rests on.
Repeated failures: what occurrence_count and the captured evidence answer, and
don’t. An operator asking “how many times did this node die in the last hour” reads it
off the fault record itself, not off this detector: GRAPH_NODE_DISAPPEARED’s own
occurrence_count starts at 1 on the first raise and increments by exactly one each
time a FAILED report reactivates a record that was CLEARED - never merely re-raised while
still CONFIRMED, and never merely HEALED. Healing (the fault manager’s debounce counter
crossing healing_threshold on clean sweeps, see “Closing the loop” in the README) and
clearing are different things: DELETE
/api/v1/apps/graph_watchdog/faults/GRAPH_NODE_DISAPPEARED - the fault manager’s own
~/clear_fault underneath - is what acknowledges an occurrence and closes its cycle. A
still-CONFIRMED, or still-HEALED-but-unacknowledged, fault that fires FAILED again is the
SAME occurrence continuing, not a new one, so restarting the dead node fast enough to heal
it before anyone acknowledges it will not move the count.
The honest limits, read off this branch’s fault-manager storage rather than assumed: the
per-fault rosbag store enforces fault_code as UNIQUE, so a fault can hold at most one
recording at a time - a later confirmation’s capture replaces the earlier one on disk
rather than accumulating a history. The freeze frame is the same shape: one row per
fault_code, overwritten on every capture, so a fifth occurrence’s captured values
overwrite the first’s. What survives every occurrence by default is the count itself; a
hash-chained record of every raise/clear/heal transition also exists, but only once the
fault manager’s own audit_log.enabled is turned on, which it is not by default.
Recordings and the freeze frame do not accumulate a per-occurrence history either way.
A second death while the first is still outstanding gets no evidence of its own. This
detector folds every currently-affected node into ONE GRAPH_NODE_DISAPPEARED record via
AggregatedFault::emit() (see “The boundary with” lifecycle_expectation above): a
second node dying while the fault is still CONFIRMED is added to the description on the next
tick, but the ReportFault call that carries it lands on fault_storage.cpp’s
already-CONFIRMED, still-FAILED branch - the same re-report-not-a-new-occurrence path
“Repeated failures” above describes for one node dying twice, reached here by two different
nodes sharing one code instead. occurrence_count does not move, the fault manager
publishes EVENT_UPDATED rather than EVENT_CONFIRMED, and just_confirmed in
fault_manager_node.cpp stays false - which is what gates capture_pool_, so neither a
freeze frame nor a recording is captured for the second node. Acknowledging between the two
deaths avoids this: the DELETE route’s ~/clear_fault closes the first occurrence, so the
second node’s next FAILED report reactivates a CLEARED record instead of updating a CONFIRMED
one - a genuine new occurrence, with its own occurrence_count and its own capture.
Two fixes were measured and rejected here, so this is not open for reconsideration without new evidence:
Clear-then-raise from inside the detector does nothing.
AggregatedFaultalready has both halves available -emit()sends a PASSED-shapedclear_fault()whenaffectedis empty, a FAILED-shapedraise_fault()otherwise - so sending one right after the other on the same tick looks like a free fix. It is not:compute_debounce_status()’s hysteresis latch holds a CONFIRMED (or HEALED) status until the debounce counter itself crosses the OPPOSITE threshold, and a single PASSED event only moves that counter by one step. Unless the fault happened to sit exactly one step short ofhealing_thresholdalready, the clear changes nothing the status can see, and the raise that follows lands right back on the same already-CONFIRMED branch as before.Calling the real
~/clear_faultservice instead is worse than doing nothing.SqliteFaultStorage::clear_fault()runsDELETE FROM snapshots WHERE fault_code = ?- it deletes the per-topic readings captured for the fault it clears, keeping only the freeze-frame row. AndClearFault.srv’s ownskip_correlation_auto_cleardefaults to false, so clearing a root cause also clears every symptom the correlation engine attributes to it unless the caller explicitly opts out (the gateway’s own REST DELETE routes do; a detector reaching for this service directly would have to remember to as well). Having the detector call this automatically the moment a second node dies would destroy the FIRST node’s evidence - snapshots and correlated symptoms alike - to chase evidence for the second, before anyone has necessarily seen the first.
A correct fix does not belong in this detector at all: it needs an operation in the fault
manager that re-confirms a CONFIRMED record - bumping occurrence_count, publishing
EVENT_CONFIRMED again, and re-arming just_confirmed for a fresh capture - without
routing through CLEARED and without deleting anything the first occurrence already earned.
Test tiers. Three tiers each prove a different layer, deliberately not overlapping, plus the suppression chain the detector shares no code with any sibling for:
Unit (
test_node_liveness_tracker.cpp): the pure presence/absence state machine - a key present but never armed is never tracked; once armed, presence alone (not continued arming) keeps a tracked key’s miss counter at zero; a key is reported dead only once misses exceedmiss_grace; freshness ordering keeps a brand-new death out of a capped description’s blind spot;prune()reclaims a key only once it has been suppressed for MORE thanprune_ticksCONSECUTIVE calls, resets the streak the moment a veto lifts even once, and never reclaims an unsuppressed death no matter how long it stays dead; the cap leaves every PRESENT key individually tracked and reportable however far the map is past it, collapses only departed entries that have crossedmiss_grace, keeps the collapsed count monotone within one tracker lifetime, and - the property most at odds with a naive read of the sibling detector - reportstracking_saturatedfor a departed set that immature entries alone keep oversized, which unlike the sibling’s saturation refuses nothing and loses no report.test_suppressor.cpppins theSuppressorinterface and the freeapply_suppressors()helper (order-independent, null-safe, returns the dropped count);test_allowlist_suppressor.cppandtest_lifecycle_shutdown_suppressor.cpppin each suppressor’s own matching rules independently of any detector.Integration (
test_node_death_integration.cpp): the full config contract (miss_grace,prune_graceandtick_interval_msrange checks and their floor interaction; malformedallowlist/suppressentries;tracked_node_caprange checks, including a value far pastINT_MAXrejected rather than wrapped; an unknown top-level key), that a peer-aggregated app and anis_online: falseapp are never tracked, that the tracked count stays bounded and shrinks after a reclaim under sustained identity churn, that a clean managed-lifecycle departure is suppressed exactly AT themiss_graceboundary and stays reclaimed - not re-raised - past the retention window the plugin sizes for it, that the allowlist’s id-form suppresses only the colliding node and not its namesake, and the ungated-clear guard’s four properties: a stored fault stays unhealed while an unrelated node arms alone, a clear flows once this process instance has genuinely raised, the guard tracks DELIVERY rather than intent, and a reconfigure while the node is absent does not withhold a clear the process had already earned.E2e, four files, each launching its own real gateway + fault_manager + demo-node stack:
test/e2e/test_node_death_e2e.test.py(ten scenarios covering the raise/clear round trip, no-heal-standalone, the lifecycle-deactivate/manifest-never-online/ros2cli non-cases, a bare-name collision naming only the node that exited, a fast tick alone raising nothing, a three-cycle restart loop reachingoccurrence_count: 3, and a fault surviving a gateway restart);test/e2e/test_node_death_suppression_e2e.test.py(five scenarios covering the allowlist, its inertness when unnamed insuppress, the clean-shutdown/still-active contrast, opt-in suppression, and pruning without a false heal); andtest/e2e/test_node_death_boundary_e2e.test.py(eight scenarios proving the seam withlifecycle_expectationdirectly: a node stuck inactive but never gone raisesGRAPH_NODE_INACTIVEalone, an ARMED node killed belowgraceraisesGRAPH_NODE_DISAPPEAREDalone, a node killed AFTERGRAPH_NODE_INACTIVEhas confirmed raises BOTH at once, a healthy node that is simply killed raisesGRAPH_NODE_DISAPPEAREDalone, a restart-looping required node is caught every cycle regardless of lifecycle maturity, a NEVER-armed node killed belowgraceraisesGRAPH_NODE_INACTIVEalone instead - node_death cannot track a node the gate never admitted for ownership, so absence has to mature it here - a largemiss_gracedelays but does not swallow the report, and a restarted gateway’s warmup window never produces a spurious PASSED); andtest/e2e/test_presence_ownership_e2e.test.py(seven scenarios against the fixture whoseGetStatenever answers, so the unread state is permanent rather than the race the boundary file’s B6 row has to live with: a managed node the watcher asked and could not read, killed, raises BOTHGRAPH_NODE_DISAPPEAREDandGRAPH_NODE_UNREADABLE; the same death with norequire_activeentry anywhere - the shipped default, where nothing else is watching at all - still raisesGRAPH_NODE_DISAPPEARED, and that pair also rules out an implementation deciding ownership fromrequire_activemembership; the SAME node told to start answering, measuredactive, then killed, IS reported too, under the identical configuration as the first row so the only variable is the lifecycle state; a node measuredactivethat then LOSES its lifecycle services and is killed is still reported; and with an unmeasured managed node in the graphGRAPH_PARAM_DRIFTstill names it - the one leg of the three that passes through the app-keyed gate at all - while the topic-keyedGRAPH_QOS_MISMATCHandGRAPH_ORPHANstill report too; and the one row that asserts an ABSENCE, where the same never-answering node is owned provisionally, then announcesinactiveover ~/transition_event, then dies:GRAPH_NODE_INACTIVEnames it andGRAPH_NODE_DISAPPEAREDnever appears, which is also what rejects a detector wired to the permissivereliability_allows(); and a crash loop shorter than the warmup, where the node is killed, reported, brought back for less than one re-warm and killed again - the second death must still be reported, and no PASSED may be emitted while it is dead).
Status
The plugin loads, ticks the graph, and shuts down cleanly. The reliability core is real
and already ticking. Five silent-fault detector classes raise through it today,
qos_mismatch, orphan, param_drift, lifecycle_expectation and
node_death. Two classes remain undelivered, GRAPH_TF_STALE and
GRAPH_LATENCY_BUDGET; they land in follow-up changes, each against its own issue.