Configuring Snapshot Capture
This tutorial shows how to configure snapshot capture to automatically preserve topic data when faults are confirmed, enabling post-mortem debugging.
Overview
When a fault transitions to CONFIRMED status, the system automatically captures data from ROS 2 topics. This snapshot preserves the system state at the moment of fault occurrence, similar to:
AUTOSAR DEM freeze frames - diagnostic data captured at fault detection
SOVD environment data - system context for fault analysis
Snapshots are useful for:
Debugging intermittent faults that are hard to reproduce
Understanding system state when a fault occurred
Post-mortem analysis without real-time access to the robot
Capture works out of the box with zero configuration: when no explicit snapshot config matches a fault code, the faulting entity’s own data is captured by default (see Zero-Config Entity Freeze-Frames). Explicit configuration always overrides the zero-config fallback when present.
Note
Snapshots are automatically deleted when a fault is cleared via the
DELETE /api/v1/faults/{code} endpoint or ~/clear_fault service.
Quick Start
Start the fault manager with snapshot capture enabled:
ros2 run ros2_medkit_fault_manager fault_manager_node --ros-args \ -p snapshots.enabled:=true \ -p snapshots.default_topics:="['/odom', '/battery_state']"
Start the gateway:
ros2 launch ros2_medkit_gateway gateway.launch.py
When a fault is confirmed, query its snapshots:
curl http://localhost:8080/api/v1/faults/MOTOR_OVERHEAT/snapshots
Configuration Options
Configure snapshot capture via fault manager parameters:
Parameter |
Default |
Description |
|---|---|---|
|
|
Enable/disable snapshot capture |
|
|
Topics to capture for all faults (empty entries are ignored) |
|
|
Zero-config fallback: when no explicit config matches a fault code,
capture the reporting source node’s own published topics. Set to
|
|
|
Path to YAML config for fault-specific topics |
|
|
Timeout waiting for topic message |
|
|
Maximum message size in bytes (larger messages skipped) |
|
|
Use background subscriptions (caches latest message) |
|
|
Minimum seconds between snapshot captures for the same fault code. Prevents snapshot storms when a fault is reported repeatedly. Set to 0 to disable. |
|
|
Maximum number of snapshots stored per fault code. When the limit is reached, new snapshots for that fault are rejected. Set to 0 for unlimited. |
|
|
Max concurrent capture threads under a fault storm (>= 1). This parallelizes freeze-frame snapshot capture only; rosbag capture stays single-writer regardless of this value - correlated faults confirming inside one post-roll window share a single recording. |
|
|
Max pending captures before the full-queue policy applies (>= 1). |
|
|
Policy when the queue is full: |
Advanced Configuration
For fault-specific topic capture, create a YAML configuration file:
# snapshots.yaml
fault_specific:
MOTOR_OVERHEAT:
- /joint_states
- /motor/temperature
BATTERY_LOW:
- /battery_state
- /power_management/status
patterns:
"MOTOR_.*":
- /joint_states
- /cmd_vel
"SENSOR_.*":
- /diagnostics
Topic Resolution Priority:
fault_specific- Exact match for fault codepatterns- Regex pattern match (first matching pattern wins)default_topics- Fallback for all faultsEntity-default - zero-config fallback when nothing above matches (
snapshots.entity_default, on by default)
Launch with config file:
ros2 run ros2_medkit_fault_manager fault_manager_node --ros-args \
-p snapshots.enabled:=true \
-p snapshots.config_file:=/path/to/snapshots.yaml \
-p snapshots.default_topics:="['/diagnostics']"
Zero-Config Entity Freeze-Frames
With no snapshot configuration at all, every confirmed fault still carries
at-fault-time context: the faulting entity’s own current data. Two paths
cover the two kinds of entities, both on by default with an opt-out flag,
and explicit config (fault_specific / patterns / default_topics)
always wins when it matches.
ROS-backed entities (snapshots.entity_default, fault manager): when
no explicit config matches the fault code, the fault manager captures the
topics published by the fault’s reporting source node(s) (resolved from
source_id), excluding per-node noise (/rosout,
/parameter_events) and capped at 16 topics. These captures always sample
on demand, even when snapshots.background_capture is enabled, because
the topics are not known until the fault confirms. If the source is not a
live node (e.g. a plugin entity id) or publishes nothing, no freeze-frame
row is written.
Plugin-backed entities (entity_freeze_frame.enabled, gateway): PLC
apps bridged by protocol plugins report faults under their bare SOVD entity
id and their live values are not ROS topics, so the fault manager cannot
capture them. Instead, the gateway snapshots the entity’s data values as
served by the owning plugin (its DataProvider, i.e. the latest polled
values) when the fault confirms, and merges them into the fault detail’s
environment_data.snapshots as a standard freeze_frame entry named
after the entity - unless the fault manager already captured a freeze-frame
for that fault (explicit config wins). When the entity reports its link down
(the loss-of-comms case), the frozen values are the plugin’s last known ones
and may predate the confirmation by the length of the outage; the entry’s
x-medkit block then carries connected: false and, when the plugin’s
payload includes one, source_timestamp (the payload’s own timestamp)
alongside captured_at.
Faults that are already confirmed when the gateway starts are caught up at
startup: the gateway lists the confirmed faults and captures a frame for each
plugin-backed one, so a device standing in fault across a gateway restart
still gets a frame. Catch-up frames carry "capture_origin": "startup" in
their x-medkit block because their values were read at gateway start, not
when the fault confirmed (which may be long before, since the fault manager
persists faults); captured_at always stamps the moment the values were
read. Frames without the marker were captured on the confirm edge. Disable
with:
ros2 run ros2_medkit_gateway gateway_node --ros-args \
-p entity_freeze_frame.enabled:=false
Example plugin-entity freeze-frame in the fault response:
{
"type": "freeze_frame",
"name": "beckhoff_plc_app",
"data": {"tank_level": 87.5, "pump_running": true},
"x-medkit": {
"topic": "",
"message_type": "",
"full_data": {"tank_level": 87.5, "pump_running": true},
"captured_at": "2026-07-14T12:00:00.000Z"
}
}
Querying Snapshots
Snapshots are included inline in the fault response as environment_data:
Get fault details with snapshots:
curl http://localhost:8080/api/v1/apps/motor_controller/faults/MOTOR_OVERHEAT
Response:
{
"item": {
"code": "MOTOR_OVERHEAT",
"fault_name": "Motor temperature exceeded threshold",
"severity": 2,
"status": {
"aggregatedStatus": "active",
"testFailed": "1",
"confirmedDTC": "1"
}
},
"environment_data": {
"extended_data_records": {
"first_occurrence": "2026-02-04T10:30:00.000Z",
"last_occurrence": "2026-02-04T10:35:00.000Z"
},
"snapshots": [
{
"type": "freeze_frame",
"name": "motor_temperature",
"data": 85.5,
"x-medkit": {
"topic": "/motor/temperature",
"message_type": "sensor_msgs/msg/Temperature",
"full_data": {"temperature": 85.5, "variance": 0.1},
"captured_at": "2026-02-04T10:30:00.123Z"
}
},
{
"type": "rosbag",
"name": "fault_recording",
"bulk_data_uri": "/apps/motor_controller/bulk-data/rosbags/550e8400-e29b-41d4-a716-446655440000",
"size_bytes": 1234567,
"duration_sec": 6.0,
"format": "mcap"
}
]
},
"x-medkit": {
"occurrence_count": 3,
"reporting_sources": ["/powertrain/motor_controller"]
}
}
Snapshot Types:
freeze_frame: Data captured at fault confirmation (JSON format). Entity frames caught up for faults that predate the gateway are captured at gateway start instead, markedx-medkit.capture_origin: startup; a plugin entity that reports its link down contributes its last known values, markedconnected: falseinx-medkitrosbag: Recording file available via bulk-data endpoint (binary format)
Get snapshots from fault response using jq:
curl http://localhost:8080/api/v1/apps/motor_controller/faults/MOTOR_OVERHEAT | \
jq '.environment_data.snapshots'
Example Workflow
This example demonstrates the complete snapshot capture workflow.
1. Configure and start the fault manager:
ros2 run ros2_medkit_fault_manager fault_manager_node --ros-args \
-p snapshots.enabled:=true \
-p snapshots.default_topics:="['/odom']"
2. Start a node that publishes odometry:
ros2 topic pub /odom nav_msgs/msg/Odometry \
"{pose: {pose: {position: {x: 1.5, y: 2.0}}}}" -r 10
3. Report a fault (it will be confirmed immediately by default):
ros2 service call /fault_manager/report_fault ros2_medkit_msgs/srv/ReportFault \
"{fault_code: 'NAV_ERROR', event_type: 0, severity: 2, \
description: 'Navigation failed', source_id: '/nav_node'}"
4. Query the captured snapshot:
curl http://localhost:8080/api/v1/apps/nav_node/faults/NAV_ERROR | \
jq '.environment_data.snapshots'
The response will contain the odometry data that was captured at the moment the fault was confirmed.
Troubleshooting
No snapshots captured
Verify
snapshots.enabledistrueCheck that configured topics exist and are publishing
Increase
snapshots.timeout_secfor slow-publishing topicsCheck fault manager logs for capture errors
Empty topics object in response
The fault may have been cleared (snapshots are deleted on clear)
No topics were configured for this fault code
All configured topics timed out or exceeded size limit
Snapshot data truncated
Message exceeded
snapshots.max_message_sizeIncrease the limit or filter to smaller topics
Wrong topics captured
Check topic resolution priority (fault_specific > patterns > default)
Verify regex patterns in config file are correct
Rosbag Capture (Time-Window Recording)
In addition to JSON snapshots, you can enable rosbag capture for “black box” style recording. This continuously buffers messages in memory and flushes them to a bag file when a fault is confirmed.
Key differences from JSON snapshots:
Feature |
JSON Snapshots |
Rosbag Capture |
|---|---|---|
Data format |
JSON (human-readable) |
Binary (native ROS 2) |
Time coverage |
Point-in-time (at confirmation) |
Time window (before + after fault) |
Message fidelity |
Converted to JSON |
Original serialization preserved |
Playback |
N/A |
|
Default |
Enabled |
Disabled |
Enabling Rosbag Capture
ros2 run ros2_medkit_fault_manager fault_manager_node --ros-args \
-p snapshots.rosbag.enabled:=true \
-p snapshots.rosbag.duration_sec:=5.0 \
-p snapshots.rosbag.duration_after_sec:=1.0
This captures 5 seconds of data before the fault and 1 second after.
Rosbag Configuration Options
Parameter |
Default |
Description |
|---|---|---|
|
|
Enable rosbag capture. When enabled, the system continuously buffers messages in memory and writes them to a bag file when faults are confirmed. |
|
|
Ring buffer duration in seconds. This determines how much history is preserved before the fault confirmation. Larger values provide more context but consume more memory. |
|
|
Post-fault recording duration. After a fault is confirmed, recording continues for this many seconds to capture immediate system response. A fault confirming while this window is still running attaches to the in-flight recording and shares its bag, with one metadata entry per fault. At most 32 faults attach on top of the first; any beyond that are recorded in the bag’s data but get no per-fault entry (logged as a warning). |
|
|
Topic selection mode:
Note The default changed from Faults reported by the |
|
|
Explicit list of topics to record (only used when |
|
|
Topics to exclude from recording (applies to all modes). Useful for filtering high-bandwidth topics like camera images. |
|
|
In broad modes ( |
|
|
Subscribe with each topic’s publisher-offered QoS (reliable/transient-local where offered) for faithful capture, instead of forcing best-effort. |
|
|
In-memory ring-buffer cap; oldest buffered messages drop once exceeded, so a broad subscribe set cannot grow memory without bound. |
|
|
Bag storage format: |
|
|
Directory for bag files. Empty string uses system temp directory
( |
|
|
Automatically delete a fault’s bag when it is cleared. A recording
shared by a burst of faults is deleted when the last fault referencing
it clears. Set to |
|
|
Controls when the ring buffer starts recording. See diagram below. |
|
|
Maximum size per bag file in MB. When exceeded, rosbag2 creates additional segment files. |
|
|
Total storage limit for all bag files. Oldest bags are automatically deleted when this limit is exceeded. A recording shared by a burst of faults counts once towards the total, and eviction removes a whole burst’s bag at a time. |
Understanding lazy_start Mode
The lazy_start parameter controls when the ring buffer starts recording:
lazy_start: false (default) - Recording starts immediately at node startup. Best for development and when you need maximum context for any fault.
lazy_start: true - Recording only starts when a fault enters PREFAILED state. Saves resources but may miss context if fault confirms before buffer fills.
When to use lazy_start: true:
Production systems with limited resources
When faults have reliable PREFAILED → CONFIRMED progression
Systems where most faults are debounced (enter PREFAILED first)
When to use lazy_start: false:
Development and debugging
When faults may skip PREFAILED state (severity 3 = CRITICAL)
When maximum fault context is more important than resource usage
Note
The "mcap" format requires rosbag2_storage_mcap to be installed.
"sqlite3" (the default) is always shipped with rosbag2 and needs no extra
package.
You do not have to switch formats manually: if a configured backend’s plugin
is unavailable at startup, the FaultManager logs a warning naming the missing
package and automatically falls back to "sqlite3" for black-box capture.
If no storage backend is usable at all, rosbag capture self-disables (the node
keeps running and freeze-frame snapshots are unaffected) instead of crashing.
To use mcap (e.g. for Foxglove), install the plugin:
# Install MCAP support (optional)
sudo apt install ros-${ROS_DISTRO}-rosbag2-storage-mcap
Downloading Rosbag Files
Rosbag files are downloaded via SOVD bulk-data endpoints.
1. List available rosbags for an entity:
curl http://localhost:8080/api/v1/apps/motor_controller/bulk-data/rosbags
Response:
{
"items": [
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"name": "MOTOR_OVERHEAT recording",
"mimetype": "application/x-mcap",
"size": 1234567,
"creation_date": "2026-02-04T10:30:00.000Z",
"x-medkit": {
"fault_code": "MOTOR_OVERHEAT",
"duration_sec": 6.0,
"format": "mcap"
}
}
]
}
2. Download a specific rosbag:
Use the bulk_data_uri from the fault response, or construct from listing:
# Using bulk_data_uri from fault response
curl -O -J http://localhost:8080/api/v1/apps/motor_controller/bulk-data/rosbags/550e8400-e29b-41d4-a716-446655440000
The -J flag uses the server-provided filename from Content-Disposition header.
3. Play back the rosbag:
ros2 bag play MOTOR_OVERHEAT.mcap
Via ROS 2 service (alternative):
ros2 service call /fault_manager/get_rosbag ros2_medkit_msgs/srv/GetRosbag \
"{fault_code: 'MOTOR_OVERHEAT'}"
Example: Production Configuration
For production use with conservative resource usage:
# config/snapshots.yaml
rosbag:
enabled: true
duration_sec: 3.0
duration_after_sec: 0.5
topics: "config" # Use same topics as JSON snapshots
lazy_start: true # Save resources until fault detected
format: "sqlite3"
max_bag_size_mb: 25
max_total_storage_mb: 200
auto_cleanup: true
# Exclude high-bandwidth topics
# exclude_topics:
# - /camera/image_raw
# - /pointcloud
Example: Debugging Configuration
For development with maximum context:
rosbag:
enabled: true
duration_sec: 10.0 # 10 seconds before fault
duration_after_sec: 2.0 # 2 seconds after
topics: "config"
lazy_start: false # Always recording
format: "sqlite3"
storage_path: "/var/log/ros2_medkit/rosbags"
max_bag_size_mb: 100
max_total_storage_mb: 1000
auto_cleanup: false # Keep bags for analysis
See Also
REST API Reference - REST API reference (Bulk Data section)
Faults - Fault API requirements
Gateway README - REST API reference
config/snapshots.yaml - Full configuration reference
Migration from Legacy Endpoints
If you were using the legacy snapshot endpoints, migrate to the new SOVD-compliant API:
Snapshots:
Previous (removed) |
Current |
|---|---|
|
|
|
|
Rosbag Downloads:
Previous (removed) |
Current |
|---|---|
|
|
|
|
Key Changes:
Snapshots inline: No separate snapshot endpoint; data is in fault response
Bulk-data pattern: Rosbags use SOVD bulk-data with UUID identifiers
Entity-scoped: Bulk-data endpoints require entity path (e.g.,
/apps/motor)SOVD status: Fault response includes SOVD-compliant
statusobject