Changelog
This page aggregates changelogs from all ros2_medkit packages.
Changelog for package ros2_medkit_gateway
0.7.0 (2026-08-27)
Rosbag bulk-data is addressed by recording id instead of fault code, so a fault holding several recordings can expose each one.
GET /{entity}/bulk-data/rosbagsnow emits one descriptor per recording rather than one per fault - a burst that shares a bag used to appear as several entries each reporting the full bag size - and the covered faults move intox-medkit.fault_codes(was the scalarx-medkit.fault_code). Old URLs keep working: an id that is not a recording is resolved as a fault code and serves that fault’s newest recording, which is what it returned before. Authorization is unchanged in effect - a download is allowed when any fault the recording covers is in the entity’s source scope, which is exactly the set that could reach it previously (#623, #620)Manual asset inventory: a manifest
assets:list and a newdiscovery.inventory.csv_pathparameter declare assets that no protocol layer can describe (or fully describe). Both paths recognize the canonical namesid, manufacturer, model, serial, hardware_rev, firmware, endpoint, role, areaplus the shared aliases (serial_number,hardware_revision/hw_rev,firmware_version/fw) and keep any other column / key as an extra; RFC-4180-style quoting is honored. Each asset becomes a Component withsource = "inventory"and a structured asset identity carrying per-field provenance, appended to the base manifest on every load / reload and merged into the tree by id alongside protocol-discovered structure;areaplaces the asset under an Area, without it the asset appears only in the flat component list. CSV rows never fail the load: rows without anidare skipped with a warning, for duplicate ids the first row wins, a row whose id is already a manifest component keeps the manifest definition (the row’s identity is folded in as gap-fill), and an unknownareais dropped with a warning. The CSV is size-capped at 1 MiB before being read; a missing file is skipped with a warning (mirrorsfragments_dir), while an unreadable or malformed one fails the load / reload. Requires a manifest-backed discovery mode (manifest_only/hybridwithdiscovery.manifest_pathset); empty = disabled (default) (#493, #490)Breaking: the lifecycle status
operationIdvalues were singularized -getAppStatusandputAppStatusRestartrather than the plural collection forms they were built from before - so a generated client gets renamed methods for those operations (#497)Aggregation now separates the time budget for reading metadata from the budget for real work. One
aggregation.timeout_mswas applied as both the connect and the read timeout for every call, so a synchronous service call on a peer got two seconds end to end while the peer’s own budget for the same call was ten; a large resource fanned out to a peer could not finish inside it either. The write timeout, which was never set and stayed at the cpp-httplib default, now follows the configured budget, and the timeout values are validated and reported rather than silently clamped (#638, #528)A request that ran out of time is reported as a timeout. A peer that did not answer in the budget returns
504withERR_NOT_RESPONDINGinstead of502claiming the peer is unavailable - which it did while that peer was answering the same request - a fanned-out collection carries the per-peer failure reason instead of only a boolean, and an operation that exceeds its own service-call budget says so rather than returning a generic failure (#638, #612)GET /faults/streamon an aggregating gateway relays its peers’ fault events. It previously returned200with an open stream that only ever sent keepalive comments, which is indistinguishable from a healthy system - and on a deployment where the aggregator is the only reachable port, it was the only fault stream available (#638, #611)One addressing model for an aggregating entity’s resources, so the aggregator no longer refuses or 404s work its peers can serve (#626, #613)
Nested
plugins.<name>.*parameters are rebuilt into a nested object instead of a flat dotted key, so nested plugin configuration reaches the plugin again (#518, #520)Discovery configuration is read from the documented top-level
config:key. It was only ever read fromdiscovery.config, so the documented form was dropped without a word and unmanifested nodes leaked into the tree in hybrid mode (#609, #529)Every startup parameter that is coerced or refused is reported. A clamped thread count or keep-alive timeout used to change the value and log nothing, leaving the configuration file and the running process in silent disagreement. Integer parameters are read as the int64 a ROS parameter holds and validated before narrowing, so a value past
INT_MAXcan no longer wrap back into the legal band and pass its own range check, and range checks are written so that NaN is refused rather than accepted (#607, #603)server.executor_threadsis real rather than advisory, and a cancel that runs out of time is reported as a timeout (#593)An unresponsive parameter node no longer hangs the REST API (#532), per-node parameter caches are bounded with LRU eviction (#534), the transport’s node is released before the context dies (#567), and the parameter error surface is consistent between list and get, with a non-404 error winning over NOT_FOUND across nodes (#540, #543)
Faults from plugin-provided entities are visible in the fault list, in fault detail and in freeze-frame through fault-scope ownership (#503), an external Component owns its fault-manager faults (#530), plugin-provided entities own their bulk data and logs (#560), and the external flag survives the hybrid merge (#522)
Zero-config freeze-frame for plugin-backed entities: when a fault confirms with a plugin-owned reporting source, that entity’s current data values are snapshotted with no configuration. Plugin entities report under their bare SOVD entity id and their values are not ROS topics, so the fault manager’s own snapshot capture could never reach them. Gated on at least one plugin being loaded (#538), a fault already raised at gateway startup gets one (#563), and so does an entity serving last known values (#565)
Data triggers work on plugin-provided entities, resolving from declared topics, and fail loudly on a topic name that cannot be resolved instead of silently never firing (#592). Trigger subscriptions go through the shared subscription executor (#549)
Asset identity model with per-field provenance and merge-by-identity (#488)
Build, image and test: an arm64 multi-arch image gated on tag or dispatch (#508), the fault bridges and
ros2_medkit_fault_detectionbundled into the image (#470, #494), coverage instrumentation for every C++ package (#582), and clang-tidy analysing a package’s translation units in parallel, cutting PR feedback time from about 56 minutes to about 35 (#588, #590)Config-less fault triggers: a threshold rule declared at runtime raises a fault when a data value crosses it and clears the fault when the value comes back. Rules are managed on Apps over
GETandPOST/apps/{app_id}/fault-triggersandDELETE/apps/{app_id}/fault-triggers/{trigger_id}, and survive a restart whenfault_triggers.storage.pathnames a database. The engine needs at least one loaded plugin, sofault_triggers.enabled(default true) is necessary rather than sufficient;fault_triggers.poll_interval_ms(default 1000) is floored at 50 ms and a lower value is reported. This is a separate facility from the SOVD notification/triggers(#544)An aggregating gateway can authenticate to its peers.
aggregation.peer_auth_headercarries the credential this gateway presents on connections it opens on its own behalf - the peer health check, the entity fetch, the fault-stream relay, and any forward whose caller sent no credential to pass on. Whereforward_authis enabled and the caller did send one, that token wins, so the peer keeps being told the end user. The header is empty by default, and is deliberately withheld from peers found by mDNS discovery rather than configured explicitly. The value is redacted from the configurations API alongsideauth.jwt_secretandauth.clients, so reading configuration back cannot disclose it (#638)A secure-by-default field profile ships as
gateway_params.secure.yaml, together with a hardening checklist. It turns on JWT authentication, TLS, restricted CORS and rate limiting, so it needs certificates and credentials provisioned before use and is not a drop-in replacement for the default profile (#485)The fault SSE stream carries
auto_cleared_codes. A consumer can now see which correlated symptom faults were cleared along with their root cause, which previously happened without any event of their own (#573)Breaking: an action execution now belongs to the entity it was started on. Reading, stopping or cancelling an execution through a different entity returns
404instead of being served, and an execution listing is filtered to the executions the addressed entity owns. An execution id used to work as a global handle (#591)Breaking:
GET <entity-path>/docsis readable byviewerrather thanadmin. The permission table is now derived from the route registrations instead of a hand-maintained literal, and this route’s derived pattern places it with the other read paths. Only reachable whereauth.enabledis true (#591)The OpenAPI document is derived from the code that serves the routes rather than declared beside it: the success status and its schema come from the handler’s return type, the
Locationheader follows from that status, a feature gate declares the501it returns, the lock contract and the RBAC permission table are read out of the registrations, and each<entity-path>/docssub-document is a projection of the paths that path actually serves. The rule and what is deliberately still declared by hand are written down indesign/openapi_derivation.rst(#591, #583)A build with
BUILD_TESTING=ONrecords every status the gateway puts on the wire, and an integration sweep driven from the served document asserts that each one is declared. The sweep refuses to pass vacuously: operations it could not reach must match a declared list, and it requires a minimum spread of error sites. The recorder is compiled out of the published image (#591)Statuses a caller receives, corrected: a configuration value that cannot be converted to the parameter’s ROS type answers
400rather than500; a fault manager that refuses a fault lookup or clear answers404, with503reserved for a transport that gave no answer at all; aconfig_idpast the published bound answers400onDELETEas it already did on the other verbs; and an operation-execution listing resolves on Areas and Functions instead of reporting404(#591)POSTon triggers, cyclic subscriptions and fault triggers returns a usableLocation, and everyLocationis canonicalised - a request with a trailing slash used to yield a URI that answered404(#591)Entity detail responses advertise
data-categoriesanddata-groupson every entity type,lockson components and apps where locking is configured, andfault-triggerson apps; thecapabilitiesarray now lists the collections each entity type actually serves. Two operation ids are added,getCapabilityDescriptionandgetScopedCapabilityDescription; no existing operation id was renamed and no route was added to or removed from the served API (#591)The root endpoint list includes plugin-mounted routes, and an SSE cyclic-subscription error frame carries
vendor_codealongside itserror_code(#591)In-flight update tasks are drained before their notifier is destroyed, closing a shutdown race (#591)
Contributors: @bburda, @mfaferek93, @YueBit
0.6.0 (2026-06-22)
SOVD entity status and lifecycle control endpoints:
GET /apps/{id}/statusandGET /components/{id}/status, plus lifecycle control routes backed by a newLifecycleProviderplugin interface and plugin-manager routing. Control returns501 Not Implementeduntil a provider is registered; the routes are RBAC-gated, advertised via astatuslink on app and component detail, and declared under the OpenAPILifecycletag (#437)Accurate app and component status: app status is read from the managed-node lifecycle state through a
GetState-backed reader, and a component reportsnotReadywhen all hosted apps are offline while stayingreadyas long as it is reachable (#455)Bounded the executor and HTTP server thread pools, sized to the cold-wait plus SSE budget, with misconfiguration guards and a bounded keep-alive timeout (#457)
Startup discovery summary logged at boot, with an empty-graph warning when no entities are discovered (#438)
Bounded the unsupported-message-type cache and exposed its size; unknown message types now warn once instead of on every sample (#450)
OpenAPI query parameters are derived from a typed query contract, tightening the query-parameter schema and its regression gate (#417)
PluginContextcan aggregate peer faults across daisy-chained gateways through the SOVD service interface (#419)Single-command bringup of the local medkit stack via
bringup.launch.pyandbringup_params.yaml(#439)Docker image enables CORS for the documented web UI path and uses explicit web UI origins instead of wildcard CORS (#452)
Docker image bundles the CycloneDDS RMW as an opt-in alternative implementation (#451)
Native gateway launch enables web UI CORS by default and honors the CORS settings from a supplied
config_file(#461)Incremental, embedded-hardened entity cache:
ThreadSafeEntityCachestores entities in a fixed-capacitySlotStoreobject pool indexed by open-addressed flat hash maps, andupdate_allreconciles the discovery output by id (add / remove / change only) so steady-state refresh does zero structural allocations. Capacity is reserved viaentity_cache.capacity(default 256) and cache stats are exposed on/healthasx-medkit-entity-cache. Discovery refreshes are debounced viadiscovery.refresh_debounce_ms(default 1000) and operationtype_infoschemas resolve lazily, cutting gateway CPU under graph churn roughly 4x. Entity detailoperations[]no longer embeds schemas eagerly; the schemas remain on the/operationsresource (#462)Contributors: @bburda, @mfaferek93
0.5.0 (2026-06-08)
Breaking Changes:
Typed router refactor.
HandlerContextno longer carriessend_json/send_error/send_plugin_error/send_dto/parse_body: handlers returnhttp::Result<TResponse>and the framework owns response writing throughRouteRegistry. The rawvoid(httplib::Request, httplib::Response)RouteRegistrylambda overloads are removed - call sites must use the typedreg.get<T>/reg.post<TBody, T>/reg.del<T>overloads, the multi-shapereg.post_alternates<TBody, TAlt...>/reg.del_alternates<TAlt...>, or one of the named escape hatches (reg.sse/reg.binary_download/reg.multipart_upload<T>/reg.static_asset/reg.docs_endpoint/reg.docs_subtree).static_assert(dto::has_dto_shape_v<T>)gates every typed overload, so non-DTO return types fail at compile time. The plugin ABI is unaffected:PluginResponsekeeps itssend_json/send_errorsurface and now routes through the same internalhttp::detail::write_json_bodyprimitive as the framework, so plugin wire format is unchanged (#403)Provider ABI typed.
FaultProvider,DataProvider,OperationProvider, andUpdateProvider::get_updatereturn typed DTO envelopes (FaultListResult/FaultDetailResult/FaultClearResult/ the matchingData*ResultandOperation*Resultshapes /UpdateStatusResult) instead of rawtl::expected<nlohmann::json, ErrorInfo>. The wire bytes are byte-identical because each envelope wraps an opaquecontentobject emitted verbatim byJsonWriter; commercial and out-of-tree plugins must wrap their existing JSON in the matching envelope type (mechanical:Result.content = std::move(json_payload)). The plugin ABI itself (PluginRouteshape,PluginResponsector, plugin api version) is locked bytest_plugin_abi_conformanceand is unchanged (#403)SchemaWriteremits optional DTO fields asanyOf: [<inner>, {type: "null"}](the OpenAPI 3.1 idiom) instead ofnullable: true. Generated clients seeT | nullfor every optional field rather thanT | undefined. Wire format is unchanged - the gateway still omits absent optional fields, andJsonReadercontinues to accept absent fields; the schema change only opts the published spec into round-tripping a literalnullvalue cleanly for clients that prefer to send one. As part of this, a handful of fields that were previously emitted as an explicit JSONnullwhen absent are now omitted entirely (consistent with the optional-omission policy): the script execution fieldsprogress/started_at/completed_at/parameters/error(GET .../scripts/{id}/executions/{eid}) and the scriptparameters_schemafield (GET .../scripts/{id}). Clients that testedfield === nullor relied on the key always being present must treat an absent key the same asnull(#403)Synchronous operation-execution service-call failures (
POST /api/v1/{entity-path}/operations/{id}/executionswhen the underlying ROS 2 service call fails) now return the standard SOVDGenericErrorenvelope ({"error_code": "vendor-error", "vendor_code": "x-medkit-ros2-service-unavailable", "message": "Service call failed", ...}, HTTP status 500 unchanged) instead of the previous bespoke nested{"error": {"code", "message", "details"}}object. This aligns the one remaining non-standard error path with every other gateway error; clients that parsederror.code/error.detailsfor this specific failure must readvendor_code/parametersinstead (#403)GET /api/v1/{entity-path}/datanow publishes the opaqueDataListResultschema ({type: object, additionalProperties: true, x-medkit-opaque: true}), matching howGET .../faultsalready publishesFaultListResult. The wire payload is unchanged for runtime (ROS 2) entities - it is still{"items": [...], "x-medkit": {...}}built from the typedCollection<DataItem, DataListXMedkit>- but for plugin-owned entities the provider’s free-form per-item shape now passes through verbatim instead of being re-parsed intoCollection<DataItem>. This fixes a regression in which plugin per-item fields (the OPC-UA plugin’svalue/unit/data_type/writable) were silently dropped by the typed re-parse. Clients that generated a typedDataItemmodel from the previous spec for this route now see an opaque object instead (#403)Entity responses (areas, components, apps, functions - list items and detail) now always carry a top-level
typediscriminator (an enum ofarea/component/app/function). Previously list items had notypeand detail responses exposed it only insidex-medkit.entityType. Additive and tolerant-client-safe; consumers keying on the entity kind can now read the top-leveltype(#403)ros2_medkit_msgs/srv/ClearFaultrequest gains abool skip_correlation_auto_clearfield (see the per-entity fault scope entry below for the in-tree motivation). Adding a request field changes the service type hash, so out-of-tree callers that invoke the service directly (for exampleros2 service call /fault_manager/clear_fault ros2_medkit_msgs/srv/ClearFault ...as documented in theros2_medkit_fault_managerREADME) must rebuild against the newros2_medkit_msgsto keep talking tofault_manager. The in-tree gateway client and server are updated together (#395)Per-entity fault routes are now correctly scoped to the entity’s hosted apps.
GET /api/v1/{entity-path}/faults/{fault_code},DELETE /api/v1/{entity-path}/faults/{fault_code},GET /api/v1/{entity-path}/faults, andDELETE /api/v1/{entity-path}/faultspreviously fell back to a prefix match against the entity’snamespace_path; when that was empty (host-derived / synthetic components, manifest components without anamespacefield, Areas, Functions, and Apps with a wildcardros_binding.namespace_pattern) the scope filter was silently disabled and the routes exposed - and onDELETE, cleared - faults reported by apps that belonged to entirely different entities. All four handlers now resolve the addressed entity to its hosted-app FQN set (via the newHandlerContext::resolve_entity_source_fqnshelper) and apply a strict all-sources scope check: a fault counts as in scope only when every entry in itsreporting_sourcesis owned by the entity (exact FQN match, or strict path-child via<fqn>/...). Per-fault routes return404 Resource Not Foundfor any fault that fails the check; collection routes return an emptyitemsarray. The underlyingGetFault.srvcontract is unchanged;ClearFault.srvgains a newskip_correlation_auto_clearrequest flag so per-entity DELETE can opt out of cascade-clearing correlated symptom fault codes that may live in other entities. Per-entity collection responses no longer include the globalmuted_count/cluster_count/muted_faults/clusterscorrelation metadata; those remain on the globalGET /api/v1/faultsroute. Behavior changes visible to clients: (a) faults reported by apps outside the addressed entity are no longer returned or cleared via that entity’s route, (b) mixed-source faults that include at least one out-of-entity reporter are likewise rejected with404on per-fault routes and excluded from per-entity collection responses (use the globalGET /api/v1/faultsto see them), (c) per-entity DELETE no longer cascade-clears correlated symptoms outside the entity (#395)GET /api/v1/updates/{id}/statusno longer returns404for a registered-but-idle package;POST /api/v1/updatesnow seeds apendingstatus, so the endpoint returns200 {"status": "pending"}immediately after registration.404is reserved for packages that are not registered. Clients that used404as a signal for “registered but nothing started yet” must adapt (#378)
Features:
Typed
fan_out_collection<T>aggregating helper replaces raw-JSONmerge_peer_itemson the typed collection routes (data, operations, config, logs). Peer items are decoded viadto::JsonReader<T>; items that fail validation are removed from the mergeditemsarray, recorded inx-medkit.peer_dropped_itemswith the JsonReader error plus a best-effortsource_id, and logged atWARN. Items that parse successfully are re-serialized through the localdto::JsonWriter<T>, so any peer-supplied fields outside the local DTO schema are dropped from the merged response (the previous raw passthrough preserved unknown peer fields verbatim). Previously, malformed peer items silently disappeared into the merged response; fleet operators can now detect inter-gateway schema drift directly on the wire (#403)Collection<T, XMedkitT>is now a 2-parameter template. Domain list endpoints (faults, config, logs) reference their richer per-domain collection x-medkit struct (FaultListXMedkit,ConfigListXMedkit,LogListXMedkit) directly in the published schema instead of the genericXMedkitCollection, so generated clients see aggregation counts, peer provenance, andpeer_dropped_itemsfrom the schema. The data list builds the same typedCollection<DataItem, DataListXMedkit>internally (so the wire still carries those fields) but publishes the opaqueDataListResultenvelope, because plugin-owned data entities can return vendor per-item fields the typed item schema cannot describe (see the data-list breaking-change entry above) (#403)New
opaque_object("key", &T::field)DTO field descriptor indto/contract.hpp. Binds anlohmann::jsonmember as a typed “any JSON object” field:JsonWriteremits it verbatim,JsonReaderrejects scalars / arrays / null,SchemaWriteremits{type: object, additionalProperties: true, x-medkit-opaque: true}. Used for fields whose runtime shape is decided by an upstream component the gateway cannot introspect (live ROS message payloads, plugin-defined fault envelopes, action results) (#403)GET /api/v1/faults/streamevent payloads now carry an optionalx-medkitSOVD payload-extension object withentity_typeandentity_idfields. When the gateway can resolve the fault’s first reporting source back to a SOVD entity (via the manifest-mode linking index, or a runtime-mode last-segment match against an existing App), consumers can hit/{entity_type}/{entity_id}/bulk-data/rosbags/{fault_code}directly instead of HEAD-probing every entity. Resolution is snapshotted at event arrival, so a discovery refresh between enqueue and stream-out cannot retroactively change the entity reported to consumers. Thex-medkitobject is omitted entirely when no entity can be resolved, so existing SSE consumers ignore the addition (#380)Plugin API version bumped to v7. Adds
PluginContext::notify_entities_changed(EntityChangeScope)lifecycle hook for plugins that mutate the entity surface at runtime; default no-op keeps v6 source code compiling unchanged against v7 headers. Binary compatibility is not provided: the plugin loader uses a strict equality check onplugin_api_version(), so out-of-tree plugins must be recompiled (#376)New
discovery.manifest.fragments_dirparameter: gateway scans the directory for*.yaml/*.ymlfragment files on every manifest load / reload and merges apps, components, and functions on top of the base manifest. Fragments are forbidden from declaring top-levelareas,metadata,discovery,scripts,capabilities, orlock_overrides- those stay in the base manifest. Presence of any forbidden key (including empty-valued ones likeareas: []) is reported as aFRAGMENT_FORBIDDEN_FIELDvalidation error that fails the load / reload. Unknown top-level keys (typos such asapp:vsapps:) are ignored with a warning log. Files merged in alphabetical order for deterministic duplicate-id errors (#376)Fragment files are size-capped at 1 MiB (
ManifestParser::kMaxFragmentBytes) before being read into memory, and any symlink resolving outside the canonicalfragments_diris skipped with a warning, so misconfigurations or symlink-based escapes cannot hand arbitrary bytes to the YAML parser (#376)All-or-nothing fragment semantics: a single malformed or forbidden fragment fails the entire load / reload and keeps the previously-loaded manifest active (#376)
ManifestParser::parse_fragment_fileconvenience entrypoint that injects a syntheticmanifest_versionheader when the fragment omits oneSee
design/plugin_entity_notifications.rstfor the lifecycle, merge-rule, and plugin-side write-contract walkthroughNew
GET /api/v1/apps/{app_id}/belongs-todiscovery endpoint returning the areas and components an app belongs to; thebelongs-toURI is advertised onGET /apps/{app_id}(#196)Pool-backed
TopicDataProviderfor live topic data: a shared subscription pool owned by a single-writer executor node, with LRU and idle eviction and publisher-QoS matching, replacing per-request subscriptions. Pool and executor health are surfaced as thex-medkit-subscription-executorvendor-extension stats onGET /api/v1/health, read atomically so/healthnever blocks under load (#384)GET /api/v1/updates/{id}/statusexposes the update lifecyclephaseunder the responsex-medkitobjectgateway.launch.pyandgateway_https.launch.pyaccept aconfig_filelaunch argument pointing at an external parameter YAML. Parameters present in the file override the matching gateway defaults; parameters the file omits keep their launch defaults instead of being reset (#408)Plugin-facing headers are httplib-free across the
.soboundary: the handler-result vocabulary (Result,NoContent,Forwarded,ValidatorResult,ResponseAttachments) moved to a new leaf headerhttp/handler_result.hppso provider and DTO interfaces no longer transitively include<httplib.h>. Out-of-tree plugins built against the installed gateway (build-farm / Docker topology, where the vendored httplib is not on the include path) compile again; no ABI, wire, or behaviour change. A pre-push gate and CI scan keep the plugin-facing headers httplib-free (#411)Contributors: @bburda, @eclipse0922, @evTessellate, @mfaferek93
0.4.0 (2026-03-20)
Breaking Changes:
GET /version-inforesponse key renamed fromsovd_infotoitemsfor SOVD alignment (#258)GET /root endpoint restructured:endpointsis now a flat string array, addedcapabilitiesobject,api_basefield, andname/versiontop-level fields (#258)Default rosbag storage format changed from
sqlite3tomcap(#258)Plugin API version bumped to v4 - added
ScriptProvider, locking API, and extendedPluginContextwith entity snapshot, fault listing, and sampler registrationGraphProviderPluginextracted to separateros2_medkit_graph_providerpackage
Features:
Discovery & Merge Pipeline:
Layered merge pipeline for hybrid discovery with per-layer, per-field-group merge policies (#258)
Gap-fill configuration: control heuristic entity creation with
allow_heuristic_*options and namespace filtering (#258)Plugin layer:
IntrospectionProvidernow wired into discovery pipeline viaPluginLayer(#258)/healthendpoint includes merge pipeline diagnostics (layers, conflicts, gap-fill stats) (#258)Entity detail responses now include
logs,bulk-data,cyclic-subscriptionsURIs (#258)Entity capabilities fix: areas and functions now report correct resource collections (#258)
discovery.manifest.enabled/discovery.runtime.enabledparameters for hybrid modeNewEntities.functions- plugins can now produce Function entitiesGET /apps/{id}/is-located-onendpoint for reverse host lookup (app to component)Beacon discovery plugin system - push-based entity enrichment via ROS 2 topic
x-medkit-topic-beaconandx-medkit-param-beaconvendor extension REST endpointsLinux introspection plugins: procfs, systemd, and container plugins via
x-medkit-*vendor endpoints (#263)
Locking:
SOVD-compliant resource locking: acquire, release, extend with session tracking and expiration
Lock enforcement on all mutating handlers (PUT, POST, DELETE)
Per-entity lock configuration via manifest YAML with
required_scopesLock API exposed to plugins via
PluginContextAutomatic cyclic subscription cleanup on lock expiry
LOCKScapability in entity descriptions
Scripts:
SOVD script execution endpoints: CRUD for scripts and executions with subprocess execution
ScriptProviderplugin interface for custom script backendsDefaultScriptProviderwith manifest + filesystem CRUD, argument passing, and timeoutManifest-defined scripts:
ManifestParserpopulatesScriptsConfig.entriesfrom manifest YAMLallow_uploadsconfig toggle for hardened deploymentsRBAC integration for script operations
Logging:
LogProviderplugin interface for custom log backends (#258)LogManagerwith/rosoutring buffer and plugin delegation/logsand/logs/configurationendpointsLOGScapability in discovery responsesConfigurable log buffer size via parameters
Area and function log endpoints with namespace aggregation (#258)
Triggers:
Condition-based triggers with CRUD endpoints, SSE event streaming, and hierarchy matching
TriggerManagerwithConditionEvaluatorinterface and 4 built-in evaluators (OnChange, OnChangeTo, EnterRange, LeaveRange)ResourceChangeNotifierfor async dispatch from FaultManager, UpdateManager, and OperationManagerTriggerTopicSubscriberfor data trigger ROS 2 topic subscriptionsPersistent trigger storage via SQLite with restore-on-restart support
TriggerTransportProviderplugin interface for custom trigger delivery
OpenAPI & Documentation:
RouteRegistryas single source of truth for routes and OpenAPI metadataOpenApiSpecBuilderfor full OpenAPI 3.1.0 document assembly withSchemaBuilderandPathBuilderCompile-time Swagger UI embedding (
ENABLE_SWAGGER_UI)Named component schemas with
$ref, cleanoperationIdvalues, endpoint descriptions,GenericErrorschema refs,info.contact, Spectral-clean output, multipart upload schemas, static spec cachingSOVD compliance documentation with resource collection support matrix (#258)
Other:
Multi-collection cyclic subscription support (data, faults, logs, configurations, update-status)
Generation-based caching for capability responses via
CapabilityGeneratorPluginContext::get_child_apps()for Component-level aggregationSub-resource RBAC patterns for all collections
Auto-populate gateway version from
package.xmlvia CMakeNamespaced fault manager integration -
FaultManagerPathsresolves service/topic names for custom namespacesGrouped
fault_manager.*parameter namespace for cleaner configuration
Build:
Extracted shared cmake modules into
ros2_medkit_cmakepackage (#294)Auto-detect ccache for faster incremental rebuilds
Precompiled headers for gateway package
Centralized clang-tidy configuration (opt-in locally, mandatory in CI)
Tests:
Unit tests for DiscoveryHandlers, OperationHandlers, ScriptHandlers, LockHandlers, LockManager, ScriptManager, DefaultScriptProvider
Comprehensive integration tests for locking, scripts, graph provider plugin, beacon plugins, OpenAPI/docs, logging, namespaced fault manager
Contributors: @bburda
0.3.0 (2026-02-27)
Features:
Gateway plugin framework with dynamic C++ plugin loading (#237)
Software updates plugin with 8 SOVD-compliant endpoints (#237, #231)
SSE-based periodic data subscriptions for real-time streaming without polling (#223)
Global
DELETE /api/v1/faultsendpoint (#228)Return HEALED/PREPASSED faults via status filter (#218)
Bulk data upload and delete endpoints (#216)
Token-bucket rate limiting middleware, configurable per-endpoint (#220)
Reduce lock contention in ConfigurationManager (#194)
Cache component topic map to avoid per-request graph rebuild (#212)
Require cpp-httplib >= 0.14 in pkg-config check (#230)
Add missing
ament_index_cppdependency topackage.xml(#191)Unit tests for HealthHandlers, DataHandlers, and AuthHandlers (#232, #234, #233)
Standardize include guards to
#pragma once(#192)Use
foreachloop for CMake coverage flags (#193)Migrate
ament_target_dependenciesto compat shim for Rolling (#242)Multi-distro CI support for ROS 2 Humble, Jazzy, and Rolling (#219, #242)
Contributors: @bburda, @eclipse0922, @mfaferek93
0.2.0 (2026-02-07)
Initial rosdistro release
HTTP REST gateway for ros2_medkit diagnostics system
SOVD-compatible entity discovery with four entity types:
Areas, Components, Apps, Functions
HATEOAS links and capabilities in all responses
Relationship endpoints (subareas, subcomponents, related-apps, hosts)
Three discovery modes:
Runtime-only: automatic ROS 2 graph introspection
Manifest-only: YAML manifest with validation (11 rules)
Hybrid: manifest as source of truth + runtime linking
REST API endpoints:
Fault management: GET/POST/DELETE /api/v1/faults
Data access: topic sampling via GenericSubscription
Operations: service calls and action goals via GenericClient
Configuration: parameter get/set via ROS 2 parameter API
Snapshots: GET /api/v1/faults/{code}/snapshots
Rosbag: GET /api/v1/faults/{code}/snapshots/bag
Server-Sent Events (SSE) at /api/v1/faults/stream:
Multi-client support with thread-safe event queue
Keepalive, Last-Event-ID reconnection, configurable max_clients
JWT-based authentication with configurable policies
HTTPS/TLS support via OpenSSL and cpp-httplib
Native C++ ROS 2 serialization via ros2_medkit_serialization (no CLI dependencies)
Contributors: Bartosz Burda, Michal Faferek
Changelog for package ros2_medkit_fault_manager
0.7.0 (2026-08-27)
Rosbag black-box recordings are no longer limited to one per fault code. A fault that re-confirms keeps a bounded history of recordings instead of overwriting the previous one, controlled by the new
snapshots.rosbag.max_bags_per_fault(default1, which reproduces the previous behaviour exactly;0= unlimited). Retention is keep-newest and the bag is unlinked only when no fault still references it, so a burst that shares one recording behaves as before. Internally therosbag_filesgrain changed from “one row per fault” to “one row per (fault, recording) link”:recording_idis now a stored, indexed column, and the legacy column-levelUNIQUE(fault_code)is replaced by aUNIQUE INDEXon(fault_code, file_path)through an automatic, idempotent table rebuild on first open. Four latent defects are fixed on the way: quota eviction deleted by fault code rather than by recording,get_rosbag_filehad noORDER BYand would have served an arbitrary recording, the stale-row self-heals deleted a fault’s entire history because one bag had vanished from disk, and bothdelete_rosbag_file/delete_rosbag_filesread only the firstfile_pathof a fault, so deleting a fault with several recordings removed every row but left all but one bag on disk - unreachable and still charged against the quota (#623, #620)Optional append-only, hash-chained audit log of fault state transitions: each transition appends one immutable row (
record_hash = sha256(prev_hash + canonical(event))via OpenSSL EVP SHA-256) with a persisted chain head, averifyroutine, a read API, and retention that seals a segment anchor before pruning. Time-based (PREFAILED->CONFIRMED) auto-confirmations are also audited.verifyreads the chain head directly from the database, so deleting the newest row together with the head row is reported as tampering instead of silently recovering.BEFORE UPDATE/BEFORE DELETEtriggers reject out-of-band edits as defense-in-depth. The chain is unkeyed and stored in a single writable file, soverifydetects edits/deletions that did not recompute the chain (casual or accidental tampering); it is not a defence against an attacker who can rewrite the whole file. Off by default (#487, #483)Breaking: the default rosbag storage format is
mcapagain.snapshots.rosbag.formatnow defaults to"mcap", so black-box recordings land as.mcapfiles instead of.db3, and the filename a bulk-data download serves changes with them. The storage plugins are declared explicitly and the plugin loader is serialised, which is what makes the format selectable reliably rather than dependent on load order. Setsnapshots.rosbag.format: sqlite3to keep the previous on-disk format (#610)Freeze-frame: a compact JSON snapshot of the entity’s data is persisted when a fault is confirmed, so the state at the moment of confirmation survives the fault being cleared (#491)
A fault that confirms inside an active post-roll window keeps its black-box recording instead of finding the buffer already finalised (#561), and a fault landing on a recording-window boundary gets its own bag rather than none (#594)
The debounce counter is clamped and the confirmed / healed status is latched, so a counter cannot run past its threshold and a status cannot silently regress (#484)
A PASSED event no longer re-dates a fault -
first_occurredkeeps marking the start of the current occurrence - and genuine SSE loss is counted rather than absorbed (#573)The near-miss series survives a fault being cleared, so acknowledging a fault no longer discards the evidence gathered around it (#629)
The fault manager’s YAML parameter file supports launch substitutions, so a path can be composed at launch time instead of being fixed in the file (#634)
Build and test only: the package is instrumented for coverage (#582), every launch test lives under
test/integrationand takes a DDS domain at run time (#628, #551, #597)A
fault_codeup to the advertised 256 characters is accepted, where the manager previously stopped at 128, and a long code no longer costs the recording: the rosbag filename is truncated with a digest instead of exceeding the filesystem’s component limit (#591)Contributors: @bburda, @mfaferek93, @nnarain
0.6.0 (2026-06-22)
Bounded concurrent snapshot capture under fault storms with a
CaptureThreadPooland configurable capture pool / queue / overflow-policy parameters. The rosbag leg is serialized and the cooldown map is bounded, so a burst of simultaneous faults can no longer exhaust capture threads or grow memory without limit (#456)Entity-scoped rosbag capture by default (#431)
Made rosbag capture enablement crash-safe (#430)
The default rosbag storage format moved back from
mcaptosqlite3, which ships with rosbag2 and needs no extra package;mcapstayed selectable throughsnapshots.rosbag.format. This came in with the crash-safety fix above and was not recorded at the time (#430)Contributors: @bburda, @mfaferek93
0.5.0 (2026-06-08)
ClearFaulthonors the newskip_correlation_auto_clearrequest flag so per-entity fault clears can opt out of cascade-clearing correlated symptom fault codes (#395)Three-layer protection against unbounded snapshot growth (bounded buffers plus pruning)
Concurrency and lifetime hardening: serialize concurrent subscription creation in
SnapshotCapture, join capture threads in theFaultManagerNodedestructor, and defense-in-depth shutdown guards to prevent teardown crashes across distrosAggregation security hardening and improved test coverage
Build: adopt the centralized
ROS2MedkitWarningsandROS2MedkitSanitizerscmake modules andbugprone/special-member-functionsclang-tidy checksContributors: @bburda
0.4.0 (2026-03-20)
Per-entity confirmation and healing thresholds via manifest configuration (#269)
Default rosbag storage format changed from
sqlite3tomcapSupport for namespaced fault manager nodes - gateway resolves service/topic names when the fault manager runs in a custom namespace
Build: use shared cmake modules from
ros2_medkit_cmakepackageBuild: centralized clang-tidy configuration
Contributors: @bburda
0.3.0 (2026-02-27)
0.2.0 (2026-02-07)
Initial rosdistro release
Central fault management node with ROS 2 services:
ReportFault - report FAILED/PASSED events with debounce filtering
GetFaults - query faults with filtering by severity, status, correlation
ClearFault - clear/acknowledge faults
Debounce filtering with configurable thresholds:
FAILED events decrement counter, PASSED events increment
Configurable confirmation_threshold (default: -1, immediate)
Optional healing support (healing_enabled, healing_threshold)
Time-based auto-confirmation (auto_confirm_after_sec)
CRITICAL severity bypasses debounce
Dual storage backends:
SQLite persistent storage with WAL mode (default)
In-memory storage for testing/lightweight deployments
Snapshot capture on fault confirmation:
Topic data captured as JSON with configurable topic resolution
Priority: fault_specific > patterns > default_topics
Stored in SQLite with indexed fault_code lookup
Auto-cleanup on fault clear
Rosbag capture with ring buffer:
Configurable duration, post-fault recording, topic selection
Lazy start mode (start on PREFAILED) or immediate
Auto-cleanup of bag files, storage limits (max_bag_size_mb)
GetRosbag service for bag file metadata
Fault correlation engine:
Hierarchical mode: root cause to symptom relationships
Auto-cluster mode: group similar faults within time window
YAML-based configuration with pattern wildcards
Muted faults tracking, auto-clear on root cause resolution
FaultEvent publishing on ~/events topic for SSE streaming
Wall clock timestamps (compatible with use_sim_time)
Contributors: Bartosz Burda, Michal Faferek
Changelog for package ros2_medkit_fault_reporter
0.7.0 (2026-08-27)
FaultReportercan now be constructed from anrclcpp_lifecycle::LifecycleNode(and from a plainrclcpp::Nodereference or explicit node interfaces), enabling use inside lifecycle nodes (#556, #555)Build and test only: the package is instrumented for coverage (#582), and every test takes a DDS domain at run time from the shared allocator (#551, #597)
Contributors: @bburda, @mfaferek93, @zeerekahmad
0.6.0 (2026-06-22)
No functional changes; version bump for the coordinated 0.6.0 release.
Contributors: @bburda
0.5.0 (2026-06-08)
LocalFilternow debouncesPASSEDevents with the same threshold / window filtering asFAILED, preventing rapid CONFIRMED/CLEARED status cycling from triggering unbounded snapshot recapture (#308)Build: adopt the centralized
ROS2MedkitWarningsandROS2MedkitSanitizerscmake modulesContributors: @bburda, @mfaferek93
0.4.0 (2026-03-20)
Build: use shared cmake modules from
ros2_medkit_cmakepackageBuild: auto-detect ccache, centralized clang-tidy configuration
Contributors: @bburda
0.3.0 (2026-02-27)
0.2.0 (2026-02-07)
Initial rosdistro release
FaultReporter client library with simple API:
report(fault_code, severity, description) - report FAILED events
report_passed(fault_code) - report fault condition cleared
High-severity faults (ERROR, CRITICAL) bypass local filtering
LocalFilter for per-fault-code threshold/window filtering:
Configurable threshold (default: 3 reports) and time window (default: 10s)
Prevents flooding FaultManager with duplicate reports
PASSED events always forwarded (bypass filtering)
Configuration via ROS parameters (filter_threshold, filter_window_sec)
Thread-safe implementation with mutex-protected config access
Contributors: Bartosz Burda, Michal Faferek
Changelog for package ros2_medkit_fault_detection
0.7.0 (2026-08-27)
Initial release of the package: shared, protocol-agnostic fault-detection model for medkit gateway plugins. A single header-only evaluator maps a raw value read from any source (OPC UA, S7, Modbus, ADS, …) into the set of faults it implies, using one of three composable detection modes:
ThresholdRule(numeric above/below a setpoint),StatusWordRule(decode named bits of an integer status register, with optional source-width masking to drop sign-extended high bits), andEnumMapRule(map a fault-code register value to a fault code + text, with an optional catch-all for unmapped values) (#481).evaluate(value, rule)is a pure function with no ROS / protocol dependencies, so it is trivially unit-testable and safe to compile into a dlopen-loaded plugin MODULE. Undecidable input (a non-finite double, a string, a failed numeric conversion) yields an empty result so a transition tracker holds the prior state instead of clearing a standing fault - a bad read never masks a real alarm.FaultTransitionTrackerlayers stateful raise/clear edge detection on top, keyed byfault_codealone; consumers that share one tracker across many points must enforce global fault-code uniqueness at config-load time.Shipped as a header-only INTERFACE library (
cxx_std_17); theOPC UAplugin is the first consumer and migrates its threshold / status-bit / enum detection onto this module.Global fault-code uniqueness is enforced across an OPC UA node map, so two rules cannot claim the same code and leave which one raised it undefined (#486)
A tracker holds its prior state on an undecidable read instead of clearing a standing fault, and an enum value with no mapping is labelled rather than dropped (#486)
Build and test only: the package is instrumented for coverage (#582), and its test is registered through the shared macros and opted out of DDS domain allocation, the evaluator being pure logic with no ROS node (#597)
Contributors: @mfaferek93, @bburda
Changelog for package ros2_medkit_diagnostic_bridge
0.7.0 (2026-08-27)
The
hardware_idof a diagnostic message is passed through as the fault’s source id, so a fault raised from/diagnosticsis attributed to the device the publisher named rather than to the bridge (#622)A fault code can be extracted from a diagnostic message’s key/value attributes, so a publisher that already carries its own code no longer has to encode it in the message name (#527)
The Humble fallback warning builds again (#622)
Build and test only: the package is instrumented for coverage (#582), every integration launch test runs on its own DDS domain taken at run time (#551, #597), and the suite re-emits and polls until the fault surfaces instead of asserting once against a graph that may not have settled (#504)
Contributors: @bburda, @mfaferek93, @nnarain
0.6.0 (2026-06-22)
Tests: label
test_integrationas an integration test so it runs in the integration suite instead of the unit set (#443)Contributors: @bburda
0.5.0 (2026-06-08)
Build: adopt the centralized
ROS2MedkitWarningsandROS2MedkitSanitizerscmake modulesTests: use centralized
ROS_DOMAIN_IDallocation for DDS isolationContributors: @bburda
0.4.0 (2026-03-20)
Build: use shared cmake modules from
ros2_medkit_cmakepackageBuild: auto-detect ccache, centralized clang-tidy configuration
Contributors: @bburda
0.3.0 (2026-02-27)
0.2.0 (2026-02-07)
Initial rosdistro release
Bridge node converting standard ROS 2 /diagnostics to FaultManager fault reports
Severity mapping:
OK -> PASSED event (fault condition cleared)
WARN -> WARN severity FAILED event
ERROR -> ERROR severity FAILED event
STALE -> CRITICAL severity FAILED event
Auto-generated fault codes from diagnostic names (UPPER_SNAKE_CASE)
Custom name_to_code mappings via ROS parameters
Stateless design: always sends PASSED for OK status (handles restarts)
Contributors: Michal Faferek
Changelog for package ros2_medkit_log_bridge
0.7.0 (2026-08-27)
Two parameters were being swallowed.
max_tracked_nodeswas clamped into range with nothing written to the log, so a configuration file and the running process could disagree in silence.report_cooldown_secwas worse: the range check was written so that NaN passed it, then passed the “non-positive disables the cooldown” test, and reachedDuration::from_seconds, where the conversion is undefined - after which the cooldown window suppressed either every ERROR report or none, with no way to tell which from outside. Both are now validated with a positive, finite range test, and both configuration rows describe what the code does (#607)Integer parameters are read as the int64 a ROS parameter actually holds and clamped in that domain before narrowing, so an out-of-range
severity_floorormax_tracked_nodescan no longer wrap back into the legal band and pass its own range check (#607)Build and test only: the package is instrumented for coverage (#582), every integration launch test runs on its own DDS domain taken at run time (#551, #597), and the log-bridge integration test waits for
/rosoutdiscovery before emitting its first line instead of racing itContributors: @bburda
0.6.0 (2026-06-22)
Initial release: promote
/rosoutlog entries (WARN/ERROR/FATAL) to FaultManager faults, attributed to the originating node via a per-source FaultReporter, with auto-generated stable fault codes (#422)Ships a default configuration so the bridge starts out of the box (#449)
Skips the medkit stack’s own nodes by default, matching on the raw logger name (#460)
Contributors: @mfaferek93
Changelog for package ros2_medkit_action_status_bridge
0.7.0 (2026-08-27)
A deferred fault report is retried instead of dropped, and delivered promptly through a fast retry timer rather than waiting for the next ordinary cycle (#472)
Silent parameter coercion is closed off, including a NaN that passed a
FloatingPointRangedescriptor and two integer narrowing paths where an out-of-range value wrapped back into the legal band before its own check ran. A value the node refuses or corrects is now reported rather than applied quietly (#607)Build and test only: the package is instrumented for coverage (#582), every integration launch test runs on its own DDS domain taken at run time (#551, #597), and the suite synchronises on bridge discovery before raising a fault instead of assuming the graph has settled (#504)
Contributors: @bburda, @mfaferek93
0.6.0 (2026-06-22)
Initial release: generic action-status bridge. Watches every
/<action>/_action/statustopic and turns ABORTED goals into FaultManager faults (<PREFIX>_<ACTION>_ABORTED), heals on a non-failing terminal state. Captures the terminal action verdict that the diagnostic and log bridges cannot see (#423).Fault state is per-ACTION, derived from the whole
GoalStatusArrayon each message: a fault is raised only on thehealthy -> failedtransition and healed only onfailed -> healthy, so it is order-independent and resilient to dropped terminal messages. Per-goal dedup now only suppresses duplicate log lines.canceled_is_faultemits a status-aware<PREFIX>_<ACTION>_CANCELEDcode; such a fault heals on the next non-failing terminal or when the canceled goal ages out of the retained status array.Vanished actions (lifecycle deactivate, one-shot nodes) are pruned on rescan.
Parameters are range-checked/normalized at load; an incompatible action status QoS is now warned about instead of silently dropping faults.
Ships a default configuration so the bridge starts out of the box (#459)
Fault source is fixed at first report instead of being re-attributed later: faults raised before discovery settles register the server FQN, faults are never attributed to the discovery placeholder node name, and the subscriptions and timer are reset in the destructor (#466)
Contributors: @bburda, @mfaferek93
Changelog for package ros2_medkit_serialization
0.7.0 (2026-08-27)
Build and test only: the package is instrumented for coverage (#582), its ROS-dependent tests take a DDS domain at run time from the shared allocator while the pure-logic ones are opted out (#597), and clang-tidy no longer reports findings in vendored headers that the header filter matched
Contributors: @bburda, @mfaferek93
0.6.0 (2026-06-22)
TypeIntrospectionresolves service and actiontype_infoschemas lazily and caches them as shared, immutable per-type objects, so discovery no longer rebuilds and deep-copies operation schemas on every refresh (#462)Contributors: @bburda
0.5.0 (2026-06-08)
TypeIntrospectionrelocated into this package as part of the gateway core / ROS 2 layer splitBuild: adopt the centralized
ROS2MedkitWarningsandROS2MedkitSanitizerscmake modulesContributors: @bburda
0.4.0 (2026-03-20)
Enable
POSITION_INDEPENDENT_CODEfor MODULE target compatibilityBuild: use shared cmake modules from
ros2_medkit_cmakepackageBuild: auto-detect ccache, centralized clang-tidy configuration
Contributors: @bburda
0.3.0 (2026-02-27)
0.2.0 (2026-02-07)
Initial rosdistro release
Runtime JSON to ROS 2 message serialization using vendored dynmsg C++ API
TypeCache - thread-safe caching of ROS type introspection data with shared_mutex for read concurrency
JsonSerializer - bidirectional JSON <-> ROS message conversion via dynmsg YAML bridge, including CDR serialization/deserialization for GenericClient/GenericSubscription
ServiceActionTypes - helper utilities for resolving service and action internal types (request/response, goal/result/feedback)
SerializationError exception hierarchy for structured error handling
Contributors: Bartosz Burda
Changelog for package ros2_medkit_msgs
0.7.0 (2026-08-27)
GetRosbag.srvgains arecording_idrequest field (tried beforefault_codeand falling back to it when it names no recording, so a caller holding one identifier that may be either can set both;fault_codekeeps its meaning of “the newest recording of this fault”) andrecording_id/fault_codes[]response fields.ListRosbags.srvgains a parallelrecording_ids[].Snapshot.msg’sbulk_data_idnow carries a recording id for rosbag snapshots rather than the fault code; the field type is unchanged. Additive, but the service type hashes change, so the gateway and the fault manager must be deployed together (#623, #620)last_passedis exposed on the fault wire, so a consumer can tell when the monitored condition was last observed healthy without inferring it from the status (#573)Rosbag recordings are addressed by recording id rather than by fault code, so a fault holding several recordings can expose each one (#623)
The fault contract documentation now matches the implementation on three points that had drifted:
occurrence_countis an edge counter that advances on a new fault and on every re-raise after CLEARED,first_occurredmarks the start of the current occurrence rather than the first one ever, and healed clears and SSE loss accounting are described as the code implements them (#618)Contributors: @bburda, @mfaferek93
0.6.0 (2026-06-22)
No functional changes; version bump for the coordinated 0.6.0 release.
Contributors: @bburda
0.5.0 (2026-06-08)
New service definitions
ListEntities.srv,GetEntityData.srv,GetCapabilities.srvand theEntityInfo.msgtype for exposing the entity tree over ROS 2 services (#330)ClearFault.srvrequest gains abool skip_correlation_auto_clearfield so callers can opt out of cascade-clearing correlated symptom fault codes; out-of-tree callers must rebuild against the new message (#395)Contributors: @bburda, @mfaferek93
0.4.0 (2026-03-20)
MedkitDiscoveryHintmessage type for beacon discovery publishersContributors: @bburda
0.3.0 (2026-02-27)
0.2.0 (2026-02-07)
Initial rosdistro release
Fault management messages:
Fault.msg - Core fault data model with severity levels (INFO/WARN/ERROR/CRITICAL) and debounce-based status lifecycle (PREFAILED/PREPASSED/CONFIRMED/HEALED/CLEARED)
FaultEvent.msg - Real-time event notifications (EVENT_CONFIRMED/EVENT_CLEARED/EVENT_UPDATED) with auto-cleared correlation codes
MutedFaultInfo.msg - Fault correlation muting metadata
ClusterInfo.msg - Fault clustering information
Fault management services:
ReportFault.srv - Report fault events with FAILED/PASSED event model
GetFaults.srv - Query faults with filtering by severity, status, and correlation data
ClearFault.srv - Clear/acknowledge faults by fault_code
GetSnapshots.srv - Retrieve topic snapshots captured at fault time
GetRosbag.srv - Retrieve rosbag recordings associated with faults
Contributors: Bartosz Burda, Michal Faferek
Changelog for package ros2_medkit_integration_tests
0.7.0 (2026-08-27)
Every integration launch test runs on its own DDS domain, taken when the test starts and held through an open socket for exactly as long as it runs, so a crash or a SIGKILL releases it the same way an ordinary exit does. This replaces the hand-maintained per-package domain pools, which had to stay pairwise disjoint because colcon runs one ctest per package in parallel and a CTest
RESOURCE_LOCKbinds only inside a single ctest run (#551, #597)Feature tests are registered by glob: a file dropped into
test/features/*.test.pyis registered with its port and domain assigned automatically, so no explicit registration is needed and none can bypassGATEWAY_TEST_PORTInline gateways and fault managers go through the shared launch helpers rather than each test hand-building its own, which also fixed the SIGKILL these tests hit under coverage instrumentation (#554)
The Python in this package is linted (#598), and the package is instrumented for coverage (#582)
Flaky suites are fixed at their cause rather than retried: the OPC UA Alarms and Conditions test, the action-status and diagnostic-bridge fault paths, the operation-handlers fixture,
type_infocaching and the documentation linkcheck (#504, #636)Contributors: @bburda, @mfaferek93, @YueBit
0.6.0 (2026-06-22)
0.5.0 (2026-06-08)
New peer-aggregation suites: peer aggregation, cross-ECU fan-out across all resource types, daisy-chain hierarchical aggregation, leaf-collision aggregation, and cross-ECU log aggregation
New SOVD-aligned entity-model suites: runtime entity model, flat entity tree without areas, hybrid suppression, and per-entity fault scope isolation
New OpenAPI conformance suites:
test_openapi_callabilityandtest_openapi_response_driftvalidate the published spec against live responsesNew graph-event discovery suite covering rclcpp graph-event driven discovery refresh
Coverage for the
GET /apps/{id}/belongs-todiscovery endpoint, nestedx-medkitvendor payloads (#385), version-info aggregation, and pending-after-register update status (#378)Tests: centralized
ROS_DOMAIN_IDallocation and widened shutdown timeouts for sanitizer overheadContributors: @bburda, @eclipse0922
0.4.0 (2026-03-20)
Integration tests for SOVD resource locking (acquire, release, extend, fault clear with locks, expiration, parent propagation)
Integration tests for SOVD script execution endpoints (all formats, params, output, failure, lifecycle)
Integration tests for graph provider plugin (external plugin loading, entity introspection)
Integration tests for beacon discovery plugins (topic beacon, parameter beacon)
Integration tests for OpenAPI/docs endpoint
Integration tests for logging endpoints (
/logs,/logs/configuration)Integration tests for linux introspection plugins (launch_testing and Docker-based)
Port isolation per integration test via CMake-assigned unique ports
ROS_DOMAIN_IDisolation for integration testsBuild: use shared cmake modules from
ros2_medkit_cmakepackageContributors: @bburda
0.3.0 (2026-02-27)
Changelog for package ros2_medkit_cmake
0.7.0 (2026-08-27)
ROS_DOMAIN_IDisolation is allocated at run time instead of from a hand-maintained table. A test registered throughmedkit_add_gtest,medkit_add_gmock,medkit_add_launch_test,medkit_add_pytest_testormedkit_add_wrapped_testtakes a domain when it starts and holds it through an open socket for exactly as long as it runs, so a crash or a SIGKILL releases it the same way an ordinary exit does. The isolation reaches across packages, which a CTestRESOURCE_LOCKcannot, because colcon runs one ctest per package in parallel.medkit_test_needs_no_domain()opts out a test that creates no ROS node. The usable band excludes the domains whose RTPS port slice falls inside the kernel ephemeral port range (#551, #597)Launch tests are no longer registered in a system package build, where they failed as a group and were reported as passing anyway because bloom runs the test step permissively; the macro now owns their properties instead of each package setting them (#608, #602)
clang-tidy analyses a package’s translation units in parallel, with the job count capped for an 8 GB machine and packages serialised so the cap holds, cutting PR feedback time from about 56 minutes to about 35 (#588, #590)
Every C++ package is instrumented for coverage, and every package that lints exports a compile database (#582)
Sanitizer builds keep their asserts, and the build reports what ccache did (#590)
Contributors: @bburda
0.6.0 (2026-06-22)
ROS2MedkitTestDomain: carve dedicatedROS_DOMAIN_IDranges for the new log bridge (210-214) and action-status bridge (215-219) test suites out of the integration-tests range (#422)Contributors: @mfaferek93
0.5.0 (2026-06-08)
ROS2MedkitWarnings.cmakemodule centralizes compiler warning flags across all packages, with selective-Werror(namespacedMEDKIT_ENABLE_WERROR, defaults OFF) applied only to flags safe against external headersROS2MedkitSanitizers.cmakemodule adds ASan/UBSan and TSan support for sanitizer CI jobsROS2MedkitTestDomain.cmakecentralizesROS_DOMAIN_IDallocation for per-test DDS isolationVendored cpp-httplib 0.14.3 as a build-farm fallback (
VENDORED_DIRparameter), marked as a SYSTEM include to suppress third-party warningsmedkit_find_cpp_httplibcaps cpp-httplib at< 0.20across both the pkg-config andfind_package(httplib)tiers, so distros shipping 0.20+ (Ubuntu 26.04 ships 0.26, which dropped the multipartRequest::has_fileAPI the gateway uses) fall through to the vendored 0.14.3 header instead of failing the build;ROS2MedkitCompat.cmakeextended to cover ROS 2 Lyrical / Ubuntu 26.04 Resolute (rclcpp 32,yaml_cpp_vendortarget export,ament_target_dependenciesremoval in ament_cmake 2.8.5+), which replaces Rolling in CI (#405)Contributors: @bburda, @mfaferek93
0.4.0 (2026-03-20)
Initial release - shared cmake modules extracted from gateway package (#294)
ROS2MedkitCcache.cmake- auto-detect ccache for faster incremental rebuildsROS2MedkitLinting.cmake- centralized clang-tidy configuration (opt-in locally, mandatory in CI)ROS2MedkitCompat.cmake- multi-distro compatibility shims for ROS 2 Humble/Jazzy/RollingContributors: @bburda
Changelog for package ros2_medkit_linux_introspection
0.7.0 (2026-08-27)
Container CPU and memory limits are reported on every common cgroup layout, not just one. The reader parses the cgroup v1 line format alongside v2, and resolves the limit files under both
cgroupns=hostandcgroupns=private- the latter being the Docker default, where the container sees its own cgroup mounted directly at/sys/fs/cgroupand the previously built path did not exist. A limit that could not be read is now distinguished from a container that genuinely has no limit, instead of both being reported as unlimited (#637, #604)Limits are read from the cgroup that owns them rather than from the process’s own leaf, a legacy controller’s limit is no longer outranked by a hierarchy that does not set one, and a limit file is read to its end instead of to the first short read (#637)
Containers are detected per process rather than once for the whole node, and the ctest suite exercises every supported cgroup layout against synthetic hierarchies (#637)
Build and test only: the package is instrumented for coverage (#582), and its tests are opted out of DDS domain allocation, reading synthetic /proc and cgroup trees rather than creating a ROS node (#597)
Contributors: @bburda
0.6.0 (2026-06-22)
No functional changes; version bump for the coordinated 0.6.0 release.
Contributors: @bburda
0.5.0 (2026-06-08)
Migrated the introspection plugins to the
get_routes()plugin APIDeclared
pkg-configas abuildtool_dependand fixed a route-separator bugBuild: adopt the centralized
ROS2MedkitWarningsandROS2MedkitSanitizerscmake modulesContributors: @bburda
0.4.0 (2026-03-20)
Initial release - Linux process introspection plugins for ros2_medkit gateway
procfs_plugin- process-level diagnostics via/procfilesystem (CPU, memory, threads, file descriptors)systemd_plugin- systemd unit status and resource usage via D-Buscontainer_plugin- container runtime detection and cgroup resource limitsPidCachewith TTL-based refresh for efficient PID-to-node mappingproc_readerandcgroup_readerutilities with configurable proc rootCross-distro support for ROS 2 Humble, Jazzy, and Rolling
Contributors: @bburda
Changelog for package ros2_medkit_beacon_common
0.7.0 (2026-08-27)
0.6.0 (2026-06-22)
No functional changes; version bump for the coordinated 0.6.0 release.
Contributors: @bburda
0.5.0 (2026-06-08)
Updated include paths for the gateway core / ROS 2
PluginContextlayer split and removal of backwards-compat shim headersBuild: adopt the centralized
ROS2MedkitWarningscmake moduleContributors: @bburda
0.4.0 (2026-03-20)
Initial release - shared utilities for beacon discovery plugins
BeaconHintStorewith TTL-based hint transitions and thread safetyBeaconValidatorinput validation gateBeaconEntityMapperto convert discovery hints intoIntrospectionResultbuild_beacon_response()shared response builderBeaconPluginbase class for topic and parameter beacon pluginsTokenBucketrate limiter with thread-safe accessContributors: @bburda
Changelog for package ros2_medkit_param_beacon
0.7.0 (2026-08-27)
Build and test only: every integration launch test runs on its own DDS domain, taken when the test starts and released when it ends, so a crash frees the domain the same way an ordinary exit does (#551, #597). The usable band excludes the domains whose RTPS port slice falls inside the kernel ephemeral port range, where any process on the machine can steal a port and kill a node with a bind failure
Build: the package is instrumented for coverage (#582)
Contributors: @bburda
0.6.0 (2026-06-22)
No functional changes; version bump for the coordinated 0.6.0 release.
Contributors: @bburda
0.5.0 (2026-06-08)
Migrated
ParameterBeaconPluginto theget_routes()plugin APIAdded shutdown guards and
noexceptdestructors that reset rclcpp resources before member destruction, preventing teardown SIGSEGV; the graph poll now swallowsrcl“context invalid” during shutdownAdded post-shutdown guard unit tests
Build: adopt the centralized
ROS2MedkitWarningscmake moduleContributors: @bburda
0.4.0 (2026-03-20)
Initial release - parameter-based beacon discovery plugin
ParameterBeaconPluginwith pull-based parameter reading for entity enrichmentx-medkit-param-beaconvendor extension REST endpointPoll target discovery from ROS graph in non-hybrid mode
ParameterClientInterfacefor testable parameter accessContributors: @bburda
Changelog for package ros2_medkit_topic_beacon
0.7.0 (2026-08-27)
Build and test only: every integration launch test runs on its own DDS domain, taken when the test starts and released when it ends, so a crash frees the domain the same way an ordinary exit does (#551, #597). The usable band excludes the domains whose RTPS port slice falls inside the kernel ephemeral port range, where any process on the machine can steal a port and kill a node with a bind failure
Build: the package is instrumented for coverage (#582)
Contributors: @bburda
0.6.0 (2026-06-22)
No functional changes; version bump for the coordinated 0.6.0 release.
Contributors: @bburda
0.5.0 (2026-06-08)
Migrated
TopicBeaconPluginto theget_routes()plugin APIAdded shutdown guards and
noexceptdestructors (withoverride) that reset rclcpp resources before member destruction, preventing teardown SIGSEGVAdded post-shutdown guard unit tests
Build: adopt the centralized
ROS2MedkitWarningscmake moduleContributors: @bburda
0.4.0 (2026-03-20)
Initial release - topic-based beacon discovery plugin
TopicBeaconPluginwith push-based topic subscription for entity enrichmentx-medkit-topic-beaconvendor extension REST endpointStamp-based TTL for topic beacon hints
Diagnostic logging for beacon hint processing
Contributors: @bburda
Changelog for package ros2_medkit_graph_provider
0.7.0 (2026-08-27)
Breaking: the
x-medkit-graphhealth model is reworked andschema_versionis now"2.0.0".pipeline_statusis derived from data freshness and node reachability instead of the previous status model, and no longer reports"broken"for an edge that simply has no/diagnosticscoverage. Theerror_reasonvaluesnode_offline,topic_staleandno_data_sourceare gone;metrics_staleis the only reachable value. A consumer that matched on the old values needs a change (#546, #545)metrics.sourcereports the resolved publisher rather than a placeholder, falling back to the single publisher of a topic where the RMW does not expose the identity directly, and the negotiation topic filter is anchored to path segments so an unrelated topic whose name merely contains the segment is no longer matched (#546)Per-function configuration overrides take effect, an invalid override warns instead of being applied silently, and the freshness clock is monotonic with a stateless stale grace, so a wall-clock step no longer moves an edge in or out of staleness (#546)
Rate measurement is defended against multi-publisher inflation, state is bounded, and shutdown is guarded against callbacks firing on a partially destroyed provider (#546)
The plugin and its robustness fields are documented, and the never-executed stale-topic path is removed rather than left as dead code (#546)
Contributors: @bburda
0.6.0 (2026-06-22)
No functional changes; version bump for the coordinated 0.6.0 release.
Contributors: @bburda
0.5.0 (2026-06-08)
Migrated
GraphProviderPluginto theget_routes()plugin API and fixed a route-separator bugAdded shutdown guards and
noexceptdestructors that reset rclcpp resources before member destruction, preventing teardown SIGSEGVAdded post-shutdown guard unit tests; dropped the cpp-httplib source install from the Dockerfile (now vendored via
ros2_medkit_cmake)Build: adopt the centralized
ROS2MedkitWarningscmake moduleContributors: @bburda
0.4.0 (2026-03-20)
Initial release - extracted from
ros2_medkit_gatewaypackageGraphProviderPluginfor ROS 2 graph-based entity introspectionStandalone external plugin package with independent build and test
Locking support via
PluginContextAPIContributors: @bburda
Changelog for package ros2_medkit_graph_watchdog
0.7.0 (2026-08-27)
Initial release of the package: a gateway plugin that raises faults for silent failures in the ROS 2 graph - problems that leave every node alive and every topic present, so nothing in the stack reports them today. The package carries the plugin skeleton, a central reliability gate that holds raises until the graph has quiesced, and a detector registry. Detectors are configured under
plugins.graph_watchdog.<key>, each with araise/advisory/offmode, and an unknown key underdetectors.<id>is reported rather than silently ignored, so a typo cannot quietly disable a check (#571)qos_mismatchdetector: raisesGRAPH_QOS_MISMATCHfor a publisher and subscriber whose QoS profiles cannot match, which leaves the connection silently unestablished (#571)orphandetector: raisesGRAPH_ORPHANfor a one-sided topic - a publisher with no subscriber, or the reverse - but only when a complementary near-miss counterpart exists: same message type, a namespace and leaf within the configured edit distance, and the opposite side. A lone unsubscribed topic raises nothing. The pair is the signature of a remap or a topic-name typo, which is what the detector is for (#578)param_driftdetector: raisesGRAPH_PARAM_DRIFTwhen a node’s live parameter value diverges from the declared expectation (#580)lifecycle_expectationdetector: raisesGRAPH_NODE_INACTIVEfor a node named inrequire_activethat is not in theactivelifecycle state once itsgracewindow has passed. It raises two further codes for the cases where the state could not be established at all:GRAPH_NODE_UNREADABLEwhen the lifecycle service does not answer, andGRAPH_NODE_NOT_MANAGEDwhen the named node has no lifecycle to read. These are distinct codes on the wire, so a consumer filtering onGRAPH_NODE_INACTIVEalone sees neither unmeasured case. A node is matched by itsApp::id, its full FQN, or the bare leaf of that FQN, and the grace streak is counted per node, so naming one node in two documented forms cannot halve the grace that was configured (#587)node_deathdetector: raisesGRAPH_NODE_DISAPPEAREDfor a node that leaves the graph when nothing else is left to report it, with a suppression framework that decides ownership of a departure from knowledge rather than from silence. A node whose lifecycle state was never readable is admitted only provisionally, and is handed back the moment a label arrives saying the departure belonged to the lifecycle detector instead. Zero-config: there is no list of nodes to maintain (#625, #624)Two silent-fault classes are not delivered in this release.
GRAPH_TF_STALEandGRAPH_LATENCY_BUDGEThave their fault codes reserved in the frozenGRAPH_*namespace and land in later changes, each against its own issueThe plugin keeps its ROS entities off the gateway executor, so an entity created and destroyed while the gateway runs cannot have its destructor run on an executor thread concurrently with a create on the same node
Contributors: @bburda
Changelog for package ros2_medkit_opcua
0.7.0 (2026-08-27)
Config-less discovery. A read-only network scan finds the OPC UA server instead of requiring its endpoint up front (#509),
auto_browsewalks the address space recursively and builds the SOVD tree from it (#510), and identity, writability and fault triggers are read from the device itself rather than declared in a node map (#544)Native OPC UA Alarms and Conditions become medkit faults with no configuration (#511). Alarms are routed to distinct faults by message substring (#506),
severityis accepted as an alias inevent_alarms(#501), and events that are not conditions are dropped from the alarm path instead of being reported as alarms (#552)Asset identity is populated from the device nameplate, including
order_code(#492, #499)PLC_COMMS_LOSTis raised when the connection to the PLC drops, so a lost link is a reported fault rather than an absence of data (#507)Security and reconnect hardening: a secure connection profile, replay of alarm state after a reconnect, and correct handling of several simultaneous alarms (#485)
Fixed a data race between the REST thread and the poll thread on the pending-report buffer that crashed the gateway with a SIGSEGV (#519, #521)
Fault detection moved onto the shared
ros2_medkit_fault_detectionevaluator, and the node map enforces global fault-code uniqueness so two rules cannot claim the same code (#486)This package is still excluded from the rosdistro binary release pending vendoring of open62541pp (#366)
Contributors: @bburda, @mfaferek93
0.6.0 (2026-06-22)
No functional changes; version bump for the coordinated 0.6.0 release.
Contributors: @bburda
0.5.0 (2026-06-08)
Native OPC-UA Part 9
AlarmConditionTypeevent subscription. The plugin now subscribes to vendor-defined alarms (Siemens S7-1500Program_Alarm/ ProDiag, Beckhoff TF6100, CodeSys 3.5+, Rockwell via FactoryTalk Linx) and bridges each event into the SOVD fault lifecycle. Configured via a new top-levelevent_alarms:block in the node map YAML; mutually exclusive per entry with the existing threshold-basedalarmform. (issue #386)New SOVD operations on entities that host alarm sources:
acknowledge_faultinvokes the inheritedAcknowledgemethod on the liveConditionId(i=9111, EventId tracked per Part 9 §5.7.3);confirm_faultinvokesConfirm(i=9113). Both accept an optionalcommentrendered asLocalizedTexton the server.OpcuaClientgainsadd_event_monitored_item/remove_event_monitored_item/call_methodand a generation counter that filters callbacks fired from defunct subscriptions after a reconnect. Heap-ownedEventCallbackContextresolves the open62541pp / raw-C lifetime hazard.Header-only
AlarmStateMachinemappingEnabledState x ShelvingState x ActiveState x AckedState x ConfirmedState x BranchIdto SOVDCONFIRMED / HEALED / CLEARED / Suppressed. Full transition table documented indesign/index.rst.ConditionRefresh(Server method i=3875) is invoked on subscribe and on every reconnect, withRefreshStartEvent/RefreshEndEventbracketing tracked for diagnostics.New
test_alarm_serverfixture (open62541-based, full namespace 0 + alarms enabled) emits AlarmConditionType events on stdin commands; integration testrun_alarm_tests.shruns in CI alongside the existing OpenPLC threshold suite. The fixture builds by default via the workspacecolcon build(gated onMEDKIT_OPCUA_BUILD_ALARM_SERVERwhich defaults to ON;ExternalProject_Addrebuilds open62541 withUA_NAMESPACE_ZERO=FULLand alarms ON, with a serial sub-build to dodge the upstream-jrace onnamespace0_generated.c).New CTest wrapper
test_alarm_server_smokeboots the fixture on an ephemeral port and runs the asyncua smoke test against it; skips with CTest exit 77 (treated as pass) whenasyncuais not importable, so iterating on plugin code without the Python dependency does not fail the suite.Contributors: @mfaferek93, @bburda
0.4.0 (2026-04-11)
Initial release
OpcuaPluginimplementation ofGatewayPluginandIntrospectionProviderthat bridges OPC-UA capable PLCs into the SOVD entity treeREST endpoints via the new
get_routes()plugin API:x-plc-data,x-plc-operations,x-plc-statusVendor capabilities registered per entity - only PLC-backed apps and the PLC runtime component advertise the
x-plc-*endpointsFull OPC 10000-6 section 5.3.1.10 node identifier support (
i=numeric,s=string,g=GUID,b=opaque ByteString); example node maps for OpenPLC, Siemens S7-1500 TIA Portal, Beckhoff TwinCAT 3, Allen-Bradley via Kepware and KUKA KR C5NodeMapdriven by YAML configuration - same binary serves any OPC-UA compliant server by changing the node map fileDeterministic entity ordering in
IntrospectionResultoutput (entries sorted by id)Threshold-based PLC alarm detection routed to SOVD faults via
ros2_medkit_msgsservicesReportFault/ClearFaultOptional bridging of numeric PLC values to ROS 2
std_msgs/Float32topics fromset_context()Type-aware writes with per-node range validation
Robust connection-loss detection: all three OPC-UA client paths (
read_value,read_values,write_value) mark the connection as dropped on terminal status codes soOpcuaPollerreconnect kicks in without stallingPolling mode (default) and OPC-UA subscription mode, backed by
open62541ppv0.16.0Integration test suite against an OpenPLC IEC 61131-3 tank demo container
Contributors: @mfaferek93
Changelog for package ros2_medkit_sovd_service_interface
0.7.0 (2026-08-27)
A cyclic-subscription request naming a resource path the collection cannot stream is refused instead of being accepted and never delivering
The sampler declaration stays inside the gateway, so the service interface no longer carries a second copy of it
Build and test only: the package is instrumented for coverage (#582), and its tests take a DDS domain at run time from the shared allocator (#597)
Contributors: @bburda
0.6.0 (2026-06-22)
list_entity_faultshandles a faultstatusreturned as an object (not just a string), supporting peer-fault aggregation across daisy-chained gateways (#419)Contributors: @bburda
0.5.0 (2026-06-08)
Initial release - gateway plugin that exposes the medkit entity tree and fault data over standard ROS 2 services, so other ROS 2 nodes can query entities and faults without going through the HTTP API (#330)
Creates ROS 2 services under a configurable prefix (default
/medkit):list_entities,list_entity_faults, andget_capabilities; aget_entity_dataservice is also registered but currently returns “not implemented” (use the HTTP REST API for topic data) pending (#351)Runs as a gateway MODULE plugin with read-only
PluginContextaccess to the entity cache and fault manager; implements no provider interfacesShutdown guards and
noexceptdestructors reset ROS resources before member destruction to prevent teardown crashesContributors: @bburda, @mfaferek93