Skip to content

NMS-20352: Publish telemetryd connector configs once per location - #8895

Open
cgorantla wants to merge 4 commits into
foundation-2026from
cg/jira/NMS-20352
Open

cgorantla wants to merge 4 commits into
foundation-2026from
cg/jira/NMS-20352

Conversation

@cgorantla

@cgorantla cgorantla commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Publish telemetryd connector configs once per location instead of per node.
Batch service tracker changes so telemetryd builds all connector configs in a change and publishes once per location. Interpolated parameters are computed once
Fixes startup time with Telemetryd

External References

…tead of per node

Batch service tracker changes so telemetryd builds all connector configs in a change and
publishes once per location. Interpolated parameters are computed once and the duplicate
connector guard works.
…S-20352

# Conflicts:
#	opennms-dao/src/main/java/org/opennms/netmgt/dao/support/DefaultServiceTracker.java
#	opennms-dao/src/test/java/org/opennms/netmgt/dao/support/DefaultServiceTrackerTest.java

@marshallmassengill marshallmassengill left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving but there are some minor things we may want to fix since we're in here:

  • Failed config build drops the node permanently. ConnectorManager.java:116. Verified: node 0 never republished after later changes. Retry failed refs on the next update call; don't rethrow.
  • Failed publish retries only on a change in that location. ConnectorManager.java:136. Verified: no-change refreshes and other-location changes don't retry. Acceptable; a retry timer would close it.
  • HashMap order churns twin patches. LocationPublisher.java:43. Verified: adding 3 configs across a resize = 567 ops at 190 nodes, 2289 at 1530; LinkedHashMap = 3.
  • Per-service publish/remove are dead in production. OpenConfigTwinPublisher.java:33. Verified: only the openconfig itest calls them.

@cgorantla

Copy link
Copy Markdown
Contributor Author

Approving but there are some minor things we may want to fix since we're in here:

  • Failed config build drops the node permanently. ConnectorManager.java:116. Verified: node 0 never republished after later changes. Retry failed refs on the next update call; don't rethrow.
  • Failed publish retries only on a change in that location. ConnectorManager.java:136. Verified: no-change refreshes and other-location changes don't retry. Acceptable; a retry timer would close it.
  • HashMap order churns twin patches. LocationPublisher.java:43. Verified: adding 3 configs across a resize = 567 ops at 190 nodes, 2289 at 1530; LinkedHashMap = 3.
  • Per-service publish/remove are dead in production. OpenConfigTwinPublisher.java:33. Verified: only the openconfig itest calls them.

Handled these review comments. Retrying publish failures should be handled at the Twin layer. Will do that in a separate issue.

@cgorantla

Copy link
Copy Markdown
Contributor Author

@christianpape

christianpape commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor
  1. Retries may never happen. The PR has three places that say "we'll try again on the next update":
    • LocationPublisher.publishPending after a failed session.publish().
    • ConnectorManager.failedBuilds after interpolation fails.
    • DefaultServiceTracker re-sending a batch after the listener throws.

All three depend on DefaultFilterWatcher.refreshNow(), which does nothing unless the filter results actually changed
(Objects.equals(last, new) → return). There is also a second limit: a location that failed to publish is only retried when a later change touches that same location, because publish() only visits locations in changesByLocation. On a stable inventory, a failed twin publish can leave a Minion location without its connector configs until someone edits a node. The comment "sends them with its next publish" suggests a retry that may never come. A periodic retry, or retrying all pending publishers on every call, would fix this.

  1. One queue name per location (existing problem, now more likely to bite). LocationPublisher holds a single queueName, and setQueueName overwrites it on every call. With several connectors that use different queue names at the same location, whichever connector publishes last sets the queue name for everyone. Batching doesn't cause this, but it makes the result depend on which (connector, package) batch runs last.

  2. Exceptions on first registration aren't caught. DefaultFilterWatcher catches listener exceptions in notifyCallbacks, but the first callback.accept(lastFilterResults) in addCallback (around line 213) is not wrapped. If a BatchServiceListener throws there, the exception escapes through the TrackingSession constructor into trackServiceMatchingFilterRule. ConnectorManager catches everything, so it's safe today. Other code adopting the new API could hit this.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants