Skip to content

prometheus: fix scrape duration growing unbounded (#13667)#13696

Open
PrashantBhanage wants to merge 1 commit into
apache:ghi13586-prometheusDrainage-20from
PrashantBhanage:fix/13667-clean
Open

prometheus: fix scrape duration growing unbounded (#13667)#13696
PrashantBhanage wants to merge 1 commit into
apache:ghi13586-prometheusDrainage-20from
PrashantBhanage:fix/13667-clean

Conversation

@PrashantBhanage

Copy link
Copy Markdown

Types of changes

  • Breaking change (fix or feature that would cause existing functionality to change)
  • New feature (non-breaking change which adds functionality)
  • Bug fix (non-breaking change which fixes an issue)
  • Enhancement (improves an existing feature and functionality)
  • Cleanup (Code refactoring and cleanup, that may add test cases)
  • Build/CI
  • Test (unit or integration test code)

Feature/Enhancement Scale or Bug Severity

Feature/Enhancement Scale

  • Major
  • Minor

Bug Severity

  • BLOCKER
  • Critical
  • Major
  • Minor
  • Trivial

Screenshots (if appropriate):

N/A — backend-only change, no UI impact.

How Has This Been Tested?

  • Added unit tests for the TTL guard behavior.
  • Verified repeated scrapes within the refresh interval reuse cached metrics.
  • Verified scrapes after the interval trigger a new metrics recomputation.
  • Built the Prometheus integration module successfully.

How did you try to break this feature and the system with this change?

  • Sent concurrent /metrics scrape requests to verify the bounded executor and synchronization prevent overlapping recomputation.
  • Verified the executor shuts down cleanly in stop().
  • Verified changing prometheus.exporter.metrics.min.refresh.interval at runtime changes the refresh behavior without requiring a restart.

Three changes scoped exclusively to the prometheus exporter plugin:

1. Bounded HTTP executor — give the HttpServer a FixedThreadPool(2)
   instead of the JDK default single-threaded executor; shut it down
   cleanly in stop().

2. TTL guard on updateMetrics() — add a synchronized guard with a
   last-run timestamp so scrapes arriving faster than the configurable
   minimum interval (prometheus.exporter.metrics.min.refresh.interval,
   default 5 s) reuse the previously computed metrics instead of
   triggering a new recomputation.

3. Timing instrumentation — wrap updateMetrics() body with
   System.nanoTime() start/end and log elapsed wall-clock time at
   info level on every run (including exception path), so slow
   sub-collectors can be identified from logs.

Fixes: apache#13667
Ref: apache#13586
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant