Skip to content

feat(core): add observe(value, count) for batched observations - #2508

Closed
callumdonald96 wants to merge 3 commits into
prometheus:mainfrom
callumdonald96:batch-observe-1157
Closed

callumdonald96 wants to merge 3 commits into
prometheus:mainfrom
callumdonald96:batch-observe-1157

Conversation

@callumdonald96

Copy link
Copy Markdown

Closes #1157.

Adds observe(double value, long count) to record the same value count times as a single
operation, for callers that already have pre-aggregated data ("this value occurred n times") —
importing counts from another system, backfilling, or driving a histogram from a batch job's
tallies.

histogram.labelValues("GET", "/", "200").observe(0.012, 4_711);

The cost does not depend on count: a batch of a million costs the same as a batch of sixteen.

Why this needs a change in Buffer, not just Histogram

Worth flagging up front, because it makes the change larger than the issue suggests.

Buffer counts one ticket per observation, and collection spins until the metric's count adder
matches the ticket total. A batch that takes one ticket but adds n to count overshoots
expectedCount and never matches it again — every later scrape of that data point spins to its
5 s deadline and throws IllegalStateException. I confirmed this with the naive implementation
before designing around it (collect() spun 5012 ms, then threw).

So the first commit teaches Buffer to count weighted tickets. append(value, weight) claims
the ticket range (count - weight, count] with a single atomic add, which is what makes it safe: a
batch cannot straddle a collector's activation, because both are single atomic ops on the same word.
It is either entirely inside that collection's expectedCount (direct path) or entirely outside it
(buffered for replay) — never split. The existing late-reader guard generalises from a point test to
a range test and is identical for weight == 1.

Buffered generations carry a lazily allocated long[] weights, left null while every entry has
weight 1, so a scrape racing only single observations allocates exactly what it does today. Replay
moves from Consumer<Double> to a primitive WeightedObserver, which also stops boxing every
replayed value.

Semantics

Buckets and counts are exactly what count single calls produce. Two deliberate differences:

  • A batch is atomic with respect to scrapes — a snapshot contains all of it or none. A loop
    never was.
  • The sum is increased by the correctly rounded product value * count, not by count
    successive additions. The product rounds once where the loop rounds count times, so it is at
    least as accurate (0.1 × 10 gives 1.0; the loop gives 0.9999999999999999), and the two are
    identical for count == 1. Making it bit-for-bit equal to a sequential loop would mean O(count)
    work inside observationLock, which would let a large batch stall a scrape into its timeout — to
    preserve a reproducibility that a striped DoubleAdder does not offer in the first place.

count == 0 is a no-op; negative throws IllegalArgumentException; NaN is ignored as in
observe(double); at most one exemplar is sampled per batch. The interface method is default (a
loop), so existing implementations of this @StableApi interface keep compiling and behave
correctly. Summary batches count and sum; with quantiles configured the CKMS sketch has no
weighted insert, so that part still costs one insert per observation.

Hot path

observe(double) and Buffer.append(double) keep their existing bodies — the only single-
observation code that changed is the cold maybeResetOrScaleDown call, which gains a 1L.
LongAdder.increment() is add(1L) and value * 1L is exact for every double, so the delegation
adds a multiply by a constant the JIT folds.

Verification

  • 187 tests pass (175 existing + 12 new), and each commit passes on its own.

    • BufferWeightedAppendTest pins the ticket protocol deterministically using Buffer's existing
      test seams, including the late-reader case: tickets claimed under generation A, generation read
      after B activated → B's expectedCount is the full weight and the batch is observed directly,
      never buffered.
    • BatchObserveTest compares batch against sequential across 13 values (0, ±1e-9, -2.25,
      1e300, ±∞, …) × 5 counts, across a forced native scale-down, and across a reset (the whole
      batch is re-applied). One test runs 4 threads × 2000 batches against 200 concurrent scrapes and
      asserts every snapshot's count is a multiple of the batch size — no torn batches — with
      count, classic total and native total all reconciling.
  • japicmp vs 1.9.0: additive only, Semantic versioning suggestion: 0.1.0.

  • JMH (4 threads, each op = 10240 observations):

    Benchmark ops/s observe calls/op vs. loop
    prometheusNativeLoop16 2,212 10,240 1.0×
    prometheusNativeBatch16 29,430 640 13.3×
    prometheusNativeLoop1024 2,528 10,240 1.0×
    prometheusNativeBatch1024 2,733,829 10 ≈1,080×
    prometheusNativeBatch1M 2,768,002 10 —

    Batch1024 and Batch1M both make ten calls per op and run at the same speed, which is the point:
    per-call cost is independent of count.

Caveats, honestly

  • Benchmarks are from a laptop (JDK 21, macOS), not a quiet host — error bars are wide and this
    machine is ~5× slower than the Ryzen 9 7900 the reference numbers in HistogramBenchmark come
    from. The hot-path A/B measured at parity over two interleaved rounds (single-thread classic
    −0.1%/−0.6%, native −2.1%/−1.6%, all inside error), but that comparison deserves a re-run on your
    benchmark host before you trust it.
  • I could not run mise run lint locally. Changed files are formatted with the pinned
    google-java-format 1.36.1 and verified clean, but checkstyle and the rest of flint have not run.
  • CONTRIBUTING suggests discussing involved changes on the mailing list first. Efficient Batch Observation Support for Native Histograms #1157 was closed as
    stale rather than by a maintainer decision, so there is no recorded objection to point at — happy
    to take this to prometheus-developers, or to close it if the API is not wanted.

Commits

  1. refactor(core): weighted tickets in Buffer (no public API change)
  2. feat(core): the observe(value, count) API
  3. perf: benchmarks

Possible follow-ups, none included here: observeWithExemplar(value, count, labels); a weighted
CKMS insert so Summary quantiles batch too; multi-value batches; and replacing the frexp loop in
findBucketIndex with Math.getExponent, which would speed up every native observation.

🤖 Generated with Claude Code

callumdonald96 and others added 3 commits September 29, 2026 00:00
Buffer counts one ticket per observation, and collection spins until the
metric's count adder matches the ticket total. That invariant makes a
batched observation ("this value occurred n times") impossible to add
without desynchronising collection permanently: the count adder would
overshoot the expected count and never match it again, so every
subsequent scrape of that data point would spin to its deadline and
throw.

Teach Buffer to count weighted tickets. append(value, weight) claims the
ticket range (count - weight, count] with a single atomic add, so a batch
cannot straddle a collector's activation -- it is either entirely inside
that collection's expected count or entirely outside it, never split.
The late-reader guard generalises from a point test to a range test and
is identical for weight 1. Buffered generations carry a lazily allocated
weights array, null while every entry has weight 1 so that scrapes
racing only single observations allocate exactly as before, and replay
passes the weight through a primitive WeightedObserver rather than
boxing each value into a Consumer<Double>.

Histogram.doObserve and Summary.doObserve gain a private multiplicity
parameter to match. No public API change: observe(double) and
append(double) take the same path they did before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Callum Donald <callumdonald96@protonmail.com>
Recording the same value n times required a loop over observe(double),
which repeated work that does not depend on the multiplicity: the buffer
ticket, the classic bucket scan, findBucketIndex, the bucket map lookup,
the schema-maintenance check and the exemplar-sampler call. Since all n
values are identical they land in exactly one classic bucket and one
native bucket, so the whole operation can be done once with a count
passed to the four adders.

Add observe(double value, long count) to DistributionDataPoint with a
default implementation that loops, so existing implementations of this
@stableAPI interface keep compiling and behave correctly. Histogram and
Summary override it: one bucket lookup, then add(count) on the bucket,
zero-count, sum and count accumulators. The cost does not depend on
count -- a batch of a million costs the same as a batch of sixteen.

Semantics: buckets and counts are exactly what count single calls
produce, and the batch is atomic with respect to scrapes, which a loop
never was. The sum is increased by the correctly rounded product
value * count, which is at least as accurate as count successive
additions and identical for count == 1; making it bit-for-bit equal to a
sequential loop would require O(count) work inside the collector's lock,
for a reproducibility that a striped DoubleAdder does not offer anyway.
count == 0 is a no-op, a negative count throws IllegalArgumentException,
and NaN is ignored as in observe(double). At most one exemplar is
sampled per batch.

Summary batches count and sum in constant time; with quantiles
configured the CKMS sketch has no weighted insert, so that part still
costs one insert per observation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Callum Donald <callumdonald96@protonmail.com>
LoopN records runs of identical values with N calls to observe(value),
BatchN with one observe(value, N); the ratio is the speedup. Batch1024
and Batch1M both make ten calls per op, so comparing them shows whether
the cost of a call depends on the count.

Also add prometheusNativeSingleThread, mirroring the existing
prometheusClassicSingleThread. Single-threaded numbers are much less
noisy than the four-thread ones and are what exposes a change in
per-call overhead on the native path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Callum Donald <callumdonald96@protonmail.com>
@callumdonald96

Copy link
Copy Markdown
Author

Closing — I opened this against the wrong base by mistake, it was meant for my own fork. Apologies for the noise.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Efficient Batch Observation Support for Native Histograms

1 participant