| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
CI Test ResultsRun: #32263379800 | Commit: deac396 | Duration: 28m 21s (longest job)
Status Overview
Legend: ✅ passed | ❌ failed | ⚪ skipped | 🚫 cancelled Summary: Total: 32 | Passed: 32 | Failed: 0 Updated: 2026-08-19 14:51:30 UTC |
Sorry, something went wrong.
Reliability & Chaos Results❌ 1 failure(s) detected Pipeline: https://gitlab.ddbuild.io/DataDog/java-profiler/-/pipelines/127431974 ❌ chaos: profiler gmalloc aarch64 21 0 3 temXchaoschaos.jar unavailable |
Sorry, something went wrong.
Benchmark Results (commit cfec595)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/123822759 Commit: cfec5956f438cc43b62cd2c81cbe46077ef66168 ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
…racking Addresses review findings from PR #644 (reference chains for surviving live-heap samples): FrontierTable's shared-lock mutation race, a single-byte JFR event size prefix that silently truncates chains longer than 255 bytes, blocking sample-lock retries in writeReferenceChain() now bounded by a shared per-batch deadline with a drop counter, unvalidated referencechains sub-option values, a startThread()/ pthread_kill() race publishing _running before the thread handle is initialized, a null-JNIEnv leak in threadLoop(), releaseSearchTags() now surfacing GetObjectsWithTags() failures so restartSearch() never resets tag state prematurely, resolveLoadedClasses() skipping its per-class scan only when the loaded-class count is unchanged (not just non-decreasing), a spurious COMPLETED state after a failed first-pass FollowReferences call, and removal of a leftover debug helper in ExternalProcessReferenceChainTest. Adds regression test coverage for release-failure and negative-value option paths. Verified via the full ddprof-lib gtestDebug suite (149/149 tasks, 57/57 referenceChains_ut tests). Environment: Datadog workspace Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Benchmark Results (commit 0714bd3)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124004132 Commit: 0714bd3166261f2710c22414dbfa68c11c699bdc ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit 700d838)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124275960 Commit: 700d838fc33298d5789fd9e1b29c8a142c11f41e ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
…racking Addresses review findings from PR #644 (reference chains for surviving live-heap samples): FrontierTable's shared-lock mutation race, a single-byte JFR event size prefix that silently truncates chains longer than 255 bytes, blocking sample-lock retries in writeReferenceChain() now bounded by a shared per-batch deadline with a drop counter, unvalidated referencechains sub-option values, a startThread()/ pthread_kill() race publishing _running before the thread handle is initialized, a null-JNIEnv leak in threadLoop(), releaseSearchTags() now surfacing GetObjectsWithTags() failures so restartSearch() never resets tag state prematurely, resolveLoadedClasses() skipping its per-class scan only when the loaded-class count is unchanged (not just non-decreasing), a spurious COMPLETED state after a failed first-pass FollowReferences call, and removal of a leftover debug helper in ExternalProcessReferenceChainTest. Adds regression test coverage for release-failure and negative-value option paths. Verified via the full ddprof-lib gtestDebug suite (149/149 tasks, 57/57 referenceChains_ut tests). Environment: Datadog workspace Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Benchmark Results (commit 9972bff)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124511782 Commit: 9972bff9464d4d2a8ee4a279dc8f8d35c88dd2e9 ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit a636398)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124700626 Commit: a636398f93c3fd15fb9fd3b6253da66f36e56742 ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit f2d8978)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124747648 Commit: f2d8978c23b568eb4ce8006e8e9b901cdaed11c3 ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit a48899b)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124773626 Commit: a48899b982ca5139081f665d6ffe1b8b5c3790fa ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit 787f29e)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124811024 Commit: 787f29e84914d8da397ed2a644c91d59a2f1e08e ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit fede524)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124833099 Commit: fede5244169eddd7f6280cd0a9fe39b618080bd7 ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit ef0f5c6)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124961539 Commit: ef0f5c6cd6793c440715d8f51a20218afa8cd188 ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit 1299316)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124967829 Commit: 1299316723d6feb8baf7e573b7925f816b772b79 ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
…racking Addresses review findings from PR #644 (reference chains for surviving live-heap samples): FrontierTable's shared-lock mutation race, a single-byte JFR event size prefix that silently truncates chains longer than 255 bytes, blocking sample-lock retries in writeReferenceChain() now bounded by a shared per-batch deadline with a drop counter, unvalidated referencechains sub-option values, a startThread()/ pthread_kill() race publishing _running before the thread handle is initialized, a null-JNIEnv leak in threadLoop(), releaseSearchTags() now surfacing GetObjectsWithTags() failures so restartSearch() never resets tag state prematurely, resolveLoadedClasses() skipping its per-class scan only when the loaded-class count is unchanged (not just non-decreasing), a spurious COMPLETED state after a failed first-pass FollowReferences call, and removal of a leftover debug helper in ExternalProcessReferenceChainTest. Adds regression test coverage for release-failure and negative-value option paths. Verified via the full ddprof-lib gtestDebug suite (149/149 tasks, 57/57 referenceChains_ut tests). Environment: Datadog workspace Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Benchmark Results (commit baba070)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124983919 Commit: baba0702198bb700cacfe9ab28c900eaf7d3f729 ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit 35c49c2)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/124997895 Commit: 35c49c2663a2e6e270a0287e4e059b8dab047faa ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit a9fb05f)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125008807 Commit: a9fb05fe90d68c1097e48893b1256661ecdfce30 ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit 2033006)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125072159 Commit: 20330063bf66bf581cf9ae47f94cff709e5aa6be ⚠️ Significant outliers
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Benchmark Results (commit 02a0a93)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125255579 Commit: 02a0a93ab9ef8ac64aa3c023b521fb75c90101c0 ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Sorry, something went wrong.
Bits has a CI fix ready🟢 Investigated · 🟢 Fix prepared · ⚪ Validation skipped · 🟠 Ready symbols_linux.cpp performed pointer arithmetic on a null ELF load base for non-PIE images while generating PLT and relocation addresses. The calculations now use integer address arithmetic before conversion back to pointers. View in Datadog | Reviewed commit ed6525a · Any feedback? Reach out in #deveng-pr-agent |
Sorry, something went wrong.
Logs why a klass fails the qualifying-growth check: INSUFFICIENT_FILL (ring not full enough), SLOPE_TOO_SMALL (mean growth below threshold), or floor_rising=false (minimum not rising). This helps diagnose why a leak with a clear upward population trend isn't being flagged as a candidate.
The thirds approach (comparing earliest third vs recent third of the ring) only uses 2/3 of the data and is noisy for oscillating-but-growing trends. A klass that grows steadily over 30+ GC epochs but oscillates within each 30-sample window can show a negative third-vs-third slope even though the overall trend is clearly up. Replace ringThirdsStats with a full-window least-squares linear regression (y = a + b*x, x = sample position 0..fill-1). This uses all samples and produces a proper slope that matches what a dashboard graph shows. Updated SecondsToOOMTest.RisingFloorProjectsExpectedSeconds: time_delta changed from 7s (thirds: 8-1) to 9s (full window: 9-0). Expected secondsToOOM changed from 10.0 to 9.0.
Add TEST_LOG to: - recordHeapFloorSample: logs used bytes + timestamp for each sample - onGC: logs when heap floor is recorded + GC epoch - secondsToOOM: logs ring fill, bytes_delta, time_delta, means, and why it returns -1 (insufficient fill, not rising, etc.) - heapFloorRising: logs mean/floor values + why it returns false - selectLeakCandidates: logs summary (candidate count, required hysteresis, heapFloorRising state) - hasLeakSignal: logs final verdict (secondsToOOM + candidates) This gives full visibility into why the leak detector isn't firing, in one deployment.
…or ring The heap floor ring shows fill=0 despite onGC recording a sample. Added TEST_LOG to recordHeapFloorSample to see whether it's entered and whether _heap_floor_recording_disabled_for_test is incorrectly true on the hotdog.
On JDK 26 with generational ZGC, _used_at_last_gc (read from CollectedHeap via VMStructs) reports the committed heap size, not the used bytes. The heap floor would appear flat (7.5GB constant) even as the real used bytes grow (2.19GB and rising). Fix: in resolvePostGcHeapUsage(), when UseZGC is enabled and hotspot_version >= 26, fall back to HeapUsage::get(false)._used (the JMX path) instead of _used_at_last_gc. Scoped to ZGC on JDK 26+ only — other GC/JDK combinations keep the existing _used_at_last_gc path.
VMFlag::find is a VMStructs table walk — don't call it on every GC. Cache the result (UseZGC && hotspot_version >= 26) as _is_zgc_jdk26_plus in initialize(), read it in resolvePostGcHeapUsage().
LEAK_GROWTH_ABS_MIN was 5 (minimum 5 instances/epoch slope). For large-population klasses (e.g. byte[] with thousands of instances), a slope of 0.5 instances/epoch is significant (~5% growth) but far below 5. The absolute floor of 5 was too high, rejecting real leaks in large-population klasses. Lowered to 1 so only klasses with very small populations (< 7 instances) need an absolute floor. This still rejects single-instance oscillation noise while allowing real leaks in large-population klasses. Also: cache ZGC JDK26+ check in initialize() with debug logging.
selectLeakCandidates() was reading cached_slope directly without calling hasQualifyingGrowth(), so consecutive_positive was not updated for the current scan. The cached_slope could be stale (set by a previous push with a negative slope), causing selectLeakCandidates to skip the klass even though the current trend is positive. Fix: call hasQualifyingGrowth(entry) in selectLeakCandidates() so consecutive_positive is updated and the slope is fresh. Made hasQualifyingGrowth() take const ref (cached_slope is mutable).
Instead of pre-tagging candidate objects with negative marker tags (MARKER_TAG_BASE - i), match by class_tag in heapReferenceCallback(). This avoids SetTag entirely: the walk visits every object normally, and when it encounters one whose class matches a candidate klass ID, it records the chain link and skips expanding that object's children. - _candidate_tags[i] now holds the klass ID (not marker tag) - _candidate_frontier_tags[i] holds the frontier tag assigned by nextTag() when pruning - buildCanaryChainEvent() uses _candidate_frontier_tags[i] as both the frontier table key and the target_tag - pollWatchedTargets() canary path checks tag > 0 (frontier tag from heapReferenceCallback) instead of tag <= MARKER_TAG_BASE Tests need updating: PollWatchedTargetsTest expects the old marker-tag approach. 3 tests fail.
Tests now set _candidate_frontier_tags via setCandidateFrontierTagForTest() so buildCanaryChainEvent() can reconstruct the chain. Added setCandidateFrontierTagForTest() test accessor.
… shouldRunPass MIN_EFFECTIVE_BUDGET was 50, which the PID controller used as a floor when throttling. For canary search, 50 edges/pass is too low — the walk never reaches deep candidates. Raised to 500 so the PID controller can let each pass explore more edges while still bounding STW pause. Also moved the cadence sleep to after shouldRunPass() so passes run back-to-back when canary search is active (no sleep when a pass is about to run). Updated PacingGrowsBudgetBackAndRelaxesCadenceWhenUnderCeiling test to use setEffectiveBudget(600) (above the new floor).
Canary-matched candidates were never getting a JVMTI tag set, so buildCanaryChainEvent() could never see them as found. Also add a TEST_LOG line to trace klass_id -> class name resolution.
resolveClassMap only logs a class the first time it's tagged, so a class already loaded before referenceChains' own scan first ran never gets logged. Resolve the name live off the candidate's representative object in pollWatchedTargets instead.
Class-tag matching let heapReferenceCallback record a chain for any instance of a candidate's class, not the specific representative LivenessTracker flagged as growing - for common classes (e.g. byte[]) this meant the recorded chain almost never belonged to the actual leak. Revert to pre-tagging each candidate's representative with a distinct marker tag and matching by identity. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Update pacing tests' configured budget/iteration count to match. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ents Integer division in autoTuneDefaults() truncated the budget-to-cap ratio, undershooting the frontier cap. Comments describing FRONTIER_CAP_HIT as abandoning the search outright didn't match runPass()'s actual behavior.
Compares the JVM heap limit against the container/cgroup memory limit (unavailable treated as unbounded) and runs the trend extrapolation once against whichever boundary is tighter, so container-memory-driven OOM kills are projected even when heap occupancy alone looks safe.
…udget for non-safepoint pass cost Split pass timing into safepoint_ticks (real JVMTI/FollowReferences duration, fed to updatePacing()) and non_safepoint_ticks (bookkeeping/dispatch, spent into a new independent _cpu_pain_budget gating shouldRunPass()). Renamed _pain_budget to _safepoint_pain_budget for clarity. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ling forever Once the frontier table saturates, no new entries can ever be admitted, so the no-progress detector can never actually fire -- the frontier-cap branch matched every subsequent pass, permanently blocking abandonment. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Replaces the binary 5ms/50ms pause-target step with a continuous exponential ramp (5ms->100ms, cadence 1s->10ms) as secondsToOOM() falls within OOM_RAMP_START_S, so the search spends progressively more STW budget the closer the app gets to dying. Budget bump is held for the whole ramp window rather than a one-shot at the old threshold. OOM_RAMP_START_S is temporarily set to 6000s (100 min, TESTING ONLY) to observe ramp behavior live in the hotdog pod; must be reverted to 1800.0 before merging.
Guards against a one-time step change (e.g. cache prefill) that plateaus immediately: the full-window regression alone stays "rising" until the step's samples age out of the ring, but a second fit over just the most recent half collapses to flat as soon as growth stops.
cadence_ns was only used for the loop's own sleep calls and never written back to _effective_cadence_ns, so shouldRunPass()'s cadence gate stayed pinned at the 1s baseline during urgency instead of ramping down.
Without it, the sweep always rescanned from tag 1, so a large permanent population of low-tag EXPANDED entries could fill its per-pass cap forever, starving any higher-tag entry (e.g. a static field's collection admitted after startup) from ever being re-queued. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…latch, completion gating, generation sync Rediscovering new elements in an already-discovered static-field collection was failing end-to-end. Root causes fixed together (verified via ExternalProcessReferenceChainTest): - classMap dictionary-id instability -> shared ClassTagAllocator with stable class_tag matching, replacing referrer_klass - klassPopulationSetRepresentativeForTest() test seam bypassed stable class-tag minting -> routed through mintStableClassTagIfNeeded() - isUrgent()/hasLeakSignal() flapped near the OOM threshold, causing a search restart almost every tick -> latch + release hysteresis - runPass() could reach terminal SearchState::COMPLETED while a leak klass was still watched -> gated on _watched_leak_klass_count == 0 - LivenessTracker::_last_class_map_generation wasn't synced in initialize(), spuriously wiping _klass_population on the first cleanup_table() call -> synced to the real classMap generation Also reverts OOM_RAMP_START_S from its temporary 6000.0 (100 min, used to observe ramp behavior live) back to the production 1800.0 (30 min). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds Component 4 to the reference-chains design summary, describing the growing-collection fix set from 209336e. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rgent-independent stuck detector pollWatchedTargets() indexed candidate arrays by loop position instead of the tag-decoded slot, so canary chains were never resolved. Pre-tagging was also one-shot, letting the tracked candidate set desync from selectLeakCandidates(). Additionally add a canary-specific no-progress abandon path that isn't suppressed by isUrgent(), so a livelocked urgent search can still terminate.
…ount Queue abandoned-search events at the moment of abandon so dump() no longer races restartSearch()'s state reset. Escalate the non-safepoint pain budget while a canary search has unresolved candidates, and drop a leftover unconditional sleep in threadLoop() that blocked back-to-back canary passes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
| Back | FazBrowse Home | New Git URL |
What does this PR do?:
Implements reference-chain reconstruction for live-heap samples that survive
past their allocation window (PROF-15341). A ReferenceChainTracker runs a
dedicated BFS thread that tags reachable objects with JVMTI object tags,
walks the heap incrementally across GC epochs, and records the referrer-type
chain back toward a GC root (labelled with its GC-root kind). It is bridged
to LivenessTracker::selectLeakCandidates()'s population-slope leak
detection: when a still-live sampled object starts to look leak-shaped, its
chain is reconstructed from the frontier table and emitted to JFR.
flowchart TD LT["LivenessTracker::selectLeakCandidates (population-slope ranking)"] -->|"ranked klass candidates"| PWT["ReferenceChainTracker::pollWatchedTargets"] BFS["BFS thread: threadLoop"] -->|"shouldRunPass gate: GC epoch advanced or cadence elapsed"| RP["runPass"] RP -->|"every pass"| MW["runPassManualWalk: IterateOverReachableObjects seeds roots"] MW --> EF["expandFrontier: batched array-holder FollowReferences"] EF --> FT["FrontierTable: FRONTIER / EXPANDED / EDGE / ABANDONED"] FT -->|"tag already set on a leak-candidate instance"| PWT PWT -->|"buildChainEvent"| RC["cacheResolvedChain: one entry per klass id, cap 256"] RC -->|"Profiler::dump, snapshot without clearing"| DR["drainPendingChainEvents"] DR --> JFR["datadog.ReferenceChain / datadog.ReferenceChainAbandoned"]Heap-walk mechanism (the main design decision in this branch): every pass
takes a bounded manual walk (runPassManualWalk()) — pure JVMTI
(IterateOverReachableObjects for root enumeration + expandFrontier() for
incremental frontier expansion), resumable across passes, and bounded so its
safepoint pauses stay small on every collector, including ZGC. The walk
issues only JVMTI heap calls, which run inside the VM_HeapWalkOperation
safepoint and honor ZGC's load barriers, so concurrent relocation cannot
corrupt it — it reads no raw oop.
expandFrontier()'s batching, concretely: instead of one FollowReferences
call per frontier entry, up to budget pending tags are resolved in one
GetObjectsWithTags call, packed into a single JNI array (holder), and
expanded via exactly one FollowReferences(initial_object = holder) call —
so one BFS level is discovered per VM-safepoint operation rather than one
safepoint per entry:
sequenceDiagram participant BFS as BFS thread participant JVMTI as JVMTI participant JNI as JNI holder array participant CB as heapReferenceCallback BFS->>BFS: pull up to budget tags from front of _pending_expand BFS->>JVMTI: GetObjectsWithTags, resolve which tags are still live JVMTI-->>BFS: live jobject references BFS->>JNI: EnsureLocalCapacity, then NewObjectArray to build holder alt exception, EnsureLocalCapacity failure, or null holder BFS->>BFS: ctx.truncated = true, retry this batch next pass else holder built successfully BFS->>JNI: SetObjectArrayElement per resolved object BFS->>JVMTI: FollowReferences, initial_object = holder JVMTI->>CB: heapReferenceCallback per outgoing edge CB-->>JVMTI: descend only for batch_tags boundary objects JVMTI-->>BFS: one BFS hop expanded for the whole batch BFS->>BFS: markExpanded, admitObject appends children to _pending_expand endexpandFrontier()'s JNI/JVMTI error handling is defensive by construction: a
null holder, a pending JNI exception after NewObjectArray/
SetObjectArrayElement, or an EnsureLocalCapacity failure all set
ctx.truncated = true (retry next pass) instead of marking the batch
permanently EXPANDED; a failed java/lang/Object class resolution with
pending work also forces truncated = true, so runPass() cannot mistake
it for SearchState::COMPLETED.
Termination and pacing — an unbounded traversal could otherwise stall a
GC safepoint or run forever:
stateDiagram-v2 direction LR [*] --> RUNNING RUNNING --> COMPLETED: frontier drained,<br/>no truncation this pass,<br/>no leak klass still watched RUNNING --> ABANDONED: frontier-size cap hit,<br/>or wall-clock TTL exceeded<br/>with work still pending COMPLETED --> RUNNING: restartSearch ABANDONED --> RUNNING: restartSearch note right of ABANDONED SearchAbandonReason (FRONTIER_CAP/TTL) records which cutoff fired, for the T_REFERENCE_CHAIN_ABANDONED JFR event end note note right of RUNNING restartSearch only fires once a leak candidate is seen and PainBudget allows it end notetreated as immediate search abandonment, not per-pass truncation) —
configured via referencechains=true:hops=N:budget=N:ttl=N:framecap=N
(plus firstpassbudget=N for the first pass's own budget). Negative
hops/budget/framecap values are floored (an unfloored negative
hops would otherwise wrap to ~4e9 as a u32, silently disabling the cap
it's meant to enforce); all three are also ceiling-clamped.
adapts the effective per-pass budget and cadence toward a configured
safepoint-pause SLO.
restarted walk is very likely to still find the object) plus a
PainBudget cooldown (painbudget=N, 0-100) — a leaky bucket over
cumulative safepoint time spent, so a cheap search can restart sooner than
an expensive one. This closes a structural gap where a one-shot walk could
finish before population-trend detection had accumulated enough GC epochs
to flag a candidate, leaving anything allocated afterward permanently
undiscoverable.
a klass's population growth to clear both a magnitude bar and a floor-rise
bar for LEAK_TREND_HYSTERESIS_BASE (5) consecutive qualifying epochs —
lowered to LEAK_TREND_HYSTERESIS_CORROBORATED (3) when the aggregate
post-GC heap floor is also rising — rather than trusting a single epoch's
positive slope. This closes a false-positive gap where an oscillating
("see-saw") population could otherwise trigger a search with no real
longer-term growth.
JFR persistence across dumps: a resolved chain is cached per klass id
(_resolved_chains, capped at 256 entries, drop-not-evict once full — see
REFERENCE_CHAIN_EVENTS_DROPPED) and re-stamped into every subsequent dump
the sample survives into (drainPendingChainEvents() snapshots without
clearing, mirroring how LivenessTracker re-emits live-object samples), so a
long-lived leak's chain is present in each JFR chunk rather than only the
chunk active when it was first reconstructed. Verified end-to-end in
ExternalProcessReferenceChainTest.
Rediscovering growth in an already-visited container: the frontier walk
above visits each object once. That is insufficient for the actual leak shape
this feature targets in production — a static final collection field that
is appended to, not reassigned — because the container is already
EXPANDED long before its element klass earns a leak signal, and a one-time
visit never sees elements added afterward. runPass()'s completion branch
now also requires _watched_leak_klass_count == 0 (no klass currently under
active leak watch) before moving to SearchState::COMPLETED; while a klass
is watched, the search stays RUNNING and a two-tier rotation mechanism
re-queues the container holding the watched klass's elements for
re-expansion, so newly appended elements are picked up on a later pass:
flowchart TD LT2["LivenessTracker::topKlassesByGenerationCount"] -->|"refreshed once per tick,<br/>only after hasLeakSignal() fires"| WK["_watched_leak_klass_ids<br/>(max 5)"] WK -->|"klass_id newly watched"| SEED["seedLeakAccumulationForNewlyWatchedKlass:<br/>one-time scan of already-EXPANDED<br/>frontier entries by class_tag"] ADM["admitObject: ADMITTED"] -->|"class_tag of new object"| TLA["trackLeakAccumulation"] SEED --> TLA WK -->|"class_tag match?"| TLA TLA --> T1["Tier 1: _leak_signature_totals<br/>(leaf_klass_id, parent_class_id) -> count"] TLA --> T2["Tier 2: _leak_parent_fanout<br/>parent_tag -> count, within winning signature"] T1 -->|"delta vs previous pass's snapshot"| RANK["collectLeakAccumulationCandidatesForRotation:<br/>pick winning signature, then its top parent_tag(s)"] T2 --> RANK RANK -->|"re-queue for re-expansion,<br/>budget 16/pass"| EF2["expandFrontier"] EF2 -->|"new elements admitted"| ADMMatching a newly-admitted object against a watched klass id uses a stable
JVMTI class tag (classTagAllocator.h, shared between
ReferenceChainTracker and LivenessTracker) rather than the classMap
dictionary id, since that id is not guaranteed stable if the dictionary is
compacted/regenerated mid-search.
A related fix latches isUrgent()'s seconds-to-OOM projection, which
otherwise flaps by orders of magnitude between consecutive readings of the
same growing heap (observed in one run: 128s, then 52769s, then back) — each
flap back to "urgent" used to bypass the per-klass hysteresis gate and
restart the search, discarding the Tier 1/Tier 2 state above before it could
converge:
stateDiagram-v2 direction LR [*] --> NOT_URGENT NOT_URGENT --> LATCHED: secondsToOOM() < OOM_URGENT_THRESHOLD_S LATCHED --> LATCHED: still below OOM_URGENT_RELEASE_S,<br/>or below release-consecutive count LATCHED --> NOT_URGENT: at/above OOM_URGENT_RELEASE_S for<br/>URGENT_RELEASE_CONSECUTIVE (5) observations note right of LATCHED _urgent_search_spent authorizes at most one restart per latched episode, independent of the per-klass hysteresis gate end noteSee doc/reference-chains-collection-summary.md's Component 4 for the full
mechanism.
Motivation:
PROF-15341 — give live-heap samples that survive long enough to look
leak-shaped an actual referrer chain, not just "this object is still alive",
without regressing safepoint-pause behaviour on low-pause collectors.
Additional Notes:
See doc/reference-chains-design.md and doc/reference-chains-collection-summary.md
for the full design and leak-detection mechanism. Per-pass
budget/hop-cap/TTL/frontier-cap/pause-target/pain-budget defaults are round,
unbenchmarked placeholders pending a measurement pass — no JMH harness for
this feature exists yet, though utils/ has repro, parameter-sweep, and
JFR-report shell/Python tooling used to characterise pause behaviour ad hoc.
Earlier planning/proposal docs (implementation plan, remaining-work plan,
benchmark plan, and an alternative VMStructs-walk design) described approaches
or work that were superseded or shipped differently than planned, so they were
dropped from this PR (kept locally, not committed) rather than left to drift
further from the code.
How to test the change?:
ReferenceChainsTest/ReferenceChainsBfsTest/ReferenceChainsTagTest/
FrontierTableTest/PollWatchedTargetsTest/ResolvedChainCacheTest,
PainBudgetTest, SearchRestartTest (restart gate),
ReferenceChainJfrRoundtripTest (JFR encode/decode), ArgumentsTest
(referencechains= sub-option parsing, including hops/budget/framecap
clamping behaviour), and the LivenessTracker/SelectLeakCandidates
bridging tests.
genuinely separate-process end-to-end test: runs a real leaking-cache Java
app in a child JVM, asserts the reconstructed chain's leaf class, and
confirms the chain re-emits into three fresh JFR dumps (across-dumps
persistence). Includes
shouldReconstructReferrerChainForGrowingStaticFieldCollection()
(StaticFieldGrowingCollectionScenario), which appends new elements to an
already-discovered static final collection after its first expansion and
asserts the chain is still reconstructed for one of the later elements.
in-process coverage for the walk engine, target-selection bridging, and
abandonment reporting.
For Datadog employees:
credentials of any kind, I've requested a security review (run the dd:platform-security-review
skill, or file a request via the PSEC review form).
bewaire also runs automatically on every PR.
Unsure? Have a question? Request a review!
🤖 Generated with Claude Code