Skip to content

Memory Leak Switching Activities #12253

Description

@dodgei

Android framework version

net11.0-android (Preview)

Affected platform version

Visual Studio 2026 v18.7.3

Description

Having experienced the same issue as #10989 I tested the fix in .NET11 Preview 6 but it still seems to be present.

Created a small app which auto switches between 2 activities the native memory increases until the application crashes. Memory viewed in Android Studio Live Telemetry.

Image

Steps to Reproduce

Install application on an Android device from https://github.com/dodgei/android_leak/tree/main/LeakTest

Witness the climbing native memory in a profiler.

Did you find any workaround?

I found that calling GC.Collect(); in OnDestroy(); or OnCreate(); keeps the memory stable

Relevant log output

Activity

jonathanpeppers commented on Jul 28, 2026

@jonathanpeppers
Member

Thanks for the small, self-contained repro — that made this quick to measure.

I profiled it on a physical Pixel 5 with .NET 11 Preview 6 (Microsoft.Android.Sdk 37.0.0-preview.6.59), on both CoreCLR and Mono. There is no leak here. Nothing is retained — but the app will still OOM as written, and I want to explain exactly why, because the distinction matters for what we do about it.

The measurement

As written, over 9 minutes / ~1,000 activity transitions:

minute Activities Views JNI grefs GC.CollectionCount(0) Native heap
1 113 1,582 246 0 96 MB
3 341 4,774 702 0 149 MB
5 570 7,980 1,158 0 196 MB
7 800 11,200 1,618 0 243 MB
9 1,030 14,420 2,082 0 290 MB

Dead linear — no plateau, no inflection. The important column is GC.CollectionCount(0): not a single garbage collection ran in over 1,000 transitions. The managed heap only ever reached ~3 MB.

Each transition allocates roughly 4 grefs and ~6 KB of managed memory, while pinning an entire Java Activity plus its view hierarchy and the native memory behind it. The GC decides when to run based on managed bytes allocated, and 6 KB/transition never comes close to the gen0 budget.

With a deterministic collection

Adding GC.Collect() in OnDestroy, over 448 transitions / 8 minutes on CoreCLR:

minute transitions JNI grefs Activities Views Native heap
1 55 19 5 70 69.4 MB
3 167 19 4 56 69.7 MB
5 279 19 4 56 70.0 MB
8 448 19 5 70 70.4 MB

Perfectly flat — GC.GetTotalMemory returned the identical byte count at every checkpoint. Mono is the same story (grefs steady at 22-24, Activities 2-4, native heap 15.5 MB → 16.5 MB over 431 transitions).

Every peer and every global reference is fully reclaimable the instant a GC runs. So this is a GC scheduling problem, not a retention bug.

I opened dodgei/android_leak#1 against your repro showing the change, with the full numbers.

Why this is not the fix from #11134

Your instinct to test against that fix was reasonable, but it does not apply. #11134 (and the equivalent in dotnet/java-interop/main for .NET 11) fixed a genuine global-reference leak in ConstructPeer / Dispose for Invalid references, which was hit when LayoutInflater.Inflate inflated custom views. Your repro inflates only stock RelativeLayout and TextView, so that code path is never entered. Both bugs produce "grefs go up and memory climbs", which is why they look alike.

What we should do about it

The real gap is that the GC budget is denominated in the wrong unit. A managed peer is a few dozen bytes but pins hundreds of KB of Java heap and native memory, and no managed-byte budget — at any size — can see that.

There is an existing safety valve that forces a collection at 90% of the max gref count, but on a physical device that is 46,080 grefs. At ~4 grefs per transition this app would need over 11,000 transitions to trip it, by which point it would need multiple GB of native memory. It is roughly two orders of magnitude too high to ever help here.

GC.AddMemoryPressure / RemoveMemoryPressure on peer construction and disposal is the most promising direction — it is precisely the API for "this small managed object owns a large unmanaged allocation", and it would let the GC account for the Java and native memory a peer pins. It needs care (the pressure has to be released correctly on dispose and on finalization, and mis-estimating the size can cause over-collection), so I want to prototype and measure it rather than turn it on blindly. A much lower gref-delta-based trigger is a cheaper interim option.

I will retitle and repurpose this issue to track that work.

In the meantime, the GC.Collect() workaround you found is correct and not just papering over a leak — it is genuinely reclaiming everything. I would put it after tearing down a large view hierarchy rather than on every transition, since a collection is not free.

jonathanpeppers commented on Jul 28, 2026

@jonathanpeppers
Member

FYI above is initial investigation from copilot, I'm having it drill into these better:

Image

Seems like it's growing by 0.4 each time, "perfectly flat"? LOL!

jonathanpeppers commented on Jul 28, 2026

@jonathanpeppers
Member

Full investigation summary

Profiled on a physical Pixel 5 (Android 14, arm64), .NET 11 Preview 6, Microsoft.Android.Sdk 37.0.0-preview.6.59, using the repro from this issue. Tested both CoreCLR (the .NET 11 default) and Mono (-p:UseMonoRuntime=true -p:_DisableCheckForUnsupportedMonoMobileRuntime=true).


1. There is no leak

As written, no explicit GC — CoreCLR, 9 minutes

minute Activities Views JNI grefs GC.CollectionCount(0) Native heap
1 113 1,582 246 0 96 MB
3 341 4,774 702 0 149 MB
5 570 7,980 1,158 0 196 MB
7 800 11,200 1,618 0 243 MB
9 1,030 14,420 2,082 0 290 MB

Dead linear, no plateau. GC.CollectionCount(0) is 0 the entire time — not a single collection in over 1,000 activity transitions. Managed heap peaked at ~3 MB. Each transition costs ~4 grefs and ~6 KB managed, while pinning a whole Java Activity + view hierarchy.

Mono is identical: 122 transitions, gc0=0, grefs 14 -> 506, 244 Activities, 3,416 Views.

With GC.Collect() in OnDestroy — CoreCLR, 448 transitions / 8 min

minute transitions JNI grefs Activities Views GC.GetTotalMemory
1 55 19 5 70 80,840
3 167 19 4 56 80,840
5 279 19 4 56 80,840
8 448 19 5 70 80,840

Flat, to the byte. Mono equivalent (431 transitions): grefs steady 22-24, Activities 2-4, GC.GetTotalMemory oscillating in noise around 4,314,7xx.

Every peer and gref is fully reclaimable the instant a GC runs. Nothing is retained.

This is also not the leak fixed by #11134 / the ConstructPeer + Dispose-on-Invalid-ref change in dotnet/java-interop. That one fired when LayoutInflater.Inflate inflated custom views; this repro inflates only stock RelativeLayout/TextView, so that path is never entered.

Sample PR showing the workaround: dodgei/android_leak#1


2. The residual native heap growth is Android's, not ours

With the explicit GC in place, native heap still crept ~1.2 KB per transition. I chased this down because it looked like a second, smaller leak.

It is not fragmentation (HeapFree pinned at ~9.8 MB while HeapAlloc climbed) and not JIT warm-up (linear for 28 min, no plateau). It scales with transitions, not wall-clock:

Task.Delay transitions HeapAlloc delta per transition
500 ms 679 +871 KB 1.28 KB
2000 ms 356 +369 KB 1.04 KB

Removing SetContentView entirely did not help (1.48 KB/transition), so it is not proportional to peer count.

Root-caused with Perfetto heapprofd (works on production builds for debuggable apps; malloc_debug needs root, and this device plus both installed emulator images are non-rootable). Top retained libc.malloc callsites:

net bytes stack
62 KB Activity.setContentView
53 KB RenderThread -> EglManager::createSurface -> eglCreateWindowSurface -> libGLESv2_adreno.so
47 KB RenderThread -> CanvasContext::create (operator new, libhwui)
41 KB setContentView -> PhoneWindow.installDecor -> DecorView.onResourcesLoaded -> LayoutInflater.inflate -> ActionBarContextView.<init> -> NinePatchDrawable.inflate -> ImageDecoder -> Bitmap::allocateHeapBitmap

Filtering every retained callstack for any frame in libmonodroid.so, libcoreclr.so, libmonosgen-2.0.so, libclrjit.so, libSystem.Native.so, or libxamarin-app.so returns zero rows. Stacks enter via art::JNI::CallNonvirtualVoidMethodA (our managed->Java setContentView call); everything past that is hwui / EGL / Adreno / decoded nine-patch bitmaps.

At ~113 transitions/min that is ~140 KB/min, versus +23 MB/min for the gref issue — about 170x smaller. Android's per-Activity native cost, not ours.


3. What I think the actual product problem is

An app reached 290 MB and was heading for an OOM kill while its managed heap sat under 3 MB and the GC never ran. That is not ordinary GC non-determinism — "you don't know when" is fine, "never, while the process dies" is not. Requiring app authors to know where to put GC.Collect() pushes JNI peer/gref internals into user code.

Two concrete gaps:

3a. The gref safety valve is calibrated ~2 orders of magnitude too high

AndroidRuntime.CreateGlobalReference already forces a collection when grefs pile up:

src/Mono.Android/Android.Runtime/AndroidRuntime.cs:237-241

if (gc >= JNIEnvInit.gref_gc_threshold) {
    Logger.Log (LogLevel.Warn, "monodroid-gc", gc + " outstanding GREFs. Performing a full GC!");
    System.GC.WaitForPendingFinalizers ();
    System.GC.Collect ();
}

The threshold comes from:

  • src/Mono.Android/Android.Runtime/JNIEnvInit.cs:46 — internal static int gref_gc_threshold, set at line 207 from args.grefGcThreshold, or lines 88-90 for NativeAOT
  • src/native/mono/runtime-base/android-system.cc:406-439 — AndroidSystem::get_max_gref_count_from_system() returns 2000 in an emulator, 51200 on a device
  • src/native/clr/runtime-base/android-system-shared.cc:53-62 — same values for CoreCLR
  • src/native/mono/runtime-base/android-system.cc:441-447 and src/native/clr/include/runtime-base/android-system.hh:44-50 — get_gref_gc_threshold() takes 90% of that
  • Plumbed through src/native/clr/host/host.cc:437, src/native/mono/monodroid/monodroid-glue.cc:834, src/native/nativeaot/host/host.cc:89

So on a real device the valve fires at 46,080 grefs. At ~4 grefs/transition this repro needs 11,520 transitions to trip it — by which point it would need multiple GB of native memory. It dies thousands of transitions earlier. The mechanism is right; the number never fires in practice.

A gref-delta-based trigger (collect after N grefs created since the last collection, N in the low thousands) would be a small, targeted change in CreateGlobalReference and is by far the cheapest option.

3b. onTrimMemory / onLowMemory are NOT a viable trigger (tested — do not build on this)

I originally proposed hooking onTrimMemory here. I then tested it, and it does not work. Leaving the correction in place so nobody implements it.

It is true that these appear in src/ only as public bindings for user code (Android.Content.IComponentCallbacks2.OnTrimMemory, IComponentCallbacks.OnLowMemory) with no hook in src/native/, src/java-runtime/, or Android.Runtime. But wiring one up would not have helped this issue at all.

I added an [Application] subclass plus Activity overrides of both callbacks and ran the leaking repro for 12 minutes:

minute OnTrimMemory calls OnLowMemory calls Activities grefs Native heap MemoryInfo.AvailMem LowMemory
2 0 0 226 429 124 MB 4001 MB False
4 0 0 455 939 180 MB 3798 MB False
8 0 0 915 1,857 279 MB 3654 MB False
12 0 0 1,376 2,775 375 MB 3686 MB False

Zero callbacks, at 375 MB native heap with 1,376 leaked Activities.

Two controls confirm the overrides themselves are wired up correctly, so this is a real negative and not a binding bug. Pressing HOME:

Application.OnTrimMemory level=Background (40) grefs=2883 gc0=0

And adb shell am send-trim-memory <pid> <level>:

Application.OnTrimMemory level=RunningModerate (5)   Activity1.OnTrimMemory level=RunningModerate (5)
Application.OnTrimMemory level=RunningLow (10)       Activity1.OnTrimMemory level=RunningLow (10)
Application.OnTrimMemory level=RunningCritical (15)

Why it never fires: onTrimMemory is a system-wide pressure signal, not a per-process one. Throughout the run ActivityManager.MemoryInfo reported ~3.6 GB available against a 216 MB threshold with LowMemory=False. A single app can grow until it hits its per-process dalvik.vm.heapgrowthlimit and gets OOM-killed while an 8 GB device is nowhere near system pressure, so ART never broadcasts a trim. That is exactly this repro's shape, and it is the common case: one leaky app on an otherwise healthy device.

Consequence: 3a is the only proposal that addresses the reported problem. The trigger has to come from something we control — grefs created since the last collection — because Android will not tell us.

A trim hook is still worth adding as a secondary, opportunistic trigger: note grefs=2883 gc0=0 at the HOME press above, i.e. we currently ignore a genuine invitation to free memory. But it is a nice-to-have, not a fix for this issue, and it must not be the primary mechanism.

3c. On GC.AddMemoryPressure

Worth prototyping in JniRuntime.JniValueManager.ConstructPeer / peer disposal, but I would not start here. Per-peer footprint is wildly variable — an Activity pins megabytes, a String peer pins bytes — and a bad estimate risks GC thrashing. It also needs the pressure released correctly on both Dispose and finalization or it degrades over time. Try 3a first and measure (3b is a dead end — see above).


4. Open question I could not answer

Several reports here and in #10989 claim .NET 8/9 were fine and .NET 10/11 are not (see the .NET 8 vs .NET 10 comparison in #10989, and @Cheesebaron's "we cannot reproduce on .NET 9"). I only tested .NET 11 P6, so I can neither confirm nor refute this.

This matters a lot for severity. If there is a real regression, something changed that used to cause collections to happen, and that is more urgent than the general ergonomics gap — it deserves a bisect across .NET 8 / 9 / 10 / 11 using this repro and GC.CollectionCount(0) as the metric, which is unambiguous and takes about two minutes per SDK to measure.


Repro recipe for whoever picks this up

dotnet build -c Debug -f net11.0-android -t:Install
adb shell am start -n com.companyname.LeakTest/crc648ee606ad49166d9e.Activity1
adb shell dumpsys meminfo com.companyname.LeakTest   # watch Activities:, Views:, Native Heap

Log Java.Interop.Runtime.GlobalReferenceCount and GC.CollectionCount(0) per transition. The Activities: and Views: counters in dumpsys meminfo are the clearest signal; GC.CollectionCount(0) staying at 0 is the smoking gun.

jonathanpeppers commented on Jul 29, 2026

@jonathanpeppers
Member

@dodgei I didn't come up with a concrete thing to do here, except this:

Running the sample for a long time, the GC never triggered on its own. This behavior is the same as Mono was, though. I do think it is reasonable to just call GC.Collect() in certain activity's OnDestroy().

Long term, we should investigate calling GC.Add/Remove pressure APIs for Java objects.

Stensan commented on Sep 23, 2026

@Stensan

We can confirm this on .NET 10 (net10.0-android36.0, Android workload 36.1.69, Mono runtime, arm64) in a production app with many Activities.

Production symptom: Zebra TC52 (Android 11). The process ends with ApplicationExitInfo.REASON_ANR after ~80 Activity starts, with PSS of 500–765 MB. This happens regardless of uptime (3 hours or 3 days). A watchdog thread that normally reports a blocked UI thread after 5 s logged nothing, so the whole process appears to be frozen (long stop-the-world GC?).

Repro on a Zebra TC53 (Android 14): open and close the same Activity (~530 views) 100 times, then read dumpsys meminfo:

Activity starts Activities Views PSS
0 9 2,419 395 MB
40 48 23,313 604 MB
50 31 14,850 533 MB
100 56 28,125 670 MB (swap 124 MB)

Destroyed Activities are only collected occasionally and partially, and the lower bound keeps rising.

Workaround: with GC.Collect() after leaving an Activity, the numbers stay flat at 6 Activities, 1,846 views and ~375 MB for 100 iterations.

This was measured with a Debug build. The production ANRs were on a Release build.

Is a fix for .NET 10 servicing planned, e.g. a backport of #11112?

jonathanpeppers commented on Sep 23, 2026

@jonathanpeppers
Member

#11112 is already backported and available in 36.1.69:

The title is slightly different, because the fix was in the now archived dotnet/java-interop repo.

Note above, "I didn't come up with a concrete thing to do here".

Stensan commented on Sep 24, 2026

@Stensan

Thanks, that matches what we see, the GC.Collect() workaround is in place

jonathanpeppers commented on Sep 24, 2026

@jonathanpeppers
Member

.NET 11 is now in RC, it uses CoreCLR by default but we should check if it has the same behavior.

dodgei commented on Sep 24, 2026

@dodgei
Author

@Stensan

We can confirm this on .NET 10 (net10.0-android36.0, Android workload 36.1.69, Mono runtime, arm64) in a production app with many Activities.

Production symptom: Zebra TC52 (Android 11). The process ends with ApplicationExitInfo.REASON_ANR after ~80 Activity starts, with PSS of 500–765 MB. This happens regardless of uptime (3 hours or 3 days). A watchdog thread that normally reports a blocked UI thread after 5 s logged nothing, so the whole process appears to be frozen (long stop-the-world GC?).

Repro on a Zebra TC53 (Android 14): open and close the same Activity (~530 views) 100 times, then read dumpsys meminfo:

Activity starts Activities Views PSS
0 9 2,419 395 MB
40 48 23,313 604 MB
50 31 14,850 533 MB
100 56 28,125 670 MB (swap 124 MB)
Destroyed Activities are only collected occasionally and partially, and the lower bound keeps rising.

Workaround: with GC.Collect() after leaving an Activity, the numbers stay flat at 6 Activities, 1,846 views and ~375 MB for 100 iterations.

This was measured with a Debug build. The production ANRs were on a Release build.

Is a fix for .NET 10 servicing planned, e.g. a backport of #11112?

That is quite a coincidence because that is exact hardware that I am using to host my application. I think the issue may be more obvious in a scanner type environment that is locked to one application and is always kept on a charging dock compared to a normal mobile application where by there will be more opportunities for the app to get restarted.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions