Open Bug 2072217 Opened 17 days ago Updated 16 days ago

Rethink how often MemoryTelemetry gathers memory measurements

Categories

(Core :: XPCOM, task)

task

Tracking

()

People

(Reporter: jstutte, Unassigned)

References

Details

MemoryTelemetry::Poke() runs after every main-thread task (from XPCJSContext::AfterProcessTask, https://searchfox.org/firefox-main/rev/6c74efe2fcddf84b6f320959064a66946b4a1759/js/xpconnect/src/XPCJSContext.cpp#1542). Ten tasks within one second count as "active" and arm a one-shot timer of kTelemetryIntervalS (60 s, https://searchfox.org/firefox-main/rev/6c74efe2fcddf84b6f320959064a66946b4a1759/xpcom/base/MemoryTelemetry.cpp#44). A content process trivially produces ten tasks per second even while idle in the background, so in practice every live process gathers once a minute for as long as it exists, and the parent additionally reads the resident-unique size of every child in GatherTotalMemory on the same cadence.

The gather is not cheap, because the resident-unique measurement walks the whole page table of the process. Bug 2069230 switched Linux to smaps_rollup, which removed the text formatting and parsing but not the walk. Weighted medians of memory.collection_time from Nightly telemetry, three days before versus three days after that landing:

population before after
Fenix Nightly 146 ms (p90 414 ms) 73 ms (p90 268 ms)
Desktop Nightly Linux 52 ms 37 ms
Desktop Nightly Windows 6.5 ms 6.5 ms

The cost was large enough to show up as a 69% improvement in background-resource cpuTime-tab on android-hw-a55 (perf alert 52774): before bug 2069230 roughly 70% of the CPU time of an idle background tab in that ten-minute test was spent gathering memory telemetry. Fenix Nightly clients record about 200 tab-process samples per foreground hour.

Bug 2034081 wants to add per-process thread statistics to the same gather.

Note on channels: on Android this only affects Nightly (and the Nightly-flavour CI builds), because Fenix release does not run unified telemetry and Telemetry::CanRecordReleaseData() stays false there, so MemoryTelemetry never gathers (2 of 1.18 M Fenix release clients had a memory.unique sample on 2026-09-13). Desktop is different: unified telemetry keeps recording on for every channel, and on 2026-09-13 90-99% of desktop release clients gathered, about 3200 (Windows) to 5000 (Linux) times per client-day. At the current Linux median of 37 ms per gather that is roughly three minutes of CPU per Linux release user per day.

You need to log in before you can comment on or make changes to this bug.