Profile Compose Performance with Layout Inspector, Macrobenchmark, and Baseline Profiles

Quick answer: Use Layout Inspector to find where Compose recomposes or skips during a specific interaction. Use Macrobenchmark on a release-like build and physical device to measure whether that journey is actually slow or janky. Use Baseline Profiles to precompile code used by the critical journeys you have measured, then benchmark again to prove the improvement.

These tools answer different questions. A recomposition count is not a frame-time measurement, and a Baseline Profile does not repair expensive work on the main thread. Treat performance as a loop: observe a real symptom, form a hypothesis, measure a representative journey, make one targeted change, and compare the same measurement.

Give each tool one job

ToolPrimary questionBest time to use it
Layout InspectorWhich composables recompose or skip while I reproduce this interaction?Early diagnosis in a debuggable build.
MacrobenchmarkHow does a complete user journey perform for startup time or frame timing?Before and after a change, on a physical device.
Baseline ProfilesWhich critical startup and runtime paths should ART precompile for a fresh install?After critical journeys are defined; verify with Macrobenchmark.

The Compose performance tooling guide makes the same distinction: recomposition inspection narrows the search, while system traces and benchmarks provide timing evidence.

Start with a reproducible user journey

Do not benchmark an abstraction such as “the feed is slow.” Write down the action a person takes and what completion means:

  • Cold-launch the app until the home content is ready.
  • Open a detail screen, then scroll its image-rich list.
  • Type a search query and wait for the first result state.
  • Expand a section and run the animation to completion.

Keep the data, network state, account state, device, and build variant consistent enough that the same journey is comparable before and after a change. A benchmark that measures an empty list one day and a populated list the next does not answer a useful question.

The work inside the journey matters more than a synthetic composable call. Macrobenchmark is designed for these larger app interactions, including startup, scrolling a LazyColumn, and animations; Android recommends it for Compose UI measurement.

Use Layout Inspector to form a hypothesis

Run the app, reproduce the problem once, then open Tools → Layout Inspector and attach to the process. In the Component Tree, enable Show Recomposition Counts if the composition and skip columns are hidden. Reset the counters, perform one known interaction, and inspect the affected subtree.

Layout Inspector can show both composition and skip counts for individual composables. On API 29 or later with Compose 1.2.0 or later, those counters make it easier to see which parts of the hierarchy responded to the interaction. The current Compose debugging documentation also notes that selecting a node exposes its parameters and dimensions when available.

The useful question is not “did this recompose?” Recomposition is normal. Ask instead:

  • Did an unrelated, large subtree recompose when only a small state value changed?
  • Is a list item reading broad screen state instead of its own item state?
  • Is a composable unexpectedly never recomposing when visible state changes?
  • Are skips rare in a hotspot where stable inputs should allow them?

A high count is a clue, not a verdict. A tiny composable can recompose often with no visible cost, while a single composition can block the main thread long enough to cause jank. Reset counts for each focused interaction, then confirm the hypothesis with timing data.

For the Compose-side causes and fixes—state-read scope, stable UI models, derivedStateOf, and lazy-list identity—see Jetpack Compose Recomposition: Debug & Optimize Performance Guide and 10 Jetpack Compose Performance Tips. Make only the change that your observation supports.

Set up a Macrobenchmark module

Macrobenchmarks run separately from the app being measured. Android’s current guidance uses a dedicated com.android.test module; Android Studio’s New Module → Benchmark template creates the normal structure and a starter benchmark.

The target app should be as close to production as possible: non-debuggable, profileable, and preferably minified. A debug build changes execution costs enough to make performance conclusions unreliable. The Macrobenchmark setup guide also requires ProfilerInstaller so the benchmark can capture profiles and reset compilation and shader state as needed.

Run the benchmark on a physical device. An emulator shares CPU, memory, and scheduling with the host machine, so its numbers are not representative of an end user’s device. Keep the same physical device, Android version, thermal state, and power conditions when comparing revisions.

Measure startup with a small benchmark

The benchmark owns the measured journey through MacrobenchmarkRule.measureRepeated. It supplies a package name, metrics, iterations, a startup mode when relevant, and the actions to perform.

import androidx.benchmark.macro.StartupMode
import androidx.benchmark.macro.StartupTimingMetric
import androidx.benchmark.macro.junit4.MacrobenchmarkRule
import androidx.test.ext.junit.runners.AndroidJUnit4
import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith

private const val TARGET_PACKAGE = "com.example.app"

@RunWith(AndroidJUnit4::class)
class StartupBenchmark {
    @get:Rule
    val benchmarkRule = MacrobenchmarkRule()

    @Test
    fun coldStartup() = benchmarkRule.measureRepeated(
        packageName = TARGET_PACKAGE,
        metrics = listOf(StartupTimingMetric()),
        iterations = 10,
        startupMode = StartupMode.COLD,
        setupBlock = {
            pressHome()
        },
    ) {
        startActivityAndWait()
    }
}

Cold startup is appropriate when process creation is part of the user experience you want to study. For a scroll or animation benchmark, use a FrameTimingMetric, navigate to the same content in setupBlock, and place only the gesture in the measured block. The Macrobenchmark guide documents StartupTimingMetric, FrameTimingMetric, TraceSectionMetric, startup modes, compilation modes, JSON output, and trace files.

When UI Automator must locate a Compose node, expose a stable resource ID through testTagAsResourceId on the relevant top-level semantics modifier. Avoid selecting by visible text if localization or dynamic content will make the benchmark fragile.

Read the result and the trace together

Macrobenchmark reports measurements in Android Studio and writes detailed JSON output and a trace per measured iteration. The trace is where a suspicious number becomes an actionable lead: it can reveal work on the main thread, repeated composition, slow layout, image decoding, binder activity, or a blocking call during the journey.

Use the result to compare the same journey across changes, not to chase a universal number. If the result improves but the user-visible symptom remains, revisit the journey definition. If the number regresses, open the associated trace before making another speculative Compose optimization.

Useful controls include:

  • Compare release-like variants, not a debug build with a release build.
  • Keep the iteration count and target journey stable across the comparison.
  • Leave CompilationMode explicit only when the question is about a compilation state, such as the effect of a Baseline Profile.
  • Record device model, OS version, app revision, build variant, and data setup with the result.

Generate Baseline Profiles from critical journeys

A Baseline Profile records code used during important flows so Android Runtime can precompile it at installation time. It can improve startup and runtime responsiveness, but it does not make network calls, inefficient algorithms, or oversized images cheaper. First choose the journeys that matter to users; then collect the profile from those journeys.

The easiest current setup is Android Studio’s Baseline Profile Generator module template. It creates the generation and benchmark structure, and Android’s Baseline Profile creation guide recommends extending the generated startup test with real critical user journeys.

At the core is BaselineProfileRule:

import androidx.benchmark.macro.junit4.BaselineProfileRule
import androidx.test.ext.junit.runners.AndroidJUnit4
import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith

private const val TARGET_PACKAGE = "com.example.app"

@RunWith(AndroidJUnit4::class)
class BaselineProfileGenerator {
    @get:Rule
    val baselineProfileRule = BaselineProfileRule()

    @Test
    fun startup() = baselineProfileRule.collect(
        packageName = TARGET_PACKAGE,
    ) {
        startActivityAndWait()
    }
}

Add navigation, list scrolling, or a primary action only when they are true critical journeys. Do not generate a profile from every debug-only route; extra rules consume profile budget without necessarily helping the path users care about. If a journey is part of startup, configure the generated rule to include it in the Startup Profile; non-startup interactions should not be added to that startup subset.

Baseline Profiles are installed for release builds, not ordinary development builds. Generate them using the module template or the matching Gradle task, commit the generated profile through the project’s normal review process, and regenerate it when critical flows or code shape changes significantly.

Prove the profile helps

Do not stop after a Baseline Profile generates successfully. Benchmark the same journey with the profile enabled and disabled. Macrobenchmark’s CompilationMode makes that comparison explicit.

@Test
fun startupWithBaselineProfile() = benchmarkRule.measureRepeated(
    packageName = TARGET_PACKAGE,
    metrics = listOf(StartupTimingMetric()),
    compilationMode = CompilationMode.Partial(
        baselineProfileMode = BaselineProfileMode.Require,
    ),
    iterations = 10,
    startupMode = StartupMode.COLD,
    setupBlock = { pressHome() },
) {
    startActivityAndWait()
}

Use a corresponding no-profile or different compilation-mode benchmark as the comparison, then review the numbers and traces on a physical device. Android’s Baseline Profile measurement guide explains these compilation states and cautions against drawing conclusions from emulator results.

The snippet intentionally omits imports because the generated Macrobenchmark module supplies the exact benchmark dependencies and versions for the project. Use that template rather than pasting version numbers from an article into a production build.

Turn it into a sustainable loop

Performance work should not depend on somebody remembering a one-off manual test.

  1. Keep one or two benchmarked critical journeys for startup, scrolling, or a high-value animation.
  2. Run the benchmarks before and after meaningful UI or startup-path changes.
  3. Review a trace when a result changes enough to matter; do not optimize from counts alone.
  4. Generate Baseline Profiles from the same critical journeys and verify their effect.
  5. Add stable benchmark execution to CI only after the device and environment can produce trustworthy numbers.

Compose performance improves fastest when every optimization has an observed cause and a measured outcome. Layout Inspector directs attention, Macrobenchmark turns a real journey into evidence, and Baseline Profiles optimize the code paths that evidence says matter.