App performance has published thresholds, so "it feels slow" is never the right level of detail. On the web, Core Web Vitals set the bar: Largest Contentful Paint under 2.5 seconds, Interaction to Next Paint under 200 milliseconds, Cumulative Layout Shift under 0.1. On mobile, Google Play flags an app when its crash rate or ANR rate exceeds the bad-behaviour thresholds, and users abandon a cold start that runs past a few seconds.
This guide covers what to measure, where the bottlenecks usually are, and how to run an improvement plan that produces a number rather than an impression.
Key takeaways
- Measure at the 75th percentile, not the average: Averages hide the quarter of sessions that are actually bad.
- Field data beats lab data: Real devices on real networks; your laptop is not a representative device.
- Most wins are in three places: Payload size, database queries, and doing work that could wait.
- Set a budget and enforce it in the pipeline: Performance regresses silently unless a build fails.
Measure the right numbers
| Surface | Metric | Target |
|---|---|---|
| Web | Largest Contentful Paint | Under 2.5s |
| Web | Interaction to Next Paint | Under 200ms |
| Web | Cumulative Layout Shift | Under 0.1 |
| Mobile | Cold start | Under 2s |
| Mobile | Crash-free sessions | Above 99.5% |
| Mobile | ANR rate (Android) | Below Play's bad-behaviour threshold |
| API | Server response, 95th percentile | Under 500ms |
| API | Error rate | Below 0.5% |
Track these at the 75th percentile. An average response time of 300ms can conceal a quarter of users waiting three seconds, and those are the users who leave.
Web vitals also affect search: they are part of Google's page experience signals, so this is not only a usability question.
Use field data, not just lab tests
Lab tools run on a fast machine on a fast connection and tell you what is possible. Field data tells you what is happening.
Collect real-user monitoring from production, segmented by device class, connection type and geography. A product that performs well on a recent phone in a city and badly on a three-year-old device on a rural connection has a problem that no lab test will reveal, and in most markets that second group is a large share of users.
For mobile, the store consoles already provide crash, ANR and start-up data segmented by device. It costs nothing and is routinely unread.
Find the bottleneck before optimising
Optimising without measuring is how teams spend a fortnight on something that was never the constraint. Profile first, then fix the largest contributor, then measure again.
The bottleneck is usually in one of three places:
Payload. Uncompressed images, unnecessary fonts, JavaScript bundles containing libraries used on one screen. This is the most common web problem and often the cheapest to fix.
Database. Missing indexes, and the N+1 query pattern where a list of fifty items triggers fifty-one queries. This is the most common backend problem and typically produces the largest single improvement.
Work that could wait. Anything computed during a request that could be precomputed, cached, or moved to a background job: thumbnail generation, report aggregation, third-party calls in the response path.
The fixes that usually pay first
| Fix | Typical impact | Effort |
|---|---|---|
| Compress and correctly size images | Large on web LCP | Low |
| Add missing database indexes | Often dramatic | Low |
| Eliminate N+1 queries | Large under load | Low to medium |
| Cache expensive reads | Large | Medium |
| Move third-party calls out of the request path | Removes an external dependency from your latency | Medium |
| Code-split the front end | Improves first load | Medium |
| Reserve space for images and ads | Fixes layout shift | Low |
Third-party scripts deserve particular attention on the web. Analytics, chat widgets and tag managers are frequently the largest contributor to interaction delay, and they are added without anyone measuring the cost.
Set a performance budget
A budget turns performance from an occasional project into a constraint. Pick a small number of limits, bundle size, LCP, API response at the 95th percentile, and fail the build when a change exceeds them.
Without this, performance decays release by release, because no individual change is obviously responsible. Each addition is small; the accumulation is what users feel.
Run the check in continuous integration so the feedback arrives while the change is still being written.
Treat perceived performance as real
Users judge responsiveness rather than elapsed time. A screen that shows structure immediately and fills in feels faster than one that waits and appears complete.
Skeleton screens, optimistic updates that assume success and correct on failure, and prefetching the likely next screen all improve the experience without changing a single backend metric. They are frequently cheaper than the equivalent real improvement and sometimes matter more.
The exception is anything involving money or irreversible action, where an optimistic update that later fails is worse than an honest wait.
Make it someone's job
Performance regresses because nothing prevents it. Assign an owner, review the field metrics on a fixed cadence, and treat a threshold breach as a defect rather than a discussion.
Include performance in the definition of done for new work. Retrofitting is always more expensive than not regressing in the first place, and by the time it is visible to users it has usually accumulated across dozens of releases.
Related guides
- The groundwork for this is set out in AI Workflow Automation: Use Cases, Risks, and Roadmap.
- A useful companion to this one is Admin Dashboard Development: Features, Architecture, and Cost Drivers.
If App Performance: Metrics, Bottlenecks, and an Improvement Plan has you costing a migration off something old, Discuss Your Modernization Plan.
Ali Boran Gazel