The performance score and Core Web Vitals measure different things
A Lighthouse performance score of 98 describes a simulated test. Core Web Vitals uses real visitors’ experiences. A good test run can coexist with slow interactions, layout shifts, or long loading times for actual visitors. Neither result cancels the other.
PageSpeed Insights contains both kinds of data. Its real-user section uses Chrome User Experience Report data, usually called CrUX. The performance score further down comes from Lighthouse. Compare the field metrics at the top with Search Console’s Core Web Vitals report; use the lab test to investigate possible causes.
Google explains this distinction in About PageSpeed Insights. A green performance score is not a certificate that your real-user metrics pass.
An example: the page loads quickly, but the filter freezes
Imagine a shop category with a mobile Lighthouse score of 98. A customer opens the category, selects a size, and waits while the filter blocks the page. A fast initial load does not tell you whether that interaction feels responsive.
| Metric | Result | Assessment |
|---|---|---|
| Largest Contentful Paint (LCP) | 2.1 seconds | Good |
| Interaction to Next Paint (INP) | 420 milliseconds | Needs improvement |
| Cumulative Layout Shift (CLS) | 0.04 | Good |
The next task is to investigate the slow interaction. Compressing an already small logo to push the score from 98 to 100 would not address this example. Test the size filter, sorting, menu, and cart controls; identify which interaction causes the delay before changing code.
This is a made-up scenario, not a claim that every high-score mismatch is caused by filters. The same approach applies to the metric that is actually failing on your site.
What counts as good?
| Field metric | Good | Needs improvement | Poor |
|---|---|---|---|
| LCP: loading | ≤ 2.5 s | > 2.5–4 s | > 4 s |
| INP: responsiveness | ≤ 200 ms | > 200–500 ms | > 500 ms |
| CLS: visual stability | ≤ 0.1 | > 0.1–0.25 | > 0.25 |
These thresholds apply at the 75th percentile, separately for mobile and desktop. For a metric to be good, at least 75% of the measured experiences must meet its good threshold. The assessment is not an average of the three metrics: good loading cannot compensate for poor responsiveness. See Google’s explanation of the thresholds.
A missing metric is also not a zero. Read the report’s data-availability message rather than declaring a pass from the numbers you happen to have.
Check four labels before comparing reports
- Device. Compare mobile with mobile. A good desktop experience does not establish a good mobile experience.
- Field or lab. Compare real-user LCP with real-user LCP. The Lighthouse score is a different measurement.
- URL or origin. An origin covers a host such as
https://example.com. Its data can include many pages. Check whether PageSpeed is showing this URL’s field data or falling back to the origin. - URL group. Search Console groups pages with similar experiences. One tested URL may be better or worse than its group. Check several representative pages, especially those sharing the affected template.
The grouping and device rules are documented in the Search Console Core Web Vitals report. An origin, a URL group, and an individual URL should not be treated as interchangeable samples.
CrUX also has eligibility and data-volume requirements. A small or new page may have no field data even when it works well. Running Lighthouse repeatedly does not create real-user observations. See the CrUX methodology.
Choose a fix from the failing experience
| Example problem | Start the investigation here |
|---|---|
| Poor LCP on product pages | Identify the main content element. Check when its image or text starts loading, the server response, and anything that delays its display. |
| Poor INP after opening a menu | Reproduce the interaction and inspect the work triggered by it. A page-load test alone will not reproduce every menu action. |
| High CLS after a banner appears | Watch the page through loading and interaction. Check whether banners, images, ads, or fonts move content after it is visible. |
Give the person fixing the site an example URL, the device, the failing metric, and a way to reproduce the problem. “Make the score green” is much less useful than “the mobile category jumps when the promotion banner loads.” These are starting checks, not diagnoses made from a score.
Why yesterday’s fix is not fully reflected yet
PageSpeed’s field data covers a trailing 28-day collection period. A lab test can reflect a deployment immediately, while field data still contains visits from before the change. The 75th percentile can move unevenly as new observations arrive; it is not a linear countdown. PageSpeed’s reporting periods explain the window.
For example, if you fix the shop filter on Monday, test the interaction that day and record the deployment date. Review the mobile field trend over subsequent days. Do not expect Tuesday’s report to contain only post-fix visits, and do not assume it must pass exactly 28 days later if some visitors still experience the problem.
Search Console’s validation process has its own monitoring period. Starting validation does not force a recrawl or replace old observations with a fresh lab test. Follow the report’s validation instructions after checking the fix across the affected pages.
If the concern is missing search traffic rather than page experience, compare the evidence separately with the traffic-change analyzer. A failing metric alone does not establish why clicks changed.