WPO: a year measuring this site's performance
On this page
The idea in one sentence
Measuring performance only pays off if the measurement is cheap to maintain.
This started in August 2025 as a draft about Google’s Core Web Vitals and Lighthouse. It stalled at the “plan of attack”. A year later, this is the ending: what was built, what was thrown away and which numbers are left.
Starting point
After building this web app, erades.com, I am going to focus on Google’s Core Web Vitals. For this task I am going to focus on Lighthouse.
The first measurements (August 2025, mobile) scored 94 on “About”, 93 on a post and 69 on search, with a 6.8 s LCP.
First thoughts
The data shows good results; partly thanks to the tech stack (server side rendering and static site generation) and the chosen meta-framework, Astro.
But to get acceptable numbers the developer has to be an expert and to have fought plenty of front-end battles to understand the best practices.
Even so, things always slip through, and there should be a way to keep the Core Web Vitals under control from day one so the numbers can keep improving. On the other hand, achieving that means running these checks programmatically.
Nowadays unit tests and e2e or functional tests are already part of developers’ DNA. However, I see that automated visual regression and performance tests have not caught on, partly, I think, because they make pipelines take longer.
I consider automated visual regression tests essential: if I had to choose between unit or component tests and visual regression tests, I would go for the regression ones.
As for WPO tests, I think developers do not realise how much value they bring to the business, and think more about the work they take than about the benefit.
What was built: LHCI with a server
The plan was the textbook one: Lighthouse CI locally and in CI, with budgets that break the build and an LHCI server keeping the history.
It got built: an LHCI server in Docker, with its SQLite database versioned in Git LFS. It was used for ten days and abandoned. When it was picked up again a year later, it took installing git-lfs, downloading 124 MB and digging a token out of the SQLite file itself before anything could be measured.
- ~20 MB of LFS per run to store about 40 numbers.
- The same numbers as plain text take ~1.8 KB.
The measuring traps
Worse than the cost was that part of what was measured meant nothing:
- Local and production mixed up. The suite measured
localhostURLs next to production ones. Half the measurements “improved” on deploy, not on code changes, and the 15-22 s local LCPs were the laptop, not the site. - A “desktop” that wasn’t. The desktop config only changed the screen and inherited mobile throttling: a desktop viewport over slow 4G with the CPU at 1/4. Once fixed, the home page went from perf 74 to 93 without touching a line of the site.
- Runner noise. On GitHub Actions, two identical runs vary by ±5
perfpoints. A timing budget that breaks the build with that noise fails at random.
What stayed: a text file on a branch
The server was replaced with an NDJSON file: one line per URL and form factor, with the median of 3 runs against production. It lives on an orphan metrics branch (on master, every weekly bot commit would trigger a deploy), and Git is the database: git log -p is the chart.
The lesson that pays off most: bytes are deterministic, timings are not. A rise in JS weight is always a real regression; a 3-point drop in perf says nothing. Timings are read as trends over weeks.
Results
With a reliable measurement, the improvements came on their own, one per PR: hero images resized by Astro, avatar and logo at their real size, fonts subset to Latin (237 → 80 KB preloaded), inline CSS to remove the only render-blocking resource, gtag.js after load and a prerendered home page preloading the LCP image.
/es home |
August 2026 | October 2026 |
|---|---|---|
| Total weight (mobile) | 2.74 MB | 0.39 MB |
| Mobile LCP | 12.1 s | 2.3 s |
| Desktop perf | 93 | 100 |
| Mobile perf | 75 | 92 |
What was not fixed: on the runner, some mobile loads stay blank for ~1 s longer with no cause reproducible locally, and that row reports a ~4 s LCP. It was accepted rather than chased.
Opinion
I stand by what I wrote a year ago: performance has to be measured from the start and automatically. What I change is how. The textbook plan (server, budgets, dashboards) died under its own weight; what has lasted is a text file anyone can read with jq.
And I don’t miss the budgets that break the build: with ±5 points of noise, a timing-based performance test is a flaky test. If they ever come back, they will be about bytes.
Bibliography:
- Web Vitals. web.dev, Google
- Introduction to Lighthouse. Chrome for Developers
- Lighthouse CI. GoogleChrome
- Lighthouse Variability. GoogleChrome