Metrics before dashboards: where a BI project starts
Digital

Metrics before dashboards: where a BI project actually starts

The page for my data analytics service carries a line: “I do not build dashboards before the metric definitions are agreed: otherwise the argument about numbers moves to a new screen.” This article is about where that rule came from, and why a BI project starts not with choosing a tool but with a document that says, for every number, what it means.

First, the frame. There will be no client stories here about three departments counting revenue three different ways: the ones I have I cannot describe, and inventing them is not an option. Every example below comes from my own measurement setup. It is small but real: a data mart with its own sources, a build schedule and an owner. It covers a website and infrastructure rather than revenue and margin, but the mistakes in it are exactly the ones every BI project makes. In the four months since June it has shown me something untrue at least seven times. All figures were checked on 26 September 2026.

A screen does not settle an argument about numbers

Metrics before dashboards: two people pull a sheet reading by invoice and by payment in front of a screen showing 53 with no date

The usual story: a company argues about how many customers it has or what the quarter’s revenue was, because sales has one number and finance has another. The fix looks obvious: put everything into one dashboard and the argument ends. It does not end. It moves to a new screen, where there is now one number, and both sides still think it is wrong.

The reason is that people are not arguing about the visualisation but about the definition. Revenue by invoice date or by payment date, with or without tax, before or after refunds, in which currency and at which rate. Until those questions have a written answer, everyone has their own number, and each is right in its own way. A dashboard built before that conversation simply picks one of the definitions silently, and the argument becomes an argument with the dashboard.

My setup is small but real

For this site I keep a data mart: a table with one row per pair of Russian and English pages. It currently has 151 rows and 78 columns: addresses, dates, presence in the sitemap, indexing in Google and Yandex, SEO scores, internal links, text structure, and dozens more attributes. A script builds it from several sources: WordPress itself, the translation plugin, the sitemap, Google Search Console, Yandex Webmaster and the Rank Math plugin. The result syncs to Google Sheets.

Next to it run automated checks, 39 critical ones, from markup integrity to whether pages are actually served, plus a weekly report sent to me on Telegram. So this system has everything a corporate BI setup has: several sources, a regular build, a data mart, a report and an owner. Below is what it showed wrongly, and why.

The table is fresh, and a column has been frozen since July

This is the main example, and I found it while gathering the figures for this article. The number of Russian pages indexed in Google, according to the table, is 53. It has been 53 since 23 July 2026. Since then the table has been rebuilt dozens of times, most recently yesterday, and the number has not moved by one. The English figure sits at 48 and Yandex at 138, also since 23 July.

The cause is mundane. Access to Google Search Console was lost when the working environment was moved, and the table build is designed sensibly: if a source is unavailable, it does not wipe the column with zeros but keeps the previous value. That is the right decision, or a single network error would erase the whole history. But it has a consequence: the build date of the table stopped matching the date of the data in the column. The table says “built yesterday”, and that is true. The indexing column, meanwhile, is two months old, and nothing says so.

It gets worse. Over that time the site gained pages: there were 139 rows, now there are 151. New rows arrive with no indexing data, the numerator stays still, the denominator grows, and the share of indexed pages drifts down on its own, from 38 percent to 35. Looking at that, it is easy to conclude that indexing is getting worse and to start fixing it. In fact nothing at all is known about indexing since July.

The conclusion I draw: every metric needs its own “as of” date, shown next to the number rather than once for the whole table. My table does not have that yet, and the figure 53 is exactly the consequence.

A zero that meant “no data”

In early June the Yandex indexing column showed zero out of 131 for a week. Not because Yandex had indexed no pages, but because the source had not yet been connected. It was connected on 8 June, and within a day the number became 130.

Zero and “no data” are different values, and in a table they look identical. If a column that should hold numbers shows zero because a source is not connected, anyone looking at it reads it as a fact. In a BI project this is the most common cause of panic over nothing: the report shows zero sales for the day, when the data for that day simply did not load. An empty value should be empty and look empty.

One word, three different numbers

The table has three columns that in conversation all get called “indexed”. The sitemap lists 151 pages out of 151. The Yandex index holds 138. The Google index holds 53. All three numbers are correct and none contradicts the others, because they answer three different questions: what I told the search engines, what Yandex took, and what Google took according to its console.

The argument starts when someone says “150 pages are indexed” and someone else says “53 are indexed”, and both are right. It is the same story as revenue by invoice versus revenue by payment. A metric needs not a name but a definition: where it comes from, what exactly is counted, as of which date. “Indexed” without a source is not a metric, it is a reason to argue.

A green response that did nothing

For a while I submitted new pages to Google through the Indexing API. Every request got a success response: “accepted”. Record that in the table as “submitted for indexing” and the table turns green. Checking showed otherwise: immediately after such an “accepted”, a status query for the same page answered “not found”.

Google’s documentation says so directly: the Indexing API can only be used for pages with job posting or livestream markup. For ordinary articles the request is accepted and triggers nothing. Since then the table records that status as “accepted”, not as “indexed” and not as “submitted”. The difference is one word, but it is the word that separates a fact from a hope.

The rule that grew out of it: a success response only means the request arrived. Everything that was supposed to happen afterwards is checked separately. In corporate BI this looks like “the load job completed successfully” when zero rows were loaded.

Even the word “day” needs a definition

The same API has a daily quota. I scheduled a one-off run at 04:30 UTC to send the remaining pages the next day, and got 139 rate-limit refusals and not one successful request. The reason: Google’s daily quota resets at midnight Pacific Time, not at UTC midnight and not at local midnight. At 04:30 UTC Google was still in the previous day, and that day’s limit was used up.

That seems trivial until it lands in a report. “Sales for the day” in a system that counts days in UTC and in an accounting team that counts in local time differ by every order placed overnight. The boundary of a day, a week and a month is part of a metric’s definition, not a server setting.

A number needs one address

The SEO score that the Rank Math plugin shows in the editor is computed in the browser and stored in the plugin’s own table, not in the post metadata where people usually look for it. Read it from anywhere else and you get a different number or none at all. The table takes the score from exactly where the plugin itself stores it, and that location is written down next to the code.

In a company the same thing looks like revenue that exists in the accounting system, in the CRM and in an export for investors, and differs slightly in each. Every metric needs one canonical place it comes from. All other copies are derived, and when they disagree with the canonical one, the canonical one wins.

A correct number that arrives late

In June the site had 96 orphan pages that no internal link pointed to. I placed about a hundred and fifty links, published them and recounted straight away: there were now 89 orphans. It looked as though the work had achieved almost nothing. In fact the plugin’s link graph updates with a delay rather than immediately. Today there are 58.

The figure 89 was correct for that moment and wrong for a decision. Had I concluded from it that the method did not work, I would have abandoned a method that does. Every metric has a delay between an event and its reflection, and you have to know it before you look at the number. A decision made on a metric that has not caught up with reality is a decision about the past.

The definition changes the number more than the data does

Once a check reported that 29 pages had no call to action. I started digging and found the check was wrong: exactly one link was genuinely broken. Twenty-nine was not a property of the site but a property of the counting rule.

A second case from the same series. The duplicate-link check first counted how many distinct pages linked to a target, and saw few problems. When the rule changed to “how many times one article links to the same target”, 270 surplus links came out of the texts. Not a byte of data changed. The definition changed, and the number grew by a factor of hundreds.

This is the central point of the whole article: a metric’s definition affects the number more than the data does. That is why the argument about definitions cannot be skipped by leaving it to whoever builds the dashboard.

A metric that is wrong half the time

One of the checks compares the number of sections in the Russian and English versions of an article. A mismatch is supposed to flag that something was lost in translation. When I went through every mismatch by hand, 16 turned out to be genuine gaps, and about twenty were deliberate: different examples for different markets.

A metric like that must not become a performance target, “get mismatches to zero”. Its proper place is a list for manual review: it is good at finding candidates and bad at passing verdicts. A BI project is the same: some metrics are signals to look at with your own eyes, not targets to hit. Confuse the two and people start fixing the number instead of what it was supposed to show.

A green screen over dead pages

The most expensive case. After one plugin updated itself automatically, every page on the site failed with a critical error when rendered normally. For two days nobody noticed: the watchdog polled ten addresses, and all of them answered 200. The page cache was answering, holding copies made before the breakage.

The “site is up” metric was measuring the cache, not the site. It was accurate, fresh and completely useless. The fix was not in the metric but in its definition: the check now bypasses the cache and looks for the critical error text on the page. Once a week the entire sitemap is crawled separately. The question “what exactly are we measuring” turned out to matter more than “how often”.

What this means for a BI project

Every case above is about the same thing: the number was correctly computed and wrongly understood. That is why a BI project starts for me with a document, not a tool. For each metric it records:

  • what exactly is counted, in words that both sales and finance will understand;
  • one source the number comes from, and where the derived copies live;
  • the formula and the time boundary: which day, which time zone, by which date;
  • the “as of” date, shown next to the number;
  • the delay between an event and its appearance in the metric;
  • what “no data” looks like, so it cannot be mistaken for zero;
  • whether it is a target or a signal: something to hit or something to look at;
  • the owner of the definition, the person who settles arguments about it.

The tool is chosen after that document, and then the choice becomes easy. Power BI, Tableau or a plain spreadsheet solve the same problem if the definitions are written down, and are equally useless if they are not.

A data project starts with a person, not a data warehouse

The second rule on the service page: “I do not leave behind a system nobody can run: if the company has no one responsible for data, I suggest starting with that person rather than with a warehouse.” It follows from the first. Metric definitions change as the business changes, sources drop away the way access to Search Console dropped away, and somebody has to notice.

My small setup works because it has an owner and checks, not because it is well built. And even with an owner, the indexing column sat frozen for two months until I started gathering figures for this article. A system without an owner does the same thing, only nobody ever finds out.

What I deliberately do not do

  • I do not build dashboards before the definitions are agreed. Otherwise the argument just moves to a new screen.
  • I do not show a number without its “as of” date. The table’s build date is not the date of the data in a column.
  • I do not record “accepted” as “done”. A success response means the request arrived and nothing more.
  • I do not turn signal metrics into targets. Otherwise people start fixing the number.
  • I do not leave a system without an owner. If nobody is responsible for the data, that person is where to start.
  • I do not invent client stories for examples. Every case in this article comes from my own setup.

The stack

Why this matters if you are paying

An argument about numbers costs money, and it is counted in executive hours. The meeting where sales and finance spend half an hour working out whose revenue is right happens every month. A dashboard bought before the definitions are agreed does not save those hours: it adds the cost of licences and of the people who built it.

Second: decisions made on a misunderstood number cost more than having no number. Had I believed indexing was falling, or that the linking work was not paying off, I would have spent time fixing what was not broken. In a company the price of that mistake is higher: budget moved away from a channel that works, or people penalised for a metric they do not control. A definitions document costs a few days of work. Not having one costs something every month.

Frequently asked questions

Where should a BI project start?

With a document that records, for each metric, what exactly is counted, where it comes from, which time boundary it uses, what delay it has and who owns the definition. The tool is chosen after that document, not before.

Why does a dashboard not settle an argument about numbers?

Because people are not arguing about how to show a number but about what it means. A dashboard built before the definitions are agreed silently picks one of them, and the argument becomes an argument with the dashboard.

What is dangerous about a metric without a date?

The table’s build date may not match the date of the data in a column. When a source is unavailable, the previous value is often kept, and the table looks fresh while the data is stale. An “as of” date should sit next to every number.

Why not show zero when there is no data?

Because zero reads as a fact. A report showing zero sales for a day when the data simply failed to load causes panic over nothing. Missing data should look like missing data.

Do I need Power BI or Tableau to begin?

No. If the definitions are written down, Power BI, Tableau and a plain spreadsheet solve the same problem, and the choice comes down to convenience and cost. If there are no definitions, every tool is equally useless.

Who should be responsible for metrics?

A specific person, the owner of the definition, who settles arguments about it and notices when a source stops updating. A system without such a person is better not built: it will start showing stale numbers, and nobody will know.

Sources

Need a consultation?

If your company argues about numbers, or you are about to buy a BI tool, book a conversation. We will start with the definitions of your key metrics and choose the tool afterwards. More about how I work with data on the data analytics service page.

Rate article