A baseline has a shelf life.

A free evaluation is accurate on the day it runs. What it cannot tell you is which of its findings will still be true next month, because that depends on things nobody controls.

How long is a one-time AI visibility report good for?

Run the free evaluation and you get an honest picture of a single moment: what search and AI systems can access, understand and verify about the business today. That picture is genuinely useful and it is genuinely perishable.

It perishes at different speeds in different places, which is the part a single score cannot express.

What ages fastest

  • Anything a third party controls. A review platform changing what it displays, a directory dropping a field, a competitor publishing the better answer. None of it touches the website and all of it changes the evidence.
  • Retrievability. One line in a robots file, one redirect, one template change that moves the answer behind a script, and a page that was quotable this morning is not this afternoon. Nothing on the page looks different to a person.
  • Reputation recency. Review evidence does not decay because the reviews were removed. It decays because they got older while nobody added to them.

What holds longer

Semantic clarity and entity consistency usually hold, and when they break it is almost always because the business changed and the pages did not. A new service line, a new territory, a rename, a merged location. The drift is created internally, which makes it both the most common finding and the most preventable one.

The question a baseline cannot answer

The useful question is not "what is our score", it is "what changed, and does it matter". Answering it requires two readings taken the same way, and that is the whole difference between an evaluation and AI visibility tracking: one describes a state, the other describes movement and hands somebody the next action.

This is also why re-running a free evaluation weekly is not tracking. Without a stored history, a comparable method and a recorded engine version, two readings are two opinions. Compare them and the difference is as likely to be a change in how the evaluation ran as a change in the business.

So when is a re-read worth it

When something happened. A site rebuild, a rename, a new location, a migration, a period of sustained publishing, a competitor visibly moving. Those are the moments a fresh reading earns its keep.

The rest of the time, cadence is a ceiling rather than a target. Measuring more often than the underlying signals move produces a longer history that says the same thing, and the extra rows make it harder, not easier, to see the movement that mattered.

What makes the second reading comparable

Two readings only become a trend if three things held between them. The method was identical, including the exact wording of anything asked. The scale did not change, so a score out of ten still means what it meant. And the earlier reading was kept rather than replaced, with the version of the engine that produced it recorded alongside.

That last point is why measurement here is append-only. When an engine improves, later readings are better and are also not directly comparable to earlier ones, and the only way anybody can tell is if each row says which engine produced it. A history that quietly restates old numbers under a new method is worse than no history, because it looks like evidence.

This is one question beneath AI visibility tracking, the ongoing work Digilu does through The Observatory. The free point-in-time baseline is AIOInsights. Digilu cannot make a private AI model recommend a business. It can make the public evidence clearer, stronger and easier to verify, then track whether visibility improves.