OptFor.AI Consulting / Transformation / Development

Day 04 · 30-day GEO plan

How OtterlyAI shows a company’s visibility in AI

A report can show whether a brand appears for important customer questions, how it compares with competitors and which sources AI cites. It cannot measure every ChatGPT conversation or the number of customers acquired.

Author
Michal Shyjatyj · CEO 1xTeam | Co-founder OptFor.AI
Published
Time
10 min read
An analyst observes controlled measurements of brand visibility in AI answers

A company buys AI visibility monitoring to answer a business question: will a potential customer see its brand when asking ChatGPT for a recommendation? An OtterlyAI report provides important signals, but it does not directly measure reach across all users or resulting sales.

The tool runs a prepared set of questions, records the answers and checks them for brands and sources. This reveals which customer needs bring up the company alongside competitors, how AI describes it and whether its domain appears as a source. Those findings can guide content, SEO and digital PR decisions.

A search rank tracker is the closest comparison. OtterlyAI observes a fixed basket of prompts under controlled conditions. It does not receive logs of real customer conversations or see whether an answer led someone to contact the company. The report is valuable for detecting changes and problems to investigate, not for promising a number of future customers.

Mechanism

The measurement OtterlyAI performs

The researcher defines the brand, name variants, domains, competitors, prompts and country. OtterlyAI runs tracked prompts in supported engines every day. According to OtterlyAI’s documentation, most measurements use public web interfaces, while Claude is monitored through an API.

One prompt, engine and day produces one recorded answer. The platform stores its content, detected brands and displayed source URLs. The result can be described as:

selected prompt + country + neutral session + engine + day = one observation

This is a real answer obtained in a controlled session, but the prompt itself is synthetic. It does not come from customer query history. OtterlyAI creates a neutral baseline, so the result may differ from an answer received by a person with account memory, conversation history, custom instructions and another location.

I would not call this a simulation of ChatGPT. A more accurate description is a synthetic measurement of a real interface.

Research sample

What determines result quality

The prompt list, not the chart, is the most important part of the study. A result applies specifically to the set of questions entered into the system.

If a brand appears in 42 of 100 prompts, we can say it appeared in 42% of the completed trials for that basket. We cannot conclude that the company has 42% visibility across all ChatGPT answers.

The basket should cover real customer situations, including problem discovery, supplier search, comparison and selection. Countries, languages, categories and engines also need to remain separate. One hundred near-identical questions may produce a stable number and still provide a poor picture of the market.

OtterlyAI also reports Intent Volume. The company explains that this is an estimate created because AI platforms do not disclose usage data. Its algorithm is based on Google search volume and presents a five-point scale. It is not a count of ChatGPT queries. It may help prioritize topics, but it cannot establish a brand’s share of demand among AI users.

Source data

How to read mentions and citations

The raw answer is the most reliable part of the report. It lets us inspect four different events:

  • the brand appeared as text,
  • the company was described accurately,
  • the system actually recommended it in the given context,
  • the answer displayed a company-owned or third-party source.

OtterlyAI detects defined name variants. Each execution receives a value of 1 or 0, regardless of how often the name occurs in the answer. This method can miss an unspecified abbreviation or product name. It can also count another company with the same name or a negative reference. A mention is therefore not a recommendation.

A domain citation is usually less ambiguous, but it answers only: “Which link was shown with this answer?” It does not prove that the link caused the brand selection. A source may support a secondary detail, while parts of the answer may come from model knowledge or other, undisclosed retrieval steps.

Likewise, a sentence stating that a company is recommended for its integrations and price shows the explanation presented to the reader. It does not reveal source weights, system instructions, document rankings or why competitors were rejected.

Metrics

How to understand the main indicators

The OtterlyAI metric definitions can be translated into simpler questions:

MetricWhat it tells usWhat it does not prove
Brand MentionsHow many executions detected the brandHow often users asked about the company
Brand CoverageThe percentage of tracked trials that included the brandThe company’s share of all AI answers
Share of VoiceThe brand’s share of detected mentions in the studied basketMarket share, traffic or sales
Average Brand PositionHow high the brand appeared in answers where it was foundThe chance of a click or purchase
Domain CoverageThe percentage of trials displaying the studied domainWhether the domain caused the recommendation

Likelihood to Buy requires particular care. OtterlyAI converts average position to a fixed scale. First position gives 100%, second 77.5% and third 55%. This is not a probability estimated from purchasing behaviour. A value of 77.5% does not mean that this proportion of users will buy the product.

Every metric inherits the limitations of its input. An exact formula cannot repair an unrepresentative prompt set or an incorrectly detected brand.

Uncertainty

Why one result is not enough

Generative answers vary between runs. A brand present today can disappear tomorrow without any change made by the company.

OtterlyAI runs prompts once per day. For one question, a single day tells us only what happened in that trial. The study “Don't Measure Once” shows that one-off observations can overestimate or underestimate visibility and argues that visibility should be treated as a distribution from repeated measurements, not a single point.

Thirty days provide history, but conditions do not remain perfectly fixed. Web content, search indexes, the model, routing and answer generation may all change. The trend is operationally useful, yet it is not a laboratory repetition of an identical trial.

I therefore assess separately:

  • the broad trend across a well-designed prompt basket,
  • the stability of the most important prompts,
  • differences between platforms and countries,
  • individual answers that require manual review.

Decisions

What the data can and cannot support

OtterlyAI monitoring can detect a change worth investigating. It shows which questions lose the brand, where a competitor appears more often, which domains receive citations and whether the system describes the offer accurately. This can guide content-gap analysis, digital PR, brand-entity checks and the selection of answers for manual auditing.

I would not use the data to claim a “number one position in ChatGPT”, forecast sales or prove the effect of a single publication. A rise after a website change does not establish causation. The engine or third-party sources may have changed at the same time.

The report is an alert and a map of observations. Diagnosis begins only after opening the prompt, full answer and sources.

Practice

How to build reliable monitoring

Before starting, I record which decisions the monitoring should support and which intent groups the prompts represent. I do not change the basket after seeing an inconvenient result without marking a new version of the study.

My working rules are:

  1. Keep engines, countries and languages separate.
  2. Base assessment on a trend, not one day.
  3. Retest high-value prompts in several independent runs.
  4. Record mention, recommendation, description accuracy and citation as separate signals.
  5. Verify every material metric change in the raw answers.

OtterlyAI scales observation well. It does not remove the need to design the sample and interpret the result. With that distinction in place, the tool can track brand exposure and identify issues for further investigation. Without it, a precise chart can easily become an answer to a question the system never measured.

OptFor.AI

Want a reliable measurement of your brand’s AI visibility?

We can help select the prompts, assess raw answers and build a measurement that separates a trend from a chance result.

Let’s talk