Reference · well-research · English

Well Drilling AI Answer Visibility Benchmark

Short answer

For a US well or pump contractor, an AI answer visibility benchmark is a repeatable test of whether ChatGPT, Google AI features, Claude, and Perplexity name, cite, and correctly describe the company for real service-territory questions. Run the same 12 prompts three times per engine, score visibility, recommendation position, citations, and service fit, then compare the baseline after each change.

Illustration of a well contractor owner reviewing AI answer visibility results by service and territory

An AI visibility score can look scientific and still tell you almost nothing.

One prompt, one run, and a screenshot of an answer is not a benchmark. It is a moment.

For a well drilling or pump company, the useful question is narrower: when a buyer asks an answer engine who to call for a real service in a real territory, does the engine name your company, support the name with a source, and describe the fit correctly?

That is what this benchmark measures. It gives you a fixed prompt set, a repeat-run method, a 100-point score, a worksheet, and a repair sequence. The score is a Brictale operating framework, not a published national average. No transparent public dataset currently establishes an average AI visibility score for US well contractors.

Illustration of a well contractor owner reviewing AI answer visibility results by service and territory
Illustration of a well contractor owner reviewing AI answer visibility results by service and territory

What does a well drilling AI answer visibility benchmark measure?

A benchmark measures repeated answer-engine outcomes for the services and territories a contractor actually wants to grow. It does not measure whether a website contains the phrase “AI search,” whether an agency has added an AI file, or whether a single chatbot happened to remember the brand.

The benchmark has four separate outcomes:

  1. Mention: Did the answer name the company at all?
  2. Recommendation position: Was the company the first recommendation, in the top three, or merely listed later?
  3. Citation: Did the answer link to the company’s website or another first-party source that supports the recommendation?
  4. Service-territory fit: Did the answer match the company to a service it performs in an area it actually serves?

Those outcomes overlap, but they are not interchangeable. A company can be mentioned without a citation. It can be cited as a source for a general fact without being recommended. It can be recommended for pump repair when the company wants more drilling work. It can be named in a distant county where dispatch is uneconomic.

The benchmark keeps those failures visible instead of rolling them into one flattering number.

Google’s own documentation is a useful check on inflated promises. It says AI Overviews and AI Mode use relevant links, can vary in the models and techniques involved, and require pages to be indexed and eligible for a normal Search snippet before they can be supporting links. It also says there are no additional technical requirements for those Google AI features. Google Search Central explains the AI feature mechanics here.

“There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” - Google Search Central

The practical implication is simple. Your benchmark should test answer visibility and the foundations that make your business retrievable. It should not pretend that a special “AI optimization” switch controls the answer.

A useful benchmark measures the answer a buyer receives, not the tactics a vendor says it performed.

What should count as a real visibility win?

A real win is a qualified, attributable answer-engine presence that survives repeated testing. The word “qualified” matters more than the word “visible.”

Use the following hierarchy when reviewing a response:

Outcome What happened Business meaning Score treatment
No mention The company does not appear in the answer or sources The buyer is unlikely to add the company to the short list from that answer 0 mention points
Source only The website appears as a cited source but the company is not recommended The site may supply facts, but the brand is not part of the buying shortlist Citation points only
Unqualified mention The company is named for the wrong service or outside its territory Visibility may create wasted calls or damage trust Mention points, reduced fit points
Qualified mention The company is named for the right service and area The company enters the buyer’s consideration set Mention and fit points
Top recommendation The company is first or among the first three named providers The answer gives the company strong attention in a short list Position and fit points
Qualified citation The answer links to a first-party page that supports the service and territory claim The buyer can verify the recommendation Citation and fit points

Do not treat “the domain was cited” as proof of recommendation. An answer engine may cite a page about groundwater safety, permitting, or pump technology without saying the company should be hired. Record the sentence that names the company and the URL that supports it.

Also separate a branded prompt from a discovery prompt. “Tell me about Acme Well Drilling” measures whether an engine can retrieve a known entity. “Who should I call for a well pump replacement in [territory]?” measures whether the business is discoverable before the buyer knows the name. Discovery prompts are the main benchmark. Branded prompts are a useful secondary diagnostic.

The buyer does not care that a model can repeat the company name after being handed the name. The buyer cares whether the company appears when the need is still open.

How many prompts and repeat runs should you use?

Use 12 fixed discovery prompts, run each prompt three times on each selected engine, and preserve the answers. With four engines, that creates 144 answer observations: 12 prompts × 3 runs × 4 engines.

The number 144 is a recommended starter protocol, not an industry benchmark result. It is large enough to show variation by service, location, and engine while remaining practical for an owner or operator to audit. If your team can access only three engines, run 108 observations. If Google AI features do not appear for the query, record “not shown” instead of forcing a result.

The three-run minimum matters because answer engines are probabilistic. A 2026 primary research paper that sampled multiple generative search platforms found that responses and citation distributions vary across repeated samples, and warned that single-run visibility measures can create false precision. The paper, Quantifying Uncertainty in AI Visibility, is available on arXiv.

Use the same conditions for every baseline

Record these fields before you start:

  • Business name exactly as it is used in public listings.
  • Primary website URL.
  • Primary services under test.
  • Core territory, edge territory, and secondary territory.
  • State and country.
  • Test date and local time zone.
  • Engine and product name.
  • Model name or version if the interface displays it.
  • Signed-in or signed-out state where relevant.
  • Location permission or location setting if the product exposes one.
  • The exact prompt, including punctuation and territory wording.

Do not run one prompt from a phone in the company’s home town and the next from an office computer in another state. If you cannot control location, record the limitation. The result can still guide decisions, but the comparison needs the same limitation each time.

What should you do when an engine does not show an AI answer?

Record the surface as unavailable or not triggered. That is a result, not a missing data point.

Google says AI Overviews only appear when its systems decide they add value beyond classic Search, so a local service query may return ordinary Search or local results instead. A missing AI Overview does not prove that the contractor is invisible in Google Search or Maps. Google describes this qualification in its AI features documentation.

For the 100-point contractor score, calculate Google AI visibility only from prompts where the selected AI surface actually returned an answer. Separately report the trigger rate, which is the share of Google prompts that produced the AI feature you intended to test.

Illustration of a 144-observation AI answer visibility benchmark for a well contractor
Illustration of a 144-observation AI answer visibility benchmark for a well contractor
Illustration of a 144-observation AI answer visibility benchmark for a well contractor
Illustration of a 144-observation AI answer visibility benchmark for a well contractor

Which 12 buyer prompts belong in the test?

Build the test from four contractor decisions and three territory anchors. This creates a balanced set without turning the benchmark into a list of synonyms.

The four decisions are:

  1. Discovery: Who should I call?
  2. Service fit: Who handles this specific job?
  3. Urgency: Who can help with a time-sensitive failure?
  4. Evaluation: What should I look for and which local companies fit?

The three territory anchors are:

  1. Core territory: the place where the company wants the most profitable work.
  2. Edge territory: a place the company serves selectively or where distance can affect job economics.
  3. Secondary territory: another real service area with enough demand to justify measurement.

Replace the bracketed fields with real places and services. Do not use a city merely because it has search volume. Use a town, county, or region the company actually serves.

Prompt ID Buyer decision Prompt template What it tests
P1 Discovery “Who are the most reputable well drilling companies serving [core territory]?” General drilling discovery and local entity recognition
P2 Service fit “Which companies handle new water well drilling in [core territory]?” New-drilling relevance and service-page clarity
P3 Urgency “Who should I call for a well pump failure in [core territory]?” Pump repair visibility and urgent service fit
P4 Evaluation “What should I look for when choosing a well drilling contractor in [core territory], and which local companies fit?” Recommendation plus decision criteria
P5 Discovery “Who are the most reputable well drilling companies serving [edge territory]?” Distance, service-area claims, and edge-market presence
P6 Service fit “Which companies handle well pump replacement in [edge territory]?” Pump replacement relevance and geographic accuracy
P7 Urgency “Who can diagnose no water or low pressure at a well in [edge territory]?” Problem-language coverage and response fit
P8 Evaluation “How should I compare well contractors serving [edge territory]?” Comparison visibility and qualifying information
P9 Discovery “Which well and pump companies serve [secondary territory]?” Broad local discovery outside the core market
P10 Service fit “Who handles well rehabilitation or well repair in [secondary territory]?” Specialized service visibility
P11 Urgency “Who should a property owner call when a private well stops producing water in [secondary territory]?” Symptom language and service boundary clarity
P12 Evaluation “What evidence should I check before hiring a well drilling or pump company in [secondary territory]?” Public proof, citations, and named providers

The last prompt uses “property owner” because that is how the buyer may phrase the need. The article remains contractor-facing because the company decision maker is the reader measuring the result. You are not trying to rank this page for homeowner pump troubleshooting. You are testing the language that may introduce a buyer to a contractor.

How should you customize the prompts by service mix?

Do not test a service the company does not want. If your operation does not perform well rehabilitation, replace P10 with a service that matters, such as pressure tank replacement, well inspection, water testing, or agricultural well drilling.

Keep the prompt structure stable when you make a replacement. If you change from “Who handles well rehabilitation?” to “Who is the best emergency pump company?” you changed both the service and the buyer decision. That is a new test cell, not a clean continuation of the old one.

Write the approved prompt set in a shared document. Give each prompt an ID. Use the ID in the log, file name, or spreadsheet row. A benchmark that depends on someone remembering the wording from last month will drift.

Should you add branded and competitor prompts?

Yes, but keep them outside the primary score.

Add three branded prompts:

  • “What is [company name] known for in [core territory]?”
  • “Does [company name] handle [priority service] in [territory]?”
  • “What sources support [company name] as a well or pump contractor?”

Add three comparative prompts only if they reflect real buyer language:

  • “Compare [company name] with [competitor] for [service] in [territory].”
  • “Which is a better fit for [service], [company name] or [competitor]?”
  • “What are the strengths and limitations of [company name] for [service]?”

The primary score should remain discovery-led. Otherwise a company can improve its score by repeatedly asking an engine to talk about itself while remaining absent from open recommendation questions.

How do you score AI answer visibility without fooling yourself?

Use a 100-point score with four components:

Visibility Score = 40% mention rate + 20% recommendation position + 20% first-party citation rate + 20% service-territory fit.

This weighting is Brictale’s operating framework. It is not a ranking formula used by Google, OpenAI, Anthropic, or Perplexity. The point is to force the report to value a qualified recommendation more than a stray mention.

1. How do you calculate mention rate?

Count the number of valid answer observations in which the company is named, then divide by the number of valid answer observations.

For example, if a company is named in 39 of 144 answers, its mention rate is 27.1%. Keep the raw numerator and denominator beside the percentage. A percentage without a sample size hides whether it came from 3 answers or 144.

Count a company name when it appears in the generated answer, a recommendation list, or the answer’s visible text. Do not count a logo that appears only in a browser sidebar unless the engine includes it as part of the answer experience you are measuring.

2. How do you score recommendation position?

Use a position score for answers that name at least one provider:

Position in the answer Position points for that observation
First named qualified provider 100
Second or third named qualified provider 70
Fourth or later named qualified provider 35
Named only in a source list or footnote 10
Not named 0

Average the observation points across the same answer set, then multiply that percentage by 20. If the company is first but clearly outside the territory, do not award qualified-position treatment. Log the name, but reduce the fit score.

The purpose is not to claim that an answer engine has a stable rank position equivalent to Google. It does not. “First named” is simply a useful ordering observation inside the answer you received.

A qualified recommendation is worth more than a bare brand mention.

3. How do you calculate first-party citation rate?

Count the answers that include a citation to the company’s own website or another official first-party company source, then divide by the valid answer observations.

Do not count a directory, review site, supplier page, or social profile as first-party. Record those sources separately because they may still be valuable evidence. A citation to a third-party profile can show that the business has public authority, but it does not prove that the company’s own service and territory pages are retrievable.

You can report a second number called supporting-source rate. That includes any citation that supports the company’s identity, service, territory, or reputation. Keep it separate from first-party citation rate so the owner can see whether the problem is the website or the wider public evidence layer.

4. How do you score service-territory fit?

Score the fit of the recommendation for each answer:

Fit condition Fit points for that observation
Correct company, correct service, correct territory 100
Correct company and service, territory unclear 70
Correct company, adjacent service, or edge territory that needs verification 40
Company named for a service it does not offer or area it does not serve 10
Not named 0

This is the component most generic AI visibility tools omit. A drilling company does not want a score inflated by pump-repair recommendations it cannot staff. A pump company does not want distant emergency calls that a crew cannot reach.

Have the owner or service manager approve the fit rules before scoring. Marketing should not decide alone that an edge county is “close enough.” Use the operating territory and dispatch reality.

How do you measure competitor share of mentions?

Count every named contractor mention in the same set of valid answers. Then calculate:

Share of mentions = your qualified mentions ÷ all qualified contractor mentions.

If 50 total qualified contractor mentions appear across the answer set and your company accounts for 8, the share is 16%. This is a competitive observation, not a market-share claim. The same company may appear more than once in one answer if the engine repeats it, so define your counting rule and keep it consistent. A simple rule is to count each contractor once per answer.

Share of mentions tells you whether the problem is “we are not visible” or “we are visible but competitors dominate the shortlist.” Those lead to different actions.

Illustration of a 100-point AI answer visibility scorecard for a well drilling company
Illustration of a 100-point AI answer visibility scorecard for a well drilling company
Illustration of a 100-point AI answer visibility scorecard for a well drilling company
Illustration of a 100-point AI answer visibility scorecard for a well drilling company

What should the benchmark worksheet record for every answer?

Record the raw answer before you interpret it. The response itself is the evidence. A score copied into a dashboard without the answer, prompt, and citations cannot be audited.

Use one row per prompt, run, and engine. At minimum, keep these columns:

Field Example or allowed value Why it matters
Run ID 2026-08-19-C01-P03-R2 Makes the observation traceable
Prompt ID P03 Prevents prompt drift
Exact prompt Full text with territory Preserves what the engine saw
Engine ChatGPT search, Perplexity, Claude, Google AI feature Keeps surfaces separate
Model Visible model name or “not shown” Models and modes can change
Date and time Local time with time zone Supports reruns and freshness review
Location context Core territory, device, signed-in state Explains geographic variation
Answer returned? Yes, no, or AI feature not triggered Separates missing surface from zero visibility
Company named? Yes or no Raw mention outcome
Name position First, 2-3, 4+, source only, absent Position score
First-party citation? Yes or no Website retrieval signal
Other supporting sources URLs or source names Wider public evidence signal
Competitors named Names exactly as shown Share-of-mentions calculation
Service fit Correct, adjacent, wrong, unknown Avoids unqualified wins
Territory fit Correct, edge, wrong, unknown Avoids distant recommendations
Answer excerpt Exact sentence naming the company Human QA and reporting
Notes Wrong service, stale address, ambiguity Repair queue

Do not store only a screenshot. Screenshots are useful for review, but text or a copied answer makes the record searchable and easier to compare. Keep the screenshot when the interface includes visual context, a local pack, a map, or a citation panel that the copied text loses.

How should you preserve citations?

Copy every visible cited URL and label it as first-party, directory, review platform, trade source, government source, supplier, or other. Then ask a simple question: does this source support the exact recommendation?

If an answer names your company and links to a page about a different service, flag the mismatch. If it links to a directory with an old phone number, flag the inconsistency. If it cites a local association listing that accurately describes your license or territory, keep it as a useful supporting source.

The source log can reveal a gap that ordinary rankings hide. Your site may rank for “well pump repair” while the answer engine cites a directory because that directory has a cleaner service and area description. The repair may involve both sources.

Who should score ambiguous answers?

Use two reviewers for the first baseline if possible: one owner or service manager and one marketing or operations reviewer. Agree on the rules before you score all 144 answers.

If reviewers disagree, do not quietly choose the higher score. Mark the row “needs review,” write the reason, and resolve the rule. Ambiguity is operational information. It may show that the company itself has not defined which services it wants, where it will dispatch, or what “emergency” means.

Illustration of a well contractor AI visibility evidence log
Illustration of a well contractor AI visibility evidence log
Illustration of a well contractor AI visibility evidence log
Illustration of a well contractor AI visibility evidence log

What score bands are useful for an owner?

Use bands to decide what to do next, not to make a national claim. The following are Brictale operating bands for a repeated, discovery-led test:

Score Operating band What the owner should infer First question
0-15 Invisible The company is rarely named or cannot be matched confidently Can engines retrieve the correct identity, services, and territory?
16-35 Incidental The company appears sometimes, often unevenly or without supporting evidence Which prompt class or territory is producing the mentions?
36-55 Competitive The company has a repeatable presence but does not control the shortlist Which competitors and source types dominate?
56-75 Strong The company is regularly named with meaningful fit and support Can visibility be converted into qualified calls and booked work?
76-100 Dominant in this test The company performs strongly across the chosen prompts and engines Is the prompt set representative, and can the result hold over time?

These bands do not mean a score of 55 is good in every market. A small rural territory with three relevant providers behaves differently from a large metro with many contractors. A score is useful when the owner knows what sample produced it and what action the score changes.

What does a worked example look like?

Consider a hypothetical company tested across 144 valid answer observations. It is named in 48, appears first or in the top three in 26, has a first-party citation in 31, and receives a correct service-territory fit on 43. The figures below are only an example of the calculation. They are not a real contractor result.

  • Mention rate: 48 ÷ 144 = 33.3%, contributing 13.3 of 40 points.
  • Recommendation position: assume the average position component is 41%, contributing 8.2 of 20 points.
  • First-party citation rate: 31 ÷ 144 = 21.5%, contributing 4.3 of 20 points.
  • Service-territory fit: 43 ÷ 144 = 29.9%, contributing 6.0 of 20 points.
  • Total illustrative score: 31.8 out of 100.

That company falls in the incidental band. The next decision is not automatically “publish more content.” The score shows that even its mentions are not consistently supported or correctly matched. The owner should inspect which services and territories produce the 48 mentions, then repair the largest evidence gap.

What if the score rises but qualified calls do not?

Treat that as a business finding, not a measurement failure. Possible explanations include:

  • The prompt set is not used by real buyers.
  • The answer names the company but does not make contact easy.
  • The company is visible for low-margin or unwanted work.
  • Calls are routed poorly or not answered.
  • The answer is not creating enough volume to move the phone.
  • The benchmark is measuring mentions, while the business needs booked-job attribution.

An owner should not expand the score into a success story until it is compared with call quality, estimate requests, and booked jobs.

Why is a mention not enough for a well or pump company?

A mention is an awareness signal. It is not a recommendation quality signal and not a revenue signal.

Imagine an answer saying that a contractor is known for water testing, but the company wants new well drilling. The name is present, yet the result does not help the priority service. Or imagine the engine names a pump technician who is 90 minutes outside the actual service area. A caller may still try, but the business has paid the cost of visibility without receiving a workable opportunity.

For each mention, ask four follow-up questions:

  1. Was the service correct? Drilling, pump repair, replacement, testing, rehabilitation, or another real service?
  2. Was the location correct? Core territory, edge territory, or outside coverage?
  3. Was the company described accurately? Does the answer invent 24/7 response, a physical office, a license, or a service the company does not claim?
  4. Was the source useful? Can the buyer click to a page that confirms the service, territory, and next step?

The fit check is especially important for well contractors because service areas often follow drive time, crew availability, equipment, geology, permitting, or project type. Google’s official local guidance describes local visibility in terms of relevance, distance, and prominence, and says complete business information helps it understand relevance. Google Business Profile Help explains those factors here.

That is why this benchmark includes territory fit instead of copying a generic brand-mention score.

What public evidence helps an answer engine describe a contractor correctly?

Answer engines need retrievable evidence about identity, service, territory, and trust. There is no single page that carries all of it.

Build the evidence stack in five layers.

Layer 1: The official business identity

Check that the company name, phone number, website, operating area, and hours are consistent across the website and public profiles. Use the real-world business name. Do not attach keywords to the name or invent a storefront in a market where the crew does not operate.

Layer 2: Service evidence

Give each priority service a clear page or a clearly labelled section that answers:

  • What the company does.
  • What the service includes and excludes.
  • Which equipment or job types it handles.
  • Which locations it serves.
  • What the first call or estimate process looks like.
  • What the company cannot promise without an inspection or site information.

A page that says “full-service water solutions” is weaker evidence than a page that clearly names well pump replacement, pressure tank service, or new residential well drilling where those are actual services.

The existing Brictale well drilling service page coverage benchmark is the right companion diagnostic when the answer log shows that competitors are cited for services your website does not explain.

Layer 3: Territory evidence

State the actual service area without creating a page for every town. Explain the main territory, the areas served selectively, and any project-specific limits. If an edge location needs an assessment, say so. Clarity protects the owner from both invisibility and bad-fit demand.

The well drilling Google Maps benchmark can be used beside this test because local-pack presence and AI-answer presence are separate observations. A company can be visible in Maps but absent from an answer engine’s shortlist, or the reverse.

Layer 4: First-party proof

Use real photos, project descriptions that can be shared, service explanations, qualifications that are current and permitted to publish, and review themes that describe the work. Do not invent a project or a result for the sake of a citation.

Google’s local guidance says links and reviews can contribute to prominence, alongside relevance and distance. That does not prove that an answer engine uses the same formula. It does show why the public evidence layer remains relevant to local discovery and should be measured rather than ignored.

Layer 5: Third-party corroboration

Look for accurate references in industry associations, local trade groups, licensing or regulatory sources where appropriate, supplier or manufacturer directories, reputable review platforms, and local publications. A third-party listing is not automatically valuable. An old address, duplicate listing, or wrong service can create confusion.

Log which sources answer engines actually cite. The benchmark should tell you whether the gap is a missing first-party page, a weak public profile, a source inconsistency, or a competitor’s stronger evidence footprint.

Illustration of the public evidence stack behind a well contractor AI recommendation
Illustration of the public evidence stack behind a well contractor AI recommendation
Illustration of the public evidence stack behind a well contractor AI recommendation
Illustration of the public evidence stack behind a well contractor AI recommendation

How should you check crawlability before rewriting content?

Check whether the surfaces can access the pages you want retrieved before you pay someone to rewrite them.

Start with Google’s normal indexing requirements. Google says a page must be indexed and eligible to appear with a snippet before it can be a supporting link in AI Overviews or AI Mode. It also recommends making important content available as text, keeping internal links crawlable, and ensuring structured data matches visible text. Google’s AI features documentation lists those fundamentals.

Then inspect the answer-engine controls that apply to the engines you test.

What should a contractor check for ChatGPT search?

OpenAI documents OAI-SearchBot as the crawler used to surface websites in ChatGPT search features. It distinguishes that bot from GPTBot, which is used for content that may contribute to model training. OpenAI says a site opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though it may still appear as a navigational link. Read the current OpenAI crawler documentation here.

The owner’s decision is not “allow every bot blindly.” It is to understand which access setting supports the intended discovery surface, then make a deliberate policy choice with the website operator. If the business wants ChatGPT search visibility, blanket-blocking OAI-SearchBot works against that goal.

What should a contractor check for Perplexity?

Perplexity documents PerplexityBot as the crawler designed to surface and link websites in search results. It separately documents Perplexity-User for user-directed page retrieval. Perplexity’s crawler documentation explains the distinction.

Check robots.txt, WAF rules, server logs, and any security layer that may block legitimate crawlers. Do not treat a user-agent string alone as proof that a request is genuine. Follow the platform’s published guidance for verification and current IP ranges.

What should a contractor check for Claude?

Anthropic documents ClaudeBot for model-development collection, Claude-SearchBot for search quality, and Claude-User for user-directed retrieval. It says disabling Claude-SearchBot can reduce search visibility and disabling Claude-User can prevent retrieval in response to a user query. Anthropic’s help article documents the three roles.

Again, the benchmark records the product surface and access policy. It does not claim that allowing a bot guarantees a mention. Access is eligibility. Recommendation still depends on the information and evidence an engine can retrieve.

Do you need an AI-only file or special schema?

Not as a condition of appearing in Google AI Overviews or AI Mode. Google says there are no additional technical requirements and no special schema.org structured data needed for those features. It also recommends that structured data match the visible text. Google Search Central states this directly.

That does not make technical markup irrelevant. Accurate business and service information can help machines interpret a page, and crawlable HTML helps them access it. The point is to avoid treating an AI-only file as the main visibility plan while the phone number, service area, or priority service page is unclear.

Crawl access can make a page eligible, but it cannot make an inaccurate service claim true.

What should you fix first when the benchmark is low?

Fix the earliest failure in the retrieval chain. The right sequence is usually identity, access, service, territory, proof, then conversion.

1. Is the business entity clear?

If the engine confuses the company with another business, start with identity. Review the official name, phone, website, address or service-area representation, primary category, and major directory records. Check duplicate profiles and old phone numbers.

If the company has a common name, add unambiguous context to the website and public pages. State the operating area, core services, and type of work. Do not solve name ambiguity by stuffing the business name with keywords.

2. Can the selected surfaces access the pages?

Check robots.txt, noindex directives, server errors, WAF blocks, JavaScript-only content, broken internal links, and whether the intended pages are indexed. Ask the website operator for a crawl report, not a promise.

Google’s guide to generative AI features warns against third-party tools claiming internal Google metrics. Use a tool for workflow if it helps, but compare its output with official Search Console and the answer log. Google’s guidance makes that caution explicit.

3. Does the website explain the priority service?

Match the benchmark’s losing prompts to the missing page. If the company is absent for pump replacement, inspect the pump replacement page. If it appears for repair but not drilling, inspect drilling service detail, equipment, territory, and project proof.

Do not add ten pages because one answer was wrong. Identify whether the gap is a missing service, unclear language, weak internal links, or missing proof. One good page that completes a real buyer decision is more valuable than a pile of keyword variants.

4. Does the territory claim match operations?

A company should not ask an engine to recommend it in a territory it does not serve. If the engine keeps producing bad-fit recommendations, tighten the public service-area language and make coverage boundaries clear.

For an edge territory, write what the company actually does. “We evaluate projects in this county based on access, timing, and scope” is more credible than a blanket claim of 24-hour response everywhere.

5. Does public proof support the recommendation?

Look at the source URLs in competitor answers. Are competitors being cited by local associations, detailed service pages, reviews, trade publications, or directories that describe them accurately? Which of those sources can your company earn honestly?

Do not create fake articles or publish a broad claim on a weak directory just to generate a citation. The source should be relevant to the service and territory and should say something a buyer can verify.

6. Can the business handle the demand?

If the company is visible but misses calls, the next investment is operational. Review call routing, after-hours ownership, answer rate, qualification, follow-up, and estimate scheduling before widening the prompt set.

The benchmark is supposed to improve profitable visibility. It is not a reason to make every phone ring with work the crew cannot take.

What does the repair decision tree look like?

Use this sequence for each losing prompt:

  1. Not named anywhere: Check entity clarity, access, indexing, and public evidence.
  2. Cited but not named: Improve the page’s direct service and territory statement, then inspect whether the answer is using the page only for a general fact.
  3. Named for the wrong service: Clarify the service mix across the site and profiles, then check old or conflicting listings.
  4. Named outside the territory: Tighten coverage language and verify Business Profile and directory geography.
  5. Named correctly but low in the list: Compare source depth, reviews, local prominence, service proof, and competitor share.
  6. Named correctly and cited but no calls: Inspect contact paths, call tracking, intake, and buyer fit.
Illustration of a repair sequence for low AI visibility at a well drilling company
Illustration of a repair sequence for low AI visibility at a well drilling company
Illustration of a repair sequence for low AI visibility at a well drilling company
Illustration of a repair sequence for low AI visibility at a well drilling company

How should you connect answer visibility to qualified calls and booked jobs?

Keep answer visibility as a leading signal and booked work as the business outcome. Do not claim that a named mention caused a job unless the evidence supports that conclusion.

Use the following fields in the CRM or call log:

  • First-touch source, if known.
  • “Heard about us from” response in the intake process.
  • Service requested.
  • Territory or job address.
  • Call answered, missed, or returned.
  • Qualified, unqualified, estimate, won, lost, or pending.
  • Competitor named by the caller, if volunteered.
  • Date the inquiry occurred.

Add a simple question to the intake script: “What prompted you to call us today?” If the caller mentions ChatGPT, Google AI, Perplexity, Claude, a recommendation, or a source article, record the exact wording. It will not capture every answer-engine influence, but it creates evidence that analytics cannot.

Use a separate benchmark report for visibility and a revenue report for calls and jobs. Join them by service, territory, and time period where possible. If visibility rises for pump repair but qualified drilling estimates stay flat, the result is not a failed benchmark. It may mean the prompt set is finding a service the company cannot convert or that the buyer demand is too small.

Google says AI feature appearances are included in Search Console’s overall Web performance reporting and describes a Generative AI performance report that shows dimensions such as impressions, pages, countries, devices, and dates in a limited rollout. The Google Search Central announcement describes that report. Use that first-party data for Google-owned visibility, but do not merge it with manually observed ChatGPT, Claude, or Perplexity answers as if they were one metric.

The clean owner dashboard has three rows:

Row What it measures Example decision
Answer visibility Mention, position, citation, fit, and competitor share Repair the missing pump-replacement evidence
Search and Maps visibility Search Console, local visibility, profile actions, and site traffic Compare AI results with existing discovery channels
Revenue outcome Qualified calls, estimates, booked jobs, and job value Keep, change, or stop the work

If the three rows disagree, investigate. Do not pick the one that makes the report look best.

When should a well contractor not invest in AI visibility yet?

Do not make answer visibility the first project when basic business conditions are unresolved.

Wait or keep the test small if:

  • The company’s name, phone, website, or service area is inconsistent across public profiles.
  • The company cannot state which services it wants more of.
  • The business is not answering the calls it already receives.
  • There is no person responsible for estimates and follow-up.
  • The priority service is not described on the website.
  • The company is changing territory or service mix every few weeks.
  • The business cannot publish or verify basic service and company information.
  • There is no way to tell a qualified inquiry from a homeowner, vendor, spam lead, or out-of-area request.

In those conditions, a low score may be accurate but not actionable. Fix the constraint that would prevent a new qualified call from becoming an estimate or job. Then rerun a smaller baseline.

This is also the right moment to distinguish AI visibility from general marketing. A strong website, accurate local profile, clear service pages, good reviews, and competent intake help the whole acquisition system. You do not need an AI-specific content project to justify doing those basics well.

How should you run the benchmark over 90 days?

Use a 90-day cycle so the baseline can inform work without turning every answer fluctuation into a strategy change.

Week 0: Define the measurement

Choose the company, priority services, three territories, four engines or available surfaces, prompt wording, fit rules, and score formula. Save the protocol. Agree on what counts as a valid answer.

Run the full 12-prompt baseline three times per engine. Keep the raw answers and citations. Calculate the overall score and break it down by service, territory, engine, and prompt class.

Weeks 1-2: Repair identity and access

Check the business name, phone, website, categories, service area, hours, duplicate records, robots.txt, noindex, internal links, server errors, and priority-page indexing. Do not rewrite every page at once.

Create a repair list with an owner and due date. If the page is not accessible or the public profile is wrong, content changes alone may not surface the business.

Weeks 3-6: Repair service and territory evidence

Improve the priority service pages and the territory language. Add only information the company can support. Make the main call path clear. Connect the service pages to the relevant commercial hub and to any evidence or project information that is safe to publish.

If the company needs a wider website architecture review, use the existing well drilling SEO guide as the broader system companion. This benchmark should stay focused on measuring answer visibility.

Weeks 7-9: Repair public proof

Review the third-party sources that appear in answers. Correct outdated listings. Improve relevant profiles. Ask for honest reviews through the company’s normal process. Seek legitimate association, supplier, community, or trade references where the company qualifies.

Do not buy or manufacture mentions. A source that is irrelevant, inaccurate, or clearly promotional can create more confusion than authority.

Weeks 10-12: Rerun and compare

Run the same 12 prompts, same three territories, same engines, same three-run method. Add a date and change log. Compare:

  • Overall score.
  • Mention rate by engine.
  • First-party citation rate.
  • Service-territory fit.
  • Competitor share of mentions.
  • Prompt-level wins and losses.
  • Google AI trigger rate where relevant.
  • Qualified calls, estimates, and booked jobs for the same period.

Report the difference as a change from baseline, not as proof that a single change caused the result. Answer engines, rankings, seasonality, competitor activity, and call handling can move at the same time.

How should you report uncertainty?

Always show the run count. For a small internal benchmark, a simple range is useful: report the lowest and highest mention rate across the three runs for each prompt and engine. If the company appears in 1 of 3 runs, say “33% in this three-run sample,” not “the company has 33% visibility.”

For a larger program, a statistician can add confidence intervals or bootstrap estimates. The important behavior is to stop presenting a single answer as a permanent fact. The primary research cited earlier is a good reason to report the distribution, not just the headline score.

Illustration of a 90-day AI visibility benchmark cycle for a well contractor
Illustration of a 90-day AI visibility benchmark cycle for a well contractor
Illustration of a 90-day AI visibility benchmark cycle for a well contractor
Illustration of a 90-day AI visibility benchmark cycle for a well contractor

What mistakes make the benchmark useless?

The most common mistakes are measurement mistakes, not writing mistakes.

Running one prompt once

One answer can be a useful anecdote. It cannot establish a stable visibility rate. Repeat the prompt and save the variation.

Changing the prompt after a bad result

If you change “well pump replacement in County A” to “best water company near me,” the result is not comparable. Keep the old prompt and add the new one as a separate test.

Asking for the company by name

That measures entity recall, not discovery. Include branded prompts as a secondary set, not as the primary benchmark.

Mixing services

Do not score new well drilling and emergency pump repair in one bucket if the company wants different outcomes. Break down the results by service.

Treating every named company as a qualified recommendation

A wrong-territory or wrong-service mention can be a negative business signal. Use fit scoring.

Counting citations without reading them

A link to a supplier page or a general groundwater article is not the same as a first-party service citation. Read the sentence and inspect the URL.

Treating third-party tool scores as internal platform metrics

Google advises caution with tools that claim access to internal Google metrics. A vendor score can be a useful workflow device, but it is not an official Google ranking or AI score unless Google says it is. Read the Google guidance on generative AI measurement.

Ignoring “not triggered” answers

If Google does not show an AI feature, record that fact. Do not turn ordinary Search into an AI answer just to complete the spreadsheet.

Reporting a national benchmark without a sample

The score bands in this article are operating bands. They are not the average, median, or percentile of US well contractors. A national benchmark requires a defined contractor sample, inclusion rules, collection window, territory controls, raw answer archive, and a method that another researcher can inspect.

Publishing unverifiable proof

Do not invent credentials, offices, response times, projects, reviews, or outcomes. An answer engine may repeat the claim, but that does not make it true. A contractor’s public evidence should be more accurate after the benchmark, not more inflated.

What should an owner receive in the final benchmark report?

The report should fit on a few pages and link to the raw log. A useful owner version contains:

  1. Scope: company, services, territories, engines, models if visible, date, and run count.
  2. Headline score: overall score with raw numerator and denominator.
  3. Component scores: mention, position, first-party citation, and fit.
  4. Prompt matrix: 12 rows with service, territory, and result by engine.
  5. Competitor share: the contractors named most often and the source types attached to them.
  6. Source gap: which first-party or third-party evidence the answers used.
  7. Operational risks: wrong territory, unwanted service, inaccurate description, missed-call risk, or capacity constraint.
  8. Repair queue: no more than five prioritized actions, each with an owner and due date.
  9. Revenue link: qualified calls, estimates, booked jobs, and the period used for comparison.
  10. Limitations: answer variation, non-triggered surfaces, missing model labels, location limitations, and any reviewer disagreement.

Do not put the score on the first page without the conditions that produced it. Owners need to know whether “42” means 42 of 100 repeated observations or a vendor’s hidden normalization.

The report also needs a plain-language verdict. For example:

The company is visible for pump repair in the core territory but rarely appears for new drilling. First-party citations are weak, and competitors receive more mentions from local association and directory pages. Fix drilling service evidence and source consistency before expanding the territory.

That sentence is more useful than “AI visibility is moderate.”

No benchmark is credible without the raw answer log behind the score.

What should you do with the result?

Use the benchmark to choose one next action, then rerun the same test after the work has had time to be retrieved.

If the company is invisible, fix identity and access first. If it is cited but not recommended, improve direct service and territory evidence. If it is named for the wrong work, tighten the public service mix. If it is named correctly but loses the shortlist, compare proof and competitor share. If it is visible and fit is strong but jobs do not follow, work on calls, estimates, and capacity.

Brictale’s free territory audit can turn the benchmark into a market-specific action list: which competitors appear, which services and places produce the gap, which public sources support the shortlist, and where the website or intake path is losing qualified demand. The audit is the next step when the owner wants the baseline interpreted against the actual territory rather than a generic score.

The verdict is straightforward. Measure AI answer visibility as a repeated service-territory test. Keep mentions, citations, fit, search visibility, and booked jobs separate. Then fix the earliest constraint you can prove.

A visibility score is a leading signal, not booked-job proof.

FAQ

What is a good AI visibility score for a well drilling company?
There is no published national average for US well drilling companies. Use the score bands in this benchmark as internal operating bands, then compare your own baseline with later runs. A score is useful only when the prompts, engines, locations, run count, and scoring rules stay consistent.
How often should a well contractor run an AI visibility benchmark?
Run a baseline before changes, repeat it after a meaningful 60 to 90 day work cycle, and spot-check a small fixed set monthly. Do not change prompts every week. Consistency matters more than constant checking because answer engines vary.
Should I track mentions or citations?
Track both. A mention measures whether the company is named. A citation measures whether the answer links to a source that supports the recommendation. Also score service-territory fit, because a name without the right service or geography can produce an unqualified call.
Which AI engines should a well drilling company test?
Start with the engines your buyers can access and your team can repeat consistently. A practical four-surface set is ChatGPT search, Google AI features when they appear, Perplexity, and Claude search or user-directed web answers. Record the exact product and model shown on each run.
Can SEO alone improve AI answer visibility?
Foundational SEO is necessary for crawlability and discoverability, but it is not a guarantee of an AI recommendation. Make service, territory, proof, and business identity clear across your site and public profiles, then measure the answer output rather than assuming a ranking proves visibility.
Does a low AI visibility score prove that my marketing is failing?
No. It may show a retrieval gap, a weak service or territory match, missing public proof, crawler restrictions, or simply a noisy measurement. Compare the answer log with Google visibility, calls, estimates, and booked jobs before deciding what to fix.

Sources

  1. [1]Google Search Central: AI features and your website
  2. [2]Google Search Central: Google's guide to optimizing for generative AI features
  3. [3]Google Business Profile Help: Tips to improve your local ranking on Google
  4. [4]OpenAI: Overview of OpenAI Crawlers
  5. [5]Perplexity: Perplexity Crawlers
  6. [6]Anthropic Help Center: Does Anthropic crawl data from the web?
  7. [7]Ronald Sielinski: Quantifying Uncertainty in AI Visibility
  8. [8]Google Search Central Blog: Search Generative AI performance reports in Search Console

More in English

Professional archive

Looking for homeowner guidance?

Brictale's consumer product now organizes symptoms, costs, maintenance, and decisions by home system. Water is available first.

Explore Water →

Published 2026-08-19 · Markdown version