Back to blog
GEOJuly 23, 20268 min read

Eight questions to ask your current agency about GEO

Your agency has never brought up GEO. Before you replace anyone, ask these eight questions — starting with whether your site lets AI crawlers in at all.

If your marketing agency has never brought up GEO, the fastest way to find out whether that matters is to ask them eight specific questions. Start with the most concrete one: does your site currently allow AI crawlers at all? A site can rank perfectly well on Google and still be invisible inside ChatGPT for a reason as mundane as one line in robots.txt. The questions below are the ones we would send an account manager, along with what a competent answer sounds like.

None of this assumes your agency is doing bad work. Most were hired to win Google rankings and are still winning them. The point is to find out whether anyone is watching the other surface — the one where the answer appears above the links.

Are AI crawlers even allowed on our site?

This is the first question because it is binary, cheap to check, and it silently undoes everything else. The engines that answer questions do not use Googlebot. They run their own agents, and each vendor separates them by purpose:

  • OpenAI runs OAI-SearchBot to surface sites in ChatGPT's search features, GPTBot for model training, and ChatGPT-User for pages fetched live when someone asks a question (OpenAI crawler documentation).
  • Anthropic runs ClaudeBot, Claude-User, and Claude-SearchBot, controlled per-agent through robots.txt (Anthropic crawler documentation).
  • Perplexity runs PerplexityBot to surface and link sites in its results, and Perplexity-User for live user requests (Perplexity crawler documentation).

Note what OpenAI's own documentation says about the search agent: sites excluded from OAI-SearchBot will not appear in ChatGPT search answers. That is a switch, not a strategy. Plenty of sites had these agents blocked during the 2023–2024 wave of "block the AI scrapers" advice, and nobody ever went back to reconsider it. Ask your agency to read you the current robots.txt line by line.

Is Google-Extended blocked, and do we know what that costs?

This one is worth separating out, because the tradeoff is genuinely subtle and it is where a thoughtful agency shows its work. Google-Extended is the control publishers use to manage whether crawled content may be used to train Gemini models and for grounding — supplying search index content to improve factual accuracy. Google's documentation is explicit that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal" (Google crawler documentation).

So blocking it costs you nothing in Google Search and may cost you presence in Gemini's answers. That is a real business decision with arguments on both sides. The wrong answer is not "we block it" — the wrong answer is "we don't know."

Which engines mention us today, and who gets named instead?

Nobody improves what they have not measured. A serious answer looks like a fixed set of category questions, asked in a clean session across ChatGPT, Perplexity, Gemini, and Claude, with three things recorded each time: were you mentioned, were you cited with a link, and who was named in your place.

That last column is the useful one. The brands that show up when a buyer asks your category question are your actual competitive set for AI-referred demand, and they are frequently not the competitors on your quarterly slide.

If your agency proposes a content plan before showing you that snapshot, they are guessing.

How will you measure traffic that comes from AI answers?

Ask them to pull the report in front of you. In GA4 this lives in the traffic acquisition report, segmented by session source (Google Analytics documentation), and the hostnames worth watching are chatgpt.com, perplexity.ai, gemini.google.com, claude.ai, and copilot.microsoft.com.

Two failure modes to listen for. One is an agency that cannot name the sources it would look for — nothing is being tracked. The other is an agency that reports AI referrals as a single blended "referral" number, which tells you a channel exists but nothing about which engine to work on.

Does our content answer the question in the first paragraph?

Generative answers changed what the top of a page is for. Google now places AI-generated overviews above the classic list of links for many queries (Google's announcement of AI Overviews), and the passage a model lifts tends to be a short, self-contained one that answers the question directly.

Pull up any article your agency published this quarter and check three things:

  • Does the first paragraph answer the title's question, or does it open with company history and industry context?
  • Are the headings phrased as questions a buyer would actually type?
  • Does every non-obvious claim carry a link to a source someone can open?

The third one is where most agency content — including plenty of ours before we tightened our own rules — falls down. An unsourced statistic reads as unverifiable, and unverifiable text is a weak thing to build an answer on.

What are we doing with structured data?

GEO is not only writing. Engines lean on machine-readable signals: structured data marked up in JSON-LD so the page states plainly what it is and who published it (Google structured data documentation), and consistent entity information — the same organization name, address, and description everywhere, using vocabulary like schema.org/Organization.

You do not need to understand the implementation. You need to confirm that someone on their side does, and that it is on a roadmap with a date.

What is your position on llms.txt?

This question is a calibration test more than a technical one. llms.txt is a proposal published by Jeremy Howard on 3 September 2024 for a markdown file at the domain root that gives models a curated, context-window-friendly map of a site (the llms.txt proposal). It is still an open community proposal — no major engine has announced that it treats the file as a ranking or citation requirement.

That makes both extreme answers informative. An agency selling llms.txt as a mandatory fix is overclaiming on something unproven. An agency that has never heard of it is not reading the space it would be charging you to work in. The answer you want sits in the middle: they know what it is, they have a view on whether it is worth the hour it takes, and they can say why.

Who actually writes our English?

If you sell in Mexico and to the U.S. Hispanic market, your content competes for citations in two languages. Translating one version into the other produces a second-hand text that a native reader notices immediately. Ask directly who writes each language, and whether the two versions are written or converted.

We go deeper on this — and on the rest of the evaluation — in our guide to choosing a GEO agency, which is the companion piece to this one for when the conversation moves from fixing to replacing.

What does a good answer sound like?

A good answer is specific and unflattering in places. It sounds like: here is what your robots.txt says today, here is where you are mentioned and where you are not, here is the report we would track, here is what we would fix first and roughly what it costs. It does not sound like a promise that you will appear in ChatGPT — nobody controls that, and anyone who says otherwise is describing an outcome they cannot deliver.

If the conversation stalls, get the snapshot independently first. We run one at no cost: a 30-minute audit that shows you, with screenshots, what each engine answers when someone asks about your category. Take it back to your agency as a baseline. A good one will be glad to have it.

Frequently asked questions

What should I ask my agency if they have never mentioned GEO?

Start with the concrete, checkable questions rather than strategy. Ask what your robots.txt currently says about AI crawlers, which engines mention your brand today and who is named instead, how they would measure traffic arriving from AI answers in GA4, and what is on the roadmap for structured data. Specific answers with dates are the signal; general enthusiasm about artificial intelligence is not.

Can my site rank well on Google but still be invisible in ChatGPT?

Yes. AI engines do not use Googlebot — they operate their own agents, such as OpenAI's OAI-SearchBot for ChatGPT's search features, Anthropic's Claude-SearchBot, and PerplexityBot. OpenAI's documentation states that sites excluded from OAI-SearchBot will not appear in ChatGPT search answers. If those agents are blocked in robots.txt, your Google rankings will look completely healthy while you are absent from AI answers.

Does blocking Google-Extended hurt my Google rankings?

No. Google's crawler documentation states that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal. What it does control is whether your crawled content may be used to train Gemini models and for grounding. Blocking it is a legitimate business decision — but it is a decision, and your agency should be able to tell you which way it is currently set and why.

How do I measure traffic that comes from ChatGPT or Perplexity?

In GA4, open the traffic acquisition report and segment by session source, then look for the hostnames chatgpt.com, perplexity.ai, gemini.google.com, claude.ai, and copilot.microsoft.com. Ask your agency to pull this live rather than describe it. Reporting these visits blended into a single referral number tells you a channel exists but not which engine to work on.

Is llms.txt required for AI engines to cite my site?

No. llms.txt is an open proposal published by Jeremy Howard on 3 September 2024 for a markdown file at the domain root that gives models a curated map of a site. No major engine has announced that it treats the file as a citation or ranking requirement. Treat an agency that sells it as mandatory with the same caution as one that has never heard of it.

Do I need to replace my agency to start doing GEO?

Not necessarily. Most agencies were hired to win Google rankings and are still doing that job well; GEO is usually a gap in scope, not in competence. Get an independent baseline of where the engines mention you, hand it to your current agency, and judge them on the response. If the answer is specific and includes what they would fix first, keep them. If it stays abstract after a second conversation, then compare alternatives.

About the author

Diego CaballeroFounder of HopperCat. Builds AI-first systems for Spanish-speaking businesses.