Skip to content
    topical maptopicalmap.app
    FeaturesSolutionsPricingGuidesBlogAboutContact
    Launch app
    1. Home
    2. Blog
    3. Semantic SEO
    4. Getting cited by ChatGPT, Gemini and Perplexity
    AI SearchSemantic SEO

    Getting cited by ChatGPT, Gemini and Perplexity

    AI assistants such as ChatGPT, Gemini and Perplexity now answer questions with links to the pages they used. Each has its own crawlers and its own rules. This article covers how they find pages, which crawlers you should allow, and what makes a page worth citing, using only what each company documents.

    Mohamed YounsSemantic SEO Engineer · Author & system developerOctober 8, 20267 min read
    On this page — 5 sections
    01

    How do AI assistants find web pages?

    Quick answer

    Through their own crawlers, search indexes they build or license, and live fetches when a user asks. ChatGPT search uses OpenAI's crawler plus third-party search providers. Gemini grounds answers in Google Search. Perplexity runs its own crawler and index.

    When you ask an assistant a question that needs fresh information, it searches the web, reads a handful of pages and writes an answer with citations. The pages it can read depend on its sources. OpenAI says ChatGPT search draws on its own crawler and third-party search providers. Google grounds Gemini answers in Google Search. Perplexity documents its own crawler, PerplexityBot.

    The practical result: a page that is well indexed in Google and Bing, and not blocked to the assistants' crawlers, is a candidate everywhere. A real estate agency in Dubai that blocks every unfamiliar bot "to be safe" may be invisible to half of these answers.

    02

    Which crawlers should you allow in robots.txt?

    Quick answer

    Separate search crawlers from training crawlers. OAI-SearchBot lets pages appear in ChatGPT search; GPTBot is for model training. PerplexityBot indexes for Perplexity answers. Google-Extended controls Gemini training and grounding, not Google Search. Allow the search ones if you want citations.

    Each company documents several user agents with different jobs. Blocking a training crawler does not remove you from search answers, and the reverse is also true. Decide each one on purpose:

    User agentCompanyDocumented purpose
    OAI-SearchBotOpenAISurfacing and linking pages in ChatGPT search. OpenAI says sites that block it will not be shown in ChatGPT search answers
    GPTBotOpenAICollecting content that may be used to train models
    ChatGPT-UserOpenAIFetching a page when a user asks ChatGPT to visit it
    PerplexityBotPerplexityIndexing pages to surface and link in Perplexity answers
    Google-ExtendedGoogleUse of content for Gemini training and grounding; does not affect Google Search
    AI crawlers and what their owners say they do. Check each company's docs before changing robots.txt.

    A sensible default for a business that wants to be cited: allow the search crawlers, then decide separately on training. Also check your CDN or firewall. Some security settings block AI crawlers by default, whatever robots.txt says.

    03

    What kind of page gets cited?

    Quick answer

    Pages that answer one question directly and specifically, near the top, with facts the assistant can quote. Assistants cite the passage that supports a sentence in their answer, so a clear definition, a number, or a step list is easier to cite than a long story.

    An assistant writes its answer, then attaches sources to the claims. The easiest page to attach is one where a single passage supports a single claim. "Tinting car windows in Saudi Arabia is allowed up to a set percentage on side windows" is citable if your page states the rule and its source plainly. A page that buries it in paragraph six is not.

    • Question as heading, answer first: the passage under the heading should stand on its own.
    • Specific facts: numbers, dates, names and conditions, with sources for anything official.
    • Information gain: something other pages do not have, such as your own prices, cases or test results.
    • Clean, readable HTML: text in the page, not only in images or behind scripts that crawlers may not run.
    04

    Do assistants cite Arabic pages?

    Quick answer

    Yes, when the question is asked in Arabic and good Arabic pages exist. In many Gulf and Egyptian niches, few Arabic pages answer questions clearly, so a well-structured Arabic page faces less competition for citations than its English equivalent.

    Assistants generally answer in the user's language and prefer sources in that language when they are good enough. This is an opportunity. Ask an assistant in Arabic about end-of-service pay in Saudi Arabia or rental contracts in Egypt, and you often see the same few sources, sometimes translated from English.

    Write the Arabic page natively rather than translating it. Use the words people actually search with, including Gulf or Egyptian terms where they differ, and keep Modern Standard Arabic for the explanation. An assistant can only cite what is on the page.

    05

    How do you track whether assistants cite you?

    Quick answer

    There is no complete report. Check referral traffic from chatgpt.com, perplexity.ai and gemini.google.com in your analytics, run a fixed list of real questions in each assistant every month, and record which pages are cited. Watch the trend, not single answers.

    Assistant answers vary from one run to the next, so a single check proves little. A simple routine works better: pick 20 to 30 questions your customers really ask, run them in each assistant on the same day each month, and note which sources appear. Pair that with referral data in your analytics tool.

    The same work, a new audience

    Getting cited by assistants rarely needs new tricks. It needs pages that are crawlable, clear and specific. The work that earns rankings earns citations too, but only if you have not blocked the crawlers.

    This article is part of the Semantic SEO series — Writing for meaning: entities, definitions, attribute-value facts, semantic distance and pages that actually answer.

    About the author

    MY

    Mohamed Youns

    Semantic SEO Engineer · Author & system developer

    Mohamed Youns writes about how search engines understand content — the same standards he applies when building semantic systems at Nut Hub. nut-hub.org

    FacebookXnut-hub.org
    NewerQuery fan-out: writing for the questions behind the questionOlderHow Google AI Overviews choose the pages they cite

    Related reading

    AI SearchHow Google AI Overviews choose the pages they cite7 min readAI SearchQuery fan-out: writing for the questions behind the question7 min readSERP ClusteringSERP clustering: let the search results decide which keywords share a page6 min read
    All articlesOpen the app

    On this page

    Reading progress

    7 min · 0% read

    topical map

    Topical Map plans, writes and checks every page your site needs to build topical authority, in Arabic and English, from your own brand facts.

    Free during early access · Your own API keys

    Product

    • Overview
    • Features
    • Solutions
    • Open the app
    • Pricing
    • Roadmap
    • Settings

    Resources

    • Guides
    • Blog
    • Help center

    Company

    • About
    • Contact

    Legal

    • Privacy
    • Terms
    • Refund policy

    © 2026 topical map — a Nut Hub product. All rights reserved.