# ── Search + AI-knowledge crawlers (Google, Bing, GPTBot, ClaudeBot, …) ── # Allowed everywhere except the low-value sections below. Crawl-delay is # honoured by Bing, Yandex, and most minor bots (Google ignores it and uses the # Search Console rate setting) — it spaces out crawl hits to curb edge-request load. User-agent: * Allow: / Crawl-delay: 10 # Sections not worth crawling. Private pages are deliberately absent from this # list — they are gated at the edge (middleware.ts) and noindex'd via # X-Robots-Tag, and naming them here would only publish where they are. Disallow: /api/ Disallow: /geophanies/ Disallow: /job/ # ── Abusive / zero-value scrapers ─────────────────────────────────────── # SEO-backlink and aggregation crawlers that hammer sites for data we gain # nothing from. Fully disallowed to cut parasitic edge-request load. User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: PetalBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / Sitemap: https://globaia.org/sitemap-index.xml