# Search engine crawlers welcome. The rules below stay directly under this # group, above any comment blocks, so even strict parsers that treat a blank # line as a record separator still apply them to User-agent: *. User-agent: * Allow: / Disallow: /*?cat= Disallow: /*&cat= Disallow: /*?q= Disallow: /*&q= Disallow: /*?n= Disallow: /*&n= Disallow: /*?sort= Disallow: /*&sort= Disallow: /api/ # Filter, search and show-all views. Every one of them canonicalises to the # page it was reached from, so none is meant to be indexed, but a crawler has # to fetch a URL before it can learn that. The category filter is the reason # this matters: each toggle link is built from the current selection and keeps # its order, so ?cat=A&cat=B and ?cat=B&cat=A are both reachable, and with 11 # categories offered on all 335 video pages that is a combinatorial space no # crawl budget survives. # # Two parameters are deliberately NOT blocked. ?p= is the pagination that makes # 62 mid-archive videos reachable at all, and ?t= carries the timestamp shares: # social scrapers honour robots.txt too, so blocking it would break every # shared-moment link preview. # # The legal pages carry meta noindex, which is the directive that actually # keeps them out of the index. Disallowing them as well would be self # defeating: a page that is never fetched is a page whose noindex is never # read, and it could still be indexed URL-only from an inbound link. # AI training crawlers blocked User-agent: CCBot Disallow: / User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: anthropic-ai Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: Google-Extended Disallow: / User-agent: FacebookBot Disallow: / # Meta's AI crawler. It ignores the ?cat= disallow group above and walks the # category-filter permutations on every video page, so it is the single largest # source of load on this site. meta-externalfetcher is its link-preview fetcher. User-agent: meta-externalagent Disallow: / User-agent: meta-externalfetcher Disallow: / User-agent: Bytespider Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Amazonbot Disallow: / # Crawlers that walk every shared-moment URL (/v//t/) although the # page canonicalises to the video and is never worth indexing on its own. They # keep the rest of the site, and lose only those URLs. Search engines and the # link-preview fetchers are deliberately not listed: Google and Bing follow the # canonical, and a preview fetcher has to read the moment page to show a card. # A named group replaces the wildcard group, so the default rules are repeated. User-agent: Amzn-SearchBot User-agent: PetalBot User-agent: LinkupBot User-agent: Reflectionbot Disallow: /v/*/t/ Disallow: /*?cat= Disallow: /*&cat= Disallow: /*?q= Disallow: /*&q= Disallow: /*?n= Disallow: /*&n= Disallow: /*?sort= Disallow: /*&sort= Disallow: /api/ # High-volume bulk crawler with no search or AI-answer surface that would send # visitors here. It honours robots.txt, so it is blocked outright rather than # left to hammer the archive. User-agent: SleepBot Disallow: / Sitemap: https://askhosk.com/sitemap.xml