# This site's content is provided for human readers and disclosed search/AI-answer citation only. # Commercial scraping, bulk extraction, or use of this content to train AI models is not authorized. # See /terms for our full Terms of Service. User-agent: * Allow: / # ===================================================== # AI training crawlers — disallowed. We do not permit # our content to be copied into AI model training sets. # ===================================================== User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Google-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: FacebookBot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Amazonbot Disallow: / User-agent: cohere-ai Disallow: / User-agent: Diffbot Disallow: / User-agent: Omgilibot Disallow: / User-agent: img2dataset Disallow: / User-agent: HuggingFaceBot Disallow: / User-agent: Ai2Bot Disallow: / User-agent: MistralAI-Bot Disallow: / User-agent: xAI-Bot Disallow: / User-agent: Timpibot Disallow: / # AI retrieval/citation crawlers — explicitly allowed. # These answer live user queries and cite back to us. User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Claude-Web Allow: / # Traditional search engines — explicitly allowed. User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / Sitemap: https://newsanarchist.com/sitemap.xml Sitemap: https://newsanarchist.com/sitemap-news.xml