User-agent: * Allow: / # Cloudflare's email-obfuscation landing page. Scrape Shield rewrites mailto: # links to this URL; with the feature off it 404s, and Google has been holding # one since a 2026-08-07 crawl. Scoped to the one endpoint on purpose: blocking # all of /cdn-cgi/ would also hide email-decode.min.js, which Googlebot needs to # render the page if Scrape Shield is ever switched back on. Disallow: /cdn-cgi/l/email-protection Sitemap: https://novamsl.com/sitemap.xml # Text and data mining rights are RESERVED. See /.well-known/tdmrep.json for the # machine-readable declaration and /terms for the policy. # # The distinction below is the whole point: AI providers run separate crawlers # for search and for training. Crawlers that cite and link back are welcome. # Crawlers that absorb the content into a model are not. # ---- Allowed: these cite and link back ---- User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # ---- Disallowed: model training ---- User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / # Google-Extended has no search/training split: it governs Gemini grounding AND # training together. Disallowing it reserves training at the cost of Gemini # citation. Allow it if that reach matters more than the reservation. User-agent: Google-Extended Disallow: / # ---- Disallowed: bulk dataset crawlers, no citation and no referral ---- User-agent: CCBot Disallow: / User-agent: Amazonbot Disallow: /