# robots.txt — dkharlanau.github.io # Canonical: https://dkharlanau.github.io/robots.txt # Policy: https://dkharlanau.github.io/legal/ai-crawler-policy/ # IndexNow: public key verification is published at /31ff0ffa67bc49f3bca0a4a719e30fa2.txt; production submissions for public changes run from GitHub Actions on main. # # AI Crawler Policy (documented choice) # See /legal/ai-crawler-policy/ for the public explanation. # ------------------------------------- # OAI-SearchBot → ALLOW (ChatGPT Search visibility desired) # ChatGPT-User → ALLOW (User-facing ChatGPT browsing) # GPTBot → BLOCK (OpenAI training crawler; search bot is separate) # Claude-SearchBot → ALLOW (Claude search visibility desired) # Claude-User → ALLOW (User-facing Claude browsing) # ClaudeBot → BLOCK (Anthropic training crawler) # Google-Extended → BLOCK (Google training crawler) # PerplexityBot → ALLOW (Perplexity search visibility) # CCBot → ALLOW (Common Crawl retrieval) # FacebookBot → ALLOW (Meta search/retrieval) # # Content-Signal is informational, not enforceable: # ai-train=no → do not use for model training without permission # search=yes → search indexing welcome # ai-input=yes → AI retrieval and user-directed answering allowed User-agent: * Allow: / Content-Signal: ai-train=no, search=yes, ai-input=yes Disallow: /DAMA/ Disallow: /agentic-bytes/ Disallow: /TRIZ-bytes/ Disallow: /LLM-prompts/ User-agent: OAI-SearchBot Allow: / User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: PerplexityBot Allow: / User-agent: CCBot Allow: / User-agent: FacebookBot Allow: / # Sitemap references — canonical sitemap index first, then section sitemaps Sitemap: https://dkharlanau.github.io/sitemap.xml Sitemap: https://dkharlanau.github.io/sitemap-pages.xml Sitemap: https://dkharlanau.github.io/sitemap-data.xml Sitemap: https://dkharlanau.github.io/sitemap-atlas.xml