# arvaken.com # Publications are licensed under CC BY 4.0. This file blocks dedicated training crawlers and states a no-training # preference; both are requests the licence does not enforce. See https://arvaken.com/publications#reuse User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / # Crawlers whose operators document them as collecting AI training data, named as the operators name them. # Common Crawl's CCBot is not listed: it is a public archive, and the ai-train=no signal carries the preference. # Names reviewed 2026-09-14. Operators rename crawlers: review this list whenever security.txt is renewed. User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended Disallow: / Sitemap: https://arvaken.com/sitemap.xml