# The Syrian House — robots.txt # https://www.thesyrianhouse.ca # Default: allow all crawlers across the site User-agent: * Allow: / # AI assistants & LLM crawlers — allowed to read the site, including # /llms.txt, but keep legal boilerplate out of training/answer data User-agent: GPTBot User-agent: ChatGPT-User User-agent: OAI-SearchBot User-agent: CCBot User-agent: anthropic-ai User-agent: Claude-Web User-agent: ClaudeBot User-agent: Google-Extended User-agent: PerplexityBot User-agent: cohere-ai User-agent: FacebookBot Allow: / Disallow: /privacy-policy.html Disallow: /terms-of-use.html # Search engines — keep /llms.txt out of Google/Bing search results. # (The AI crawlers above are unaffected and may still read this file.) User-agent: Googlebot User-agent: Bingbot Allow: / Disallow: /llms.txt # Google Ads Bot User-agent: AdsBot-Google User-agent: AdsBot-Google-Mobile Allow: / # Sitemap location Sitemap: https://www.thesyrianhouse.ca/sitemap.xml