# robots.txt — taginsight.com (site vitrine) # Objectif : couper les crawlers IA / scrapers qui consomment de la bande passante # sans valeur SEO. Les moteurs classiques (Google, Bing) restent autorisés. # NB : seuls les bots "polis" respectent ce fichier. Les scrapers malveillants # l'ignorent — pour ceux-là il faut un WAF (cf. migration Cloudflare Pages). # --- Crawlers IA / entraînement de modèles --- User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: OAI-SearchBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: anthropic-ai Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: meta-externalagent Disallow: / User-agent: FacebookBot Disallow: / User-agent: Diffbot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / # --- Tout le reste : autorisé (Googlebot, Bingbot, etc.) --- User-agent: * Allow: / Sitemap: https://www.taginsight.com/sitemap.xml