# robots.txt for www.tmt.org # Updated 2026-07 in response to crawler-induced load spikes. # Updated 2026-08: Baiduspider added to the blocked list (also enforced # at the load balancer since 2026-08-04). # Keeps article/page/image content crawlable for mainstream search engines, # but blocks the expensive dynamic endpoints (search, downloads, # parameterized listing pages) and asks crawlers to pace themselves. User-agent: * Crawl-delay: 10 Disallow: /search Disallow: /download/ Disallow: /news? Disallow: /in-the-news? Disallow: /images? Disallow: /multimedia? Disallow: /users/ Disallow: /admin/ Disallow: /collaborate/ # SEO-index bots that send no visitor traffic; full block. User-agent: MJ12bot Disallow: / User-agent: SemrushBot Disallow: / # These largely ignore robots.txt (enforced at the load balancer # instead), but the policy is stated here for completeness. User-agent: YisouSpider Disallow: / User-agent: Bytespider Disallow: / User-agent: Baiduspider Disallow: /