All Wiki
Tag
#crawlers (4)
robots.txt — Robots Exclusion Protocol
analysisThe original machine-readable web standard: a plain-text file at a site's root telling crawlers which paths they may visit. Advisory, not enforced.
robots.txt — Robots Exclusion Protocol
sourceThe original machine-readable web standard: a plain-text file at a site's root telling crawlers which paths they may visit. Advisory, not enforced.
Sitemaps Protocol (sitemap.xml)
analysisThe positive complement to robots.txt: an XML file declaring what URLs exist on a site and what the operator wants crawled first.
Sitemaps Protocol (sitemap.xml)
sourceThe positive complement to robots.txt: an XML file declaring what URLs exist on a site and what the operator wants crawled first.