Unfortunately, not all search bots and spiders comply with robots exclusion rules; nor do they have to either. While we’re not lawyers (and we could be wrong), as far as we’re aware, there is no U.S. law prohibiting search engines from ignoring robots.txt exclusions on websites. However, that doesn’t mean there’s no point in using them; as most of the major search engines comply with the robots.txt exclusions, including Google and Bing/Yahoo!

What search engines do not comply with robots.txt exclusions?

We have suspicion to believe Baidu, a popular search engine in China, does not comply with robots.txt exclusions. However, some smaller search engines worldwide may also not comply with robots.txt exclusions.

How do I prevent non-abiding search engines from crawling my website or a specific folder?

The only plausible way to prevent is likely to blacklist their IPs on your Dedicated Server or your individual hosting account. You may also need to blacklist an IP range in order for it to be more effective. Otherwise, password-protect the specific directory.

Popular Posts You May Read

Explore more hosting insights, tips and industry updates.

Xen Dedicated Server

On Linux, there are applications like kernel (Xen, KVM) allows you to do virtualization on…

Apache-Vs-NGINX-Which-Is-The-Best-Web-Server-for-You-BLOG

NGINX Vs Apache – Which Is The Best Web Server for You?

The Internet was born in the 1990s. The entire “Web” protocol can be summarised as…

VPS Shutting Down

VPS Shutting Down? Don’t Worry bodHOST to Your Rescue

THE BUZZ AROUND!! Is your current VPS shutting down or leaving you stranded? Don’t worry…