apache禁止搜索引擎收录、网络爬虫采集的配置方法

时间：2022-03-07 08:52:32|栏目：Linux|点击：次

Apache中禁止网络爬虫，之前设置了很多次的，但总是不起作用，原来是是写错了，不能写到Dirctory中，要写到Location中

 
 <Location />
 
 SetEnvIfNoCase User-Agent "spider" bad_bot
 
 BrowserMatchNoCase bingbot bad_bot
 
 BrowserMatchNoCase Googlebot bad_bot
 
 Order Deny,Allow
 
 #下面是禁止soso的爬虫
 
 Deny from 124.115.4. 124.115.0. 64.69.34.135 216.240.136.125 218.15.197.69 155.69.160.99 58.60.13. 121.14.96. 58.60.14. 58.61.164. 202.108.7.209
 
 Deny from env=bad_bot
 
 </Location>

这是禁止了所有包含spider字符的爬虫。
如果要针对性的禁止爬虫，改成精确匹配的爬虫字符串，如果bingbot、Googlebot等等

上一篇：101个脚本之建立linux回收站的脚本

栏目：Linux

下一篇：Linux中无法远程连接数据库问题的解决方法

本文标题：apache禁止搜索引擎收录、网络爬虫采集的配置方法

本文地址：http://www.codeinn.net/misctech/195499.html

更多Linux

Linux

apache禁止搜索引擎收录、网络爬虫采集的配置方法

阅读排行

推荐教程