hetacv.be
robots.txt

Robots Exclusion Standard data for hetacv.be

Resource Scan

Scan Details

Site Domain hetacv.be
Base Domain hetacv.be
Scan Status Ok
Last Scan2024-09-19T17:46:52+00:00
Next Scan 2024-10-19T17:46:52+00:00

Last Scan

Scanned2024-09-19T17:46:52+00:00
URL https://hetacv.be/robots.txt
Domain IPs 194.78.53.107
Response IP 194.78.53.107
Found Yes
Hash 915c5e34f5ac29bc7d4a8935c39d9d5c56f0d68fcd13ab5bed409fd39e9144b8
SimHash 3a1f71d1c7c1

Groups

*

No rules defined. All paths allowed.

Other Records

Field Value
crawl-delay 10

googlebot

No rules defined. All paths allowed.

Other Records

Field Value
crawl-delay 10

slurp

No rules defined. All paths allowed.

Other Records

Field Value
crawl-delay 10

msnbot

No rules defined. All paths allowed.

Other Records

Field Value
crawl-delay 10

googlebot/2.1 (+http://www.google.com/bot.html)

Rule Path
Disallow

googlebot-image/1.0

Rule Path
Disallow

googlebot-image

Rule Path
Disallow

mediapartners-google*

Rule Path
Disallow

mozilla/5.0 (compatible; googlebot/2.1; +http://www.google.com/bot.html)

Rule Path
Disallow

yahoofeedseeker/1.0 (compatible; mozilla 4.0; msie 5.5; http://my.yahoo.com/s/publishers.html)

Rule Path
Disallow

yahoo-mmcrawler/3.x (mms dash mmcrawler dash support at yahoo dash inc dot com)

Rule Path
Disallow

mozilla/5.0 (compatible; yahoo! slurp;http://help.yahoo.com/help/us/ysearch/slurp)

Rule Path
Disallow

mozilla/5.0 (compatible; yahoo! slurp china; http://misc.yahoo.com.cn/help.html)

Rule Path
Disallow

mozilla/3.0 (slurp/si; slurp@inktomi.com; http://www.inktomi.com/slurp.html)

Rule Path
Disallow

scooter-3.2.ex

Rule Path
Disallow

mozilla/2.0 (compatible; ask jeeves/teoma)

Rule Path
Disallow

msnbot/1.0 (+http://search.msn.com/msnbot.htm)

Rule Path
Disallow

baiduspider ( http://www.baidu.com/search/spider.htm)

Rule Path
Disallow

ia_archiver

Rule Path
Disallow

ia_archiver/1.6

Rule Path
Disallow

alexibot

Rule Path
Disallow

mj12bot

Rule Path
Disallow

w3c-checklink

Rule Path
Disallow /

w3c_validator/1.432.2.10

Rule Path
Disallow /

xenu's link sleuth 1.3.8

Rule Path
Disallow /

xenu's

Rule Path
Disallow /

black hole

Rule Path
Disallow /

titan

Rule Path
Disallow /

webstripper

Rule Path
Disallow /

netmechanic

Rule Path
Disallow /

cherrypicker

Rule Path
Disallow /

emailcollector

Rule Path
Disallow /

emailsiphon

Rule Path
Disallow /

webbandit

Rule Path
Disallow /

emailwolf

Rule Path
Disallow /

extractorpro

Rule Path
Disallow /

copyrightcheck

Rule Path
Disallow /

crescent

Rule Path
Disallow /

wget

Rule Path
Disallow /

sitesnagger

Rule Path
Disallow /

prowebwalker

Rule Path
Disallow /

cheesebot

Rule Path
Disallow /

mozilla/4

Rule Path
Disallow /

mozilla/5

Rule Path
Disallow /

mozilla/4.0 (compatible; msie 4.0; windows nt)

Rule Path
Disallow /

mozilla/4.0 (compatible; msie 4.0; windows 95)

Rule Path
Disallow /

mozilla/4.0 (compatible; msie 4.0; windows 98)

Rule Path
Disallow /

teleport

Rule Path
Disallow /

teleportpro

Rule Path
Disallow /

miixpc

Rule Path
Disallow /

telesoft

Rule Path
Disallow /

website quester

Rule Path
Disallow /

webzip

Rule Path
Disallow /

moget/2.1

Rule Path
Disallow /

webzip/4.0

Rule Path
Disallow /

websauger

Rule Path
Disallow /

webcopier

Rule Path
Disallow /

netants

Rule Path
Disallow /

mister pix

Rule Path
Disallow /

webauto

Rule Path
Disallow /

thenomad

Rule Path
Disallow /

www-collector-e

Rule Path
Disallow /

rma

Rule Path
Disallow /

libweb/clshttp

Rule Path
Disallow /

asterias

Rule Path
Disallow /

httplib

Rule Path
Disallow /

turingos

Rule Path
Disallow /

spanner

Rule Path
Disallow /

infonavirobot

Rule Path
Disallow /

harvest/1.5

Rule Path
Disallow /

bullseye/1.0

Rule Path
Disallow /

mozilla/4.0 (compatible; bullseye; windows 95)

Rule Path
Disallow /

crescent internet toolpak http ole control v.1.0

Rule Path
Disallow /

cherrypickerse/1.0

Rule Path
Disallow /

cherrypickerelite/1.0

Rule Path
Disallow /

webbandit/3.50

Rule Path
Disallow /

nicerspro

Rule Path
Disallow /

microsoft url control - 5.01.4511

Rule Path
Disallow /

dittospyder

Rule Path
Disallow /

foobot

Rule Path
Disallow /

webmasterworldforumbot

Rule Path
Disallow /

spankbot

Rule Path
Disallow /

botalot

Rule Path
Disallow /

lwp-trivial/1.34

Rule Path
Disallow /

lwp-trivial

Rule Path
Disallow /

wget/1.6

Rule Path
Disallow /

bunnyslippers

Rule Path
Disallow /

microsoft url control - 6.00.8169

Rule Path
Disallow /

urly warning

Rule Path
Disallow /

wget/1.5.3

Rule Path
Disallow /

linkwalker

Rule Path
Disallow /

cosmos

Rule Path
Disallow /

moget

Rule Path
Disallow /

hloader

Rule Path
Disallow /

humanlinks

Rule Path
Disallow /

linkextractorpro

Rule Path
Disallow /

offline explorer

Rule Path
Disallow /

mata hari

Rule Path
Disallow /

lexibot

Rule Path
Disallow /

web image collector

Rule Path
Disallow /

the intraformant

Rule Path
Disallow /

true_robot/1.0

Rule Path
Disallow /

true_robot

Rule Path
Disallow /

blowfish/1.0

Rule Path
Disallow /

jennybot

Rule Path
Disallow /

miixpc/4.2

Rule Path
Disallow /

builtbottough

Rule Path
Disallow /

propowerbot/2.14

Rule Path
Disallow /

backdoorbot/1.0

Rule Path
Disallow /

tocrawl/urldispatcher

Rule Path
Disallow /

webenhancer

Rule Path
Disallow /

tighttwatbot

Rule Path
Disallow /

suzuran

Rule Path
Disallow /

vci webviewer vci webviewer win32

Rule Path
Disallow /

vci

Rule Path
Disallow /

szukacz/1.4

Rule Path
Disallow /

queryn metasearch

Rule Path
Disallow /

openfind data gathere

Rule Path
Disallow /

openfind

Rule Path
Disallow /

zeus

Rule Path
Disallow /

repomonkey bait & tackle/v1.01

Rule Path
Disallow /

repomonkey

Rule Path
Disallow /

zeus 32297 webster pro v2.9 win32

Rule Path
Disallow /

webster pro

Rule Path
Disallow /

erocrawler

Rule Path
Disallow /

linkscan/8.1a unix

Rule Path
Disallow /

keyword density/0.9

Rule Path
Disallow /

kenjin spider

Rule Path
Disallow /

cegbfeieh

Rule Path
Disallow /

ninjabot

Rule Path
Disallow /

googlebot$

Rule Path
Disallow *.css$
Disallow *.js$
Disallow *.pdf$
Disallow *.doc$
Disallow *.docx$
Disallow *.xls$
Disallow *.xlsx$
Disallow *.ppt$
Disallow *.pptx$

googlebot-image/1.0

Rule Path
Disallow *.jpg$
Disallow *.jpeg$
Disallow *.gif$
Disallow *.png$
Disallow /Restricted/
Disallow /App_Data/
Disallow /App_Readme/
Disallow /bin/
Disallow /ClientBin/
Disallow /Content/
Disallow /Customization/
Disallow /Mvc/
Disallow /ResourcePackages/
Disallow /Sitefinity/
Disallow /Global.asax
Disallow /Silverlight.js
Disallow /packages.config
Disallow /web.config

Other Records

Field Value
sitemap https://hetacv.be/sitemap.xml

Comments

  • All
  • Specific
  • Sitemap
  • Allowed User-agents
  • Googlebot
  • Googlebot-Image
  • Google Image
  • Google AdSense
  • Googlebot
  • YahooFeedSeeker
  • Yahoo
  • Yahoo Slurp
  • Yahoo Slurp
  • Inktomi Slurp
  • Scooter
  • Ask Jeeves Teoma
  • Bing
  • Baidu
  • ia_archiver
  • ia_archiver
  • Alexibot
  • MJ12bot
  • Disallowed User-agents
  • Start validation user-agents
  • Rogue user-agents & spam bots
  • Disallowed Googlebot wildcards
  • Disallowed Googlebot-Image wildcards
  • Directories & files
  • Sitefinity
  • Files
  • Default.aspx is the only file
  • that can be accessed by the spiders
  • Even if they are coming from Mars.

Warnings

  • `noindex` is not a known field.