unctad.org
robots.txt

Robots Exclusion Standard data for unctad.org

Archived Snapshots

Resource Scan

Scan Details

Site Domain	unctad.org
Base Domain	unctad.org
Scan Status	Ok
Last Scan	2025-11-01T06:37:59+00:00
Next Scan	2025-12-01T06:37:59+00:00

Last Scan

Scanned	2025-11-01T06:37:59+00:00
URL	https://unctad.org/robots.txt
Domain IPs	104.20.27.93, 172.66.174.241, 2606:4700:10::6814:1b5d, 2606:4700:10::ac42:aef1
Response IP	172.66.174.241
Found	Yes
Hash	f94c1e3378b6b454992b2402c3ebf30100169e9f15b28bfc60eea09e8a24ff4d
SimHash	3d96bd11c64c

Groups

addsearchbot
ai2bot
ai2bot-dolma
aihitbot
amazonbot
andibot
anthropic-ai
applebot
applebot-extended
awario
bedrockbot
bigsur.ai
brightbot 1.0
bytespider
ccbot
chatgpt agent
chatgpt-user
claude-searchbot
claude-user
claude-web
claudebot
cloudvertexbot
cohere-ai
cohere-training-data-crawler
cotoyogi
crawlspace
datenbank crawler
deepseekbot
devin
diffbot
duckassistbot
echobot bot
echoboxbot
facebookbot
facebookexternalhit
factset_spyderbot
firecrawlagent
friendlycrawler
gemini-deep-research
google-cloudvertexbot
google-extended
google-firebase
googleagent-mariner
googleother
googleother-image
googleother-video
gptbot
iaskspider/2.0
icc-crawler
imagesiftbot
img2dataset
isscyberriskcrawler
kangaroo bot
linerbot
meta-externalagent
meta-externalagent
meta-externalfetcher
meta-externalfetcher
meta-webindexer
mistralai-user
mistralai-user/1.0
mycentralaiscraperbot
netestate imprint crawler
novaact
oai-searchbot
omgili
omgilibot
openai
operator
pangubot
panscient
panscient.com
perplexity-user
perplexitybot
petalbot
phindbot
poseidon research crawler
qualifiedbot
quillbot
quillbot.com
sbintuitionsbot
scrapy
semrushbot-ocob
semrushbot-swa
shapbot
sidetrade indexer bot
terracotta
thinkbot
tiktokspider
timpibot
velenpublicwebcrawler
wardbot
webzio-extended
wpbot
yak
yandexadditional
yandexadditionalbot
youbot

Rule	Path
Disallow	/

Rule

Path

Disallow

/

*

Rule	Path
Allow	/core/*.css$
Allow	/core/*.css?
Allow	/core/*.js$
Allow	/core/*.js?
Allow	/core/*.gif
Allow	/core/*.jpg
Allow	/core/*.jpeg
Allow	/core/*.png
Allow	/core/*.svg
Allow	/profiles/*.css$
Allow	/profiles/*.css?
Allow	/profiles/*.js$
Allow	/profiles/*.js?
Allow	/profiles/*.gif
Allow	/profiles/*.jpg
Allow	/profiles/*.jpeg
Allow	/profiles/*.png
Allow	/profiles/*.svg
Disallow	/core/
Disallow	/profiles/
Disallow	/README.md
Disallow	/composer/Metapackage/README.txt
Disallow	/composer/Plugin/ProjectMessage/README.md
Disallow	/composer/Plugin/Scaffold/README.md
Disallow	/composer/Plugin/VendorHardening/README.txt
Disallow	/composer/Template/README.txt
Disallow	/modules/README.txt
Disallow	/sites/README.txt
Disallow	/themes/README.txt
Disallow	/web.config
Disallow	/admin/
Disallow	/comment/reply/
Disallow	/filter/tips
Disallow	/node/add/
Disallow	/search/
Disallow	/*/search/
Disallow	/user/register
Disallow	/user/password
Disallow	/user/login
Disallow	/user/logout
Disallow	/media/oembed
Disallow	/*/media/oembed
Disallow	/index.php/admin/
Disallow	/index.php/comment/reply/
Disallow	/index.php/filter/tips
Disallow	/index.php/node/add/
Disallow	/index.php/search/
Disallow	/index.php/*/search/
Disallow	/index.php/user/password
Disallow	/index.php/user/register
Disallow	/index.php/user/login
Disallow	/index.php/user/logout
Disallow	/index.php/media/oembed
Disallow	/index.php/*/media/oembed

Rule

Path

Allow

/core/*.css$

Allow

/core/*.css?

Allow

/core/*.js$

Allow

/core/*.js?

Allow

/core/*.gif

Allow

/core/*.jpg

Allow

/core/*.jpeg

Allow

/core/*.png

Allow

/core/*.svg

Allow

/profiles/*.css$

Allow

/profiles/*.css?

Allow

/profiles/*.js$

Allow

/profiles/*.js?

Allow

/profiles/*.gif

Allow

/profiles/*.jpg

Allow

/profiles/*.jpeg

Allow

/profiles/*.png

Allow

/profiles/*.svg

Disallow

/core/

Disallow

/profiles/

Disallow

/README.md

Disallow

/composer/Metapackage/README.txt

Disallow

/composer/Plugin/ProjectMessage/README.md

Disallow

/composer/Plugin/Scaffold/README.md

Disallow

/composer/Plugin/VendorHardening/README.txt

Disallow

/composer/Template/README.txt

Disallow

/modules/README.txt

Disallow

/sites/README.txt

Disallow

/themes/README.txt

Disallow

/web.config

Disallow

/admin/

Disallow

/comment/reply/

Disallow

/filter/tips

Disallow

/node/add/

Disallow

/search/

Disallow

/*/search/

Disallow

/user/register

Disallow

/user/password

Disallow

/user/login

Disallow

/user/logout

Disallow

/media/oembed

Disallow

/*/media/oembed

Disallow

/index.php/admin/

Disallow

/index.php/comment/reply/

Disallow

/index.php/filter/tips

Disallow

/index.php/node/add/

Disallow

/index.php/search/

Disallow

/index.php/*/search/

Disallow

/index.php/user/password

Disallow

/index.php/user/register

Disallow

/index.php/user/login

Disallow

/index.php/user/logout

Disallow

/index.php/media/oembed

Disallow

/index.php/*/media/oembed

Back to top

Comments

robots.txt
This file is to prevent the crawling and indexing of certain parts
of your site by web crawlers and spiders run by sites like Yahoo!
and Google. By telling these "robots" where not to go on your site,
you save bandwidth and server resources.
This file will be ignored unless it is at the root of your host:
Used: http://example.com/robots.txt
Ignored: http://example.com/site/robots.txt
For more information about the robots.txt standard, see:
http://www.robotstxt.org/robotstxt.html
Politely ask AI rebots to leave.
From https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.txt
Default Drupal rules
CSS, JS, Images
Directories
Files
Paths (clean URLs)
Paths (no clean URLs)

Back to top

unctad.orgrobots.txt

Resource Scan

Scan Details

Last Scan

Groups

*

Comments

unctad.org
robots.txt