antonsmindstorms.com
robots.txt

Robots Exclusion Standard data for antonsmindstorms.com

Archived Snapshots

Resource Scan

Scan Details

Site Domain	antonsmindstorms.com
Base Domain	antonsmindstorms.com
Scan Status	Ok
Last Scan	2025-10-06T13:18:28+00:00
Next Scan	2025-10-13T13:18:28+00:00

Last Scan

Scanned	2025-10-06T13:18:28+00:00
URL	https://antonsmindstorms.com/robots.txt
Domain IPs	104.21.16.238, 172.67.216.219, 2606:4700:3032::ac43:d8db, 2606:4700:3037::6815:10ee
Response IP	104.21.16.238
Found	Yes
Hash	9126f6c19e816c5d5ef6d75dbe4293b5a4bc3dbd0a2ef9a6c3552a06413f9f8a
SimHash	4674cb13c494

Groups

*

Rule	Path
Allow	/

Rule

Path

Allow

amazonbot

Rule	Path
Disallow	/

Rule

Path

Disallow

applebot-extended

Rule	Path
Disallow	/

Rule

Path

Disallow

bytespider

Rule	Path
Disallow	/

Rule

Path

Disallow

ccbot

Rule	Path
Disallow	/

Rule

Path

Disallow

claudebot

Rule	Path
Disallow	/

Rule

Path

Disallow

google-extended

Rule	Path
Disallow	/

Rule

Path

Disallow

gptbot

Rule	Path
Disallow	/

Rule

Path

Disallow

meta-externalagent

Rule	Path
Disallow	/

Rule

Path

Disallow

googleother
img2dataset
petalbot
chatgpt agent
cotoyogi
echobot bot
echoboxbot
gemini-deep-research
quillbot
sbintuitionsbot
wpbot
yak
applebot
googleagent-mariner
googleother-video
meta-externalfetcher
pangubot
oai-searchbot
qualifiedbot
quillbot.com
awario
diffbot
duckassistbot
meta-externalagent
mycentralaiscraperbot
bigsur.ai
claude-user
cohere-training-data-crawler
panscient.com
perplexity-user
datenbank crawler
mistralai-user/1.0
perplexitybot
phindbot
youbot
ai2bot
amazonbot
chatgpt-user
claudebot
mistralai-user
omgili
andibot
bytespider
google-firebase
iaskspider/2.0
novaact
yandexadditionalbot
facebookbot
google-cloudvertexbot
operator
panscient
semrushbot-ocob
anthropic-ai
ccbot
gptbot
semrushbot-swa
yandexadditional
googleother-image
openai
ai2bot-dolma
claude-searchbot
cloudvertexbot
friendlycrawler
poseidon research crawler
wardbot
brightbot 1.0
firecrawlagent
imagesiftbot
isscyberriskcrawler
linerbot
factset_spyderbot
icc-crawler
meta-externalagent
addsearchbot
applebot-extended
bedrockbot
cohere-ai
crawlspace
netestate imprint crawler
webzio-extended
velenpublicwebcrawler
google-extended
omgilibot
sidetrade indexer bot
thinkbot
tiktokspider
aihitbot
scrapy
shapbot
claude-web
devin
kangaroo bot
meta-externalfetcher
timpibot

Rule	Path
Disallow	/

Rule

Path

Disallow

Comments

As a condition of accessing this website, you agree to abide by the following
content signals:
(a) If a content-signal = yes, you may collect content for the corresponding
use.
(b) If a content-signal = no, you may not collect content for the
corresponding use.
(c) If the website operator does not include a content signal for a
corresponding use, the website operator neither grants nor restricts
permission via content signal with respect to the corresponding use.
The content signals and their meanings are:
search: building a search index and providing search results (e.g., returning
hyperlinks and short excerpts from your website's contents). Search does not
include providing AI-generated search summaries.
ai-input: inputting content into one or more AI models (e.g., retrieval
augmented generation, grounding, or other real-time taking of content for
generative AI search answers).
ai-train: training or fine-tuning AI models.
ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF
AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET.
BEGIN Cloudflare Managed content
END Cloudflare Managed Content

Warnings

`content-signal` is not a known field.

antonsmindstorms.comrobots.txt

Resource Scan

Scan Details

Last Scan

Groups

*

amazonbot

applebot-extended

bytespider

ccbot

claudebot

google-extended

gptbot

meta-externalagent

Comments

Warnings

antonsmindstorms.com
robots.txt