antonsmindstorms.com
robots.txt

Robots Exclusion Standard data for antonsmindstorms.com

Resource Scan

Scan Details

Site Domain antonsmindstorms.com
Base Domain antonsmindstorms.com
Scan Status Ok
Last Scan2025-10-06T13:18:28+00:00
Next Scan 2025-10-13T13:18:28+00:00

Last Scan

Scanned2025-10-06T13:18:28+00:00
URL https://antonsmindstorms.com/robots.txt
Domain IPs 104.21.16.238, 172.67.216.219, 2606:4700:3032::ac43:d8db, 2606:4700:3037::6815:10ee
Response IP 104.21.16.238
Found Yes
Hash 9126f6c19e816c5d5ef6d75dbe4293b5a4bc3dbd0a2ef9a6c3552a06413f9f8a
SimHash 4674cb13c494

Groups

*

Rule Path
Allow /

amazonbot

Rule Path
Disallow /

applebot-extended

Rule Path
Disallow /

bytespider

Rule Path
Disallow /

ccbot

Rule Path
Disallow /

claudebot

Rule Path
Disallow /

google-extended

Rule Path
Disallow /

gptbot

Rule Path
Disallow /

meta-externalagent

Rule Path
Disallow /

googleother
img2dataset
petalbot
chatgpt agent
cotoyogi
echobot bot
echoboxbot
gemini-deep-research
quillbot
sbintuitionsbot
wpbot
yak
applebot
googleagent-mariner
googleother-video
meta-externalfetcher
pangubot
oai-searchbot
qualifiedbot
quillbot.com
awario
diffbot
duckassistbot
meta-externalagent
mycentralaiscraperbot
bigsur.ai
claude-user
cohere-training-data-crawler
panscient.com
perplexity-user
datenbank crawler
mistralai-user/1.0
perplexitybot
phindbot
youbot
ai2bot
amazonbot
chatgpt-user
claudebot
mistralai-user
omgili
andibot
bytespider
google-firebase
iaskspider/2.0
novaact
yandexadditionalbot
facebookbot
google-cloudvertexbot
operator
panscient
semrushbot-ocob
anthropic-ai
ccbot
gptbot
semrushbot-swa
yandexadditional
googleother-image
openai
ai2bot-dolma
claude-searchbot
cloudvertexbot
friendlycrawler
poseidon research crawler
wardbot
brightbot 1.0
firecrawlagent
imagesiftbot
isscyberriskcrawler
linerbot
factset_spyderbot
icc-crawler
meta-externalagent
addsearchbot
applebot-extended
bedrockbot
cohere-ai
crawlspace
netestate imprint crawler
webzio-extended
velenpublicwebcrawler
google-extended
omgilibot
sidetrade indexer bot
thinkbot
tiktokspider
aihitbot
scrapy
shapbot
claude-web
devin
kangaroo bot
meta-externalfetcher
timpibot

Rule Path
Disallow /

Comments

  • As a condition of accessing this website, you agree to abide by the following
  • content signals:
  • (a) If a content-signal = yes, you may collect content for the corresponding
  • use.
  • (b) If a content-signal = no, you may not collect content for the
  • corresponding use.
  • (c) If the website operator does not include a content signal for a
  • corresponding use, the website operator neither grants nor restricts
  • permission via content signal with respect to the corresponding use.
  • The content signals and their meanings are:
  • search: building a search index and providing search results (e.g., returning
  • hyperlinks and short excerpts from your website's contents). Search does not
  • include providing AI-generated search summaries.
  • ai-input: inputting content into one or more AI models (e.g., retrieval
  • augmented generation, grounding, or other real-time taking of content for
  • generative AI search answers).
  • ai-train: training or fine-tuning AI models.
  • ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF
  • RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT
  • AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET.
  • BEGIN Cloudflare Managed content
  • END Cloudflare Managed Content

Warnings

  • `content-signal` is not a known field.