sipoonsanomat.fi
robots.txt

Robots Exclusion Standard data for sipoonsanomat.fi

Resource Scan

Scan Details

Site Domain sipoonsanomat.fi
Base Domain sipoonsanomat.fi
Scan Status Ok
Last Scan2025-12-12T15:47:26+00:00
Next Scan 2025-12-19T15:47:26+00:00

Last Scan

Scanned2025-12-12T15:47:26+00:00
URL https://sipoonsanomat.fi/robots.txt
Redirect https://www.sipoonsanomat.fi:443/robots.txt
Redirect Domain www.sipoonsanomat.fi
Redirect Base sipoonsanomat.fi
Domain IPs 54.246.245.212
Redirect IPs 18.161.111.36, 18.161.111.40, 18.161.111.94, 18.161.111.96
Response IP 13.226.2.112
Found Yes
Hash c38ea49bb367fa5306619c15c9946d414d41eae3f1809603570b5b1a2affc816
SimHash 7230f15d3554

Groups

amazonbot

Rule Path
Disallow /

anthropic-ai

Rule Path
Disallow /

claudebot

Rule Path
Disallow /

claude-web

Rule Path
Disallow /

bytespider

Rule Path
Disallow /

gptbot

Rule Path
Disallow /

chatgpt-user

Rule Path
Disallow /

cohere-ai

Rule Path
Disallow /

ccbot

Rule Path
Disallow /

diffbot

Rule Path
Disallow /

facebookbot

Rule Path
Disallow /

google-extended

Rule Path
Disallow /

imagesiftbot

Rule Path
Disallow /

meta-externalagent

Rule Path
Disallow /

omgilibot

Rule Path
Disallow /

omgili

Rule Path
Disallow /

oai-searchbot

Rule Path
Disallow /

perplexitybot

Rule Path
Disallow /

youbot

Rule Path
Disallow /

googlebot

Rule Path
Disallow /kaupalliset/*.jpg$
Disallow /kaupalliset/*.Jpg$
Disallow /kaupalliset/*.jPg$
Disallow /kaupalliset/*.jpG$
Disallow /kaupalliset/*.jPG$
Disallow /kaupalliset/*.JPg$
Disallow /kaupalliset/*.JpG$
Disallow /kaupalliset/*.JPG$
Disallow /kaupalliset/*.png$
Disallow /kaupalliset/*.Png$
Disallow /kaupalliset/*.pNg$
Disallow /kaupalliset/*.pnG$
Disallow /kaupalliset/*.pNG$
Disallow /kaupalliset/*.PNg$
Disallow /kaupalliset/*.PnG$
Disallow /kaupalliset/*.PNG$
Disallow /kaupalliset/*.gif$
Disallow /kaupalliset/*.Gif$
Disallow /kaupalliset/*.gIf$
Disallow /kaupalliset/*.giF$
Disallow /kaupalliset/*.gIF$
Disallow /kaupalliset/*.GIf$
Disallow /kaupalliset/*.GiF$
Disallow /kaupalliset/*.GIF$

Other Records

Field Value
sitemap https://www.sipoonsanomat.fi/sitemap.xml

Comments

  • Scraping is not allowed for training AI language models, or selling to AI companies
  • Amazon: used to improve/enable Alexa to answer questions
  • Anthropic/Claude: provides no documentation whether these are effective
  • Anthropic/Claude
  • Anthropic/Claude
  • ByteDance LLMs, including Doubao
  • ChatGPT crawler
  • ChatGPT plugins
  • Cohere: associated with Cohere's chatbot
  • Common Crawl
  • Diffbot: collects data to train LLMs
  • Facebook: crawls to improve language models
  • Google: Bard and Vertex AI generative APIs
  • ImagesiftBot: associated with a company that produces models for image generation
  • Meta
  • Omgilibot/webz.io: sells data for training LLMs
  • OpenAI Search
  • Perplexity AI
  • SuSea
  • Disable indexing of native ad images
  • Sitemap