About our crawler
We visit public business websites to work out what a business does and where it works. Here is exactly how it behaves — and how to stop it.
How to identify it
It sends this user agent string:
cornwall_hairbot/0.1 (+https://104-248-161-77.sslip.io/about/bot; contact hello@cornwallhairdressers.co.uk)
What it does
- Reads and honours
robots.txtbefore fetching anything. - Fetches your home page, and sometimes your contact or about page. That is the whole visit.
- Waits at least 2.0 seconds between requests, and longer if
your
robots.txtsets aCrawl-delay— we use whichever is greater. - Keeps a short text extract, your council area, and whether the site displays a company registration number. Not the whole page.
- Deletes stored text after 90 days.
What it does not do
- Submit forms, or collect email addresses for marketing.
- Ignore
robots.txt, or crawl behind a login. - Crawl your whole site, or come back repeatedly.
If your site is behind a CDN — Cloudflare, or anything your host puts in front of it — our requests almost certainly never reach your server at all.
To block it
Add this to your robots.txt and we will stop:
User-agent: cornwall_hairbot
Disallow: /
Or ask us to remove your listing and we will not come back.