Skip to main content

About the SHC-Scanner crawler

The SHC-Scanner you see in your server logs is the scanner behind Sag3o's website health check. It only visits when someone pastes a URL into Sag3o and asks for a check; it never crawls a site on its own.

How it identifies itself

User-Agent for plain requests (robots.txt, sitemaps, llms.txt, link checks):

SHC-Scanner/1.0 (+https://sag3o.com/bot)

When rendering the page in a browser (Playwright), the User-Agent is browser-shaped but still names the scanner:

Mozilla/5.0 (compatible; SHC-Scanner/1.0; +https://sag3o.com/bot)

1. What it is

Sag3o checks a site's SEO (search engines), GEO (generative engines) and AEO (answer engines) health for its owner. Every scan is triggered by a person entering one URL; the scanner looks at that page and a few of the site's public configuration files, and does not follow links to other pages.

2. Public pages only

The scanner does not sign in, submit forms or carry cookies; it reads only what anyone could open in a browser. What it fetches is used to produce that one report, never to train models or to resell.

3. What one scan fetches

A scan takes about 20–60 seconds and makes these requests to your site:

  • The page being checked: rendered once in headless Chromium (with the images, scripts and styles that page loads), plus one Lighthouse run of the same page.
  • The http:// version of the same page once (no redirects followed), only to see whether it redirects to https.
  • /robots.txt, /llms.txt and /llms-full.txt, once each.
  • The sitemaps listed in robots.txt, else /sitemap.xml and /sitemap_index.xml; for a sitemap index, up to 5 child sitemaps.
  • Up to 20 links on the page, one HEAD request each to see whether they are broken; when HEAD is refused (405 / 403 / 501), a GET asking for the first byte only; a 429 is not retried.
  • Real-user speed data comes from Google's public Chrome UX Report API, which sends no extra request to your site.

4. Frequency

Only when a user asks. Anonymous users have a daily limit, and paid monitoring rescans at most once a week. Each fetch (robots.txt, sitemaps, llms.txt, link checks) times out after 10 seconds, the page render after 30 seconds and the http:// probe after 5 seconds; nothing is retried.

5. How to block it

Add this group to your robots.txt; the scanner honours it (parsed per RFC 9309). Once blocked, Sag3o's report shows the page could not be checked.

User-agent: SHC-Scanner
Disallow: /

6. Contact

Questions, or want us to stop scanning a site? Email support@sag3o.com.