What this tool is actually checking
Most SEO tools read your sitemap and tell you what is in it. That hides the most expensive problem a site can have, because a sitemap is a list you write about yourself. It is not proof that anything links to those pages.
Google follows links. A page listed in your sitemap but linked from nowhere gets discovered, gets crawled occasionally, and receives almost no authority from the rest of your site. It sits at position 20 or worse and quietly earns nothing, which looks like a content problem but is really a plumbing problem.
This crawler starts at your homepage, follows every internal link it finds, then compares what it reached against what your sitemap declares. The gap between those two numbers is usually the most useful thing you will learn about your website this month.
The numbers that matter, and what they mean
- Pages reached vs pages declared. If your sitemap lists 800 URLs and the crawl reaches 30, then roughly 770 pages are effectively invisible to the part of Google's system that decides what ranks.
- Click depth. How many clicks from the homepage to the deepest page found. A max depth of 1 means your site is completely flat, with no topic hierarchy for Google to read. Three to four is healthy for most sites.
- Pages with no inbound links. Pages the crawl reached but which nothing else points at. These are dead ends holding authority that never circulates.
- Duplicate titles. Two pages with the same title are competing with each other. Google picks one, usually not the one you wanted.
- Thin pages. Under about 300 words. Sometimes fine, often a page that was started and never finished.
- Broken links and redirect chains. Every one wastes crawl budget and, when a visitor hits one, a customer.
What to do with the results
In order of return:
- Fix broken links first. They cost you visitors today, not in three months.
- Then connect the orphans. Build hub pages that group your content by topic, link those hubs from your main navigation, and link every article to the hub and to two or three related articles. This is what turns a pile of pages into a structure.
- Then fix duplicate titles, so your pages stop competing with each other.
- Then the thin pages. Either finish them or remove them. A half-written page helps nobody.
If you want the reasoning behind the internal linking part, I have written it up in the blog, and the other free tools cover the rest of the technical checks.
Frequently asked questions
How is this different from Screaming Frog?
Screaming Frog is a desktop application that crawls without a page limit and is the industry standard. This runs in your browser with no install and no licence, caps at 40 pages, and focuses on the single check most people never run: comparing what your site links to against what your sitemap claims exists. If you want unlimited crawling, use Screaming Frog. If you want the orphan answer in 30 seconds, use this.
Why only 40 pages?
Because it runs on my server and fetches your site politely, with a pause between requests. Forty pages is enough to establish your site's structure, your click depth and your sitemap gap, which is where the useful findings are. The pattern rarely changes at page 400.
Does it change anything on my site?
No. It only reads pages, exactly as a search engine would. Nothing is modified and nothing is stored beyond the crawl you just ran.
My sitemap gap is huge. Is that always bad?
Not always. Large ecommerce and directory sites legitimately rely on category pages and pagination, and some pages are intentionally unlinked. But if your gap is most of your site and your rankings sit around position 20, that gap is very likely your main problem.
Can I crawl a competitor's site?
Yes. It only reads publicly available pages. Looking at how a competitor structures their internal links, and how deep their site goes, is one of the more useful competitive checks available.