SEO Audit

Enter a URL and the crawler checks up to 200 pages the way a search engine's first pass would: status codes, noindex and canonical directives, missing or duplicate titles and descriptions, missing H1s, thin pages, images without alt, redirect chains, broken links, orphan pages, robots.txt and sitemap coverage. Every page gets the same on-page checks as the single-page checker; the report groups them by issue and by page. No score, no fabricated priorities — findings with counts. Download as CSV. Nothing is stored beyond 24 hours.

HEAD request per unique external URL, up to 100. Untick for a faster crawl.

The crawl runs on our server: one request per second, robots.txt obeyed, no JavaScript. Only the start URL and this summary are kept, for 24 hours.

How it works

1. Our crawler (InsighthackerzBot, one request per second, no JavaScript) fetches robots.txt and any sitemaps it declares, then starts at the URL you give and follows internal links breadth-first — same host only, www and the bare domain counted as one site
2. Limits: 50, 100 or 200 pages per run, 2 MB per page, 15 s per request, 5 redirects per URL; URLs your robots.txt disallows for our user-agent are skipped and counted, pages marked nofollow are fetched but their links are not followed
3. Every HTML page goes through the same on-page engine as the on-page SEO checker: title, description, H1, canonical, robots directives, word count, images without alt — the warnings and fails are kept per page, the HTML is not
4. Internal links found broken during the crawl (4xx, 5xx, timeouts) are listed with the pages that link to them; external links are checked with HEAD only (one GET when the server refuses HEAD), at most 100 per run, nofollow links excluded
5. The result is one job on our own server: it holds only the start URL and the summary, expires after 24 hours and is never used for anything else
6. Per page: the on-page checker's warnings and fails (title, description, H1, canonical, robots, lang, viewport, word count, images, links, social tags) plus the page's HTTP status and any fetch error. Site-wide: duplicate titles and descriptions across indexable pages, redirect chains, orphan pages, robots.txt presence, sitemap presence and the two coverage gaps
7. Counts are over crawled pages only; 'thin' means fewer than 300 words on an indexable page, which is a convention, not a rule. There is no score: a number that weighs a missing alt against a noindex would be an opinion presented as a measurement

About site audits and what this one reports

A technical SEO audit answers three questions before anything about content: can search engines reach the pages, are they allowed to index them, and does each page say clearly what it is. The crawler follows internal links the way Googlebot's discovery does — same host, robots.txt respected, no JavaScript — so what it finds is close to what a first-pass crawler finds. Pages it cannot reach are the first finding; pages it reaches but is told not to index are the second; pages it indexes with a missing or duplicated title are the third.

The report is deliberately flat. Commercial audit tools produce a health score and a colour-coded priority list; both are opinions dressed as measurement, and they push people to fix what moves the score rather than what matters for their site. Here each issue is a count and a list of pages. A missing meta description on 40 blog posts and a noindex on your pricing page both appear; which one is urgent depends on the site, and you know the site.

Site-wide checks are what a single-page checker cannot do. Duplicate titles across pages usually mean a template that ignores the page name or a paginated series without page numbers. Redirect chains inside the site mean links point at old URLs. Orphan pages mean the sitemap and the navigation disagree. Pages in the crawl but not in the sitemap, and pages in the sitemap the crawl never reached, show the two directions of that disagreement. None of these are visible from one page.

The 200-page limit makes this a sample for larger sites. A sample is enough to find template-level problems — a missing H1 on every product page shows up in the first twenty — and not enough to find the one broken URL in a 5,000-page catalogue. Start from the home page for structure, from a section index for that section's template, and read counts as 'at least'.

Frequently asked questions

Why is there no score?
Because any score is a weighting somebody chose. Is a missing alt attribute worth 1 point and a noindex on a money page worth 20? Different sites, different answers. The report gives you the counts and the pages; deciding what to fix first is the part that needs to know your business, and the tool does not.
What does the audit not check?
Anything that needs a browser: rendered content, Core Web Vitals, JavaScript-inserted links or tags, mobile layout. Anything that needs external data: backlinks, rankings, search volume. Anything off-site: hreflang reciprocity across domains, structured data validity beyond presence. For speed use PageSpeed Insights; for rendering use the SSR checker on a single page.
Why does it say a page is thin?
It counted fewer than 300 visible words on an indexable page. That is a convention from content SEO, not a Google threshold; a contact page with 80 words is fine. The flag exists so you can find pages that are accidentally empty — a template with no content, a category with no products — not to push every page past a word count.
The crawler reports pages I have deleted. Why?
Because something still links to them: internal navigation, the sitemap, or a redirect target. The broken-link and orphan lists say which. Deleted pages should return 404 or 410 and not be linked; if they are still in the sitemap, that is the sitemap's error.
My site has a sitemap but the audit says it does not.
The crawler looks for Sitemap: lines in robots.txt and, failing that, /sitemap.xml at the root. A sitemap at another path that robots.txt does not declare is invisible to it — and to search engines that have not been told about it in Search Console. Add the Sitemap: line and rerun.
Why are duplicates counted only on indexable pages?
Two pages with the same title where one is noindex or canonicalises to the other are not competing in search — that is the point of the directive. Duplicates that matter are between pages that both claim to be indexable. The audit compares only those.
Is 200 pages enough?
For most small business sites, yes — it covers the site. For larger sites it is a sample that reliably finds template-level issues and unreliably finds page-level ones. The audit says when it stopped at the limit. A full crawl of a large site needs desktop crawler software or a paid service.
Sitemap GeneratorEnter a URL and the crawler follows internal links (up to 200 pages), keeps only the pages a search engine could index — status 200, no noindex, canonical pointing at itself — and writes a plain XML sitemap you can download or copy. It also reads the sitemap the site already publishes and shows both gaps: indexable pages missing from it, and sitemap URLs the crawl never reached. No lastmod, priority or changefreq are invented. Nothing is stored beyond 24 hours.Broken Link CheckerEnter a URL and the crawler follows internal links (up to 200 pages) and reports every link that fails: 404s and other 4xx, 5xx, timeouts and DNS errors — internal links found during the crawl, external links checked with a HEAD request (up to 100). Each broken link comes with the pages that point to it and the anchor text, so you can fix the link or the page, not just know that something is wrong. Download as CSV. Nothing is stored beyond 24 hours.Internal Link CheckerEnter a URL and the crawler maps the internal link graph of up to 200 pages: how many pages link to each URL, how many links each page sends out, how many clicks each page sits from the start, and which pages nobody links to — orphans that only the sitemap knows about. The table sorts by incoming links so the pages your own site treats as unimportant are at the bottom. Download as CSV. Nothing is stored beyond 24 hours.On-Page SEO CheckerEnter a URL and, optionally, a target keyword. The checker fetches the page as Chrome and as Googlebot and reports what the raw HTML says: title and meta description with character and pixel length, robots meta and X-Robots-Tag, canonical, H1 count and heading outline with skipped levels, image alt coverage, internal and external links and nofollow share, Open Graph and Twitter card, hreflang validity, JSON-LD types, lang, viewport, charset, word count and keyword density and placement. No score — a list of what passed, what to look at and what is only information. Nothing stored.SSR CheckerCheck whether a URL delivers its content in the HTML or builds it with JavaScript. Fetches the page as a browser and as Googlebot without running JS, detects the framework and rendering mode (Next.js, Nuxt, React SPA, WordPress…), counts the words, links and headings a crawler sees, and flags bot blocking or dynamic rendering. Free, no signup.