Sitemap Generator

Enter a URL and the crawler follows internal links (up to 200 pages), keeps only the pages a search engine could index — status 200, no noindex, canonical pointing at itself — and writes a plain XML sitemap you can download or copy. It also reads the sitemap the site already publishes and shows both gaps: indexable pages missing from it, and sitemap URLs the crawl never reached. No lastmod, priority or changefreq are invented. Nothing is stored beyond 24 hours.

HEAD request per unique external URL, up to 100. Untick for a faster crawl.

The crawl runs on our server: one request per second, robots.txt obeyed, no JavaScript. Only the start URL and this summary are kept, for 24 hours.

How it works

1. Our crawler (InsighthackerzBot, one request per second, no JavaScript) fetches robots.txt and any sitemaps it declares, then starts at the URL you give and follows internal links breadth-first — same host only, www and the bare domain counted as one site
2. Limits: 50, 100 or 200 pages per run, 2 MB per page, 15 s per request, 5 redirects per URL; URLs your robots.txt disallows for our user-agent are skipped and counted, pages marked nofollow are fetched but their links are not followed
3. Every HTML page goes through the same on-page engine as the on-page SEO checker: title, description, H1, canonical, robots directives, word count, images without alt — the warnings and fails are kept per page, the HTML is not
4. Internal links found broken during the crawl (4xx, 5xx, timeouts) are listed with the pages that link to them; external links are checked with HEAD only (one GET when the server refuses HEAD), at most 100 per run, nofollow links excluded
5. The result is one job on our own server: it holds only the start URL and the summary, expires after 24 hours and is never used for anything else
6. Sitemap rule: a URL is listed when it returned 200, carries no noindex (meta or X-Robots-Tag) and its canonical is absent or points at itself; redirects, errors and pages that canonicalise elsewhere are left out. Output is <loc> only, sorted alphabetically

About XML sitemaps and what the generator includes

An XML sitemap is a list of the URLs you want search engines to know about. It does not make a page rank and it does not force indexing; it is a hint that helps crawlers find pages they might otherwise reach late or not at all — deep pages, new pages, pages with few internal links. Google reads the <loc> element and, when it trusts it, <lastmod>; it has said publicly that it ignores <priority> and <changefreq>. This generator writes <loc> only, because the crawl cannot know when a page last changed and guessing would make the field worthless.

Which pages belong in a sitemap is the real question, and it is where most generated sitemaps go wrong. A sitemap should list canonical, indexable URLs and nothing else: a URL that redirects, returns an error, is marked noindex or declares another URL as its canonical sends Google a contradictory signal — 'index this' from the sitemap and 'do not' from the page. The generator applies exactly those four filters and shows which pages were excluded and why, so the list you download is the list a search engine would accept.

The comparison with the site's existing sitemap is often more useful than the new file. Indexable pages that are not in the published sitemap are usually pages the CMS forgot — a category added later, a landing page outside the blog; sitemap URLs the crawl never reached are either unlinked (orphans that only the sitemap knows about) or gone. A crawl of 200 pages cannot see a large site whole, so treat both lists as samples, not audits.

For sites that change often, a generated sitemap goes stale the day after you upload it. The right long-term answer is a sitemap the CMS or framework produces itself; this tool is for the one-off cases — a static site, a migration check, a client site you do not administer — and for seeing what a crawler actually finds when it follows your links.

Frequently asked questions

Why is a page missing from the generated sitemap?
Four reasons, all deliberate: the page returned something other than 200, it carries noindex, its canonical points at another URL, or the crawl never reached it within the page limit. The site audit tab lists every crawled page with its status and directives; check there before assuming the crawler missed it.
Why is there no lastmod?
Because a crawl does not know when a page changed — the HTTP Last-Modified header is missing or wrong on most sites, and inventing a date would teach Google to distrust your sitemap. Google uses lastmod only when it has been consistently accurate. If your CMS knows the dates, its own sitemap should carry them; this file does not pretend to.
Does Google use priority and changefreq?
No. Google has said since 2015 that it ignores both, and in 2023 it removed them from its documentation. Bing says much the same. They are left out here so the file stays honest and small.
What does 'not in the existing sitemap' mean?
The crawler fetched the sitemap(s) your robots.txt declares (or /sitemap.xml) and compared them with the pages it found. Pages it could index that are absent from your sitemap are listed there — typically pages added after the sitemap was generated, or sections the CMS does not include. The reverse list, sitemap URLs the crawl did not reach, points at orphaned or removed pages.
My site has more than 200 pages. What now?
The crawl stops at the limit and says so. The generated file is then a sample, not a complete sitemap — useful for checking structure, not for uploading. Large sites need a sitemap produced by the CMS, split into files of at most 50,000 URLs and 50 MB each with a sitemap index on top.
Where do I put the file?
At the root of the site is the convention (/sitemap.xml), and the location must be on the same host as the URLs it lists. Then declare it in robots.txt with a Sitemap: line and submit it once in Search Console; after that Google refetches it on its own schedule.
Is the crawl safe for my site?
It makes one request per second, identifies itself as InsighthackerzBot, obeys your robots.txt, reads at most 200 pages of 2 MB each and does not execute JavaScript. A site that cannot take that load has a bigger problem than its sitemap. If you want to keep it out, disallow InsighthackerzBot in robots.txt and the crawl stops at the first page.
SEO AuditEnter a URL and the crawler checks up to 200 pages the way a search engine's first pass would: status codes, noindex and canonical directives, missing or duplicate titles and descriptions, missing H1s, thin pages, images without alt, redirect chains, broken links, orphan pages, robots.txt and sitemap coverage. Every page gets the same on-page checks as the single-page checker; the report groups them by issue and by page. No score, no fabricated priorities — findings with counts. Download as CSV. Nothing is stored beyond 24 hours.Broken Link CheckerEnter a URL and the crawler follows internal links (up to 200 pages) and reports every link that fails: 404s and other 4xx, 5xx, timeouts and DNS errors — internal links found during the crawl, external links checked with a HEAD request (up to 100). Each broken link comes with the pages that point to it and the anchor text, so you can fix the link or the page, not just know that something is wrong. Download as CSV. Nothing is stored beyond 24 hours.Internal Link CheckerEnter a URL and the crawler maps the internal link graph of up to 200 pages: how many pages link to each URL, how many links each page sends out, how many clicks each page sits from the start, and which pages nobody links to — orphans that only the sitemap knows about. The table sorts by incoming links so the pages your own site treats as unimportant are at the bottom. Download as CSV. Nothing is stored beyond 24 hours.robots.txt TesterTest a robots.txt against any URL path and crawler: fetch it from a site or paste it, pick Googlebot, Bingbot, GPTBot, ClaudeBot, PerplexityBot or a custom user-agent, and see the exact rule that allows or blocks the path — plus a per-crawler summary, syntax warnings and Sitemap lines. RFC 9309 matching, runs in your browser, nothing stored.