Sitemap URL Extractor
Get the full list of pages a website wants search engines to find. Enter a domain and the tool reads robots.txt for Sitemap lines, tries the usual locations such as /sitemap.xml, follows sitemap index files to their child sitemaps and collects every URL with its last-modified date, change frequency, priority and image count. Filter the list by any part of the address and download it as plain text or CSV.
- Encrypted connection
- No sign-up
- Free to use
How to use Sitemap URL Extractor
- Enter a website or the address of a sitemap.
- Click “Extract URLs”.
- Filter the list, for example by /blog/.
- Download the URLs as TXT or CSV.
Sitemap URL Extractor features
Automatic discovery
robots.txt Sitemap lines and common names.
Sitemap indexes
Up to 20 child sitemaps are followed.
Gzip support
.xml.gz sitemaps are read.
Dates
Oldest and newest lastmod at a glance.
Filter and export
TXT or CSV of the filtered list.
Safe fetching
Only public http(s) addresses; private networks are blocked.
When to use Sitemap URL Extractor
- Building a crawl list for an audit.
- Comparing the pages in the sitemap with those in analytics.
- Checking which sections a competitor publishes.
- Preparing redirect maps for a migration.
Sitemap URL Extractor FAQ
How many URLs can it read?
Up to 50,000 URLs from up to 20 child sitemaps, and the first 20,000 are listed. The download contains every listed URL.
What if the site has no sitemap?
The tool reports which locations it tried. Many small sites have none; search engines then find pages through links.
Is the XML parsed safely?
Yes. Sitemaps are parsed without DTDs or external entities and downloads are size-capped, so malformed or hostile files cannot cause harm.
Why are some lastmod dates missing?
lastmod is optional. Google uses it only when it is consistently accurate, so a sitemap without dates is not an error.
Why extract sitemap URLs
A sitemap is the site owner’s own list of the pages that matter. Comparing it with what is actually indexed, linked or visited is one of the quickest ways to find problems: pages that are listed but return errors, sections that are missing, or old URLs that should have been removed.
The extractor saves you from opening large XML files by hand. It handles sitemap indexes, compressed files and the image extension, and gives you a clean list you can paste into a spreadsheet or a crawler.
To check whether the listed pages are healthy, use the Sitemap URL Validator, which tests a sample for status codes, redirects, noindex and canonical tags.
Results reflect the page at the moment of the check. Pages that serve different content to different visitors, countries or devices may show different results in a crawler; when in doubt, compare with the URL Inspection tool in Google Search Console, which shows what Google itself fetched.
Privacy: the address you enter is fetched by our server only to run this check. The downloaded HTML is passed to your browser for analysis and is not stored, and the results are not saved on our side. Requests are rate-limited to keep the service fair for everyone, so if you check many pages in a row you may need to wait a few minutes.
Limits: pages are read up to 2 MB of HTML, redirects are followed up to eight hops and each request has a short timeout. Very large pages, servers that block automated requests or pages behind a login cannot be analysed this way; for those, open the page in your browser, copy the page source and use the paste option where it is available.