URL Path Extractor
Strip the scheme and domain from a list of URLs and keep just the paths. Choose whether to keep the query string, decode percent-encoded characters and remove duplicates. A table shows each path’s depth, the file name and its extension, which is handy for sitemaps, redirect planning and log analysis.
- Runs in your browser
- No sign-up
- Free to use
| Path | Depth | File | Extension |
|---|
How to use URL Path Extractor
- Paste full URLs, one per line.
- Choose whether to keep the query, decode and remove duplicates.
- Read the depth, file name and extension of each path.
- Copy or download the paths.
URL Path Extractor features
Clean paths
Scheme, host, port and fragment removed.
Optional query
Keep or drop everything after ?.
Decoding
Shows %20 and similar as readable characters.
Path analysis
Depth, file name and extension per URL.
De-duplication
One line per distinct path.
Private
Runs in your browser.
When to use URL Path Extractor
- Preparing the source column of a redirect map.
- Grouping crawl results by section of a site.
- Finding file types in a list of URLs.
- Converting absolute links to site-relative ones.
URL Path Extractor FAQ
What is a URL path?
The part after the host and before the query: /blog/2026/hello-world/ in https://example.com/blog/2026/hello-world/?x=1.
What does depth mean?
The number of path segments: /blog/2026/post has depth 3; the home page / has depth 0.
Why decode?
Paths often contain encoded characters such as %20 for a space. Decoding makes them readable; keep it off if you need the exact encoded form, for example for server rules.
Is the fragment kept?
No. The part after # is not sent to servers and is never part of the path.
What if a line is not a URL?
It is skipped and counted in the status message.
Is anything sent?
No.
Paths are the structure of a site
A website’s paths reflect how its content is organised: sections, categories, dates and file names. Looking at paths without the domain makes that structure visible and is often the first step in planning a migration, a redirect map or a sitemap.
Depth is a simple but useful measure. Pages buried many levels deep tend to receive fewer internal links and less attention from search engines. Extensions reveal what kinds of resources a list contains, HTML pages, PDFs, images, scripts, which helps when filtering crawl exports.
Whether to keep the query string depends on the job. For redirects and content inventories, the path alone usually identifies the page. For analytics and debugging, the parameters can matter. The option lets you switch between the two views of the same list.
Decoding affects matching. Server rules usually work on the encoded form, while people read the decoded one. Extract both when building redirect rules: the decoded list for review and the encoded list for the configuration.