PDF Link Extractor
Collect all links from a PDF in one list. The extractor reads every clickable link – web pages, e-mail addresses, files – and, if you like, internal links and web addresses that are only written in the text. Duplicate addresses are merged with the pages they appear on, and the list downloads as CSV for spreadsheets or as plain text, one address per line.
- Files stay on your device
- No sign-up
- Free to use
How to use PDF Link Extractor
- Drop a PDF onto the page.
- Choose whether to include text addresses and internal links.
- Extract the links.
- Download CSV or a plain list.
PDF Link Extractor features
Clickable links
Web, e-mail and file links from link annotations.
Text addresses
URLs written in the text but not clickable.
Internal links
Table of contents and cross references, optional.
De-duplicated
One row per address with all its pages.
Two formats
CSV with type and pages, or plain text.
Private
Read in your browser.
When to use PDF Link Extractor
- Auditing links in reports and e-books.
- Collecting references from research papers.
- Preparing a link check before publishing.
- Moving sources into a reference manager.
PDF Link Extractor FAQ
What is the difference between clickable and text links?
Clickable links are annotations with a target. Text links are addresses written in the page text; they may or may not be clickable depending on the reader.
What are internal links?
Links that jump to another place in the same document, such as table of contents entries.
Does it check whether links work?
No; use the PDF Link Checker, which tests each address.
Why are some addresses incomplete?
Addresses split across two lines in the text cannot always be joined. Clickable links are always complete.
Does it work with scans?
Only clickable links. Text addresses need a text layer; run PDF OCR first.
Is my PDF uploaded?
No. Everything happens in your browser.
Links inside PDF documents
Links in PDFs come in two forms. Real links are link annotations: a rectangle on the page with an action that opens a web address, starts an e-mail, opens a file or jumps to a destination in the document. Many PDFs also contain addresses printed as text, which some readers detect and make clickable on their own.
The extractor uses PDF.js to read the link annotations of every page and classifies them as web, e-mail, file or internal. It then scans the page text for web addresses and adds those that are not already covered by a clickable link, so nothing is counted twice.
The result is grouped by address: a link that appears on twenty pages is one row with twenty page numbers. The CSV includes the type, which makes it easy to filter e-mail addresses or internal links in a spreadsheet.
To find broken links, pass the list to the PDF Link Checker, which tests each web address from our server.
The plain text download is handy as input for other tools: paste it into the Broken Link Checker or a spreadsheet, or use it to update links in the source document before exporting the PDF again.