Skip to content
Edit & Text

Extract text from a PDF — PDF to TXT

Take every word out of a PDF as plain text, in reading order: columns read one after the other, wrapped lines rejoined into paragraphs, words split by a hyphen at a line break put back together. It even works on PDFs that block copying — and nothing is uploaded.

Settings appear once you add a PDF.

Lines

Keep the lines for poems, addresses and code, where the line breaks matter.

Mark where each page starts
Advanced 2
Line endings

Done on your device No upload, no queue, no sign-up.

Bytes sent anywhere: 0 B

How to extract text from a PDF

Your PDF goes into this browser tab, is read in reading order on your device, and comes back as a plain text file. 0 bytes uploaded
  1. Add your PDFs

    Drop the PDF. It is opened in your browser and its text layer is read page by page.

  2. Check the settings

    Choose whether wrapped lines are rejoined into paragraphs or kept exactly as they sit on the page, and whether each page is marked.

  3. Save the result

    Save the .txt. It opens in any text editor, pastes into anything, and holds nothing but the words.

Why it doesn’t upload

People pull the text out of contracts, medical letters, research drafts and statements to quote or search them. None of that needs to be sent to a converter: the text is already inside the PDF, and your browser can read it where it is.

Check it in your browser’s Network tab

Worth knowing

  • This reads the text that is actually in the PDF. A scanned page is a photograph of words with no text in it, so it comes out empty — the tool says when a file looks scanned.
  • Tables come out as lines of text, not rows and columns. For tables, a spreadsheet converter is the better tool.
  • Reading order follows the page layout: two columns are read left then right. Unusual layouts — text wrapped around pictures, magazine spreads — can come out in a slightly different order than the eye follows.
  • Headers, footers and page numbers are part of the text, so they appear on every page unless you remove them afterwards.
  • A PDF that needs a password to open has to be opened with it first. Copy and print restrictions, on the other hand, do not stop this tool.

Under the hood

What happens to your file

pdf.js reads the characters on each page with their positions and sizes. Lines are rebuilt from shared baselines, spaces are put back from the gaps between letters, wide gaps split columns, and lines are grouped into paragraphs by their spacing — with words hyphenated at a line break joined again. The result is written as UTF-8 plain text.

Runs on pdf.js

Reference: pdf.js, Mozilla

Questions

How do I extract text from a PDF for free?

Drop the PDF here and press Extract text. You get a .txt file with every word in reading order. No account, no page limit, and the PDF never leaves your device.

Can I copy text from a PDF that won’t let me?

Usually, yes. “Copying not allowed” is a permission flag that viewers choose to respect; the text is still in the file. This tool reads it directly, so a copy-restricted PDF gives up its text like any other. It cannot open a PDF that needs a password to open.

Why is the text file empty?

The PDF is almost certainly a scan: each page is a picture of words, with no actual text inside. Text recognition (OCR) is needed to read pictures of text, and this tool does not do that.

Why do some words run together or break oddly?

Some PDFs place each letter separately with no spaces stored between words; the tool puts spaces back from the gaps it measures, which is right almost every time. Words hyphenated at a line break are joined back up when paragraphs are rejoined.

What’s the difference between “rejoin into paragraphs” and “keep as on the page”?

Rejoining turns the wrapped lines of a paragraph back into one line, which is what you want for pasting into a document or a translator. Keeping the lines preserves every line break, which matters for poems, addresses, song sheets and code.

More edit & text tools

All edit & text tools