↑ ↓ to move · Enter to open · Esc to close

Extract text from a PDF

Pull the words out of a PDF, then copy them or save them as a plain .txt file. Extraction happens in your browser, so the document stays on your device.

Runs on your device — files are never uploaded. How to check

Drop a PDF here or choose files One PDF at a time

About PDF to Text

Most PDFs created from a word processor, a web page or a spreadsheet contain real text: characters with positions, stored alongside the visual layout. This tool reads that text layer with pdf.js, Mozilla’s open-source PDF library, and hands it back to you as plain text. You get an on-screen preview with a Copy button, and you can save everything as a UTF-8 .txt file that opens in any text editor.

Before the text of each page, the output inserts a marker line such as --- Page 3 ---, so you can tell where a passage came from or delete the pages you don’t need. Because the file is UTF-8, accented letters, Greek, Cyrillic, Chinese, Japanese and other scripts come through intact, provided the PDF stores them as genuine characters. Document properties like the title and author aren’t included; only the page text is.

Scanned PDFs come out empty

A PDF produced by a scanner or a phone camera is usually just a stack of page photos. There are no characters inside to extract, so the result is blank. Reading text from images takes OCR (optical character recognition). This tool doesn’t do that itself, but OCR PDF on this site does, also inside your browser. A quick test before you start: open the PDF in any viewer and try to highlight a word. If the highlight snaps to the letters, the text layer is there.

How to use PDF to Text

  1. Add your PDFSelect a PDF from your device or drag it onto the tool. A file that asks for a password has to be unlocked before its text can be read.
  2. Check the extracted textRun the extraction and the text appears in the preview, page by page, separated by --- Page N --- lines. Skim it to make sure the reading order makes sense.
  3. Copy it or save a .txt filePress Copy to put all the text on your clipboard, or download it as a UTF-8 text file that you can search, edit or paste somewhere else.

When people use it

  • Quote a passage without retyping it

    Need a paragraph from a contract, a study or a user manual for an email or a report? Extract the text, find the right page by its marker and paste only the part you need.

  • Search or compare long documents

    A .txt file can be searched in any editor, compared line by line with an earlier draft using a diff tool, or loaded into a spreadsheet or script for counting and sorting.

  • Hand the words to another app

    Translation tools, note apps and chat assistants accept plain text more readily than a PDF. Pasting the extracted text lets you share only the passage you choose instead of uploading the whole original file.

Reading order and layout

Text comes out in the order it is stored inside the PDF, which usually, but not always, matches the order a person would read it. Single-column documents such as letters, reports and most ebooks extract cleanly. Layouts with two or three columns, sidebars, captions or floating text boxes may interleave: a line from the left column followed by a line from the right, or a caption landing in the middle of a paragraph.

Formatting isn’t part of plain text. Bold, italics, font sizes, colors and pictures are dropped, and tables turn into runs of words separated by spaces or line breaks rather than tidy cells. When the look of the page matters more than the characters, an image from PDF to PNG keeps the layout exactly, though as pixels.

Why there’s no resolution setting

The page-to-image converters on pd00 ask for 72, 150 or 300 DPI because they draw each page as a grid of pixels. Text extraction draws nothing at all. It reads the stored characters directly, so there is no resolution to choose and no quality loss: what you get is exactly the text the PDF contains. The resulting file is also light, since a .txt holds characters only, with no fonts or images, and is often far smaller than the PDF it came from.

When the output looks wrong

  • Nothing comes out. The PDF is a scan or an image export. Only OCR can help here. Run the file through OCR PDF first, which adds a hidden text layer, then extract the text with this tool.
  • Strange symbols instead of words. Some PDFs embed fonts without a map from the drawn shapes back to real characters. The page looks fine on screen, but the stored codes are meaningless outside that font, so the words can only be recovered with OCR.
  • Words split apart or run together. PDFs place text in fragments by position, so extra or missing spaces can appear, especially in justified paragraphs or letter-spaced headings.
  • The file is locked. Remove the password with Unlock PDF (you have to know it), then extract.

Only need part of the document?

The text tool works on the whole file. For a 400-page manual where you want a single chapter, cut it out first with Extract PDF Pages or Split PDF, then run that shorter PDF through here.

Your text stays private

The PDF is parsed in your browser tab and the text is produced on your device; neither the file nor the extracted words are uploaded anywhere. That matters for contracts, medical letters and other confidential paperwork. Follow the steps on the privacy page to confirm it.

Frequently asked questions

How do I copy all the text from a PDF?

Add the PDF here and press Copy once the text has appeared. Everything goes to your clipboard, page markers included.

Why is the extracted text empty?

The PDF almost certainly has no text layer, which is typical of scans and photographed pages. Try highlighting a word in your PDF viewer: if nothing gets selected, there is no text to extract. Run it through the OCR PDF tool first to add a searchable text layer, then come back here.

Why are the two columns mixed together?

The tool follows the order in which text is stored in the file. In multi-column layouts that order can jump between columns, so lines from different columns may alternate.

Are tables, images and formatting kept?

No. A .txt file holds characters only, so styling and pictures are left out and table cells become words separated by spaces or line breaks.

What are the “--- Page N ---” lines for?

They mark where each page starts, which helps you trace a quote back to its page. If you don’t want them, delete them with your editor’s find-and-replace.

Does it handle non-English text?

Yes, as long as the PDF stores proper Unicode characters. The output is UTF-8, which covers every major writing system, from accented Latin letters to Chinese and Japanese.

Is my document uploaded when I extract text?

No. Extraction runs entirely in your browser, and the text never leaves your device unless you paste it somewhere yourself. See how to verify this.

Last updated October 8, 2026