About PDF to Text
Most PDFs created from a word processor, a web page or a spreadsheet contain real text: characters with positions, stored alongside the visual layout. This tool reads that text layer with pdf.js, Mozilla’s open-source PDF library, and hands it back to you as plain text. You get an on-screen preview with a Copy button, and you can save everything as a UTF-8 .txt file that opens in any text editor.
Before the text of each page, the output inserts a marker line such as --- Page 3 ---, so you can tell where a passage came from or delete the pages you don’t need. Because the file is UTF-8, accented letters, Greek, Cyrillic, Chinese, Japanese and other scripts come through intact, provided the PDF stores them as genuine characters. Document properties like the title and author aren’t included; only the page text is.
Scanned PDFs come out empty
A PDF produced by a scanner or a phone camera is usually just a stack of page photos. There are no characters inside to extract, so the result is blank. Reading text from images takes OCR (optical character recognition). This tool doesn’t do that itself, but OCR PDF on this site does, also inside your browser. A quick test before you start: open the PDF in any viewer and try to highlight a word. If the highlight snaps to the letters, the text layer is there.
How to use PDF to Text
- Add your PDFSelect a PDF from your device or drag it onto the tool. A file that asks for a password has to be unlocked before its text can be read.
- Check the extracted textRun the extraction and the text appears in the preview, page by page, separated by --- Page N --- lines. Skim it to make sure the reading order makes sense.
- Copy it or save a .txt filePress Copy to put all the text on your clipboard, or download it as a UTF-8 text file that you can search, edit or paste somewhere else.
When people use it
Quote a passage without retyping it
Need a paragraph from a contract, a study or a user manual for an email or a report? Extract the text, find the right page by its marker and paste only the part you need.
Search or compare long documents
A .txt file can be searched in any editor, compared line by line with an earlier draft using a diff tool, or loaded into a spreadsheet or script for counting and sorting.
Hand the words to another app
Translation tools, note apps and chat assistants accept plain text more readily than a PDF. Pasting the extracted text lets you share only the passage you choose instead of uploading the whole original file.
Reading order and layout
Text comes out in the order it is stored inside the PDF, which usually, but not always, matches the order a person would read it. Single-column documents such as letters, reports and most ebooks extract cleanly. Layouts with two or three columns, sidebars, captions or floating text boxes may interleave: a line from the left column followed by a line from the right, or a caption landing in the middle of a paragraph.
Formatting isn’t part of plain text. Bold, italics, font sizes, colors and pictures are dropped, and tables turn into runs of words separated by spaces or line breaks rather than tidy cells. When the look of the page matters more than the characters, an image from PDF to PNG keeps the layout exactly, though as pixels.
Why there’s no resolution setting
The page-to-image converters on pd00 ask for 72, 150 or 300 DPI because they draw each page as a grid of pixels. Text extraction draws nothing at all. It reads the stored characters directly, so there is no resolution to choose and no quality loss: what you get is exactly the text the PDF contains. The resulting file is also light, since a .txt holds characters only, with no fonts or images, and is often far smaller than the PDF it came from.
When the output looks wrong
- Nothing comes out. The PDF is a scan or an image export. Only OCR can help here. Run the file through OCR PDF first, which adds a hidden text layer, then extract the text with this tool.
- Strange symbols instead of words. Some PDFs embed fonts without a map from the drawn shapes back to real characters. The page looks fine on screen, but the stored codes are meaningless outside that font, so the words can only be recovered with OCR.
- Words split apart or run together. PDFs place text in fragments by position, so extra or missing spaces can appear, especially in justified paragraphs or letter-spaced headings.
- The file is locked. Remove the password with Unlock PDF (you have to know it), then extract.
Only need part of the document?
The text tool works on the whole file. For a 400-page manual where you want a single chapter, cut it out first with Extract PDF Pages or Split PDF, then run that shorter PDF through here.
Your text stays private
The PDF is parsed in your browser tab and the text is produced on your device; neither the file nor the extracted words are uploaded anywhere. That matters for contracts, medical letters and other confidential paperwork. Follow the steps on the privacy page to confirm it.