About PDF to Markdown
A PDF records where each run of text sits on the page and how large it is, but it rarely says which line is a heading or which lines belong to a list. This converter reads the text layer with pdf.js and makes those calls itself, then writes the result in Markdown:
- Headings are guessed from text set in a larger font than the body.
- Lists come from lines that start with a bullet symbol or a number, which turn into bulleted or numbered items.
- Paragraphs are rejoined: lines the PDF broke at the right margin flow back together, so you don’t get a hard line break every dozen words.
- Page breaks are marked with a horizontal rule,
---.
The Markdown appears in a preview on the page. A Copy button puts all of it on your clipboard, or you can save it as a .md file. Being plain text, the output opens in any editor and drops straight into wikis, note apps and documentation sites that use Markdown.
Treat the structure as a first draft
Every heading and list in the output is an inference from visual clues. A heading set in bold at body size may come out as an ordinary line, while a large pull quote can be promoted to a heading. Skim the preview before you rely on it, especially for long or heavily designed documents.
How to use PDF to Markdown
- Load a PDF that contains textPick a PDF or drop it onto the tool. It needs a text layer: if you can’t highlight words in your PDF viewer, the file is a scan and should go through OCR PDF first.
- Read through the previewThe rebuilt Markdown shows up on the page. Check that headings and lists landed where you expect, particularly in documents with columns or tables.
- Copy it or keep the .mdPress Copy to paste the Markdown into an editor, wiki or chat, or download it as a .md file to keep alongside your other notes.
When people use it
Moving an old manual into a wiki
A product guide that survives only as a PDF can become Markdown pages with its section headings and step lists intact, ready to split into topics and edit in your documentation system.
Editing a report whose source file is gone
When the original word-processor file is lost, Markdown gives you editable text with its outline in place. Make your changes, then produce a fresh PDF with Markdown to PDF.
Feeding a long document to an AI assistant
Chat assistants read Markdown easily, and the heading markers tell them which section a paragraph sits in. Paste in the sections you need rather than attaching the full PDF.
Tables, images and multi-column pages
Tables aren’t rebuilt as Markdown tables. Inside a PDF a table is just words placed at coordinates, with no record of rows or cells, so cell contents come out as separate lines of text. Recreate important tables by hand, or keep a faithful picture of the page with PDF to PNG.
Images are left out of the Markdown. To keep photos, charts and figures, save them with Extract Images from PDF and link them into your Markdown yourself.
Multi-column layouts such as newsletters and academic papers may interleave, with lines from neighbouring columns alternating. Single-column documents exported from a word processor or a web page give the cleanest result.
Scans need a text layer first
A scanned PDF contains page images and no characters, so this tool has nothing to read and the Markdown comes out empty. Run the file through OCR PDF, which recognises the words and adds an invisible text layer, then bring the new file back here. OCR has its own error rate, so proofread names and figures in what you get.
Markdown, plain text or images: which output to pick
- Markdown (this tool) when you want the document’s outline: headings, lists and clean paragraphs you can keep editing.
- Plain text from PDF to Text when you want the words exactly as stored, with page markers and no guesses about structure.
- Images from PDF to PNG when the precise look of each page matters more than editable text.
What carries over, and where it happens
Characters come across exactly as the PDF stores them, so nothing is lost to resolution or compression. What Markdown can’t describe is left behind: fonts, colors, spacing and the exact position of things on the page. The resulting .md file is small, since it holds text and a few formatting symbols only. If the file is locked, clear its password with Unlock PDF (you’ll need to know it) and convert the unlocked copy.
The conversion happens inside this browser tab, and neither the PDF nor the Markdown is uploaded; you can confirm that yourself in a few steps.