How to Convert PDF to Markdown: 3 Methods
September 15, 2026 · 7 min read
How to Convert PDF to Markdown: 3 Reliable Methods
Converting PDF to markdown extracts text, tables, and structure from fixed-layout PDF files into editable, version-controlled markdown format. You can do it with our free online tool, the Marker CLI library, or Poppler's pdftohtml piped through Pandoc. Each method handles text extraction differently, and results vary depending on your PDF's complexity. This guide walks through each approach.
Why Convert PDF to Markdown?
PDFs lock content into a fixed layout that is hard to edit, search, or version-control. Converting to markdown unlocks your content for editing in any text editor, tracking changes in Git, publishing on static site generators (Hugo, Jekyll), migrating to documentation platforms (Notion, Confluence), and feeding into AI/LLM workflows that process markdown better than PDF.
The challenge is that PDFs store visual layout, not document structure. A heading in a PDF is just larger text at specific coordinates. The converter must figure out what is a heading, what is a paragraph, and where tables begin and end. This is why conversion quality varies by tool and PDF type.
Method 1: Convert PDF to Markdown Online (Free)
Our online converter handles the most common PDF-to-markdown scenarios with no installation needed.
Step 1: Open the PDF converter.
Step 2: Upload your PDF file or paste content.
Step 3: The tool extracts text and converts it to markdown with headings, paragraphs, lists, and basic table structure preserved.
Step 4: Copy the markdown output or continue editing in our editor.
Try free today
3 PDF conversions/day, no sign-up required. Upgrade anytime for unlimited.
The tool works best with text-based PDFs (not scanned images). For scanned documents, you need OCR preprocessing (see the troubleshooting section below).
Method 2: Convert with Marker (Best Quality)
Marker is an open-source Python library specifically designed for high-quality PDF-to-markdown conversion. It uses machine learning models to understand document layout, making it more accurate than rule-based tools.
Step 1: Install Marker:
pip install marker-pdf
Step 2: Convert a single file:
marker_single input.pdf --output_dir output/ --output_format markdown
Step 3: Check the output/ directory for your markdown file.
Marker handles complex layouts including multi-column PDFs, academic papers with equations, tables with merged cells, and documents with images. In our testing, Marker produced the best heading hierarchy and table formatting compared to other tools.
Batch conversion:
marker input_directory/ --output_dir output_directory/ --output_format markdown
Marker runs fastest on a GPU but also works on CPU (slower). Its own benchmarks report several pages per second on GPU hardware; expect CPU runs to take noticeably longer.
Method 3: Convert with pdftohtml and Pandoc
Pandoc has no PDF reader (PDF is an output format only in Pandoc), so it cannot open a PDF directly. The workaround is a two-step pipeline: Poppler's pdftohtml turns the PDF into HTML, then Pandoc converts that HTML into markdown:
pdftohtml -s -i input.pdf temp.html
pandoc temp.html -o output.md
The -s flag writes a single HTML file and -i skips images. Install Poppler with brew install poppler (macOS) or sudo apt install poppler-utils (Ubuntu).
This route extracts text linearly and does its best to identify paragraphs and headings. It struggles with complex layouts, multi-column text, and tables. Use it for simple, single-column PDFs where text order matches reading order, and expect less structure than Marker produces.
Handling Scanned PDFs (OCR)
Scanned PDFs contain images of text, not actual text data. These require Optical Character Recognition (OCR) before markdown conversion.
Using OCRmyPDF (Tesseract under the hood):
# Install Tesseract and OCRmyPDF
# macOS: brew install tesseract ocrmypdf
# Ubuntu: sudo apt install tesseract-ocr ocrmypdf
# Add a text layer to the scanned PDF
ocrmypdf -l eng scanned.pdf searchable.pdf
# Then convert searchable.pdf with the online tool, Marker, or pdftohtml
Tesseract itself reads images, not PDFs, which is why OCRmyPDF is the practical wrapper: it rasterizes each page, runs Tesseract, and writes a PDF with a searchable text layer.
Using Marker (has built-in OCR):
Marker includes OCR capabilities. It detects pages without a usable text layer and applies OCR before conversion, and --force_ocr runs OCR on every page. This makes it the simplest option for mixed PDFs (some pages scanned, some text-based).
What to Expect from PDF-to-Markdown Conversion
Not all PDF elements convert perfectly. Here is what you can expect:
| PDF Element | Conversion Quality | Notes |
|---|---|---|
| Plain text paragraphs | Excellent | All tools handle this well |
| Headings | Good (Marker) / Fair (pdftohtml + Pandoc) | Marker uses ML to detect heading levels |
| Simple tables | Good | Pipe-and-dash markdown tables |
| Complex tables (merged cells) | Fair | May need manual cleanup |
| Lists (bulleted, numbered) | Good | Most tools detect list patterns |
| Images | Extracted (Marker) / Lost (pdftohtml + Pandoc) | Marker saves images as files |
| Math equations | Good (Marker) / Poor (pdftohtml + Pandoc) | Marker converts to LaTeX notation |
| Headers and footers | Often included | May need manual removal |
| Multi-column layouts | Good (Marker) / Poor (pdftohtml + Pandoc) | Marker handles columns correctly |
Plan for 5 to 15 minutes of manual cleanup on a typical 10-page document. Complex PDFs with unusual layouts may need more editing.
Common Issues and Fixes
Problem: Headers and footers appear in the body text.
PDF converters cannot always distinguish repeating headers/footers from body content. Marker drops page headers and footers by default (pass --keep_pageheader_in_output or --keep_pagefooter_in_output to keep them); with other tools, remove them manually.
Problem: Table columns are misaligned.
Markdown tables require exact pipe alignment. Use our markdown table generator to clean up extracted tables, or reformat them manually in the editor.
Problem: Text extraction produces garbage characters.
This usually means the PDF uses custom fonts with non-standard encoding. Try a different converter, or extract text with pdftotext first and then format as markdown manually.
Problem: Scanned PDF produces no text.
The PDF contains images, not text. Use OCR (Tesseract or Marker's built-in OCR) to extract text before conversion.
Frequently Asked Questions
Summary
Converting PDF to markdown works best with the right tool for your document type. Use our online converter for quick single-file conversions. Install Marker for batch processing and complex layouts with the best accuracy. Fall back to pdftohtml and Pandoc for simple documents. Expect some manual cleanup, especially for tables and multi-column layouts. After conversion, polish your markdown in the editor or see the markdown cheat sheet for formatting reference.
Written by the Markdown Editor Online team. Last updated September 2026.