PDF to Markdown on GitHub: Choose by PDF Type
Choose a native-text converter for selectable prose, a layout-aware path for columns, and a configured OCR path for scans when choosing a PDF-to-Markdown project on GitHub. Classify pages into five cases: single-column prose, multiple columns, tables, image-only scans, and equations. Mixed documents combine these cases. Our synthetic samples show why the distinctions matter: a converter can retain table rows while losing usable headers, or produce readable Markdown while omitting scanned content.
Classify the PDF before choosing a GitHub project
Choose the extraction path from the page's structure before comparing repository feature lists.
| Page feature | How to identify it | Conversion path |
|---|---|---|
| Single-column prose | Copy a paragraph into an editor, words remain readable and consecutive | Native extraction with MarkItDown or PyMuPDF4LLM |
| Multiple columns | Copy across the column boundary, compare the pasted sequence with the visual reading order | Layout-aware extraction, PyMuPDF4LLM is the stronger candidate on our column sample |
| Tables | Compare a data row and its headers with the pasted text, mark merged or rotated headers separately | Native table reconstruction, followed by a header and column mapping check |
| Image-only scans or mixed documents with scanned pages | A visible paragraph yields no selectable words in the viewer | Open-source OCR configuration with Marker (not tested here), or hosted OCR with Markovo (tested here) |
| Equations | Compare the pasted expression with the displayed grouping, superscripts, and fractions | Text extraction for linear expressions, a math-aware path for spatial notation |
Apply the table to pages, not just filenames. A report can contain selectable prose, a native invoice table, and an image-only appendix. In our mixed sample, both Markdown converters produced readable prose and a table, but the scanned appendix was missing. The successful prose and table pages did not establish that the appendix survived.
A GitHub result can also be a wrapper or a fork rather than the package you intend to install. Follow the repository's dependency and package names to the implementation used by its conversion command. The comparisons below concern the installed libraries and explicit calls shown here, rather than every project that shares a similar name.
Run native conversion for prose and columns
Use either tested Markdown converter for our single-column sample, and prefer PyMuPDF4LLM for our column sample.
| Project or baseline | Input it serves here | Execution boundary |
|---|---|---|
| MarkItDown | Selectable prose and simple tabular documents | Core local conversion, plugins disabled 1 |
| PyMuPDF4LLM | Selectable text with layout reconstruction | Default conversion with the installed layout dependency 2 |
| pdfplumber raw-text baseline | Text extraction without Markdown table rendering | Its other extraction APIs were not evaluated |
| Marker | Document reconstruction with model-based processing | Static documentation review, not executed for this guide 8 |
Both conversion calls below were executed on the same sample. Run either call with your own input and save its result before trying the other.
from pathlib import Path
from markitdown import MarkItDown
import pymupdf4llm
pdf = "input.pdf"
text = MarkItDown(enable_plugins=False).convert_local(pdf).markdown
Path("output.md").write_text(text, encoding="utf-8")
# Alternatively:
text = pymupdf4llm.to_markdown(pdf)
Path("output.md").write_text(text, encoding="utf-8")
Both Markdown converters kept the single-column sentences intact and in order. In the double-column sample, MarkItDown alternated between unrelated paragraphs on opposite sides of the page. A sentence about invoice reconciliation breaks off, and a demand forecast interrupts it before the invoice sentence resumes.
Output from our double-column sample, unedited:
Invoice reconciliation starts with the export file, not
Demand forecasts fail upward more often than they
the ledger. Teams that reconcile weekly pull the
fail downward. Optimistic sales inputs push safety
The pdfplumber baseline joined the left and right text on each line instead. The interruption is visible without rendering Markdown.
Output from our double-column sample, unedited:
Invoice reconciliation starts with the export file, not Demand forecasts fail upward more often than they
the ledger. Teams that reconcile weekly pull the fail downward. Optimistic sales inputs push safety
payment provider's settlement report first, match it stock up for slow movers, and the excess sits in the
PyMuPDF4LLM kept the double-column sentences intact and in their expected sequence. That result supports the column choice here, but its grouped table still needs a header check.
Output from our grouped-price table sample, unedited:
# **Invoice with grouped price columns**
|**Item**|**Qty**|**Price**||
|---|---|---|---|
|||**Unit**|**Total**|
|Consulting hours|12|85.00|1020.00|
|Data export fee|3|40.00|120.00|
|Template license|1|190.00|190.00|
MarkItDown's installed converter first looks for form-like content on each page. If none qualifies, it extracts plain text across the document. If a page qualifies, it joins collected page outputs, including text from other pages. An exception triggers the plain-text fallback. Those branches explain why the converter does not always take the same extraction route. They do not establish the cause of every column failure. 3
An older MarkItDown issue reports missing tables. Our simple invoice sample instead produced a Markdown table with all row text and labels intact. We did not run the issue's original attachment, so this is evidence against a blanket claim that the tool cannot emit PDF tables, rather than a verified fix for that report. See the test table at the end. 4
Keep table headers and formulas out of false success
Accepting intact row text alone can admit a table whose headers or column meanings are unusable.
On the grouped and rotated header samples, both Markdown converters retained all four data rows. The labels did not form a usable header line. The raw-text baseline retained the same row text and scattered labels, without rendering a Markdown table.
Markovo retained the row text and leaf labels in both special-header samples, but returned HTML rather than a Markdown table. In the grouped output, the HTML shares a single text line. The labels are present, but their presence alone does not establish the parent groups or value-to-column mapping.
Output from our grouped-price table sample, unedited:
<table><tr><td>Item</td><td>Qty</td><td>Price</td><td></td></tr><tr><td></td><td></td><td>Unit</td><td>Total</td></tr><tr><td>Consulting hours</td><td>12</td><td>85.00</td><td>1020.00</td></tr><tr><td>Data export fee</td><td>3</td><td>40.00</td><td>120.00</td></tr><tr><td>Template license</td><td>1</td><td>190.00</td><td>190.00</td></tr><tr><td>Support retainer</td><td>6</td><td>55.00</td><td>330.00</td></tr><tr></tr></table>
| Output symptom | What you will see | Concrete acceptance check |
|---|---|---|
| Merged header split across output lines | Item and Qty remain above Unit and Total, rather than forming a flat label row | Associate each leaf label with its parent group and corresponding values |
| Reversed rotated labels | MarkItDown's output contains backward header words outside the rendered table | Compare each label's character order with the PDF, then map it to its data column |
| Data row promoted to header | The first invoice data row appears before the Markdown separator | Require a real label row before the separator and retain that data row as data |
| Formula text partly absent | PyMuPDF4LLM omits the energy expression while retaining other linear expressions | Compare each expected expression, then inspect its mathematical grouping |
In the rotated table, a data row occupies the header position even though its cell text survived. In the grouped table, the labels remain visible on separate lines. Neither output can be accepted by looking only for the expected words.
A separate MarkItDown issue reports reversed rotated headers. Our synthetic sample shows a related shape, but does not reproduce that attachment. The report supports a direction-sensitive text check, not a measured failure rate. 5
When headers carry meaning, use explicit labels in the normalized output. For our grouped-price sample, a flattened schema can identify the unit-price and total-price columns separately. That is an editorial normalization derived from the visible header hierarchy, not something intact row text certifies automatically.
MarkItDown and the raw-text baseline retained all four linear expressions. PyMuPDF4LLM omitted the energy expression.
Markovo recovered half of the expected expressions after mathematical markup was normalized. The energy and linear-function expressions survived. The squared-terms expression acquired double superscripts, and the summation subscript was escaped.
Output from our linear-equation sample, unedited:
$$
\[a^{\wedge}2+b^{\wedge}2=c^{\wedge}2\]
$$
The API's quality estimate does not establish that every expression survived. These results also do not establish LaTeX reconstruction for fractions, matrices, or spatial notation. In the project's equation discussion, a maintainer describes formula regions being treated as pictures or ordinary text rather than understood mathematically. See the test table at the end. 7
Route scans through a configured OCR path
Use an explicitly configured OCR path for image-only pages. Our default local runs recovered no scanned text.
Both Markdown converters and the raw-text baseline returned no scanned words from the image-only sample or the mixed report's appendix. The Markdown converters still returned the mixed report's prose and table, so a nonempty result concealed the missing page.
Markovo recovered the scanned text in both samples and retained the mixed report's page anchors. Its prose and table text matched the local paths, but its table was HTML rather than Markdown. The API distinguishes pages classified for OCR billing before conversion from pages processed by the OCR engine during conversion. In the mixed report, the engine processed the native pages as well as the image-only page. Those counts describe different stages, not a disagreement about whether OCR ran. 10
CLI equivalent for the hosted OCR path, with MARKOVO_API_KEY set in the environment: pip install "https://markovo.net/downloads/markovo-0.1.1-py3-none-any.whl#sha256=aae71e1fff133fca3a9e6c17a1fd3b1080a0bd5a05ea075eef7267491bea079d", then markovo convert input.pdf --out out --mode fast --max-credits 30. 10
The recovered image-only text reads: “This page exists only as a picture.” The raw OCR output and page counts are recorded in How we tested.
MarkItDown ran with plugins disabled and no external OCR service. PyMuPDF4LLM's default call selected no usable OCR function in this environment. Its documentation lists OCR engines and language resources as dependencies. We did not measure either converter with an operational OCR configuration. 1 2 9
The MarkItDown OCR plugin was also not executed for this guide. A plugin issue reports scanned-page fallback and page-detection failures in a different configuration from our core-only run. Installing the package, enabling the plugin, and supplying its OCR service are distinct operations. The visible page text must appear in the result. 6
For a mixed report, compare a phrase from the scanned appendix with the generated Markdown separately from the prose and table checks. If the prose is readable but the image-only phrase is absent, route that page through OCR and keep its position in the document. Changing only Markdown formatting cannot recover words that never entered extraction.
Choose the native path for readable prose and simple tables, favor the layout-aware path for the tested column shape, and rebuild special headers when their relationships matter. Route image-only content through OCR before combining it with native output.
How we tested
We compared fixed conversion calls on synthetic PDFs with paired ground truth, so the results measure these fixtures rather than general document accuracy.
Corpus: lab/corpus/ in this run:
| File | Representative shape |
|---|---|
prose-single.pdf |
Single-column prose |
two-column.pdf |
Multiple prose columns |
table-simple.pdf |
Simple invoice grid |
table-merged-header.pdf |
Grouped price header with merged cells |
rotated-header.pdf |
Rotated header text with ordinary data rows |
scan-image.pdf |
Image-only page |
mixed.pdf |
Prose, table, and scanned pages |
equation.pdf |
Linear textual expressions |
Local tests used Python 3.11.14 on macOS. PyMuPDF and its layout package were 1.28.2. The calls were core MarkItDown, default PyMuPDF4LLM, and pdfplumber extract_text, without operational OCR.
| Tool | Version identifier | Test date |
|---|---|---|
| MarkItDown | 0.1.8 | 2026-10-06 |
| PyMuPDF4LLM | 1.28.2 | 2026-10-06 |
| pdfplumber | 0.11.10 | 2026-10-06 |
| Markovo public API | rate_card_version 2026-09-19.v3-media-unified (not a pip version) |
2026-10-06 |
Run from the repository root with the lab environment:
PYTHONDONTWRITEBYTECODE=1 tools/lab-venv/bin/python tools/lab_run.py --lab runs/2026-10-06-markovo-pdf-to-markdown-github/lab --tools markitdown,pymupdf4llm,pdfplumber
PYTHONDONTWRITEBYTECODE=1 tools/lab-venv/bin/python tools/lab_score.py --lab runs/2026-10-06-markovo-pdf-to-markdown-github/lab --out runs/2026-10-06-markovo-pdf-to-markdown-github/lab/scores.json --markdown runs/2026-10-06-markovo-pdf-to-markdown-github/lab/scores.md
Submit, poll to status=succeeded, then download text/markdown:
curl -X POST https://markovo.net/v1/convert -H "Authorization: Bearer $MARKOVO_API_KEY" -F "file=@scan.pdf" -F "capability_id=pdf-to-markdown" -F "mode=fast" -F "max_credit_units=5000"
# Set job_id from the POST response; repeat GET until status=succeeded.
curl "https://markovo.net/v1/jobs/${job_id}" -H "Authorization: Bearer $MARKOVO_API_KEY"
# After succeeded, download text/markdown.
curl "https://markovo.net/v1/jobs/${job_id}/download" -H "Authorization: Bearer $MARKOVO_API_KEY" -o scan.md
Markovo is the service we build. It ran through the same corpus and scorer as the other paths; its row is reported in full, including where it scored lower.
tools/lab_score.py normalizes case, whitespace, and selected Markdown characters. Prose and scan recall measure sentence-string matches. Reading order compares adjacent paragraph-opening anchors. Rows measure same-line cell-string matches, headers measure same-line leaf-label matches, and the pipe flag detects table syntax. Mixed-page coverage checks page anchors. Equation recall uses the same markup-normalised substring matching for every tool: remove $ and \[ \] delimiters, \mathsf{} font wrappers, normalize \{ to { and \_ to _, then remove whitespace. equation_recall_exact retains the old case/whitespace-only matching for comparison, which inherently scores LaTeX output lower. Blank metrics are not applicable. Local times are single-run seconds. API elapsed_s=0.0 denotes cache reuse in lab/markovo/run.json, not zero conversion latency.
Raw outputs use lab/<tool>/<sample>.md, for example lab/markitdown/rotated-header.md and lab/pymupdf4llm/table-merged-header.md. API responses: lab/markovo/<sample>.json. Batch metadata: lab/markovo/run.json. lab/runs.json records calls and elapsed times. lab/scores.json supplies the following scores. Example outputs: lab/sample-markitdown.md, lab/sample-pymupdf4llm.md.
| tool | sample | prose_recall | reading_order | table_rows_intact | table_header_intact | table_pipe_detected | pages_covered | scan_text_recall | equation_recall | equation_recall_exact | elapsed_s |
|---|---|---|---|---|---|---|---|---|---|---|---|
| markitdown | equation | 1.0 | 1.0 | 0.1081 | |||||||
| markitdown | mixed | 1.0 | 1.0 | 1.0 | yes | 0.6667 | 0.0 | 0.0752 | |||
| markitdown | prose-single | 1.0 | 1.0 | 0.062 | |||||||
| markitdown | rotated-header | 1.0 | 0.0 | yes | 0.033 | ||||||
| markitdown | scan-image | 0.0 | 0.0227 | ||||||||
| markitdown | table-merged-header | 1.0 | 0.0 | yes | 0.0515 | ||||||
| markitdown | table-simple | 1.0 | 1.0 | yes | 0.067 | ||||||
| markitdown | two-column | 0.1111 | 0.8 | 0.1239 | |||||||
| pymupdf4llm | equation | 0.75 | 0.75 | 0.4329 | |||||||
| pymupdf4llm | mixed | 1.0 | 1.0 | 1.0 | yes | 0.6667 | 0.0 | 0.8069 | |||
| pymupdf4llm | prose-single | 1.0 | 1.0 | 0.3997 | |||||||
| pymupdf4llm | rotated-header | 1.0 | 0.0 | yes | 0.429 | ||||||
| pymupdf4llm | scan-image | 0.0 | 0.5137 | ||||||||
| pymupdf4llm | table-merged-header | 1.0 | 0.0 | yes | 0.4445 | ||||||
| pymupdf4llm | table-simple | 1.0 | 1.0 | yes | 0.4838 | ||||||
| pymupdf4llm | two-column | 1.0 | 1.0 | 0.4409 | |||||||
| pdfplumber | equation | 1.0 | 1.0 | 0.0111 | |||||||
| pdfplumber | mixed | 1.0 | 1.0 | 1.0 | no | 0.6667 | 0.0 | 0.0447 | |||
| pdfplumber | prose-single | 1.0 | 1.0 | 0.0274 | |||||||
| pdfplumber | rotated-header | 1.0 | 0.0 | no | 0.0145 | ||||||
| pdfplumber | scan-image | 0.0 | 0.002 | ||||||||
| pdfplumber | table-merged-header | 1.0 | 0.0 | no | 0.015 | ||||||
| pdfplumber | table-simple | 1.0 | 1.0 | no | 0.0145 | ||||||
| pdfplumber | two-column | 0.1111 | 0.8 | 0.0543 | |||||||
| markovo | equation | 0.5 | 0.0 | 0.0 | |||||||
| markovo | mixed | 1.0 | 1.0 | 1.0 | no | 1.0 | 1.0 | 0.0 | |||
| markovo | prose-single | 1.0 | 1.0 | 0.0 | |||||||
| markovo | rotated-header | 1.0 | 1.0 | no | 0.0 | ||||||
| markovo | scan-image | 1.0 | 0.0 | ||||||||
| markovo | table-merged-header | 1.0 | 1.0 | no | 0.0 | ||||||
| markovo | table-simple | 1.0 | 1.0 | no | 0.0 | ||||||
| markovo | two-column | 1.0 | 1.0 | 0.0 |
Details moved from the article body
The column scores were MarkItDown and pdfplumber prose_recall=0.1111, reading_order=0.8, versus PyMuPDF4LLM 1.0 on both. The single-column Markdown runs were 1.0 on both. On the mixed sample, the local paths had pages_covered=0.6667 and scan_text_recall=0.0. The table above preserves all tool/sample values, including Markovo's lower equation results and absent pipe syntax.
MarkItDown issue #1419 concerns version 0.1.3; issue #2560 concerns 0.1.8; plugin issue #2343 was not executed. The installed source branches are PdfConverter.convert and _extract_form_content_from_words. The installed PyMuPDF4LLM OCR selector returned None.
Markovo equation_recall=0.5 means 2/4 recovered after markup normalization: E=mc^2 and f(x)=3x+7. The squared-terms expression contains ^{\wedge} and the sum subscript is escaped. Its equation_recall_exact=0.0 is a comparison under the older criterion, which inherently scores LaTeX output lower. lab/markovo/equation.json records quality_score=0.9, not proof of expression recovery.
billing_metrics.pagesOcr and ocrPageIndexes identify pages classified as image-only in preflight for OCR billing. job.pages_ocr is the container-reported number actually processed by the OCR engine. Special-header and equation responses show 0/[]/1: no OCR-billed page, one page processed by OCR. In mixed.json, pagesOcr=1, ocrPageIndexes=[2], pages_ocr=3: one OCR-billed page, all three processed by OCR. The image-only sample has index [0]. The public documentation states, "The server may choose the best conversion method without changing the confirmed price". 10
Output from our image-only scan sample, unedited:
This page exists only as a picture. A scanner produced it, the PDF embeds it as a bitmap, and no amount of text extraction will recover these words without OCR.