Converting PDFs to Markdown in Python: Text, Tables, and Scans

Use MarkItDown for readable prose and a structure-aware or OCR path such as Marker for content the text baseline loses when converting PDFs to Markdown in Python. Route pages by text usability and required relationships into four cases: readable prose, tables or multi-column layouts, scans or garbled text layers, and equations; the paths include local text extraction, documented open-source OCR, and a tested hosted OCR path such as Markovo.

Classify pages by text usability and required structure

Readable prose belongs on the text path, tables and columns need relationship checks, scans need OCR, and equations need recognition plus renderer support.

What you see How to confirm Which path
Readable prose Read the output beside the source and follow the paragraphs in order. Our single-column sample kept its sentences and paragraph sequence Convert a PDF to Markdown in Python with MarkItDown
Tables or multi-column layouts Follow each value back to its row label and column heading. In columns, read consecutive sentences: our interleaved output kept paragraph openings but spliced together unrelated text Use the tested PyMuPDF4LLM call for the column fixture, repair table headings, or Marker’s documented structure path (not executed) when the baseline loses those relationships
Scanned pages or garbled text layers Search for a distinctive visible phrase from each page, including scanned attachments. Our mixed document kept native text but omitted the image-only page. For garbled text layers, compare copied characters with the page Open-source OCR configuration with Marker (not executed), or the tested hosted API path below
Equations Compare raw expressions with the page before checking the preview. Our formula sample contains linear text only. Missing symbols need extraction work, while intact expressions displayed with literal delimiters need renderer support a configured OCR path for recognition, or Repair the failed layer before ingestion for rendering

Convert a PDF to Markdown in Python with MarkItDown

MarkItDown is sufficient for readable prose when its output preserves the required paragraphs and their sequence.

Run the text baseline

MarkItDown exposes a Python API and a PDF dependency extra for text-analysis workflows, with limited fidelity for human-facing conversion. The example uses the installed MarkItDown version recorded in the test matrix. 12

python -m pip install "markitdown[pdf]==0.1.8"
        

Save the following script and run it from the directory containing your input PDF:

from pathlib import Path
        from markitdown import MarkItDown
        
        source = Path("input.pdf")
        destination = Path("output.md")
        converter = MarkItDown(enable_plugins=False)
        result = converter.convert_local(str(source))
        destination.write_text(result.markdown, encoding="utf-8")
        
python convert_text.py
        

The result's markdown field contains the converted content. The result class retains text_content as a soft-deprecated alias. 11

Output on the corpus

MarkItDown kept the sentences and paragraph sequence in our single-column sample. In the simple invoice grid, the rows and headings stayed together in a Markdown table. These results apply to our samples, not every PDF.

In our double-column sample, MarkItDown alternated fragments from opposite columns. An invoice-reconciliation sentence runs into a demand-forecast sentence before either finishes. Paragraph openings surviving in the output did not make the passage readable.

Output from our double-column sample, unedited:

Invoice reconciliation starts with the export file, not
        
        Demand forecasts fail upward more often than they
        
        the ledger. Teams that reconcile weekly pull the
        
        fail downward. Optimistic sales inputs push safety
        

PyMuPDF4LLM kept the complete sentences and their sequence on that same input. It is the tested replacement for this layout, using its default Markdown call. 14

Output from our double-column sample, unedited:

Invoice reconciliation starts with the export file, not the ledger. Teams that reconcile weekly pull the payment provider's settlement report first, match it against issued invoices by number, and only then open the accounting system. Skipping the export step leaves ghost receivables that nobody can trace back to a real transaction. 
        
        Returns processing needs a cutoff rule before it needs software. A warehouse that accepts returns within thirty days must decide whether the clock starts at shipment, delivery, or carrier scan, because each choice shifts which orders qualify by several days. The rule has to be written down before the tooling question matters. 
        
from pathlib import Path
        import pymupdf4llm
        
        source = Path("input.pdf")
        markdown = pymupdf4llm.to_markdown(str(source))
        Path("output.md").write_text(markdown, encoding="utf-8")
        

The corpus runner executed that call on each sample. We also used pdfplumber as a raw-text baseline. It retained the simple table’s rows and headings as text, without rendering a Markdown table. Its double-column output joined unrelated fragments on the same line.

Output from our double-column sample, unedited:

Invoice reconciliation starts with the export file, not Demand forecasts fail upward more often than they
        the ledger. Teams that reconcile weekly pull the fail downward. Optimistic sales inputs push safety
        payment provider's settlement report first, match it stock up for slow movers, and the excess sits in the
        

The installed MarkItDown converter looks for form content page by page. Without it, the converter uses pdfminer on the whole document. When form content is present, it joins page chunks and uses pdfplumber text for non-form pages. Exceptions and empty output also trigger pdfminer extraction. These branches do not explain every layout failure. 3

Failure conditions and replacement paths

Failure or requirement What you will see Replacement path
Column sentences interleave Unrelated sentences alternate between columns even though paragraph openings survive Try the tested PyMuPDF4LLM call and compare complete sentences as well as their sequence
Table rows survive but headings do not Data rows survive, but grouped headings split across lines or rotated headings appear backwards Repair heading-to-value relationships, or try Marker's documented structure path, which was not executed
A visible scanned page contributes no content The local text calls leave the image-only sample empty and omit the scanned page of the mixed document Use the tested Markovo OCR path described below, or Marker (not executed). MarkItDown plugins were disabled, and these results do not evaluate configured OCR
An old report describes newline-separated table text Our simple grid produced a Markdown table. The historical report does not reproduce this output 6 Compare the current rows and headings with the source instead of applying the old report to every PDF

Marker documents structure and OCR for tables, columns, scans, and equations, but it was not executed for this guide. Its package needs an inference backend: vLLM for NVIDIA GPUs, or llama.cpp elsewhere; CPU and Apple Silicon use llama-server, with brew install llama.cpp documented for macOS. Its documented fast and balanced modes include --force_ocr for unusable text; disabling OCR also disables equation VLM calls. The output helper saves Markdown, metadata, and referenced images, converting non-RGB images before JPEG saving. Metadata does not replace block-level source locations. JSON offers page-level blocks; HTML and chunks are also documented, and merged tables can remain HTML inside Markdown. These are documented interfaces, not sample results. 451213

Convert with the Markovo API from Python

Markovo recovered the scanned text and kept leaf table headings together on our corpus, but its formula output damaged superscripts and subscripts.

Run the hosted conversion

The Markovo API accepts a multipart PDF upload, returns a job ID to poll, and exposes a Markdown download after success. This standard-library example follows the same request sequence as our corpus runner. Read the account credential from the environment. 15

import json
        import os
        import time
        import uuid
        from pathlib import Path
        from urllib.request import Request, urlopen
        
        base = "https://markovo.net"
        auth = {"Authorization": f"Bearer {os.environ['MARKOVO_API_KEY']}"}
        boundary = uuid.uuid4().hex
        fields = {"capability_id": "pdf-to-markdown", "mode": "fast",
                  "max_credit_units": "30000"}
        parts = []
        for name, value in fields.items():
            parts.append((f'--{boundary}\r\nContent-Disposition: form-data; '
                          f'name="{name}"\r\n\r\n{value}\r\n').encode())
        parts.append((f'--{boundary}\r\nContent-Disposition: form-data; '
                      'name="file"; filename="input.pdf"\r\n'
                      'Content-Type: application/pdf\r\n\r\n').encode()
                     + Path("input.pdf").read_bytes() + b"\r\n")
        parts.append(f"--{boundary}--\r\n".encode())
        request = Request(f"{base}/v1/convert", data=b"".join(parts), method="POST",
                          headers={**auth, "Content-Type":
                                   f"multipart/form-data; boundary={boundary}"})
        with urlopen(request, timeout=60) as response:
            job_id = json.load(response)["job_id"]
        
        deadline = time.monotonic() + 600
        while time.monotonic() < deadline:
            with urlopen(Request(f"{base}/v1/jobs/{job_id}", headers=auth),
                         timeout=60) as response:
                payload = json.load(response)
            job = payload.get("job", payload)
            if job.get("status") == "succeeded":
                break
            if job.get("status") in {"failed", "cancelled"}:
                raise RuntimeError(f"Conversion ended: {job['status']}")
            time.sleep(2)
        else:
            raise TimeoutError("Conversion did not finish within the polling window")
        
        with urlopen(Request(f"{base}/v1/jobs/{job_id}/download?format=md",
                             headers=auth), timeout=60) as response:
            Path("output.md").write_text(response.read().decode("utf-8"),
                                         encoding="utf-8")
        

CLI equivalent, using the pinned wheel from the installation guide: pip install "https://markovo.net/downloads/markovo-0.1.1-py3-none-any.whl#sha256=aae71e1fff133fca3a9e6c17a1fd3b1080a0bd5a05ea075eef7267491bea079d" then markovo convert input.pdf --out out --mode fast --max-credits 30. 16

Output on the corpus

The image-only page and the scanned attachment both returned readable text. Native prose and table labels also survived in the mixed document, as they did in the local text outputs. The table below is HTML, not a Markdown pipe table.

Output from our mixed prose-table-and-scan sample, unedited:

<table><tr><td>Item</td><td>Qty</td><td>Unit Price</td><td>Total</td></tr><tr><td>Consulting hours</td><td>12</td><td>85.00</td><td>1020.00</td></tr><tr><td>Data export fee</td><td>3</td><td>40.00</td><td>120.00</td></tr><tr><td>Template license</td><td>1</td><td>190.00</td><td>190.00</td></tr><tr><td>Support retainer</td><td>6</td><td>55.00</td><td>330.00</td></tr><tr></tr></table>
        
        This page exists only as a picture. A scanner produced it, the PDF embeds it as a bitmap, and no amount of text extraction will recover these words without OCR.
        

In the grouped and rotated-heading samples, Markovo kept row values and leaf headings together; local outputs kept the rows but split or displaced the headings. This does not establish preservation of merged-column semantics. See the test table at the end.

Markovo retained half of the source expressions in the linear-formula sample. The squared-term expression returned double superscripts, and the sum expression had an escaped subscript. The sample does not test stacked fractions or integrals.

Output from our linear-formula sample, unedited:

$$
        \[a^{\wedge}2+b^{\wedge}2=c^{\wedge}2\]
        $$
        

Failure conditions and replacement paths

Failure or requirement What you will see Next path
Consumer accepts only Markdown pipe tables Downloaded tables use HTML tags even when row values and leaf headings survive Enable HTML rendering or retain a structured table separately
Formula recognition changes symbols Double superscripts or escaped subscripts alter the raw expression Repair extraction against the source; renderer support alone cannot restore symbols
Workflow must run offline Upload, polling, and download require a network connection and an account API key Use a local extraction or configured OCR path
OCR classification is mistaken for processing history Preflight OCR billing classification and post-conversion OCR processing can describe different page sets Read them as separate fields, not as formula or table accuracy measurements

Preflight classifies pages for OCR billing, while the conversion report records pages processed by the OCR engine; those records do not identify how each table element was recovered. 15

Repair the failed layer before ingestion

Repair table relationships, page coverage, or rendering at the layer that failed, and ingest only content with those errors resolved.

Symptom Possible cause Discriminating check and action
Table values form a vertical list Lost cell boundaries Trace one value to its source row label and column heading. Missing associations fail cell-level use
Text alternates between columns Failed reading-order reconstruction Follow consecutive source sentences through the output. Reject passages whose sequence changes
Opening content survives but later pages vanish Incomplete page processing Search output for distinctive first, middle, and last content-page phrases. Inspect page ranges and processing records when one is absent
Formula source survives but the preview shows delimiters Unsupported math notation in the renderer Compare raw formula text with supported notation. Fix the consumer instead of re-running OCR
Marker documented path (not executed) Reports describe invoice amounts interleaving with prose 8, missing MathML despite useful LaTeX 9, and rejected structured responses 10 Reject misaligned rows, change accessibility rendering, or correct the endpoint contract; repeating OCR does not fix the latter

The local text calls retained the prose and table pages in our mixed document but left out its scanned attachment. A nonempty Markdown file did not mean the whole document was present. A historical OCR-plugin report also describes missing later pages, but we did not run that configuration or attachment. 7

In the table with a grouped price heading, both Markdown converters placed Price and its leaf labels on separate lines. In the rotated-header sample, MarkItDown reversed the heading text and used a data row as the Markdown header. PyMuPDF4LLM split Unit Price across lines. The row values surviving did not preserve their heading relationships.

Output from our rotated-header sample, unedited:

ecirP tinU
        latoT
        metI
        ytQ
        | Consulting hours | 12  | 85.00  | 1020.00 |
        | ---------------- | --- | ------ | ------- |
        | Data export fee  | 3   | 40.00  | 120.00  |
        

For grouped headings, match each leaf label to its source column before flattening the table. An application could use Price / Unit and Price / Total as the resulting labels. This is a proposed repair, not a tested postprocessor.

In our linear-formula sample, MarkItDown and pdfplumber retained all the source expressions. PyMuPDF4LLM omitted an expression. This sample does not establish recognition of stacked fractions or integrals. Missing raw text needs extraction work, while intact text that fails only in the preview needs renderer work.

Record accepted content and failures

Batch records and review queues are workflow recommendations, not guarantees supplied by either library:

  • Save the source filename and hash, package version, mode or OCR configuration, output paths, and acceptance status beside each artifact
  • Keep raw extraction separate from cleanup so changed text can be traced to a processing step
  • Send unresolved coverage or relationship errors to a review queue with the failed passage identified
  • Attach source and page information to accepted RAG content so retrieved answers can be traced to the document

How we tested

We compared fixed conversion calls on synthetic PDFs with paired ground truth, so the results measure these fixtures rather than general document accuracy.

The corpus is under lab/corpus/ in this guide's run directory, with a matching .truth.json for each PDF:

File Representative shape
prose-single.pdf Single-column prose
two-column.pdf Multiple prose columns
table-simple.pdf Simple invoice grid
table-merged-header.pdf Grouped price header with merged cells
rotated-header.pdf Rotated header text with ordinary data rows
scan-image.pdf Image-only page
mixed.pdf Prose, table, and scanned pages
equation.pdf Linear textual expressions

We used MarkItDown 0.1.8, PyMuPDF4LLM 1.28.2, and pdfplumber 0.11.10 with Python 3.11.14 on 2026-10-06. Versions, Python, and date are recorded in lab/runs.json and lab/scores.json. Calls were local MarkItDown conversion with plugins disabled, default PyMuPDF4LLM to_markdown, and pdfplumber extract_text joined across pages. These local calls recovered no scan text; the API path below tested OCR routing. Marker was not executed for this guide.

Tool Version or API record Date
MarkItDown 0.1.8 2026-10-06
PyMuPDF4LLM 1.28.2 2026-10-06
pdfplumber 0.11.10 2026-10-06
Markovo rate_card_version=2026-09-19.v3-media-unified (API record, no pip version) 2026-10-06

Markovo is the service we build. It ran through the same corpus and scorer as the other paths; its row is reported in full, including where it scored lower.

The API batch used https://markovo.net/v1/convert; lab/markovo/run.json records the endpoint, date, job IDs, and cache flags. Its elapsed_s=0.0 entries are cached records, not conversion latency measurements. The equivalent request sequence is below; the download response was text/markdown; charset=utf-8. Use a PDF at scan.pdf and an API key in the environment.

curl -X POST https://markovo.net/v1/convert -H "Authorization: Bearer $MARKOVO_API_KEY" -F "file=@scan.pdf" -F "capability_id=pdf-to-markdown" -F "mode=fast" -F "max_credit_units=5000"
        # Set JOB_ID to the returned job_id; repeat GET until status=succeeded.
        curl -H "Authorization: Bearer $MARKOVO_API_KEY" "https://markovo.net/v1/jobs/$JOB_ID"
        curl -H "Authorization: Bearer $MARKOVO_API_KEY" "https://markovo.net/v1/jobs/$JOB_ID/download" -o scan.md
        

Run these commands from the repository root with the preinstalled lab environment and corpus:

PYTHONDONTWRITEBYTECODE=1 tools/lab-venv/bin/python tools/lab_run.py --lab runs/2026-10-05-markovo-pdf-to-markdown-python/lab --tools markitdown,pymupdf4llm,pdfplumber
        PYTHONDONTWRITEBYTECODE=1 tools/lab-venv/bin/python tools/lab_run.py --lab runs/2026-10-05-markovo-pdf-to-markdown-python/lab --tools markovo
        PYTHONDONTWRITEBYTECODE=1 tools/lab-venv/bin/python tools/lab_score.py --lab runs/2026-10-05-markovo-pdf-to-markdown-python/lab --out runs/2026-10-05-markovo-pdf-to-markdown-python/lab/scores.json --markdown runs/2026-10-05-markovo-pdf-to-markdown-python/lab/scores.md
        

The metrics_doc object in lab/scores.json defines the scoring. tools/lab_score.py normalizes case, whitespace, and selected Markdown characters. Prose and scan recall match normalized sentence strings containing at least four words. Reading order compares the first six words of adjacent paragraphs as anchors, counting missing anchors as failures. Row scores require each row's cell strings on an output line. Header scores require all leaf labels on the same line and do not measure merged-column grouping. The pipe flag detects table syntax. Mixed-page coverage matches page anchors. Equation recall uses the same markup-normalised substring matching for every tool: remove $ and \[ \] delimiters and \mathsf{} font wrappers, unescape \{ and \_, and remove whitespace. equation_recall_exact retains the old plain-text matching criterion for comparison; LaTeX output naturally scores lower under that criterion. Blank metrics are not applicable. Times are single-run seconds, not a throughput benchmark.

Raw outputs use lab/<tool>/<sample>.md for every tool and sample in the table below, including lab/markitdown/two-column.md, lab/pymupdf4llm/rotated-header.md, and lab/pdfplumber/scan-image.md. lab/runs.json records local output names, times, and errors; API responses are lab/markovo/<sample>.json, with batch metadata in lab/markovo/run.json. lab/scores.json supplies the full table, also saved as lab/scores.md. The earlier lab/sample.pdf and lab/markitdown-output.md remain historical evidence and are not part of this matrix.

tool sample prose_recall reading_order table_rows_intact table_header_intact table_pipe_detected pages_covered scan_text_recall equation_recall equation_recall_exact elapsed_s
markitdown equation 1.0 1.0 0.0439
markitdown mixed 1.0 1.0 1.0 yes 0.6667 0.0 0.0329
markitdown prose-single 1.0 1.0 0.0246
markitdown rotated-header 1.0 0.0 yes 0.0118
markitdown scan-image 0.0 0.008
markitdown table-merged-header 1.0 0.0 yes 0.023
markitdown table-simple 1.0 1.0 yes 0.0128
markitdown two-column 0.1111 0.8 0.046
pymupdf4llm equation 0.75 0.75 0.2054
pymupdf4llm mixed 1.0 1.0 1.0 yes 0.6667 0.0 0.4484
pymupdf4llm prose-single 1.0 1.0 0.1768
pymupdf4llm rotated-header 1.0 0.0 yes 0.185
pymupdf4llm scan-image 0.0 0.2136
pymupdf4llm table-merged-header 1.0 0.0 yes 0.2102
pymupdf4llm table-simple 1.0 1.0 yes 0.1985
pymupdf4llm two-column 1.0 1.0 0.178
pdfplumber equation 1.0 1.0 0.007
pdfplumber mixed 1.0 1.0 1.0 no 0.6667 0.0 0.0141
pdfplumber prose-single 1.0 1.0 0.0095
pdfplumber rotated-header 1.0 0.0 no 0.0044
pdfplumber scan-image 0.0 0.0006
pdfplumber table-merged-header 1.0 0.0 no 0.0045
pdfplumber table-simple 1.0 1.0 no 0.0046
pdfplumber two-column 0.1111 0.8 0.0181
markovo equation 0.5 0.0 0.0
markovo mixed 1.0 1.0 1.0 no 1.0 1.0 0.0
markovo prose-single 1.0 1.0 0.0
markovo rotated-header 1.0 1.0 no 0.0
markovo scan-image 1.0 0.0
markovo table-merged-header 1.0 1.0 no 0.0
markovo table-simple 1.0 1.0 no 0.0
markovo two-column 1.0 1.0 0.0

Detailed results and routing records

On prose-single.pdf, MarkItDown scored 1.0 for both prose_recall and reading_order. On table-simple.pdf, it scored 1.0 for rows and headers and emitted a pipe table. These are fixture scores, not general extraction accuracy.

On two-column.pdf, MarkItDown scored 0.1111 for sentence recall and 0.8 for reading order. The output interleaves fragments from opposite columns. PyMuPDF4LLM scored 1.0 on both metrics for the same input, making it the tested alternative for this layout. Its default Markdown API uses the following call. 14

The installed MarkItDown PdfConverter.convert calls _extract_form_content_from_words per page. If no page yields form content, it runs pdfminer over the whole document. Otherwise it joins page chunks, including pdfplumber text from non-form pages. Exceptions and empty output also trigger pdfminer extraction. The snapshot is lab/source/markitdown_pdf_converter.py. This establishes the branch behavior, not a causal explanation for every failed layout. 3

Markovo scored 1.0 for scan recall on scan-image and mixed, and 1.0 for mixed-page coverage, versus 0.0 and 0.6667 respectively for the three open-source calls; its mixed prose, row, and header scores tied them at 1.0, while its pipe flag was no versus yes for the two Markdown converters (lab/scores.json). The API snapshots lab/markovo/scan-image.json and mixed.json report convert_response.jobs[0].billing_metrics.ocrPageIndexes as [0] and [2], identifying the image-only pages classified by preflight for OCR billing. billing_metrics.pagesOcr counts those billed pages; job.pages_ocr counts pages processed by the OCR engine, as reported by the container after conversion. For scan-image, both counts are 1. For mixed, pagesOcr=1, ocrPageIndexes=[2], and pages_ocr=3 mean one page was billed as OCR and all three pages were processed by the OCR engine. The public documentation states: “The server may choose the best conversion method without changing the confirmed price”. 15

Marker has no output from this guide's sample because it was not executed. Its documented save_output helper writes converted/document.md, document_meta.json, and referenced images, converting non-RGB images before JPEG saving. 12

For table-merged-header and rotated-header, Markovo scored 1.0 for rows, tying the open-source calls, and 1.0 for headers versus their 0.0; its pipe flag was no versus yes for the two Markdown converters (lab/scores.json), and its downloaded outputs contain HTML tables that require renderer support. Both API snapshots report job.billing_metrics.pagesOcr=0 and empty convert_response.jobs[0].billing_metrics.ocrPageIndexes, while job.pages_ocr=1: no page was classified by preflight for OCR billing, and one page was processed by the OCR engine. These counts do not identify how the table structure was recovered. 15

The header metric checks same-line leaf labels, not merged-column semantics. A failed header score can reflect split labels rather than missing data. For grouped headings, an application can normalize the leaves to Price / Unit and Price / Total after matching them to the source columns. This repair is a workflow proposal, not a tested postprocessor.

On equation.pdf, equation-string recall was 1.0 for MarkItDown and pdfplumber and 0.75 for PyMuPDF4LLM. The fixture contains linear text expressions, so these scores do not establish recognition of stacked fractions or integrals. Missing raw text requires extraction work. Intact text that fails only in the preview requires renderer work.

Markovo scored 0.5 under markup-normalised matching (2/4 recovered: E=mc^2 and f(x)=3x+7), versus 1.0 for MarkItDown and pdfplumber and 0.75 for PyMuPDF4LLM (lab/scores.json). The OCR output for a^2+b^2=c^2 contains double superscripts (^{\wedge}), and the sum expression has an escaped subscript; both count as unrecovered. Its exact score remains 0.0 for comparison: the old plain-text matching criterion naturally scores paths that emit LaTeX lower. In lab/markovo/equation.json, job.billing_metrics.pagesOcr=0 and the conversion response has empty ocrPageIndexes, while job.pages_ocr=1: no page was classified by preflight for OCR billing, and one page was processed by the OCR engine. These fields do not measure formula correctness. 15

Sources

  1. MarkItDown official README
  2. MarkItDown 0.1.8 package metadata
  3. MarkItDown PDF converter source
  4. Marker official README
  5. Marker 2.0.0 package documentation
  6. MarkItDown issue 1419: PDF table structure
  7. MarkItDown issue 1791: scanned-page coverage
  8. Marker issue 1068: invoice table output
  9. Marker issue 563: MathML accessibility
  10. Marker issue 862: structured-response failure
  11. MarkItDown result type
  12. Marker output saving source
  13. Marker Markdown renderer source
  14. PyMuPDF4LLM official README
  15. Markovo public developer documentation
  16. Markovo installation guide