OCR Pipeline

Scanned pages.
Readable Markdown.

Frontier OCR wrapped in a reviewable, bounded pipeline.

Choose one supported file to get started.

Conversion details & limits

OvisOCR2-class OCR, inside a complete Markdown pipeline.

Searching for OvisOCR2-level document OCR? Markovo applies frontier OCR and vision-language models inside a bounded pipeline that returns clean Markdown — with source maps, quality evidence, and the original kept for side-by-side review.

OCR and formula enhancement run as signed-in routes with the same estimate-first Credit contract as every other conversion.

01 / Beyond raw OCR

Model-class OCR, pipeline-grade output.

Running a strong OCR or vision-language model is only the first step. The value is in what surrounds it: layout-aware reading order, table reconstruction, provenance, and a result you can verify.

SCAN

Pixels to structure

Frontier OCR reads scanned pages, photos, and dense layouts — then the pipeline rebuilds headings, paragraphs, lists, and common tables as real Markdown structure.

MAP

Source maps and evidence

Every result can carry source_map.json, a quality report, and warnings — so a reviewer or an agent can trace a paragraph back to the page it came from.

PAIR

Compare against the original

The Markdown Reader renders the source PDF next to the Markdown output for line-by-line verification when the source is retained.

02 / Bounded and honest

OCR that admits its limits.

Weak scans, handwriting, dense equations, and unusual layouts still produce warnings rather than silent guesses — the report tells you which pages need a human look.

EST

Estimate before spend

Every job starts with a maximum Credit estimate. You approve the ceiling; the job cannot exceed it, and failures or cancellations charge zero.

WARN

Warnings, not silent errors

Low-confidence pages are flagged in the quality report instead of being smoothed over, so downstream readers know where to double-check.

RET

Retention you control

Results live in your account history with clear expiry. Deleting a conversion removes the stored files while keeping an auditable record that the work ran.

03 / Connect in one step

One prompt connects your agent.

The copy strip above hands your AI agent a playbook at markovo.net/install.md — it picks remote MCP, local MCP, CLI, or API for your setup, guides sign-in and API-key placement, and verifies the connection before you convert anything.

WEB

Convert in the browser

Drop a scan or document above, review the rendered Markdown in the reader, and export to Word — no account needed for the preview.

MCP

Give an agent the tool

Remote MCP connects over OAuth; local stdio MCP reaches local files inside a directory boundary you choose.

API

Build it into a product

POST /v1/convert with a mandatory max_credits ceiling; poll the job; download the bundle — the same contract the web uses.

OvisOCR2 pipeline questions.

What the OCR route does and where its limits sit.

What is OvisOCR2?

OvisOCR2 refers to a class of vision-language OCR models aimed at reading document pages — text, tables, and formulas — from pixels. Markovo applies frontier OCR models of this class inside a managed pipeline rather than asking you to operate the model yourself.

How is this different from running an OCR model directly?

A raw OCR model returns text. Markovo's pipeline adds page-aware reading order, table reconstruction, source maps, quality scores, warnings, and a reviewable Markdown bundle — so output can be checked against the original instead of trusted blindly.

Which documents benefit most from OCR conversion?

Scanned PDFs, photographed pages, screenshots of documents, and image-heavy files. Digital PDFs with selectable text already convert well on the standard route — OCR matters most where no text layer exists.

Can I use OCR conversion through an API or AI agent?

Yes. The same capability is reachable through the web converter, REST API, CLI, and MCP — every call is estimate-first with a mandatory Credit ceiling, and failed jobs charge zero.

What happens when OCR is uncertain?

Low-confidence pages are flagged in the quality report with warnings. The Markdown reader lets you compare output against the retained original so uncertain regions get a human check.