We also ship smaller models for line-level text detection and ocr error detection. It works on a range of documents (see usage and benchmarks).

Are you the maintainer?
Claim this page →Surya
OCR, layout analysis, reading order and table recognition in 90+ languages.
01 / About
What Surya is.
02 / Discussion CREDIBILITY-GATED
Discussion
Reading is open to everyone. Posting and voting need a verified identity or a GitHub grade of B or higher.
- No discussions yet.
03 / Related
More around Surya.
More from the same builder
Similar tools
ToolDocling
Data Wrangling
docling-project
Parses PDFs and many other document formats into structured Markdown or JSON, with layout, table, formula, and OCR understanding
productivitydocument-processingocrragmarkdown
ToolMinerU
Data Wrangling
opendatalab
Converts PDFs and Office documents into LLM-ready Markdown and JSON with layout analysis, OCR and structure extraction
productivitydocument-processingocr
ToolOpenDataLoader PDF
Data Wrangling
opendataloader-project
Extract Markdown, JSON (with bounding boxes), and HTML from any PDF.
document-processingocraccessibilityrag
ToolPaddleOCR
Data Wrangling
PaddlePaddle
OCR toolkit that converts PDFs and images into structured Markdown or JSON, with text, table, formula, and layout recognition
productivityocrdocument-processingtable-recognitionvision-language-model
ToolDoc7
Data Wrangling
magicrew
Turn documents into AI-ready Markdown with visual understanding
document-processingocrvision-language-modelmarkdown
ToolPullMD
Data WranglingResearch
AeternaLabsHQ
Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Cl
html-to-markdownocryoutube-transcriptdocument-conversion
04 / Build
Build with Surya.
Browse the catalogue for frameworks, tools, and harnesses, each scored on real GitHub credibility.
Get Surya →