Registry/pdf-parser
Skill

pdf-parser

Robust PDF extraction skill using OCR fallback for scanned documents.

4.0

Sign in to vote

datatext

What it does

A document extraction skill that handles both digital and scanned PDFs. For digital PDFs, it uses a fast text extraction library. For scanned documents, it falls back to an OCR engine (Tesseract) with layout analysis to preserve table structures and column ordering. Returns structured content with page metadata.

How it works

Uses pdf-lib for digital PDF text extraction. If the extracted text is below a quality threshold (measured by character density), it automatically invokes Tesseract OCR on rendered page images. Table detection uses a heuristic grid analysis to infer rows and columns from the spatial coordinates of detected text blocks.