Inspect PDF structure
Identify page count, text layer, headings, bookmarks, metadata, forms, attachments, columns, scans, and likely extraction problems.
Codex Skills to Read PDFs help Codex extract, inspect, compare, and explain information from digital and scanned PDF files. Download the skills for page-aware text extraction, tables, images, OCR routing, citations, document structure, cross-file comparison, and visual verification when layout changes meaning.
task: complete pdf reading and analysis task
inspect:
- requirements and context
- existing standards
- failure and edge cases
verify: outputs + checks + handoffReliable PDF work combines programmatic extraction with page rendering when visual placement matters and OCR when a page has no usable text layer.
These skills make Codex state the question, preserve page references, distinguish extracted facts from interpretation, and verify difficult pages visually.
Identify page count, text layer, headings, bookmarks, metadata, forms, attachments, columns, scans, and likely extraction problems.
Recover text, tables, selected pages, figures, form fields, and OCR output while preserving page-level provenance.
Summarize sections, answer questions, trace definitions, compare revisions, reconcile tables, and flag conflicting statements.
Render pages to check reading order, diagrams, equations, labels, footnotes, table alignment, redactions, and blank or clipped content.
The steps keep context, implementation, and verification visible so the result can be reviewed and repeated.
Specify the question, files, page range, desired detail, citation style, language, and whether tables or visuals are central.
Check text availability and structure, choose direct extraction or OCR, and identify pages that require images.
Keep page numbers and section labels attached to facts, tables, quotes, calculations, and uncertainties.
Render difficult pages, compare extracted values against the page, distinguish inference, and report unreadable or missing material.
The workflow adjusts to the project, audience, tools, and risk while preserving the same quality standard.
Read abstracts, methods, results, references, equations, figures, and supplemental material with citations.
Analyze reports, contracts, policies, manuals, proposals, statements, and financial tables.
Route image-only pages through OCR and visually inspect low-confidence text, handwriting, stamps, and skew.
Compare editions, revisions, clauses, figures, tables, and page-level changes across several PDFs.
Start with one defined outcome and provide the source material, constraints, and checks that matter.
These skills are designed for people who need dependable pdf reading and analysis work with a visible process.
Read papers and source documents with page-aware evidence.
Locate clauses, definitions, obligations, differences, and uncertainty for qualified review.
Extract and compare reports, figures, tables, methods, and stated assumptions.
Turn manuals, forms, statements, and scanned records into structured information.
The skills do not make unreadable scans accurate, replace qualified legal or medical interpretation, or justify treating extracted text as authoritative when the visual page contradicts it.
Install the complete skill folder and add the project-specific context before beginning.
Keep extraction, OCR, table, comparison, citation, and visual-verification guidance together.
Use repository or user scope according to whether the workflow belongs to one document collection or many tasks.
Name the files, page range, desired output, citation format, and any sensitive-data handling rules.
Use page images for tables, diagrams, forms, equations, footnotes, scans, and anything with uncertain reading order.
Practical answers about capabilities, limits, setup, and review.
Yes, when OCR and page rendering are available, but low-quality scans still need confidence checks and visual review.
The workflow preserves page provenance and can provide page-level citations when the source supports them.
Yes, with visual verification when merged cells, wrapped text, ruling lines, or multi-page layouts make extraction ambiguous.
Yes. They can align sections, clauses, tables, figures, and revisions while identifying missing or conflicting material.
No. Text extraction is efficient for straightforward pages. Rendering is used where layout or extraction quality could change meaning.
Clear context. Purposeful work. Relevant checks. A result others can understand.