conversion report

How the pulled test files converted

Every source file is converted to PHP HTML and WordPress block markup. Haskell Pandoc provides the primary external reference where it can read the format; format-native references are added where Pandoc has no reader. PDFs use macOS PDFKit for independent text and line-geometry evidence, while generic XML uses libxml2/libxslt source structure. PDF HTML semantics remain explicit inference rather than fabricated source tags. This report shows the current pass/fail shape and links into a curated stress set.

109source files
40covered input formats
298/327successful conversions
29known failures
194/196text-faithful comparisons
152/152visual-structure matches
10/10Citeproc semantic comparisons
99/109import-quality passes

Success by conversion path

PHP WordPress blocks

108/109

1 failed

PHP HTML

108/109

1 failed

Haskell Pandoc HTML

82/109

27 failed

Faithful enough diff checks

These checks compare generated outputs against Haskell Pandoc when it can read the source, a format-native external reference where one is available, or PHP HTML only as a disclosed fallback. Text scores compare normalized visible words. Visual-structure scores compare the rendered document shape: headings, paragraphs, lists, tables, images, figures, captions, code, quotes, and math. Generic XML uses an independent libxml2/libxslt transform over common structural element names. PDFKit supplies text and geometry, not an HTML semantic tree, so PDF visual tag scores are explicitly excluded rather than reported as a false mismatch.

Text

Faithful enough

194

Needs review

0

Divergent or empty

2

Visual structure

Faithful enough

152

Needs review

0

Divergent or empty

0

Source semantics unavailable

44

Bibliography semantics

Standalone BibTeX, BibLaTeX, CSL JSON, EndNote XML, and RIS inputs are rendered by Haskell Pandoc through Citeproc. Their WordPress output is checked by exact reference count and Citeproc token coverage, not by wrapper tags, because the PHP port deliberately emits editable definition-list blocks.

Proven

10

Needs review

0

No baseline

0

Import quality gates

These checks evaluate the WordPress block output as an import artifact: visible text completeness, paragraph merge or split drift, Citeproc reference coverage where applicable, media extraction diagnostics, local anchor validity, and Custom HTML block share. Generic XML is checked against independent libxml2/libxslt source structure. PDFs additionally use independent native source-token coverage and line geometry because untagged PDFs do not encode HTML heading, list, table, or image semantics.

Pass

99

Review

2

Unbenchmarked

0

Fail

8

Actionable thresholds

The common-format gate is blocking and covers normal WordPress import formats. Exotic formats are tracked separately so they stay visible without diluting the common-format release signal.

SegmentStatusQuality coverageFailuresConversion failuresPolicy
commonfail54/62 pass, review, or unbenchmarked; min 397 max 51 max 0blocking
exoticpass47/47 pass, review, or unbenchmarked; min 390 max 20 max 2tracked
SampleStatusGate summary
Unstructured OCR text-layer fixture
pdf
review4 pass, 5 review, 0 unbenchmarked, 0 fail
Docling right-to-left fixture
pdf
fail6 pass, 1 review, 0 unbenchmarked, 1 fail
Docling aircraft handbook sample
pdf
fail7 pass, 0 review, 0 unbenchmarked, 1 fail
MinerU small OCR document
pdf
fail0 pass, 0 review, 0 unbenchmarked, 1 fail
Theatre script formatting example
pdf
fail6 pass, 1 review, 0 unbenchmarked, 1 fail
Grand Canyon North Rim pocket map
pdf
fail7 pass, 0 review, 0 unbenchmarked, 1 fail
Public-domain scanned book PDF
pdf
fail5 pass, 2 review, 0 unbenchmarked, 1 fail
Muir Beach illustrated brochure
pdf
review6 pass, 2 review, 0 unbenchmarked, 0 fail
QuickBooks invoice template PDF
pdf
fail7 pass, 0 review, 0 unbenchmarked, 1 fail
Ruling-free crop table PDF
pdf
fail7 pass, 0 review, 0 unbenchmarked, 1 fail

Curated stress showcase

These are representative real-world files from the pulled corpus: leaflets, brochures, a scanned book, image-heavy packages, table-heavy office documents, and rich markup fixtures.

SampleFormatSizePHP WordPress blocksPHP HTMLHaskell HTML
CDC hand hygiene brochure
Short public health brochure with flyer-like layout, large headings, side-by-side blocks, and image-heavy design elements.
pdf339.1 KBokokfailed
Grand Canyon North Rim pocket map
National Park Service leaflet-style pocket map with map panels, services, columns, and mixed visual/text layout.
pdf1.9 MBokokfailed
Muir Beach illustrated brochure
Small National Park Service brochure with photos, map-oriented visitor information, headings, and short panels.
pdf2.4 MBokokfailed
Public-domain scanned book PDF
Internet Archive public-domain scanned small book with OCR text, title pages, and book-like page flow.
pdf1.9 MBokokfailed
TraceMonkey technical PDF
Untagged two-column PDF from Mozilla pdf.js tests, with dense prose, diagrams, code listings, and figures that exercise reading-order reconstruction.
pdf992.5 KBokokfailed
QuickBooks invoice template PDF
Public invoice template with explanatory prose, shaded table headers, line items, totals, and two editable invoice layouts.
pdf138.4 KBokokfailed
Ruling-free crop table PDF
Public Tabula fixture with a data table inferred from aligned text rather than a boxed grid, followed by explanatory prose.
pdf941.3 KBokokfailed
Multi-column numeric table PDF
Public Tabula fixture with six numeric columns arranged as two adjacent visual groups, exercising coordinate-based table grouping.
pdf8.1 KBokokfailed
Illustrated Alice EPUB
Real Project Gutenberg EPUB3 with chapter XHTML, CSS, navigation, cover art, and dozens of illustrations.
epub1.4 MBokokok
EPUB 2 picture book
EPUB fixture with packaged image content from upstream Pandoc tests.
epub11.5 KBokokok
OASIS KMIP specification DOCX
Real standards-track DOCX with many sections, tables, lists, references, styles, and packaged media.
docx471.0 KBokokok
DOCX inline images
DOCX fixture with packaged image relationships from upstream Pandoc tests.
docx14.0 KBokokok
DOCX tables
DOCX table coverage from upstream Pandoc reader tests.
docx31.0 KBokokok
OASIS OpenDocument schema ODT
Real 930 KB editable ODT specification with thousands of headings, long prose, lists, tables, styles, and package metadata.
odt930.9 KBokokok
ODT table spans
ODT table fixture from upstream Pandoc tests.
odt10.4 KBokokok
CDC food safety classroom slides
Real CDC PowerPoint deck with 16 slides, many images, list-heavy slides, and a data table.
pptx842.0 KBokokok
WHO BFHI training slides
Real WHO training PowerPoint with 13 slides, photographic assets, speaker metadata, lists, and structured slide text.
pptx284.9 KBokokok
Census tax parameter workbook
Real Census Bureau workbook with 102 worksheets of federal and state tax parameters.
xlsx374.0 KBokokok
Full GitHub syntax packet
Expanded GitHub Flavored Markdown showcase covering CommonMark leaf blocks, container blocks, inline syntax, GFM tables, task lists, strikethrough, autolinks, raw HTML filtering, alerts, footnotes, emoji, math, and diagram fences.
markdown_github6.4 KBokokok
Pandoc manual Markdown
Real 293 KB Markdown manual with dense sections, tables, links, code, lists, metadata, and large generated tables.
markdown293.5 KBokokok
MediaWiki feature packet
mediawiki584 Bokokok
Generated roff manpage fixture
Checked-in roff manpage exercising TH/SH, font escapes, tagged paragraphs, indentation, no-fill code, and generated-man requests.
man453 Bokokok

Extracted media preview

The PHP path now runs an --extract-media-style pass in images we thought were important mode. Referenced package images are written beside the converted output and their <img> URLs are rewritten to hosted files; directly embeddable PDF streams are copied out, while JBIG2 streams are decoded through the bundled browser-compatible PDF.js/PDFium raster path into validated 1-bit PNG media.

222 image media entries extracted across 21 source files.

Success by input format

FormatFilesSuccessful conversionsFailures
biblatex26/60
bibtex26/60
bits26/60
commonmark13/30
commonmark_x13/30
csljson26/60
csv26/60
doc24/62
docbook26/60
docx927/270
dokuwiki13/30
endnotexml26/60
epub618/180
fb226/60
gfm13/30
html412/120
ipynb26/60
jats26/60
jira26/60
json26/60
latex39/90
man26/60
markdown13/30
markdown_github13/30
markdown_mmd13/30
markdown_phpextra13/30
markdown_strict13/30
mdoc13/30
mediawiki26/60
native26/60
odt39/90
opml26/60
pdf2344/6925
pptx412/120
ris26/60
rst13/30
rtf39/90
tsv26/60
xlsx39/90
xml24/62

WordPress block coverage

The WordPress output generated blocks for 108 samples.

BlockCount
code653
group401
heading4289
html52
image128
list1423
paragraph23597
quote56
separator66
table586
verse10