Scoring

What's New PowerPoint text is now checked for contrast in the colour PowerPoint shows it in: most text takes its colour and size from the deck’s design rather than setting its own, and that text was never checked before. Across the real decks in our test set, about half of all text is now checked, up from about 7%; links are judged in the colour PowerPoint draws links in, and highlighted text against its highlight. A colour shaded lighter or darker, or made see-through, is left unchecked rather than guessed at. Separately, a picture “described” only by a label such as “Table (page 30)”, which automatic tagging tools write, now counts as undescribed, so some grades will go down, correctly. How scoring worksUpdated October 6, 2026 · See all updates

Check your PDF, Word, PowerPoint, and Excel files for accessibility

Upload a PDF, Word, PowerPoint, or Excel file to get an instant accessibility score that counts only WCAG 2.1 A/AA — the standard ADA Title II and the Illinois IITAA 2.1 require. The audit checks up to nine categories — text extractability, heading structure, alt text, table markup, and more — plus WCAG 2.2's newer criteria — flagged for manual review on documents with interactive forms, marked as beyond the standard being measured, and never counted — and returns a detailed report with actionable findings.

Technical Details: How This Tool Analyzes Documents & Remediates PDFs

Overview: What This Tool Does

This tool checks whether a document — a PDF, Word (.docx), PowerPoint (.pptx), or Excel (.xlsx) file — can be read by people who use assistive technology — screen readers, braille displays, and other tools used by people with disabilities. It does this by examining the internal structure of the file, not just its visual appearance. A document that looks fine on screen may be completely unreadable to a screen reader if it lacks the right internal markup.

The tool evaluates documents against WCAG 2.1 Level AA (the international standard for web content accessibility) and ADA Title II digital accessibility requirements (U.S. federal law requiring state and local government digital content to be accessible; compliance due April 26, 2027 for entities of 50,000 or more and April 26, 2028 for smaller ones and special districts), as adopted in Illinois by the IITAA 2.1 standard.

What Is a PDF, Really? (And Why It's Different from Word)

To understand why some PDFs are accessible and others aren't — and why "fixing" an inaccessible PDF can be so much harder than it looks — it helps to know what a PDF actually is under the hood. Most people use PDFs every day without ever thinking about it. Here's the short version.

A PDF is an export, not a source document. Adobe created the Portable Document Format in 1993 to solve a specific problem: making a file that looks identical on every printer, every monitor, every operating system. You don't write in a PDF — you write in Word, InDesign, Pages, or Google Docs, and then you export to PDF when you want to share the finished result. PDF is the printed-and- mailed envelope at the end of the workflow, not the word-processor you used to draft the letter.

The difference between Word and PDF is about what each format stores:

Word (.docx) says:
  <h1>Annual Report 2024</h1>
  <p>In fiscal year 2024…</p>
  <img alt="Bar chart showing arrests by month" src="…" />

PDF says:
  Page 1, x=72, y=720, font=Arial-Bold, size=24pt: glyph 'A'
  Page 1, x=85, y=720, font=Arial-Bold, size=24pt: glyph 'n'
  Page 1, x=98, y=720, font=Arial-Bold, size=24pt: glyph 'n'
  Page 1, x=72, y=680, font=Arial,      size=11pt: glyph 'I'
  Page 1, x=78, y=680, font=Arial,      size=11pt: glyph 'n'
  Page 1, x=72, y=200, image XObject ref=42 (768 x 432 pixels)
  …

Word stores the meaning of your content. The <h1> tag tells any program reading the file: "this is a top-level heading." The <img> tag has an alt attribute that describes the picture. A screen reader can read a Word file and navigate it like a webpage because the meaning is right there in the file.

PowerPoint (.pptx) and Excel (.xlsx) files store meaning the same way Word does — all three are the same Office Open XML family under the hood — which is why this tool can audit all of them directly as source documents.

PDF stores where every glyph goes on the page. That's it. A PDF doesn't natively know which glyphs are a heading and which are a paragraph — only that this letter is here, that letter is there, in this font, in this color. When you read a PDF, your brain does the work of recognizing "the big bold text at the top must be a heading." A screen reader can't do that from glyph positions alone — it would just read each glyph in sequence, which sounds like gibberish.

So how can a PDF be accessible at all? Starting in 2001 (PDF version 1.4), Adobe added an optional second layer to the format called the structure tree (or "tags"). This is a separate invisible layer that runs alongside the visual content and says "the glyphs that draw 'Annual Report 2024' belong to a <H1> element. The image at x=72, y=200 is a <Figure> element with alt-text 'Bar chart showing arrests by month'." Screen readers read the structure tree first, then jump to the visual content based on what the tree tells them.

A PDF that has this layer is called a "tagged PDF." A PDF without it is "untagged." Whether a PDF gets tagged depends on how it was exported. In Word: File → Save As → PDF → Options → "Document structure tags for accessibility" (checked by default in recent versions, but commonly turned off on older Office installs or "minimum size" exports). In InDesign: File → Export → Adobe PDF (Print) → "Create Tagged PDF". Pages and Google Docs are similar. If that box is unchecked, you get an untagged PDF — visually identical, but invisible to screen readers.

The structure tree itself looks like a webpage's DOM tree, because it borrows the same ideas:

StructTreeRoot
└── Document
    ├── H1 "Annual Report 2024"
    ├── P  "In fiscal year 2024, the agency processed…"
    ├── Figure (/Alt "Bar chart showing arrests by month")
    ├── H2 "Methodology"
    ├── P  "Data was collected from…"
    └── Table
        ├── TR
        │   ├── TH (Scope=Col) "County"
        │   ├── TH (Scope=Col) "Arrests"
        │   └── TH (Scope=Col) "Year"
        └── TR
            ├── TD "Cook"
            ├── TD "12,345"
            └── TD "2024"

Standard tag types include Document, Sect, H1 through H6, P, L / LI (list / list item), Table / TR / TH / TD, Figure, Caption, Form, Link, and Artifact (used for purely decorative content that screen readers should skip). Each can carry attributes like /Alt (alt text for figures), /Lang (language declaration), and Scope (whether a TH is a row or column header).

Linking these tags back to the glyphs they describe uses Marked Content Identifiers (MCIDs). Each chunk of content in the page's drawing instructions is wrapped in a marker (/MCID 7 … /EMC), and the corresponding structure tree node points back at that marker. It's the same idea as id attributes connecting HTML elements to JavaScript handlers — a separate identifier layer that knits two parallel representations together.

This architecture is why retrofitting accessibility into an existing PDF is so much harder than getting it right at export. When Word exports a tagged PDF, it already knows your headings are headings — it just copies that semantic information into the structure tree. When somebody hands you an untagged PDF and asks you to fix it, the only thing left is the glyph positions. Reverse-engineering "what was this heading?" from "14-pt bold text at the top of page 2" is what auto-remediation tools attempt, but with the same fundamental limitation a human would have: it's a guess based on visual cues, not a recall of authorial intent.

The practical takeaway: the most reliable path to an accessible PDF is to fix accessibility issues in the source document (Word, InDesign, etc.) and re-export with tagging enabled. The next-best path — and what this tool's optional auto-remediation feature does — is to take an already-exported PDF and add structure tags after the fact. The audit results page surfaces this distinction in the "The best path to accessibility starts at the source" notice.

How It Works

When you upload a PDF, the server runs three independent, open-source checks one after another — one reads the PDF's internal structure (tags, bookmarks, form fields), one extracts text and metadata from every page, and veraPDF validates the file against the PDF/UA standard — PDF/UA-1 (ISO 14289-1), or PDF/UA-2 (ISO 14289-2) when the document declares it — and, since v1.97.0, against its machine-testable WCAG 2.2 profile as an independent second opinion; each is reported on the audit as its own panel. The combined output of the first two feeds a scorer that evaluates nine accessibility categories and produces a weighted overall score. Word, PowerPoint, and Excel files skip that two-tool step entirely — they're already ZIP archives of XML, so the server unzips them with JSZip and reads the relevant parts with fast-xml-parser inside a dedicated, short-lived child process (no external binary; killed on timeout), then scores the result against a category set adapted for that format (see "How Scores Are Calculated" below). No data is sent to third-party services or AI models — all processing happens on the server (hosted on DigitalOcean cloud infrastructure). The uploaded file is deleted (or discarded from memory) immediately after analysis — no file content is retained on the server.

PDF:
  File → [validate type & size]
       → in order: QPDF (structure) → PDF.js (content) → veraPDF (PDF/UA)
       → Scorer (9 categories) → Weighted Score → Report

Word / PowerPoint / Excel:
  File → [validate type & size]
       → JSZip + fast-xml-parser (short-lived child process)
       → Scorer (adapted categories) → Weighted Score → Report

Since v1.100.0 the web page usually drives this pipeline through a progress variant (POST /api/analyze-job + status polling) that runs the identical steps and reports each one's real state while you wait — the synchronous endpoint above remains unchanged for every other caller and as the page's automatic fallback.

An audit survives leaving the page. A single-file upload runs as a server-side job: the browser posts the file, gets back a job identifier and a one-time key, and polls for progress. The audit therefore does not depend on the page staying open — it runs to completion on the server whether or not anyone is watching, and the finished report waits in memory until a page collects it or ten minutes pass. Since v1.147.0 the page keeps that identifier and key in the browser’s per-tab session storage, so following the status link (a server route, so a real page load), reloading, or clicking any in-app link no longer discards the work: on return the page rejoins the same job and carries on, and a report that had already arrived is re-rendered from the copy the browser kept. Nothing is uploaded twice. A batch still cannot survive leaving — its queue holds the files themselves — so the warning before navigating away now appears only when leaving would genuinely lose something.

Files the tool cannot audit — legacy binary Office formats (.doc, .xls, .ppt, .rtf), CSV exports, images — are refused up front with a specific explanation instead of a generic error, even when the file has been renamed (the format is detected from content, not the name). Service health and anonymous usage totals are published on the public status page.

Audit pipeline — visual flow
Audit pipeline — visual flow

Browser uploads a file; the server validates magic bytes and size. A PDF is written to a short-lived temp copy: qpdf analyzes its structure, pdfjs then extracts its content, and veraPDF runs two checks one after the other: PDF/UA conformance (PDF/UA-1, or PDF/UA-2 when the document declares it) and its machine-testable WCAG 2.2 profile. A Word, PowerPoint, or Excel file is unzipped in memory with JSZip and parsed with fast-xml-parser instead — no temp file, no external tools. Either path feeds the scorer, which returns a grade, WCAG verdict, and findings to the browser, then discards the memory buffer.

Application Architecture

The application is a monorepo with two components, both running on the same DigitalOcean droplet:

Frontend (port 5102)

A Nuxt 4 (Vue 3) web application that provides the user interface — the upload form, progress indicators, score cards, export buttons, and shareable report pages. Styled with Tailwind CSS and Nuxt UI. Served as a server-rendered app via Nitro.

Backend API (port 5103)

An Express (Node.js/TypeScript) server that handles file uploads, runs QPDF/PDF.js analysis on PDFs and a child-process OOXML parser on Word/PowerPoint/Excel files, scores the results, and stores shared reports in a SQLite database (WAL mode). There are no accounts or sign-in — the tool is free and open to use, and stores no email addresses, IP addresses, or browser identifiers. Managed by PM2 in production.

Both processes are managed by PM2 behind an nginx reverse proxy on a single DigitalOcean droplet provisioned via Laravel Forge. The frontend proxies API requests to the backend — the user's browser never communicates directly with the API server.

Application architecture
Application architecture

Browser talks to Nginx reverse proxy. Nginx routes to either the Nuxt web app (port 5102) or the Express API (port 5103). The web app makes some API calls back to Express. For PDFs, Express shells out to qpdf, OpenDataLoader Java, and veraPDF Java; for Word, PowerPoint, and Excel files, it parses OOXML with JSZip and fast-xml-parser in a short-lived child process instead. Express reads and writes SQLite locally. No external services.

Tool 1: QPDF (PDF Structure Extraction)

QPDF is an open-source C++ command-line program for inspecting and transforming PDF files. It is maintained by Jay Berkenbilt and is widely used in PDF archival libraries, digital preservation projects, and accessibility workflows. Think of QPDF as a tool that can "open up" a PDF and read its internal blueprint — not just the words on the page, but the hidden structural information that tells assistive technology how the document is organized.

How it's called: The server invokes QPDF as a subprocess with the --json flag, which outputs the PDF's complete internal object graph as machine-readable JSON. The server writes the uploaded PDF to a temporary file, runs qpdf --json /tmp/<uuid>.pdf, parses the resulting JSON, and immediately deletes the temp file. The subprocess has a 30-second timeout and a 50 MB output buffer to handle complex documents safely.

Why QPDF? A PDF file is not a simple document — internally, it is a collection of numbered "objects" (text streams, images, fonts, bookmarks, form fields, tags) connected by cross-references. QPDF can decode and dump this entire object graph as structured data, which lets the tool inspect every accessibility-relevant feature without relying on visual rendering. No other open-source tool provides this level of structural access to PDFs.

What QPDF extracts

Data QPDF extracts from a PDF, its source, and what it's used for
DataPDF SourceUsed For
StructTreeRootCatalog /StructTreeRoot Whether the PDF is "tagged" (has a semantic structure tree)
Language declarationCatalog /LangLanguage accessibility (screen reader pronunciation)
Headings (H1–H6) Structure elements with /S = /H, /H1…/H6 Heading presence, hierarchy validation, level-skip detection
Outlines / Bookmarks/Outlines → /First//Next chain Bookmark count for the navigation advisory
Tables & structure Structure elements /Table, /TR, /TH, /TD, /Caption, /Scope, /Headers Header cells, scope attributes, row structure, nesting, captions, column consistency, header-data associations
Images & figures XObjects (/Image) + structure elements (/Figure with /Alt) Image detection and alt text presence
Form fields Widget annotations + /AcroForm /Fields + /TU tooltip Whether form fields have accessible labels
Reading order MCIDs Numeric /K values (Marked Content IDs) in structure tree Content sequence validation — detects out-of-order reading
Lists Structure elements /L, /LI, /Lbl, /LBody List detection, well-formedness (label + body per item), nesting depth
ParagraphsStructure elements with /S = /P Text organization — whether body text is structurally tagged
MarkInfo & artifactsCatalog /MarkInfo → /Marked Whether content is distinguished from artifacts (headers, footers, watermarks)
Role mapping/RoleMap on Catalog or StructTreeRoot Custom tag mappings to standard PDF roles (e.g., Title → H1)
Tab orderPage objects /TabsWhether keyboard navigation follows the structure tree
Font embedding FontDescriptor /FontFile, /FontFile2, /FontFile3 Whether the fonts used to display text are embedded (non-embedded fonts can cause garbled text; fonts no content stream can select, or that never display visible text, are exempt)
Language spansStructure elements with their own /LangInline language declarations for foreign-language content
PDF/UA identifierXMP metadata stream (pdfuaid:part) Whether the document claims PDF/UA (ISO 14289) accessibility conformance
Artifact elements Structure elements with /S = /Artifact Decorative content (headers, footers, watermarks) distinguished from real content
ActualText & expansion/ActualText and /E on structure elements Screen reader text overrides for ligatures, symbols, and abbreviation expansions
Footnotes & formulas/Note elements (/ID) and /Formula elements (/Alt) Footnote linkability advisory (Matterhorn 19); formulas without a spoken alternative reduce the alt-text score (v1.92.0)
Annotation tagging (OBJR) Structure-tree /OBJR references vs. visible widget, link, and markup annotations Form-field widgets no structure element references reduce Form Accessibility; comments and markup get tagging advisories (v1.94.0; Matterhorn 28)
Document behaviors JavaScript actions, multimedia annotations, optional-content layers (/OCG), reference XObjects, embedded files (/EF + /Desc), signature fields Disclosed as behaviors advisories so nothing a reader will encounter is silently skipped (v1.92.0–v1.94.0)

Tool 2: PDF.js (Content & Metadata Extraction)

PDF.js is Mozilla's open-source JavaScript PDF renderer — the same library that powers Firefox's built-in PDF viewer, used by hundreds of millions of people. While QPDF reads the internal blueprint, PDF.js reads the PDF the way a human would: it renders each page and extracts the actual text content, metadata (title, author, language), and interactive elements like links. It runs server-side via Node.js, processing every page of the uploaded document.

What PDF.js extracts

Data PDF.js extracts, its extraction method, and what it's used for
DataMethodUsed For
Text contentpage.getTextContent() per pageText extractability (minimum 50 chars = "has text")
Title, Author, Languagedoc.getMetadata() Title/language scoring (filename-like titles are rejected)
Links & link textpage.getAnnotations() + spatial text matching Link quality — detects raw URLs vs. descriptive text
Image count (approx.)page.getOperatorList() + image object resolution Fallback image detection when QPDF finds no tagged images — deduplicates per page, filters out images smaller than 50px (spacers, borders). Count is approximate and may include decorative graphics.
Outlinesdoc.getOutline()Bookmark detection (cross-referenced with QPDF)
Unmapped glyphsgetTextContent() character codes in the Unicode Private Use Areas or U+FFFD Text that renders fine but extracts as unpronounceable symbols — caps Text Extractability when heavy (v1.94.0; Matterhorn 10)
Untagged visible textgetTextContent({includeMarkedContent}) marked-content stream per page Visible, non-artifact text painted outside every tagged run — caps Text Extractability when heavy; affected pages are named (v1.94.0; Matterhorn 01)
Empty pagesPer-page text length < 10 chars Detects blank pages or pages with content only as images (may need OCR)

Link text extraction uses a spatial matching algorithm: for each link annotation, PDF.js finds text items whose coordinates fall within the link's bounding rectangle (±5px tolerance), then joins them to determine the visible link text. This is how the tool distinguishes descriptive links ("View the full report") from raw URLs ("https://example.com/report.pdf").

Why Two Tools?

No single open-source library can extract both the low-level PDF structure (tag trees, object references, XObjects) and the rendered text content. Each tool sees a different layer of the document:

QPDF sees:

Structure tags, heading hierarchy, table markup, image objects, form field labels, bookmark chains, reading order markers — the "skeleton" of the document.

PDF.js sees:

Rendered text content, document title and metadata, link URLs and their visible text, page count, image rendering operations — the "surface" of the document as a user would read it.

By cross-referencing both outputs, the scorer can answer questions that neither tool could answer alone. For example: "Does this image have alt text?" requires QPDF to find the image object and its Figure tag, while "Is there any readable text on this page at all?" requires PDF.js to attempt text extraction. The two run one after the other rather than at once: PDF.js works through a document page by page and holds the server's attention while it does, which used to leave qpdf waiting long enough to be cut off as if it had stalled — on documents it actually reads in a second or two. Running them in order costs almost nothing on a long report and removes that failure entirely.

Two-tool analysis (PDF)
Two-tool analysis (PDF)

For a PDF, the uploaded buffer runs through qpdf (structure tree, language, outlines, images, tables) and then pdfjs (text, metadata, content order); their results combine in the scorer for a weighted score across 9 categories. Word, PowerPoint, and Excel files don't need this two-tool split — a single JavaScript parser (JSZip + fast-xml-parser, in a short-lived child process) reads their XML directly.

How a report is presented

Every report opens in the Visual view: the grade, then a numbered action plan that walks through one fix at a time in plain language, with instructions for both the source document and Adobe Acrobat. The Detailed view holds the complete technical report — every finding, the WCAG criteria each maps to, the evidence behind it, PDF/UA signals and methodology. The chooser sits above every report and the choice is not remembered between reports: the stepper is written for document authors, so it is what greets everyone, every time, including people who prefer the detailed view and know where the toggle is.

Between the action plan and “Above and beyond” sits a Best practices section — 43 non-scored practices in all, 21 for PDF and 22 across Word, PowerPoint and Excel. Each row shows the evidence found in the document itself (the heading order it actually has, the fonts it actually carries), both routes to fix it, and a link to the rule. Nothing in the section touches the grade, and the section says so.

It is extra credit, and only extra credit (v1.148.2). A report lists a practice only when it has something a reader can act on or take credit for — worth doing, or met. Anything the grade already dealt with is left out: a defect that cost points belongs in the action plan, and a practice the checker could not judge because a scored failure got in the way (heading level order, on a document with no heading tags at all) is not something anyone could go and do. Three labels were tried on those rows in one afternoon — “not applicable” on a defect that had just cost points, “counted in your score” beside practices that are never scored, “not checked” beside a category the same page had scored zero — and each read as a contradiction, because the section was being asked to describe things that belong elsewhere. Listing fewer rows needs no label at all.

Both views end with “Still worth checking by hand”, which appears on every report at every score — including a perfect one. These checks confirm that accessibility structure is present; almost none of them can judge whether it is correct. Alt text reading “image” passes. A heading describing the wrong section passes. So each check a document passed contributes the one judgment the tool could not make, and the WCAG criteria this tool does not evaluate at all are listed by name.

Printer-friendly action steps opens the plan in a new tab as a self-contained page: every fix expanded, both routes shown, the human checks included, and nothing loaded from the network. Print it or save it as a PDF and work from it beside the document rather than behind it. The same button appears on the auto-remediation result, where it prints what the automatic fixes could not repair.

How Scores Are Calculated

For a PDF, the scorer weighs nine accessibility categories anchored to WCAG 2.1 AA and IITAA 2.1 §E205.4 — the rules that govern non-web document accessibility in Illinois. Each category receives a score from 0 to 100 (or N/A if the category doesn't apply to the document). The overall score is a weighted average across the categories. A category that doesn't apply — table markup in a document with no tables — counts as passing, because such a document has no table problem; one the tool could not evaluate is left out of the calculation entirely, because scoring it as a pass would be a claim we cannot back. Word, PowerPoint, and Excel files are graded on the same model, just with a category set adapted to each format (see below). A scanned document is the exception in the other direction: it scores zero, because nothing in it can be read by a screen reader at all.

The score is capped by the worst finding

A document's score may never outrank its worst unresolved finding — a Minor caps it at 89, a Moderate at 79, a Critical at 69. The letter then comes off the same published scale it always has (90 = A, 80 = B, 70 = C, 60 = D, below that F), so the number and the letter can never disagree. A weighted average cannot express "one thing here is disqualifying", but accessibility conformance is pass/fail per criterion, not a mean, so the average alone let four perfect categories outvote one catastrophic one.

  • A — nothing found
  • B — only minor items remain
  • C — at least one moderate issue
  • D — at least one critical issue (F if the average is also failing)

The cap only ever lowers a score, never raises one, so a document already below a ceiling keeps its own lower number. It also means two documents with the same defect get the same letter regardless of how much else was checkable in each, which was not true before tool v1.58.0: a one-page notice and a longer agenda missing the identical document title graded C and B, because the notice had only three applicable categories to average against and the agenda had seven. Where a score is sitting at its ceiling, the report says which finding is holding it there.

Scoring category weights (WCAG + IITAA §E205.4)
Category Weight WCAG + IITAA §E205.4
Text Extractability20%
Title & Language15%
Heading Structure15%
Alt Text on Images15%
Bookmarks / Navigation5%
Table Markup10%
Link Quality5%
Reading Order10%
Form Accessibility5%
Total100%

About this score

This is a WCAG-based evaluation. It aligns with WCAG 2.1 Level AA, ADA Title II, and Illinois IITAA 2.1 §E205.4 — the rules that govern non-web document accessibility in Illinois. The scorer emphasizes programmatically determinable structure (real headings, real table-header relationships, logical reading order) because that's what assistive technology can actually use.

When the veraPDF engine is configured (it is on the production deployment), every PDF audit also includes a formal PDF/UA conformance check — against PDF/UA-1 (ISO 14289-1), or PDF/UA-2 (ISO 14289-2) when the document declares it — and, since v1.97.0, a second pass against veraPDF's machine-testable WCAG 2.2 profile (the subset a dedicated checker like PAC verifies by machine, including PDF text contrast), by veraPDF, shown as its own verdict panel on the report; the optional remediation pipeline runs the same check on its output. PDF/UA is referenced by IITAA only in §504.2.2 for authoring-tool export capability, not for the final PDF artifact itself, so the WCAG-anchored score above is what governs publication decisions.

Word, PowerPoint & Excel: adapted category sets

The nine-category table above is the PDF model. Word, PowerPoint, and Excel files are scored the same way — a weighted average of 0–100 category scores, grounded in the same WCAG 2.1 AA criteria — but each format uses its own adapted category set, because not every PDF category has an OOXML equivalent:

  • Word (.docx) — 8 scored categories: Text Extractability, Title & Language, Heading Structure, Alt Text on Images, Table Markup, Color Contrast, List Structure, and Link Quality. Reading Order and Form Accessibility are shown on the report but are not automatically scored.
  • PowerPoint (.pptx) — 9 scored categories: the same eight minus Heading Structure — slides carry titles, not a heading hierarchy — plus a presentation-specific Slide Titles check (every slide needs a distinct title placeholder so screen-reader users can tell slides apart) and a Reading Order category that verifies each slide's title reads first in tab order — reported as a clearly labelled advisory, never counted. Heading Structure, Bookmarks and Form Accessibility are not scored for a presentation.
  • Excel (.xlsx) — 7 scored categories: Text Extractability, Title & Language, Sheet Names (descriptive names vs. Excel defaults like "Sheet1"), Table Markup, Alt Text on Images, Color Contrast, and Link Quality. Excel workbooks have no document-language property, so Title & Language evaluates title only. Heading Structure, List Structure, Reading Order, Bookmarks and Form Accessibility do not apply to a workbook and are not scored.
  • Empty headings: scored for Word, reported for PDF. A Heading style on a blank line — used to make space — puts a section in the outline with no content in it, so someone moving by heading lands on silence. In a Word document that is a scored WCAG 1.3.1 (Level A) failure — 10 points each, capped at 30, so it can never take Heading Structure past Minor. The identical defect in a PDF is reported and not scored: there the evidence is an estimate (pdf.js text attribution) rather than a certainty (a heading style with no text, read straight from the XML), and the PDF row says plainly that checkers disagree about it. A heading whose content is a described picture — an agency masthead — counts as a heading, using its alt text; an undescribed one is left to the alt-text check rather than charged twice.

Color contrast is one place Office formats do more than PDF: Word, PowerPoint, and Excel all read text/fill colors directly from their XML, so contrast is machine-checked wherever a resolvable color pair is set — explicit colors, theme-based colors (all three formats since v1.95.0), and Excel's legacy indexed palette. Word's and Excel's tints and shades are computed; in those two formats style-inherited and automatic colors stay unresolved. PowerPoint text that sets no color of its own is followed through its slide's layout and master text styles (since v1.166.0), while a PowerPoint color shaded lighter or darker, or made see-through, is left unassessed rather than guessed at. PDF's Color Contrast category, by contrast, remains N/A pending rendered-page analysis — see "Color contrast" under Limitations below.

Category scoring logic

Text Extractability (20% weight — highest)

What it means: Can a screen reader actually read the words in this PDF? Some PDFs are just pictures of text (scanned documents) — they look normal on screen but are completely invisible to assistive technology.

How it's scored: 100 = extractable text + structure tags. 50 = text is present but no tags (an untagged PDF). 25 = tags are present but no extractable text (partially remediated scan). 0 = no text and no tags (unremediated scanned image). Two text-layer censuses also cap this category, each a confirmed WCAG failure: a heavy share of unmapped glyphs (characters that extract as unpronounceable symbols — 1.1.1, Matterhorn 10) caps at 50, and visible text outside every tagged run (1.3.1, Matterhorn 01-005/006, with the affected pages named) caps at 50 or 85 by share; a handful of either stays an advisory. Non-embedded fonts are reported but never scored — no WCAG criterion requires embedding (a substituted font still renders and reads aloud); PDF/UA does, so they appear as a PDF/UA-only item. This category carries the highest weight because if text can't be extracted, nothing else matters.

Title & Language (15%)

What it means: The document title is the first thing a screen reader announces when a user opens the PDF. The language tag controls how the screen reader pronounces words — without it, an English document might be read with a French accent, making it incomprehensible.

How it's scored: 50 points for a document title — a title that cannot describe anything (a bare file name such as "report_final.pdf", an authoring-tool default such as "Microsoft Word - Cook.doc" or "Untitled", a placeholder, a pure timestamp or hash) earns partial credit and is a confirmed 2.4.2 failure (WCAG's own documented failure F25). A title that merely looks like a file name but carries real words ("Annual_Report_2024", a real title with an export timestamp glued on) names the document, so it earns full credit and is reported as an unscored advisory — whether it describes the document well is a judgment for a person (since 2026-09-02). Word, PowerPoint and Excel titles get the same check since 2026-10-05 — PowerPoint's own default, "PowerPoint Presentation", included. Plus 50 points for a usable language declaration. The language value is checked two ways, each a confirmed 3.1.1 failure at half credit: a declaration that is not a usable code ("english", "en_US"), and a declaration that contradicts the text's actual language (a stopword-based check with four guards against false accusations). Word and PowerPoint declarations get the same two checks since 2026-10-05, with one more guard: a language the file itself marks on some passage is never called a mismatch, because a screen reader switches to it there. And when a Word or PowerPoint file sets no document-wide language, the language marked on more than half of its text is its language (since 2026-10-06) — what selecting the text and setting its proofing language produces. The DisplayDocTitle viewer flag counts too (since 2026-09-01): a title that is set but not displayed earns half the title credit and is a confirmed 2.4.2 failure — W3C's own PDF technique for 2.4.2 (PDF18) sets the flag, and with it off every viewer shows the filename instead. Checked in QPDF's catalog /Lang and PDF.js metadata.

Heading Structure (15%)

What it means: Headings (H1, H2, H3, etc.) are how screen reader users navigate and skim documents — the same way sighted users scan bold section titles. Without headings, a blind user must listen to the entire document from start to finish to find the section they need.

How it's scored: 100 = heading tags are present — the outline exists and is programmatically identifiable. 0 = no heading tags at all while the page itself shows section headings — at least two lines that look like section headings (uniformly larger, or uniformly bold, text sitting over body text), which the report names — a confirmed 1.3.1 failure: the structure a sighted reader sees is conveyed by presentation only. A document with no heading tags and no such lines has no visual structure to convey, so it is not scored on headings (since 2026-09-02; before that the failure was inferred from page and paragraph counts). Word and PowerPoint apply the same threshold (since 2026-10-05): with no Heading styles — or no titled slide — one bold or large line is the document's title and is not scored; two or more are sections, Critical in all three formats. When heading tags do exist, a line that looks like a section heading but is tagged as an ordinary paragraph costs 15 points, at most 40 (since 2026-10-06) — Word's own rule for a paragraph formatted to look like a heading. To avoid false accusations, a PDF line counts only if it is its paragraph's whole text, doesn't end mid-sentence, isn't a caption, and such lines appear on at least two pages (so a cover page or letterhead never counts). Everything about the outline's shape — level skips (W3C's own guidance: not a WCAG failure), multiple H1s, generic /H tags, mixing conventions (PDF/UA 7.4.4 / Matterhorn 14-007), and whether the headings' text reads like headings — is reported as clearly labelled advisories and never scored.

Alt Text on Images (15%)

What it means: Every informative image in a PDF must have "alternative text" — a short description that a screen reader reads aloud. Without alt text, a blind user hears nothing when they encounter a chart, photo, or diagram.

How it's scored: The percentage of detected images that have alt text. QPDF identifies image objects (/Image XObjects) and matches them to their /Figure structure elements, then checks whether each Figure has an /Alt attribute. If QPDF finds no tagged images, PDF.js provides a fallback by counting image rendering operations — if images exist but aren't tagged, the category scores 0 (Critical) instead of N/A. N/A only if no images are detected by either tool.

Bookmarks / Navigation (5%)

What it means: Bookmarks act as a clickable table of contents in the PDF viewer's sidebar — for longer documents, a real navigation aid for every reader, screen-reader users included. No WCAG 2.1 criterion requires them inside a single document, so they can never affect the score.

How it's scored: It isn't. Under 10 pages the category is N/A; at 10+ pages a missing bookmark tree is reported as a clearly labelled advisory that never affects the grade — no WCAG 2.1 criterion requires bookmarks in a single document (2.4.5 Multiple Ways applies to sets of pages). Checked in both QPDF's /Outlines object chain and PDF.js's getOutline().

Table Markup (10%)

What it means: When a sighted user looks at a data table, they can glance at the column headers to understand what each number means. Screen readers need explicit markup to provide the same context — without it, a screen reader reads a flat stream of numbers with no structure. This category checks seven aspects of table accessibility.

How it's scored: N/A if no tables are detected; one-row and one-column constructs are layout scaffolds and are excluded (the conformance gate applies the identical rule, both halves mirrored since 2026-08-31). So is a table with no header cells that draws nothing — no ruled line and no cell fill anywhere in its area, read from the page itself (since 2026-10-06): a grid used only to line things up, which Word has never scored either. A ruled table without header cells is still a data table missing its headers. What is scored is what WCAG 1.3.1 requires: header cells (/TH present), row structure (cells grouped in /TR), a regular grid (consistent column counts after row/column-span accounting), and — only for tables whose headers run along more than one edge or contain spanned cells — a /Scope or /Headers association, without which the header-to-data relationship cannot be determined. Reported but never scored: missing /Scope on plain one-header-row tables (the shape already answers the question — a PDF/UA-only item), nested tables, captions, and — in every format — header cells that hold no text, or Excel tables still headed with its own default names ("Column1", "Column2"). Word, PowerPoint and Excel: each data table whose header row is not marked scores 45 — exactly what a PDF table with no header cells scores — so the same defect costs the same in every format. (In Word, either Table Design → Header Row or Table Layout → Repeat Header Rows marks it. A Word table counts as layout when nothing is drawn — borders switched off, or a table style that draws nothing — and as data when borders are drawn on its cells.)

Link Quality (5%)

What it means: Screen reader users often navigate by tabbing through links. Hearing "https://www.example.com/documents/2024/report-final-v3.pdf" read aloud character by character is unusable. Descriptive link text like "Download the 2024 Annual Report" tells users where the link goes.

How it's scored: N/A if no links. What is scored is untagged links — annotations no /Link structure element claims, which a screen reader following the tags never encounters (a confirmed 1.3.1 failure) — and links with no text at all, which reach a screen reader with no name to announce (a confirmed 4.1.2 Name, Role, Value failure, Level A). A link whose text this tool could not attribute — rotated text, image-only links — is flagged for hand-checking, never scored. Link wording — raw URLs as text, or vague phrases like "click here" — is reported as an advisory and never scored: WCAG 2.4.4 (Level A) allows a link's purpose to come from its surrounding context, which no text-only check can weigh; judging the text alone is 2.4.9, a AAA criterion. PDF.js extracts the visible text overlapping each link annotation using spatial coordinate matching.

Form Accessibility (5%)

What it means: If a PDF contains fillable form fields (text boxes, checkboxes, dropdowns), each field needs a label that assistive technology can read. Without labels, a screen reader user hears "edit text" or "checkbox" with no indication of what the field is for.

How it's scored: N/A if no form fields. Percentage of widget annotations (form fields) that have a /TU (tooltip) attribute, which serves as the accessible label. QPDF checks both the widget annotation and the /AcroForm fields array. Since v1.94.0, visible widgets that no structure element references (no /OBJR — a screen reader following the structure never reaches them) reduce the score proportionally as well, the same treatment untagged links get (Matterhorn 28).

Reading Order (10%)

What it means: PDFs with multi-column layouts, sidebars, or callout boxes can confuse screen readers if the reading order isn't explicitly defined. A sighted user can see that a sidebar is separate from the main text, but a screen reader reads content in the order defined by the structure tree — if that order is wrong, the document becomes a jumble of unrelated sentences.

How it's scored: only one condition scores: 0 = no structure tree at all, a confirmed 1.3.2 failure — no programmatic reading sequence exists. Everything else is measured and reported, never scored: the tagged order (structure-tree MCID sequence) is compared against the order the page's content is painted (content-stream MCID sequence) using a longest-common-subsequence match, and any divergence is disclosed with the affected pages named — but divergence proves the two orders disagree, not which side is wrong (remediated documents re-order tags away from a bad draw order on purpose), so it cannot support a deduction. A flat structure tree is likewise an advisory: a single-sequence tree is still a programmatic reading order. Image (Figure) runs are excluded from the comparison — exporters paint images by z-order, which says nothing about reading order. Forms are never measured on this metric at all, and when the sequences are too short to compare, the category reports no score rather than guessing.

Thresholds and heuristics

Every scored rule that turns on a number is listed here with the number, so a reviewer can reproduce any accusation the report makes. The values live in the analyzer and audit.config.ts; this list is kept in step with them by a guard test.

  • Visual headings (PDF, 1.3.1): a document with no heading tags is scored 0 only when at least two lines look like section headings. A line qualifies when it is 80 characters or shorter, every lettered run on it is at least 1.5 pt larger than the document's body size (the size carrying the most letters) or every run is bold at body size, its runs are contiguous (a gap wider than 1.5 × the font size marks a table row, not a heading), and a body-size, non-bold line of 40+ characters follows it on the same page. Memo header lines ("TO: Jane Smith") and table header rows are excluded by construction; one qualifying line is a title, not sections.
  • Title shape (every format, 2.4.2 / F25): confirmed only for a title that ends in a document file extension, matches an authoring-tool default or placeholder ("Untitled", "Document1", "scan_001", "Microsoft Word - X.doc", "PowerPoint Presentation"), or is a pure timestamp / digit run / hash with fewer than two real words left once file-name machinery is stripped. Anything else that looks like a file name (underscores, two or more hyphens, a 20+ character token with digits, an export timestamp or hash beside real words) is an unscored advisory.
  • Untagged visible text (PDF, 1.3.1): characters painted outside the tag structure and not marked as artifacts fail when they are at least 2 % of the visible text and at least 50 characters (score capped at 85); at 10 % and 200 characters the cap is 50.
  • Unmappable characters (PDF, 1.1.1): glyphs that extract as private-use symbols fail when they are at least 100 characters and at least 5 % of the extracted text; smaller shares are an unscored advisory.
  • Large text (Word / PowerPoint / Excel, 1.4.3): 18 pt or larger, or 14 pt bold or larger, is held to 3:1 instead of 4.5:1. For Word the size is resolved through the paragraph style and document defaults, not only the run.
  • Typed bullets (Word / PowerPoint, 1.3.1): a paragraph that starts with a hand-typed bullet or enumerator counts only when another typed or real list item sits within two non-empty paragraphs of it — a lone "1." is a label, not a list. A list typed entirely by hand floors at 45 (Moderate), the same band as an unmarked table: every word is there and in order; only the list structure is missing.
  • Word fake headings (1.3.1 / F2): a paragraph with no Heading style, 120 characters or shorter, whose own runs are bold at 14 pt or larger. Paragraphs inside data-shaped tables (two or more rows and columns) and inside text boxes are not counted — a title in a one-row banner table still is; "Title" and "Subtitle" styles are structural. With no Heading styles anywhere, a single such paragraph is the document's title and is not scored; two or more are sections.

How fix times are estimated

The action plan's minute figures count clicks, never judgment. Each estimate is the document's own defect count multiplied by a per-item rate, rounded up to a floor: headings 30 s each (floor 2 min), typed bullets 20 s (floor 2 min), contrast runs 60 s (floor 2 min), header rows 60 s per table (floor 1 min), unnamed links 30 s (floor 1 min), alt text 60 s per image to apply — writing the description is not counted, so alt-text steps never contribute to a total. Title and language are a flat 2 minutes; bookmarks a flat 5. For PDFs only the mechanical Acrobat steps (title, language, bookmarks, applying alt text) carry an estimate; tag surgery depends on the document's history, not its counts, and shows none.

Supplementary analysis

In addition to the nine PDF categories scored above, the tool appends additional findings to relevant categories. These are informational and never move a score — the alt-text quality census (boilerplate descriptions such as "image" or "picture") is reported for review only; whitespace-only alternate text is treated as missing by the alt-text check itself. They provide deeper insight into the document's accessibility posture.

Supplementary analysis checks, which category they append to, and what they report
CheckAppended ToWhat It Reports
List structureReading Order Per-list breakdown of <LI>, <Lbl>, <LBody> presence and nesting depth
Marked content & artifactsText Extractability/MarkInfo status, paragraph tag count, empty page detection
Font embeddingText Extractability Per-font embedded/not-embedded listing — reported, never scored: no WCAG criterion requires embedding (a substituted font still renders and reads aloud), so non-embedded fonts appear as a PDF/UA-only item; fonts that never display visible text are noted as harmless
Role mapping & tab orderReading Order Custom tag role mappings, per-page tab order configuration
Language spansTitle & Language Inline language declarations for foreign-language content within the document
Alt text qualityAlt Text on Images Heuristic check for non-human-readable alt text: hex-encoded data, filenames, generic placeholders, long strings without spaces
PDF/UA identifierText Extractability Checks XMP metadata for pdfuaid:part — indicates if the document claims PDF/UA (ISO 14289) conformance
Artifact taggingText Extractability Counts /Artifact structure elements — headers, footers, and watermarks should be tagged as artifacts so screen readers skip them
ActualText & expansionReading Order/ActualText for glyph/ligature overrides and /E for abbreviation expansions — help screen readers pronounce content correctly
Footnotes (<Note> IDs)Reading Order<Note> tags missing the unique /ID that lets assistive technology link a footnote to its reference — advisory with the Acrobat fix path (v1.92.0; Matterhorn 19)
Role-map validityReading Order Circular mappings, remappings of standard tags, and custom tags that never reach a standard role (resolved transitively) — advisories (v1.92.0; Matterhorn 02)
Document behaviorsReading Order JavaScript actions, audio/video annotations, optional-content layers, reference XObjects, embedded files without descriptions, and signature fields — each disclosed for human review rather than silently skipped (v1.92.0–v1.94.0)
Acrobat remediation guideAll PDF categories When a PDF category scores below "Pass", appends the exact Adobe Acrobat Accessibility Check rule names (called Full Check in older Acrobat versions), menu paths, and step-by-step fix instructions specific to that category. Word/PowerPoint/Excel findings link to the relevant Microsoft accessibility help article instead.

Categories That Don't Apply

A category that doesn't apply to a document (a text-only file has no images, tables, links, or forms) counts as passing and keeps its weight in the score's base: a document with no tables does not have a table problem. Only a category the tool could not assess — color contrast on PDFs, where rendered-page analysis is not yet implemented — sits outside the weighted score, because "we don't know" must not be scored as a pass. Through v1.58.2 the tool instead dropped non-applicable categories and renormalized the remaining weights; that magnified a simple document's single fault (a one-page notice scored worse than a longer agenda carrying the identical missing-title defect) and was removed.

Counting a non-applicable category as passing does not make the remaining categories less important. A high score can still coexist with unresolved semantic issues that matter for ADA/WCAG/IITAA review. For Illinois agency publication decisions, the score is a prioritization aid, not a substitute for the per-category findings.

A worked example makes it concrete:

Example: a 12-page PDF report — no tables, no links, no form fields

  Category               Weight   Score
  ─────────────────────────────────────
  Text Extractability      20%     100
  Title & Language         15%      75
  Heading Structure        15%      55
  Alt Text                 15%      90
  Reading Order            10%     100
  Bookmarks                 5%     100
  Table Markup             10%     n/a → passing ┐
  Link Quality              5%     n/a → passing ├─ no such content =
  Form Accessibility        5%     n/a → passing ┘  no such problem (counts as 100)

  Weighted sum = (100×20 + 75×15 + 55×15 + 90×15 + 100×10 + 100×5 + 100×10 + 100×5 + 100×5) ÷ 100
               = 8800 ÷ 100
               = 88 before the severity cap

  Severity cap: the score may never outrank the worst open finding (Minor 89 ·
  Moderate 79 · Critical 69). The broken heading hierarchy here is a Moderate
  finding, so the final score is capped at 79 · grade C — and the report says
  which finding is holding it there.

WCAG 2.2 Alignment

This tool reports against WCAG 2.1 Level AA — the version IITAA 2.1 (§E205.4) and ADA Title II both require, and the version every verdict here names. WCAG 2.2 is a strict superset of it: it adds nine success criteria (six at Level A/AA) and removes one (4.1.1 Parsing, obsolete). The automated checks are the same either way — every machine-checkable criterion is one carried forward from 2.1. The criteria 2.2 adds are interactive and manual, so they are never reported as automated failures. On a document with interactive form fields the form-relevant ones (Target Size 2.5.8, Redundant Entry 3.3.7, Accessible Authentication 3.3.8) are listed as “not assessed — manual review”, each saying plainly that it sits beyond the standard your grade measures. The rest are described on the What’s new in WCAG 2.2 page.

For a plain-language manager summary, see how WCAG 2.2 differs from 2.1. IITAA 2.1 does not yet reference WCAG 2.2, so 2.2 conformance is optional/forward-looking; WCAG 2.1 AA remains the legal minimum.

Scanned Document Detection

A PDF is flagged as a scanned image when both conditions are true: PDF.js extracts fewer than 50 characters of text content (indicating no real text layer) and QPDF finds no StructTreeRoot (indicating no semantic tags). This combination means the document is an unremediated scanned image that screen readers cannot access at all.

PDF Auto-Remediation: Pipeline Overview

As of v1.18.0, the tool also exposes an optional PDF auto-remediation feature behind the REMEDIATION_ENABLED=true env flag. When enabled, the audit results page surfaces an Attempt remediation button further down the results page. Clicking it spawns a detached worker that runs a four-stage pipeline, validates the output, and either serves the remediated file to the user (single-use download, deleted on stream close) or rejects it and surfaces a fallback message. The user re-uploads to remediate; no PDF is cached between the audit and remediation stages.

POST /api/remediate (multipart PDF)
  → [magic-byte check] → [page count cap (500)] → [pre-flight audit]
  → [job row created, sha256 content_hash recorded]
  → [spawn detached child: tsx src/jobs/remediate.ts <jobId>]
  ◄ { jobId, downloadToken }   (HTTP 202)

Worker pipeline:
  [Stage 1: preparing]   qpdf --object-streams=disable input → normalized
  [Stage 2: tagging]     OpenDataLoader convert(normalized) → tagged-pdf
  [Stage 3: validating]  qpdf --check tagged → validity verdict
                         verapdf --flavour ua1 --format json tagged → conformance verdict
  [Stage 4: comparing]   re-audit tagged → output_audit
                         guard: reject if Overall|Strict regresses

Output finalized OR job marked failed. Scratch dir wiped in `finally`.
Remediation pipeline — visual flow
Remediation pipeline — visual flow

The user re-uploads the PDF. qpdf normalizes it; original deleted with verification. OpenDataLoader adds structure tags; normalized intermediate deleted with verification. qpdf check + veraPDF validate the output. A re-audit confirms no score profile regressed. If all clear, output is held for 30 minutes; user downloads via single-use token; output deleted with verification.

Why Auditing Is Easy and Remediation Is Hard

Auditing any supported document is a read-only operation: walk the file's internal structure and report what you find. For a PDF, that means asking "does it have a tagged StructTreeRoot? Are figures marked? Is the language declared?" — the PDF specification (ISO 32000-2) is unambiguous about how to read these structures, and the libraries that parse them (qpdf, pdfjs, veraPDF) are mature and battle-tested. Word, PowerPoint, and Excel files are audited the same read-only way, walking their OOXML parts instead. Either way, the answers don't change between runs: a document can be audited a thousand times and produce the same result every time.

Remediation is a read-modify-write operation, and PDFs make that uniquely hard for several reasons that are baked into the format itself:

  1. PDF was designed for fixed-layout printing, not semantic content. Adobe published it in 1993 to make "documents that look identical on every printer." The accessibility layer (StructTreeRoot, marked content, role mapping) was bolted on in PDF 1.4 (2001) and is optional — valid PDFs can have none of it. Auto-tagging means reverse-engineering semantic meaning from raw visual presentation, which is much harder than reading existing semantic markers.
  2. There is no canonical mapping from visual layout to semantic role. Is a 14-pt bold line of text an <H2> or just emphasized body text? Is a 100×100-pixel image content (needs alt text) or decoration (mark as /Artifact)? A human reader judges from context; software guesses heuristically and is wrong some of the time.
  3. The content stream and the structure tree are coupled but separable. Every glyph and image in a PDF lives in a per-page content stream. Each one is wrapped in a "marked content" section (/MCID 7 … /EMC) that links it back to a node in the StructTreeRoot. Adding an alt-text to one image means mutating both sides coherently — write the new /Alt property on the Figure structure element AND ensure the MCID linkage stays valid. Many PDF libraries handle reading one side or the other, but not modifying both at once.
  4. The content layer can be in any of several representations. A scanned PDF has no text layer — it's just raster images, requiring OCR before any semantic remediation can happen. An optimized PDF compresses objects into "object streams" (a PDF 1.5+ feature) that some libraries can't safely round-trip. An encrypted PDF requires a password even to read. Each case is its own engineering minefield, and they layer onto each other (scanned-and-encrypted is worse than either alone).
  5. No single PDF library does everything well.pdf-lib (JavaScript, in the Node ecosystem) reads and writes metadata easily but has no StructTreeRoot builder. Apache PDFBox (Java) has the deepest structure-tree support but is Java-only. Ghostscript can rewrite PDFs but silently degrades tag structure. OpenDataLoader (Java, used here) is the only open-source tool that produces a tagged PDF from an untagged one — and even it cannot judge whether the result is meaningful.
  6. The "tagged PDF" specification is permissive. You can produce a PDF that satisfies all the technical requirements of being tagged (MarkInfo=true, StructTreeRoot exists, every page has marked content) and is still inaccessible to screen readers (e.g., every paragraph wrapped in a single <P> with no heading structure). PDF/UA-1 (ISO 14289-1) narrows this somewhat but doesn't eliminate it. Automated remediation tools often produce tagged-but-shallow output that machine validators accept but assistive technology can't navigate.
  7. Mistakes compound badly. A wrong heading level might confuse a screen reader user. A corrupted cross-reference (xref) table makes the entire PDF unreadable by any viewer. Remediation tools have to be conservative — when in doubt, don't touch. The qpdf preprocessing step in this pipeline exists precisely because OpenDataLoader's PDF writer occasionally corrupts the xref on round-trip with certain inputs (the InDesign 18.x / Word 365 case described above); we accept the cost of an extra normalization pass to avoid serving a damaged file.
  8. Round-trip fidelity is the highest bar. Remediation must add semantic markup while preserving every visual nuance: embedded fonts, raster + vector images, color spaces, ICC profiles, page labels, bookmarks, hyperlinks, form fields, digital signatures, embedded multimedia. The user doesn't want their report to look different after remediation; they want the same document with structure added. Read-modify-write while changing only the semantic layer is a class of problem the format simply wasn't designed to make easy.

The result is that PDF auto-remediation works well for the machine-checkable parts of accessibility (structure presence, metadata, language declaration, tagged content stream) and falls back to human judgment for the semantically-judged parts (alt-text quality, reading-order intent, decorative vs. informative classification). The roadmap for this tool (see docs/archive/pdf-remediation-alt-text-walkthrough-spec.md) is an interactive walkthrough that augments the machine-checkable foundation with human-authored alt text — without any AI in the loop, because the regulatory durability of agency-authored content is higher than the durability of AI-generated content.

Why OpenDataLoader Changes the Cost Equation

Until 2024–2025, programmatically tagging a PDF (auto-generating StructTreeRoot, marking figures, tables, headings) was something only a handful of commercial vendors could do, and they priced accordingly. The economics of PDF accessibility have historically been brutal for state agencies: PDF/UA expertise is rare, specialized, and was locked behind commercial walls for decades.

Commercial PDF remediation, today:

  • •Apryse / PDFTron SDK: enterprise-quoted, typically $1,500/yr minimum for the entry SDK and considerably more for the auto-tagging add-on. On-prem deployable but you pay for the privilege of running their Java/C++ binary in your own data center.
  • •Adobe PDF Services API: Accessibility Auto-Tag endpoint, free tier of 500 transactions per month (about 50 pages — exhausted by a single annual report). Beyond the free tier: enterprise-quoted, scaling per-document. Your PDF leaves your network for the API call.
  • •PDFix SDK, AbleDocs ADapi, CommonLook API: all enterprise-quoted, all opaque pricing, all aimed at large organizations.
  • •Manual remediation services: $5–$50 per page for hand-remediation of tagged-and-reviewed output. A typical 50-page agency report costs $250–$2,500 to remediate this way, and that's per document. State agencies producing dozens of reports per year face annual remediation bills in the tens of thousands.

Why so expensive? The skill is rare — there are relatively few practitioners who can read a structure tree and judge whether it's correct. The labor is real — even with good tooling, a 50-page report can require 4–8 hours of expert work. The market is small, the demand is regulated (ADA Title II, IITAA, Section 508), and the buyers are mostly governments and large organizations that aren't price-sensitive. The result is a niche industry with high prices and slow innovation.

OpenDataLoader PDF, released as Apache 2.0 in 2024 and continuously developed since, is the first credible open-source PDF auto-tagger. It does what previously required a $1,500/year SDK subscription: takes an untagged PDF and produces a tagged one. It's developed by Hancom (a Korean office-software vendor with deep PDF expertise) in collaboration with the PDF Association and Dual Lab (the same people behind the veraPDF validator). It ranks #1 overall (0.907) in 2026 PDF-extraction accuracy benchmarks — not just "as good as the commercial tools," better than them on the published metrics.

For this tool, OpenDataLoader is load-bearing. The pipeline architecture (qpdf preprocess → ODL tag → veraPDF check → re-audit) takes the most expensive part of commercial PDF remediation — the auto-tagging step — and replaces it with an apt install openjdk-17-jre-headless. The other open-source tools we pair it with (qpdf for preprocessing, veraPDF for PDF/UA-1 conformance validation) are also free and mature. Together they form a complete pipeline that until very recently did not exist in open source.

What ODL doesn't do — and no auto-tagger does — is judge whether the resulting structure is meaningful. It can mark every image as a Figure but can't write an alt-text. It can mark every table cell but can't decide which row is the header. Those remain human judgment calls. The economic shift ODL enables is from "$1,500/year + per-document manual labor" to "$0 of software + the manual labor for the parts a machine genuinely cannot do." That's an order-of-magnitude cost reduction for the agencies it serves, with no loss of output quality.

Tool 3: OpenDataLoader PDF (Auto-Tagging)

OpenDataLoader PDF (ODL) is an Apache-2.0-licensed Java application that takes an untagged PDF and writes a Tagged PDF with a populated StructTreeRoot. It is the first open-source tool to offer this transformation; it ranks #1 overall (0.907) in 2026 PDF-extraction benchmarks across reading order, table extraction, and heading detection. ICJIA maintains a fork at ICJIA/opendataloader-pdf as a hedge against future license changes upstream.

  • Invocation:@opendataloader/pdf v2.4.3 npm wrapper around a bundled JAR (lib/opendataloader-pdf-cli.jar).
  • Runtime: OpenJDK 17+ (java -version ≥ 11 required; install via apt install openjdk-17-jre-headless on Ubuntu 22.x).
  • JVM heap cap:JAVA_TOOL_OPTIONS=-Xmx768m set per-invocation by the worker as a safety rail against pathological documents.
  • Convert options used:{ outputDir, format: 'tagged-pdf', quiet: true }. Hybrid mode (docling-fast + SmolVLM) is deliberately not used in v1 — see the spike report for why.
  • Wall-clock timeout:REMEDIATION.WORKER_TIMEOUT_MS (5 min default); the JVM child is killed on overrun.

Why a Java tool in a Node.js codebase: every other open-source PDF/UA-targeted auto-tagger is either commercial (Apryse, Adobe PDF Services API), Java-only, or both. The tradeoff is one additional system dependency (JRE) on the deploy box in exchange for free, locally-hosted auto-tagging with no outbound API calls.

qpdf Preprocessing: --object-streams=disable

Stage 1 of the remediation pipeline pipes the input through qpdf --object-streams=disable INPUT NORMALIZED before ODL ever touches it. This decompresses PDF 1.5+ compressed object streams to traditional uncompressed objects. Without this preprocessing, ODL's Java PDF writer corrupts the output xref table on certain inputs — specifically, tagged PDFs emitted by modern Adobe InDesign (18.x) and Microsoft Word 365.

This bug was discovered during the OpenDataLoader feasibility spike on the FY_22_ICJIA_Annual_Report (InDesign 18.2) and 2022 SFS Process Evaluation Report (Word 365) fixtures. Without preprocessing, ODL emits a PDF that qpdf --check reports as damaged: xref num N not found, Invalid object stream, Catalog object is wrong type (null). With preprocessing, both PDFs round-trip cleanly and the score moves from F to D-grade improvement. See docs/archive/spike-remediation-results.md for the full reproducer + results.

Output Validation: qpdf --check + veraPDF

Every remediated PDF passes through two independent validators before the worker is allowed to serve it. The output is rejected (job marked failed, file deleted) on any failure, even though the upstream pipeline succeeded.

  • qpdf --check <output>: parses the entire PDF structure and reports warnings on damaged xref tables, malformed object streams, broken catalogs, etc. The worker treats "operation succeeded with warnings" as a failure — better to discard a borderline file than serve a damaged one.
  • verapdf --flavour ua1 --format json <output>: runs the veraPDF open-source PDF/UA-1 conformance validator (from the PDF Association + Dual Lab). Configured via REMEDIATION_VERAPDF_PATH; optional — when not configured, the receipt records verapdf_unavailable and skips this step; when the run itself fails (a timeout, no output) it records verapdf_error and stores no verdict at all — never a failure. veraPDF's verdict is informational, not blocking: even a PDF that veraPDF flags as non-conformant is still served if the audit score didn't regress. The result page surfaces this honestly in the IITAA compliance disclaimer.

Regression Guards

After successful tagging + validation, the worker re-audits the output and compares against the pre-flight audit stored at job creation time. Two independent comparisons run:

if (output.overallScore < input.overallScore ||
    output.scoreProfiles.strict.overallScore < input.scoreProfiles.strict.overallScore) {
  recordEvent(jobId, 'validation_failed', { regressed_profiles: [...] })
  await deleteAndVerify(jobId, taggedPath, 'cleanup')
  setFailed(jobId, `auto-remediation regressed: ${regressed.join(', ')}`)
  return
}

Why both: the headline overall score can fall back to a stored value while the strict profile is the canonical scoring methodology, so either could mask a regression in the other. Checking both ensures the user never sees a metric that decreased. The validation_failed event payload records the input/output deltas plus the regressed_profiles array, so any auditor query can identify exactly which comparison failed and by how much.

Lifecycle Audit Trail: remediation_events

Every remediation produces an append-only series of timestamped lifecycle events in the remediation_events SQLite table (apps/api/data/audit.db). The table is the canonical source for the receipt displayed on the result page, the auditor evidence trail, and any future compliance reporting. PDF content is never stored — only structural metadata.

CREATE TABLE remediation_events (
  id          INTEGER PRIMARY KEY AUTOINCREMENT,
  job_id      TEXT NOT NULL,
  event       TEXT NOT NULL,
  occurred_at INTEGER NOT NULL,
  details     TEXT,   -- JSON, content-free metadata only
  FOREIGN KEY (job_id) REFERENCES remediation_jobs(id)
);

Event vocabulary (closed set, typed at compile time):

  • received
  • processing_started
  • normalize_complete
  • input_deleted
  • tagging_complete
  • intermediate_deleted
  • validation_passed
  • validation_failed
  • verapdf_passed
  • verapdf_failed
  • verapdf_error
  • verapdf_unavailable
  • output_ready
  • downloaded
  • output_deleted
  • verified_absent
  • verify_failed
  • expired
  • error

The verified_absent event is the critical compliance signal. It is emitted after the worker (or cleanup sweep, or download handler) calls fs.unlink on a job artifact AND fs.stat returns ENOENT. The details payload contains a SHA-256 hash of the deleted path string (not file content) so auditors can reconcile event entries against expected paths without storing the paths themselves in the log.

Privacy & Retention (Remediation-Specific)

The remediation pipeline (PDF-only) maintains the same posture as the audit pipeline — no uploaded file content is persisted — with three additional rules:

  1. No between-stage cache. The just-audited PDF is not cached on disk waiting for the user to click Remediate. Clicking Remediate prompts a re-upload. UX cost: one extra upload. Privacy cost of caching: declined.
  2. Inputs deleted between pipeline stages. The worker writes work/input.pdf, normalizes it to work/normalized.pdf, then deleteAndVerify(work/input.pdf). Once ODL produces work/odl/<name>_tagged.pdf, the normalized intermediate is deleted. At any moment, at most one copy of the PDF exists on disk per job. The entire scratch dir is wiped in a finally block regardless of pipeline outcome.
  3. Output deleted on first download.GET /api/remediate/:id/download streams via createReadStream + pipe (no memory buffering); the response 'close' handler triggers deleteAndVerify(outputPath, 'download'). The job row is marked status='expired' before the stream begins, so a concurrent second download request sees 410 Gone. Files not downloaded within REMEDIATION.OUTPUT_TTL_MS (30 min default) are deleted by the cleanup sweep.

Filesystem permissions are 0700 on apps/api/data/remediation/ and 0600 on output files. Output filenames are <jobId>.pdf where jobId is a UUIDv4 (122 bits of entropy) — not derivable from the user's input. The remediation_events rows are retained per REMEDIATION.EVENT_LOG_RETENTION_DAYS (7 years default — typical state-agency records-retention schedule); the remediation_jobs row is purged separately at REMEDIATION.JOB_ROW_RETENTION_DAYS (30 days default).

Deploy Topology (Ubuntu 22.04 + PM2 + Nginx + DigitalOcean)

The API spawns the worker via spawn(process.execPath, ['--import', 'tsx', WORKER_PATH, jobId], { detached: true, stdio: 'ignore' }).unref(). PM2 does not manage the worker — it's a transient child of the API process, killed by the OS when the pipeline completes or crashes. Worker stdout is suppressed; all signals flow through the database (remediation_jobs.status, progress_pct, step) which the frontend polls via GET /api/remediate/:id/status once per second (backing off toward 8 s if requests fail).

System packages required on the deploy box:qpdf ≥ 10.x, openjdk-17-jre-headless, and (optional) the veraPDF CLI from verapdf.org. The rebuild.sh preflight verifies all three on every deploy and emits warnings if any are missing or below required version. The feature flag REMEDIATION_ENABLED is forwarded from the parent shell through ecosystem.config.cjs's env block, so the deploy idiom is:

sudo apt install -y openjdk-17-jre-headless   # one-time
echo 'REMEDIATION_ENABLED=true' | sudo tee -a /etc/environment
source /etc/environment
./rebuild.sh                                  # pulls, builds, pm2 restart

# Rollback to audit-only without redeploying:
sudo sed -i '/^REMEDIATION_ENABLED=/d' /etc/environment
pm2 restart ecosystem.config.cjs

Limitations & What This Tool Cannot Do

This tool provides a thorough automated assessment, but no automated tool can fully replace manual accessibility testing. Important limitations:

  • 1.Alt text quality: For PDFs, the tool detects whether alt text exists and runs a heuristic check for obviously poor alt text (hex-encoded strings, filenames like "IMG_001.jpg", generic placeholders like "image", and long strings without spaces); for Word, PowerPoint, and Excel, alt text is currently checked for presence only, without that heuristic quality pre-filter. In neither case can the tool evaluate whether alt text is semantically meaningful — for example, "a chart" technically passes all automated checks, but "Bar chart showing 2024 crime rates by county" is far more useful. Human review is still needed to assess alt text quality beyond the heuristic flags.
  • 2.Color contrast (PDF only): PDF color contrast analysis requires rendering each page as an image and analyzing pixel colors. This tool focuses on structural accessibility (tags, metadata, markup) and does not currently assess color contrast for PDFs. Word, PowerPoint, and Excel are the exception: their colors live in the document XML (explicit values, theme references, and Excel's legacy indexed palette are all resolved), so contrast is machine-checked for those three formats.
  • 3.Natural language clarity: The tool cannot evaluate whether the text itself is written clearly. WCAG 3.1.5 recommends content be written at a lower secondary education reading level — this requires human judgment.
  • 4.Decorative images: Not all images need alt text — decorative images should be marked as artifacts. For PDFs, the tool cannot distinguish informative images from decorative ones; it reports all images without alt text as a potential issue. Word, PowerPoint, and Excel are the exception: the tool reads each format's own "mark as decorative" flag and excludes those images from the alt-text check.
  • 5.Complex layouts: While reading order is assessed via MCID sequence analysis, extremely complex layouts (e.g., multi-column magazine spreads, nested pull quotes) may have subtle ordering issues that the 20% disorder threshold doesn't catch.

For a complete accessibility evaluation, this tool's automated analysis should be supplemented with manual testing using an actual screen reader (e.g., NVDA, JAWS, or VoiceOver) and the Adobe Acrobat Accessibility Checker for PDFs, or the Microsoft Office Accessibility Checker for Word, PowerPoint, and Excel documents.

These limitations apply to auto-remediation too. When the optional auto-remediation feature runs, OpenDataLoader can add a /Figure structure element for an image — but it cannot author a meaningful description. The same human-judgment gap applies to color contrast, reading-order ambiguity in multi-column layouts, distinguishing decorative from informative images, and writing text at a clear reading level. Auto-remediation is genuinely helpful for the machine-checkable parts of accessibility (structure, metadata, language declaration); it is not a substitute for the human-judgment parts. The result page is explicit about this in the IITAA compliance disclaimer.

The Open-Source Toolchain at a Glance

The open-source toolchain: each tool, its job, license, and pipeline stage
ToolJobLicensePipeline
qpdfStructure parsing + PDF normalizationApache 2.0Audit + Remediation (PDF)
pdfjs-distText + metadata extractionApache 2.0Audit (PDF)
jszipUnzip the OOXML package (.docx / .pptx / .xlsx)MIT / GPLv3Audit (Office formats)
fast-xml-parserParse OOXML structure & contentMITAudit (Office formats)
OpenDataLoader PDFRule-based PDF auto-taggingApache 2.0Remediation (PDF)
veraPDF PDF/UA validation — ISO 14289-1, or ISO 14289-2 when a document declares PDF/UA-2 — plus its machine-testable WCAG 2.2 profile (v1.97.0) MPL 2.0Audit + Remediation (PDF)

Privacy & Security

The application is hosted on DigitalOcean cloud infrastructure (managed via Laravel Forge). When you upload a file:

  • 1.A PDF is written to a temporary directory on the server, analyzed by QPDF, PDF.js, and veraPDF, and immediately deleted. A Word, PowerPoint, or Excel file never touches disk — it's held in server memory and parsed by JSZip + fast-xml-parser in a short-lived child process — the bytes cross over an in-memory channel, never a temp file. Either way, no file content is retained after analysis completes.
  • 2.The file exists in server memory for the duration of analysis (typically under 10 seconds); the qpdf analyzer briefly works from a randomly named temp copy that is deleted in the same request.
  • 3.No PDF data is transmitted to external APIs, cloud services, or AI models — all analysis runs on the server itself.
  • 4.Encrypted (password-protected) PDFs are rejected with a clear error before analysis begins.
  • 5.A concurrency semaphore limits the server to two simultaneous analyses to prevent resource exhaustion.

Shared reports: When you click "Share Report," the analysis results only — scores, category findings, grade, metadata (title, author, page count) — are saved to a SQLite database file on the same DigitalOcean droplet. Specifically:

  • •The original uploaded file is never saved — only the structured audit results (JSON) are stored.
  • •Shared links expire after 365 days. After expiration, the stored results are eligible for permanent deletion. The 365-day window is sized for the auditor / fleet inventory use case — fleet reports run on a multi-month cadence and reviewers need report links to stay valid for at least a year. Older results are deleted by the periodic cleanup sweep.
  • •Anyone with the link can view the report. There are no accounts on this tool at all — nothing to sign in to, for sharing or anything else.
  • •The database is stored locally on the server filesystem — it is not replicated to any external storage or cloud backup service. It is snapshotted nightly on that same machine (v1.49.0+), keeping the 5 newest snapshots, so a disk failure does not erase the audit record. A snapshot copies this database and nothing else: audit records, not audited files. Because no uploaded document is ever written to disk, no backup can contain one — why that isn't a contradiction.

When auto-remediation is enabled (the optional v1.18.0 feature behind REMEDIATION_ENABLED=true), the file lifecycle differs from a plain audit. The remediation worker needs the PDF on disk briefly to run external tools (qpdf, OpenDataLoader, veraPDF). The posture remains "as short-lived as the work requires, then deleted with verification":

  • •No between-stage cache. A PDF is never stored on disk waiting for the user to click "Remediate" after an audit. Clicking the button prompts a fresh multipart upload — the just-audited buffer is not preserved server-side.
  • •Inputs deleted between pipeline stages. After qpdf normalizes the uploaded file, the original input is deleted. After OpenDataLoader produces the tagged output, the normalized intermediate is deleted. At any moment, at most one copy of the PDF exists on disk per job. The entire scratch directory is wiped in a finally block regardless of pipeline outcome (including crashes).
  • •Output deleted on first download. The remediated PDF is served via a single-use download token. The file is deleted as soon as the response stream closes, and an fs.stat call verifies the deletion succeeded (the verified_absent event in the audit log is the auditor evidence). Concurrent or repeat download attempts return 410 Gone.
  • •Maximum 30-minute output retention. If the user never downloads, a cleanup sweep removes the file after REMEDIATION.OUTPUT_TTL_MS (default 30 minutes) and marks the job status='expired'.
  • •Lifecycle events contain no PDF content. Each step (received, normalize_complete, tagging_complete, validation_passed, output_ready, downloaded, output_deleted, verified_absent, etc.) writes a row to remediation_events with a server-side timestamp and a JSON payload of structural metadata only. File paths are recorded as SHA-256 hashes rather than literal strings.
  • •No external API calls. The remediation pipeline runs entirely on this server. OpenDataLoader and veraPDF execute locally; the file never leaves the droplet. AI-based alt text generation (which would call a hosted vision API) is explicitly not used in v1 — see the docs/archive/pdf-remediation-alt-text-walkthrough-spec.md roadmap document for the AI-free Phase 1 approach.
  • •Per-user concurrency limit. Each user can have at most one remediation job in flight at a time (REMEDIATION.MAX_CONCURRENT_JOBS_PER_USER). The 50 MB file-size cap, 500-page count cap, 5-minute wall-clock timeout, and 768 MB JVM heap cap are additional resource-exhaustion guards.

The Matterhorn Checklist: The 31 Checkpoints of PDF Accessibility — and Who Checks Each One

When you check a PDF here, how do you know the checker itself is any good? The PDF industry answers that with a published master checklist of everything a PDF accessibility checker should test. Here is that whole checklist — and who checks each item on it, in plain terms.

Why is this checklist only about PDFs?

The Matterhorn Protocol was written for exactly one file format: PDF. It is the test model for PDF/UA (ISO 14289), the accessibility standard for PDF files specifically, and its 31 checkpoints test machinery that only exists inside a PDF — the hidden tag tree, artifact marking, bookmark outlines, font embedding. Word, PowerPoint, and Excel files are built on a completely different internal format (Office Open XML), so these checkpoints have no meaning there.

Office files can absolutely still be checked here. Drop a Word (.docx), PowerPoint (.pptx), or Excel (.xlsx) file on the same tool and it gets its own audit covering the same accessibility ground — image descriptions, heading structure, table setup — tested the way those formats actually store it, with every per-format check listed on the technical details page. Only this Matterhorn checklist and the veraPDF panels are PDF-specific.

The protocol itself is published free by the PDF Association — read the Matterhorn Protocol specification if you want the full test model behind this list.

WCAG and IITAA are the law — so where does Matterhorn fit?

The legal chain, in one breath: ADA Title II (federal, in effect compliance due April 26, 2027 for entities of 50,000 or more, April 26, 2028 for smaller ones and special districts) and Illinois' IITAA 2.1 make digital accessibility a legal obligation for public bodies like ICJIA — and both name WCAG 2.1 AA as the standard to meet.

But WCAG was written mainly for web pages. It says what must be true of any content — text alternatives for images, a correct reading order, real headings, sufficient contrast — not how those requirements look inside a PDF, whose internals are nothing like a web page's.

Matterhorn is the PDF world's translation. PDF/UA (ISO 14289) is the technical standard for an accessible PDF, and the Matterhorn Protocol is its published test model: the 31 checkpoints below are the concrete, checkable form those same WCAG-style requirements take inside a PDF file.

Why you should care: when your PDF is evaluated — by an auditor, a records officer, a remediation vendor, or the professional checkers agencies use (PAC, Acrobat, veraPDF) — these checkpoints are what those tools test. Fixing your report's findings moves both needles at once: the WCAG/IITAA obligation the law names, and the PDF-specific checks evaluators actually run.

To be precise: the law requires WCAG — a PDF does not need a PDF/UA badge to be lawful. But the two overlap heavily by design, and Matterhorn is how PDF accessibility gets tested in practice.

Why "Matterhorn"?

It's the PDF industry's official test model for accessible PDFs, published by the PDF Association — the group that stewards the PDF format itself — and named after the famous Alpine mountain. Its 31 checkpoints work like the marked stops on a climbing route: clear them all and a PDF meets the formal accessibility standard (called PDF/UA). Professional checkers, like the PAC tool many agencies use, are built on this same list.

What is veraPDF?

A free, industry-standard PDF checker, created with the PDF Association — veraPDF is an independent second opinion, and it runs automatically alongside our own audit engine on every PDF you check here — once against the PDF/UA standard, and once against its machine-testable WCAG 2.2 rules (v1.97.0). You don't install or run anything: both results simply appear in your report. And if a check ever cannot run, your report says "Did not run" rather than staying silent.

Does this change my score?

No. Your score comes from the WCAG-based categories in your report, and the action plan there is still the thing to follow. These checkpoints cover the same ground — headings, image descriptions, tables, language — so fixing your report's findings improves both. The veraPDF result appears as its own informational panel on PDF reports; it informs, it never grades.

The protocol defines 31 checkpoints, made up of 136 specific ways a PDF can fail. Software can verify 87 of them; 47 need human judgment; 2 have no defined test. No tool anywhere can automate the human part — which is why every report here includes a manual-review card listing what still needs a person's eyes.

  • Audit enginechecked by this tool's own analyzer on every PDF audit (11 checkpoints)
  • Engine + veraPDFour analyzer catches the common problems; veraPDF covers the rest of what software can check (14)
  • Human reviewno software anywhere can judge these — your report's manual-review card lists them for human eyes (6)
  1. Real content tagged Engine + veraPDF

    Every piece of real content is in the tag tree; decorations are artifacts.

  2. Role Mapping Engine + veraPDF

    Custom tag names map to standard PDF structure types.

  3. Flicker Human review

    Nothing flashes or flickers — a judgment no software can make.

  4. Color and Contrast Human review

    Color is never the only carrier of meaning; contrast is sufficient.

  5. Sound Human review

    Audio content carries text alternatives.

  6. Metadata Audit engine

    XMP metadata declares the document title and any PDF/UA claim.

  7. Dictionary Audit engine

    Viewer preferences show the document title, not the filename.

  8. OCR-generated content Engine + veraPDF

    Scanned pages carry recognized text, not just pictures of words.

  9. Appropriate Tags Engine + veraPDF

    Tags mean what they say — headings are headings, figures are figures.

  10. Character Mappings Engine + veraPDF

    Every glyph maps to real Unicode text a screen reader can speak.

  11. Declared Natural Language Audit engine

    The document declares its language; foreign passages are marked.

  12. Stretchable Characters Engine + veraPDF

    Glyphs assembled from pieces carry replacement text (ActualText).

  13. Graphics Audit engine

    Figures carry alternative text; decorative graphics are artifacts.

  14. Headings Audit engine

    Headings are tagged as headings, and a document keeps to one convention — either numbered levels or generic <H> — rather than mixing them.

  15. Tables Audit engine

    Tables declare header cells and associate data cells with them.

  16. Lists Audit engine

    Lists use real list structure (L, LI, LBody).

  17. Mathematical Expressions Engine + veraPDF

    Formulas are tagged and carry a readable alternative.

  18. Page Headers and Footers Engine + veraPDF

    Running headers and footers are artifacts, not repeated content.

  19. Notes and References Engine + veraPDF

    Footnotes and endnotes use Note tags with unique IDs.

  20. Optional Content Audit engine

    Layer configurations are named and never auto-switch content.

  21. Embedded Files Engine + veraPDF

    Attachments carry filenames and descriptions.

  22. Article Threads Human review

    Reading threads, where used, follow a sensible order.

  23. Digital Signatures Engine + veraPDF

    Signature fields are labeled and reachable.

  24. Non-Interactive Forms Human review

    Print-and-fill form layouts are marked as such.

  25. XFA Audit engine

    Dynamic XFA forms — invisible to assistive tech — are flagged.

  26. Security Audit engine

    Encryption settings permit assistive technology to read the document.

  27. Navigation Engine + veraPDF

    Bookmarks and outlines support moving through long documents.

  28. Annotations Engine + veraPDF

    Links and annotations are tagged so assistive tech can reach them.

  29. Actions Human review

    Scripted actions stay accessible.

  30. XObjects Engine + veraPDF

    No prohibited reference XObjects.

  31. Fonts Audit engine

    Fonts are embedded so text renders and extracts reliably.

One more honesty rule: if veraPDF ever cannot run, PDF reports say “PDF/UA-1 machine checks (veraPDF): Did not run” instead of hiding the panel — a missing check is never presented as a passing one. The Matterhorn Protocol applies to PDF files; Word, PowerPoint, and Excel audits follow the per-format checks on the technical details page.

What This Tool Does

Audit any PDF, Word (.docx), PowerPoint (.pptx), or Excel (.xlsx) document for WCAG 2.1 AA accessibility — and (optionally) auto-remediate PDFs — all on infrastructure you control, with no AI and no per-document fees.

9
WCAG categories audited

Each document scored across up to 9 categories, counting only WCAG 2.1 Level A/AA criteria — the standard ADA Title II and the Illinois IITAA require. Critical / Moderate / Minor severity per category so you know what to fix first, and an A–F letter grade capped by the worst issue found — so a document is never graded above its most serious problem.

F → A
Auto-remediation (optional)

Tag untagged PDFs in seconds with the qpdf → OpenDataLoader → veraPDF pipeline. Output never regresses any score profile, and manual review is still recommended for IITAA compliance.

PDF/UA
Standards aligned

WCAG 2.1 Level AA, ADA Title II (compliance due April 26, 2027 for larger entities, 2028 for smaller), Illinois IITAA 2.1, and PDF/UA-1 (ISO 14289-1) via veraPDF — with PDF/UA-2 (ISO 14289-2) validated when a document declares it. Full lifecycle audit trail with deletion verification for compliance reporting.

0
Files retained

Uploaded files exist on the server only as long as the pipeline requires. Audited files: in-memory, gone in seconds. Remediated outputs: deleted on first download or 30-minute TTL, then fs.stat-verified absent.

$0
No AI, no third-party APIs

Every audit step runs on this server. No data is sent to vision models, hosted AI services, or commercial PDF/Office SDKs. The toolchain (qpdf, pdfjs, OpenDataLoader, veraPDF) is entirely open source — no per-document fees, no SDK licensing. Page views are counted by ICJIA's own self-hosted, cookie-free Plausible — never document data.

100%
Open source

Every line of code is on GitHub — fork it, audit it, run it on your own infrastructure. Underlying tools use Apache 2.0 / MIT / MPL licenses. Designed for state agencies that need control over their accessibility pipeline.