PDF← Guide

What is PDF/A, and do you need it?

A PDF is a set of instructions for drawing a page, and some of those instructions depend on things that are not inside the file. A font that lives on your computer. A colour profile your monitor supplies. A video, a JavaScript action, a link to a server that will not exist in ten years. Open such a file on a different machine now and it usually looks fine, because the missing pieces get substituted. Open it in 2050 and there may be nothing to substitute with. PDF/A is the answer to that problem: a restricted subset of PDF in which everything needed to reproduce the page exactly is inside the file, and everything that depends on the outside world is forbidden. It is not a different file format — a PDF/A file is a PDF and opens in any reader — it is a promise about what the file does not contain.

Open the PDF tools →Compress, convert and unlock PDFs in your own browser tab.

What the standard forbids

The rules follow from the goal. Fonts must be embedded, all of them, with no exceptions for the standard set that readers used to be assumed to have; a document that referred to Arial by name and let the reader find a copy is not archivable, because in fifty years the reader's substitute may have different metrics and reflow the line. Encryption is banned, since a document nobody can decrypt is not preserved. So is JavaScript, so are embedded audio and video, so are actions that launch external files or open network locations. Transparency is restricted in the older parts of the standard and permitted in the newer ones. Colour must be defined unambiguously, either through an embedded ICC profile or in a device-independent space, so that a specific red is a specific red rather than whatever the printer decides. And every file must carry XMP metadata identifying which flavour of PDF/A it claims to be.

The levels: A, B and U

Conformance comes in levels, and the letters mean different things from the version numbers. Level B, for basic, guarantees visual reproduction: the page will look right forever. Level A, for accessible, additionally requires a tagged structure — a machine-readable description of which text is a heading, which is a paragraph, which is a table cell and in what reading order — plus Unicode mappings for all text. Level U sits between them and requires the Unicode mappings without the full structure, meaning the text can be reliably searched and extracted even where the document is not fully tagged. Level A is genuinely difficult to produce from an arbitrary document because tagging is real editorial work, which is why most archives in practice ask for level B, and why an institution asking you for PDF/A-1a is asking for considerably more than one asking for PDF/A-1b.

Which version to use

PDF/A-1 came first and is the strictest: no transparency, no layers, no embedded files. PDF/A-2 relaxed those, allowing transparency, layers, JPEG 2000 images and the embedding of other PDF/A files inside a container. PDF/A-3 made one significant change by permitting arbitrary attachments of any type, which is how electronic invoicing formats work — a human-readable invoice page with the machine-readable data file attached inside it. PDF/A-4, based on the newer PDF 2.0, tidies the levels up considerably. Unless a recipient specifies otherwise, PDF/A-2b is the sensible default: modern enough not to fight your document's transparency, strict enough to be a genuine archive, and universally supported by validation tools.

Making one, and the scanned-document trap

Most office suites can export directly to PDF/A, usually as a checkbox in the export dialogue, and that is the cleanest route because the source document still knows what its headings and fonts are. Converting an existing PDF is also possible with dedicated tools, which embed missing fonts, convert colours and strip forbidden features. The trap is scans. A scanned page converted to PDF/A is perfectly valid at level B — it is a picture, it will look the same forever, the standard is satisfied. But it contains no text, so it cannot be searched, and level A is out of reach without OCR. Archiving a filing cabinet as image-only PDF/A produces a collection nobody can find anything in. If the archive is meant to be usable rather than merely preserved, OCR belongs in the workflow before conversion, and the OCR output needs proofreading.

Validate, do not assume

A file being exported with a PDF/A option ticked is not proof that it conforms, and this is where submissions to journals, courts and public archives most often bounce. Validators exist, veraPDF being the well-known open-source one, and they report precisely which clause a file violates. The common failures are dull and repeatable: a font that could not be embedded because its licence forbids embedding, a stray transparency group from a logo, a colour space with no profile, missing or malformed XMP metadata, or an annotation type the standard does not allow. Each has a specific fix. Running the validator before you submit turns a rejection weeks later into five minutes now.

When PDF/A is the wrong choice

It is an archival format, not a working one, and using it as a default has costs. It cannot be encrypted, so a confidential document cannot be a password-protected PDF/A — the two requirements are mutually exclusive and you have to protect the file some other way, through storage permissions rather than through the document. Embedding every font makes files larger, occasionally much larger. Interactive forms and multimedia are exactly what the standard removes. So the honest rule is: use PDF/A when a document must be readable long after the software that made it is gone, or when an institution requires it, and use ordinary PDF for everything you are still working on. Converting at the end, once, is much easier than living inside the restrictions the whole time.

Where compression fits

Archival and small are in tension, and it is worth being deliberate about which you want. Aggressive image compression throws away information permanently, which is precisely what an archive is supposed to prevent, so a preservation copy should be compressed conservatively or not at all. The practical arrangement is two files: a faithful PDF/A copy in the archive, and a compressed derivative for emailing and day-to-day reading, regenerated whenever it is needed. Keeping only the compressed version and calling it an archive is the mistake — it will still open in 2050, and it will still be missing the detail you threw away today.

Frequently asked questions

Is PDF/A a different file format?

No. A PDF/A file is a normal PDF that obeys extra restrictions and declares them in its metadata. Any reader opens it; only validators care about the difference.

Can a PDF/A file be password protected?

No. Encryption is forbidden, because a file nobody can decrypt is not preserved. Protect the storage location instead, or keep an encrypted non-archival copy separately.

Which PDF/A version should I choose?

PDF/A-2b unless someone specifies otherwise. It permits transparency and layers, which PDF/A-1 forbids, and level B guarantees appearance without demanding full accessibility tagging.

Does converting to PDF/A change how my document looks?

It should not, and that is the point. What changes is what travels with it — fonts get embedded, colours get profiles, forbidden features are removed. If the appearance shifts, a font could not be embedded and was substituted.

Is a scanned PDF/A file searchable?

Not by itself. It is a valid archive of images with no text layer. Add OCR before conversion if the archive needs to be searchable, and proofread the result.

How do I know a file really conforms?

Run a validator such as veraPDF. Exporting with the option ticked is not proof; validators name the exact clause that fails, which is usually an unembeddable font or missing metadata.

Compress, convert and unlock PDFs in your own browser tab. PDF re-encodes the images that make a file huge, rebuilds a Word document from the text layer, and adds or removes passwords. A twelve-page scan we measured went from 17.69 MB to 1.49 MB. The file is processed in the tab — there is no upload.

Open the PDF tools
Language: TR EN DE ES FR IT PT AR RU JA KO