Why scanned PDFs are so large
A five-page letter typed in a word processor and exported to PDF is usually under 200 KB. The same five pages, printed, signed and scanned back in, can easily be 15 MB — seventy-five times bigger for a document containing exactly the same words. Nothing has gone wrong; the two files are simply not the same kind of object. One contains text instructions and a font. The other contains five photographs of paper. Once you see a scan as a photo album rather than a document, every part of its behaviour becomes predictable: why it is enormous, why you cannot search it, why converting it to Word yields nothing, and why it compresses better than any other file you own.
Where the megabytes come from
Three scanner settings multiply against each other. Resolution comes first: an A4 page at 300 dots per inch is roughly 2480 by 3508 pixels, about 8.7 million of them. Push to 600 dpi and it becomes nearly 35 million — four times the data for detail no screen will show and no printer needs for ordinary text. Colour depth is second: full colour stores three values per pixel, greyscale one, and pure black-and-white a single bit. Scanning a black-ink-on-white-paper page in colour therefore triples the raw data before compression, and gives you nothing back except a faint beige tint where the paper was. Third is the encoding the scanner chooses. If it stores pages losslessly, every speck of dust and every fibre of the paper is preserved byte for byte, and the noise of the paper itself becomes a large part of your file.
Why the paper texture costs you
This last point surprises people. Compression works by finding repetition, and a clean field of white compresses to almost nothing. But a scanned white page is not white: it is thousands of slightly different off-white values, produced by paper grain, lamp variation and sensor noise. To a lossless compressor that field is unpredictable detail and must be preserved exactly, so a blank page costs real megabytes. Scanning in greyscale rather than colour helps, and scanning with the descreen or text mode that most drivers offer helps more, because those modes flatten near-white values to white before the page is ever encoded. If you have ever wondered why a scan of an empty sheet is bigger than a text document of forty pages, this is why.
The reason you cannot search or convert it
A scanned page contains no text layer at all. There are no glyphs and no character codes — only pixels that happen to be shaped like letters. Your reader's search function finds nothing, your copy button selects nothing, and a PDF-to-Word converter honestly produces an empty document. The only way to add text is optical character recognition, which examines the pixels and guesses characters. Modern OCR is good on clean, straight, printed pages and unreliable on handwriting, faint carbon copies, stamps, tables and anything photographed at an angle. It also fails silently: a confident wrong character looks exactly like a correct one, so OCR output on anything that matters is a draft to be proofread, not recovered text.
Fixing a file you have already been sent
You usually cannot re-scan someone else's document, so the practical fix is re-encoding. This is the case where compression shines, because a scan is pure image data and image data is precisely what a compressor can rework. In our own measurements a twelve-page scanned document dropped from 17.69 MB to 1.49 MB — 91.6 per cent — while staying perfectly legible on screen. That kind of result is normal for scans and unattainable for text documents, which have nothing comparable to give. Check the result at full zoom rather than as a thumbnail, and look specifically at the smallest marks on the page: a signature, a stamp, a footnote, a handwritten annotation. If those survive, the file is fine.
Getting it right at the scanner
When the scan is yours, the settings that matter are simple. Use 300 dpi for documents you may need to enlarge or OCR later, and 200 dpi for everyday paperwork; 600 dpi is for photographs and fine artwork, not for a signed invoice. Use greyscale unless the colour carries information — a coloured stamp, a highlighted clause, a chart. Let the driver save as JPEG-in-PDF rather than lossless if it offers the choice, and turn on any automatic deskew and blank-page removal features. A phone camera is a legitimate scanner in a hurry, but photograph the page flat, in even light, filling the frame, and use a scanning app that flattens perspective rather than the plain camera, because a page shot at an angle defeats OCR and looks careless.
When large is the right answer
Not every big scan is a mistake. Archival scanning of manuscripts, plans, medical imaging and anything with legal weight is deliberately done at high resolution in colour and kept losslessly, precisely because information that is thrown away now cannot be recovered later. The rule of thumb is to ask who reads the file and how long it must live. A scan being emailed to a colleague this afternoon should be small and readable. A scan of a title deed being stored for thirty years should be large and faithful, and the copy you email should be a compressed derivative of it rather than the archive itself. Keeping both, when the document warrants it, costs a little disk space and settles the question permanently.
A quick way to diagnose any file
If you are not sure which kind of PDF you have, two five-second tests answer it. Try to select a sentence with your cursor: if the whole page highlights as one block, it is an image and you have a scan. Then check the size against the page count — a document averaging more than about half a megabyte per page is image-heavy almost without exception. Those two checks tell you immediately whether compression will help dramatically, whether text extraction will work at all, and whether you should be asking the sender for the original digital file instead, which is very often the fastest fix of all.
Frequently asked questions
Why is my scanned PDF 20 MB when the document is only five pages?
Because each page is a photograph. At 600 dpi in colour a single A4 page can exceed 4 MB before any compression, so five pages reaching 20 MB is arithmetic, not a fault.
Does compressing a scan make it unreadable?
Not at sensible settings. Body text stays crisp because the compressor keeps enough resolution for the screen; what disappears is paper grain and colour noise you were never reading.
Can I make a scanned PDF searchable?
Only with OCR, which adds an invisible text layer over the images. It works well on clean printed pages and poorly on handwriting, faint copies and complex tables, so proofread anything important.
What scanner resolution should I use for documents?
200 dpi for everyday paperwork, 300 dpi if you plan to OCR or enlarge it. Reserve 600 dpi for photographs and detailed drawings.
Should I scan in colour or greyscale?
Greyscale unless colour carries meaning. It removes roughly two-thirds of the raw data and makes black text look cleaner, not worse.
Is a photo of a page as good as a scan?
For readability, often yes, if the page is flat and evenly lit and you use a scanning app that corrects perspective. For OCR and for anything official, a real scan is still noticeably more reliable.
Compress, convert and unlock PDFs in your own browser tab. PDF re-encodes the images that make a file huge, rebuilds a Word document from the text layer, and adds or removes passwords. A twelve-page scan we measured went from 17.69 MB to 1.49 MB. The file is processed in the tab — there is no upload.
Open the PDF tools