Most people searching for a way to compress a PDF are dealing with a problem that was created earlier, at the scanner. A 40-page contract has no business being 80 MB. It is that large because it was scanned at 600 DPI in full colour when 300 DPI in greyscale would have been indistinguishable in use and roughly a twelfth of the size.
Compression can recover some of that. Scanning correctly in the first place avoids it entirely.
The arithmetic nobody mentions
DPI means dots per inch, measured along one edge. That single word "per inch" is doing a lot of work, because the page has two dimensions. Doubling DPI quadruples the pixel count.
An A4 page is 8.27 × 11.69 inches. So:
| Resolution | Pixel dimensions | Megapixels | Relative size | |---|---|---|---| | 150 DPI | 1,240 × 1,754 | 2.2 MP | 1× | | 300 DPI | 2,480 × 3,508 | 8.7 MP | 4× | | 600 DPI | 4,960 × 7,016 | 34.8 MP | 16× | | 1200 DPI | 9,921 × 14,031 | 139 MP | 64× |
Scanning at 600 instead of 300 does not make a file "a bit bigger". It makes each page four times larger. Across a 200-page document, that is the difference between a file you can email and one you cannot.
Colour depth multiplies on top of this. The same page can be stored as:
- Bitonal (1-bit) — pure black and white, no grey. Smallest by an enormous margin.
- Greyscale (8-bit) — 256 shades. Roughly 8× a bitonal scan.
- Colour (24-bit) — full colour. Roughly 24× a bitonal scan, 3× greyscale.
A 600 DPI colour scan of a page of black text is carrying about 384 times the data of a 150 DPI bitonal scan of the same page, to convey exactly the same words.
What resolution actually buys you
Higher DPI is not wasted — it just has a ceiling past which it stops helping.
150 DPI is enough to read on a screen. Text is legible, but the edges are visibly soft when zoomed and OCR accuracy degrades noticeably. Suitable for reference copies you will never print or process.
300 DPI is the practical standard, and the right default for almost everything. It is the minimum most OCR engines want for reliable recognition of normal body text, and it is the traditional threshold for acceptable print quality. If you scan everything at 300 and never think about it again, you will rarely be wrong.
600 DPI starts to pay off with small type — footnotes below about 8 point, dense tables, engineering drawings, or degraded originals where the characters are already breaking up. It also helps OCR on poor-quality source material, where the extra pixels give the recogniser more to work with.
1200 DPI and above is for reproduction and archival preservation of physical artefacts: photographs, artwork, historical documents where the material itself matters. For a page of text it is pure waste.
OCR has different requirements than reading
This trips people up. A scan that looks perfectly clear to you can still OCR badly, because OCR engines are sensitive to things your eye compensates for automatically.
The rule of thumb is that OCR wants roughly 20 pixels of height per character to work reliably. At 300 DPI, 10-point text lands comfortably above that. At 150 DPI, the same text falls near the edge and error rates climb.
But resolution is not the only factor, and often not the limiting one:
Contrast matters more than DPI. A crisp 300 DPI scan will out-perform a muddy 600 DPI one every time. If your originals are faint, adjusting the scanner's threshold or contrast helps far more than raising resolution.
Skew is expensive. Pages fed at an angle cost accuracy disproportionately, because character segmentation assumes roughly horizontal baselines. Most scanning software can deskew automatically — turn it on.
JPEG artefacts confuse recognisers. Heavy JPEG compression creates ringing around high-contrast edges, which is exactly where letterforms live. For text you intend to OCR, use lossless or light compression at scan time and compress afterwards if needed.
Bitonal can beat greyscale. For clean printed text, a well-thresholded bitonal scan often OCRs better than greyscale, because the engine does not have to make the black/white decision itself. For faint or handwritten material the opposite is true — greyscale preserves the information the threshold would have thrown away.
A practical decision table
| What you are scanning | Resolution | Colour mode | |---|---|---| | Printed text you will read on screen | 200–300 DPI | Bitonal or greyscale | | Contracts, records for OCR and archive | 300 DPI | Greyscale | | Documents with photographs or colour charts | 300 DPI | Colour | | Small print, footnotes, dense tables | 400–600 DPI | Greyscale | | Faded, damaged, or handwritten originals | 400–600 DPI | Greyscale | | Photographs and artwork | 600–1200 DPI | Colour | | Anything going to a government portal with a size cap | 200–300 DPI | Bitonal if text only |
Fixing what you already scanned
If the scans exist and rescanning is not realistic, you are compressing rather than preventing. That works, with limits.
Downsampling a 600 DPI scan to 300 recovers most of the size and loses little you were using — the information was captured, you are simply storing less of it than the scanner produced. Our Compress PDF tool does this, and on image-heavy scans a 60–85% reduction is typical.
What you cannot do is recover detail that was never captured. A 150 DPI scan of small print will not OCR well no matter what you run it through, because the pixels that would distinguish an e from a c do not exist in the file. If accurate text extraction matters and the original is still available, rescanning at 300 DPI is faster than fighting a bad scan.
Two other things worth doing before reaching for a compressor: delete blank pages left by duplex scanners, which can each carry hundreds of kilobytes of scanned paper texture, and crop wide margins, which are pure raster data conveying nothing.
The short version
Scan at 300 DPI in greyscale unless you have a specific reason not to. Raise it for small print and damaged originals; drop to bitonal for clean printed text you need to keep small. And remember that the extra pixels from a 600 DPI colour scan are usually not buying you anything except a file you cannot email.
