• PDF file compactness (was Re: reprints of old AI memos)

    From Kragen Javier Sitaker@kragen@canonical.org to alt.sys.pdp10,comp.text.pdf,comp.compression on Thu Sep 3 12:06:18 2026
    From Newsgroup: comp.text.pdf

    Lawrence DrCOOliveiro <ldo@nz.invalid> writes:
    On Wed, 02 Sep 2026 16:35:26 -0300, Kragen Javier Sitaker wrote:
    Contrary to popular belief, PDF is a relatively compact file format.

    Only if you apply compression to it.

    This is ambiguous; you might be intending to say, rCLonly if
    you use the PDF formatrCOs compression featuresrCY or, rCLonly if
    you compress the PDF file with an additional compressor
    after creating it.rCY

    I disagree with both of these interpretations.

    For text, PDF can be a relatively compact file format even
    if you do neither of these. ThererCOs some file-format
    overhead of about a kilobyterCerCorCeyou need a header, a
    catalog, a page tree, an xrefs table, and trailer, even for
    a one-page documentrCerCorCeand then you need to specify the font
    and the coordinates where your text starts. After that, I
    think the per-line overhead is about 4 bytes per line. I
    havenrCOt tested this content-stream example, but if itrCOs
    missing something, itrCOs not missing much:

    BT
    /F0 12 Tf 50 706 Td
    (For text, PDF is a relatively compact file format even if) '
    (you do neither of these. There's file format overhead of) '
    (about a kilobyte - you need a page tree and xrefs table even) '
    (for a one-page document - and then you need to specify the) '
    (font and the coordinates where your text starts. After) '
    (that, the per-line overhead is a few bytes per line. I) '
    (haven't tested this content-stream example, but if it's) '
    (missing something, it's not missing much:) '
    ET

    The `'` PDF content-stream operator is just `T* TJ`, if
    yourCOre familiar with those. ThererCOs also a `"` shortcut
    operator that lets you set a different letter spacing and
    word spacing for each line:

    1 1.4 (for a one-page document - and then you need to specify the) "

    That costs you about 11 bytes per line instead of 4.

    So, a simply formatted PDF file without using any kind of
    compression is about 5% bigger than a plain ASCII text file,
    plus about one kilobyte.

    Perhaps you don't consider 5% overhead to be rCLrelatively
    compactrCY, but I do.

    Also, of course, PDF *does* support Deflate compression for
    content streams; and, since PDF 1.5, it also supports it for
    xrefs and object streams, so a PDF file can easily be half
    the size of a plain ASCII text file, down to a minimal size
    of a few hundred bytes. Amusingly, IrCOve even seen PDF files
    that apply Paeth compression to the xrefs table.

    Now, in practice, PDF files are often not this compact. A
    much more typical example of a PDF content-stream (from the
    Derctuo PDF: <http://canonical.org/~kragen/derctuo/>) looks
    like this (slightly reformatted):

    1 0 0 1 0 0 cm BT /F1 12 Tf 14.4 TL ET
    0 0 0 rg
    0 0 0 rg
    BT 1 0 0 1 6 747.6 Tm .533333 0 0 rg
    /F2+0 24 Tf 28.8 TL (Derctuo) Tj T* ET
    0 0 0 rg
    BT 1 0 0 1 6 714.84 Tm /F2+0 12 Tf 14.4 TL ( ) Tj T* ET
    0 0 0 rg
    BT 1 0 0 1 6 700.44 Tm /F3+0 12 Tf 14.4 TL ( ) Tj
    /F4+0 12 Tf 14.4 TL (\200\201\201\200) Tj T* ET
    0 0 0 rg
    BT 1 0 0 1 6 686.04 Tm /F3+0 12 Tf 14.4 TL ( ) Tj
    (Kragen ) Tj (Javier ) Tj (Sitaker) Tj T* ET
    0 0 0 rg
    BT 1 0 0 1 6 671.64 Tm /F3+0 12 Tf 14.4 TL ( ) Tj
    (Buenos ) Tj (Aires) Tj T* ET
    0 0 0 rg
    BT 1 0 0 1 6 657.24 Tm /F3+0 12 Tf 14.4 TL ( ) Tj
    (December, ) Tj (02020) Tj T* ET
    0 0 0 rg
    BT 1 0 0 1 6 642.84 Tm /F3+0 12 Tf 14.4 TL ( ) Tj
    (Public ) Tj (domain ) Tj (work) Tj T* ET
    0 0 0 rg
    BT 1 0 0 1 6 628.44 Tm /F3+0 12 Tf 14.4 TL ( ) Tj
    /F4+0 12 Tf 14.4 TL (\200\201\201\200) Tj T* ET
    0 0 0 rg
    0 0 0 rg
    BT 1 0 0 1 6 599.64 Tm /F2+0 12 Tf 14.4 TL ( ) Tj T* ET
    0 0 0 rg
    BT 1 0 0 1 6 581.28 Tm /F2+0 12 Tf 14.4 TL ( ) Tj
    (Derctuo ) Tj (is ) Tj (a ) Tj (book ) Tj (of ) Tj
    (notes ) Tj (on ) Tj (various ) Tj (topics, )
    Tj (mostly ) Tj (science ) Tj (and ) Tj T* ET

    ItrCOs relatively straightforward to uncompress
    content-streams like this from PDF files in Python:

    zlib.decompress(base64.a85decode(a8.removesuffix(b'~>'))
    ).decode('utf-8')

    This is obviously inefficient in many different ways:

    - Reportlab decided to Ascii85Decode the compressed data for
    no real reason, even though I specified pageCompression=True.
    - ThererCOs no need to Tj each word separately. You could Tj
    the whole line of text. I think this was my fault; I
    hacked together this PDF renderer in a week for a deadline.
    - ItrCOs unnecessary to set the transformation matrix (cm) to
    the identity matrix. ThatrCOs the default.
    - Similarly, setting the RGB color to black for each line
    (and twice for the first line) is unnecessary. Black is
    the default. I think this is ReportLabrCOs fault.
    - In the one case where the color is set to a non-default
    color, itrCOs unnecessary to specify that color to six
    significant figures.
    - Displaying runs of spaces is generally unnecessary.
    - Changing fonts to display runs of spaces is extra
    unnecessary.
    - Changing fonts twice per line is unnecessary. Most of
    this text is in a single font.
    - Chanting to the same font again is also unnecessary.
    - Creating a new text object for every line (BT ET) is
    unnecessary and also counterproductive for copy-and-paste.

    Despite all this, the 47 lines of text on the page are 8202
    bytes of uncompressed content-stream; FlateDecode encoded
    and Ascii85Decode encoded, they pack down to only 2471
    bytes, plus 149 bytes of per-stream overhead (also mostly
    unnecessary), plus 309 bytes of the Page object containing
    the content stream (also mostly unnecessary), for a total of
    3K per page, which is slightly more compact than plain ASCII
    text. (This doesnrCOt count the hyperlinks on the page,
    though.)

    ItrCOs easy to see how small inefficiencies like these can
    pile up when people (like me) who donrCOt really understand
    what theyrCOre doing get things to more or less work, and then
    stop. And that seems to be how most PDF files are built.
    Most of them are even worse than the Derctuo PDF.

    Derctuo is far from exemplary, but the PDF is 986 pages and
    5.91 megabytes, roughly 5.9KiB per page; it divides up as
    follows:

    - bytes 569 to 2.23e6: intermixed page objects and link objects
    - bytes 2.23e6 to 2.50e6: embedded fonts, covering ASCII and
    a bunch of Unicode for things like math and Greek, in
    eight display styles
    - bytes 2.50e6 to 2.62e6: more document structure, including
    outline and page tree
    - bytes 2.62e6 to 5.74e6: page content streams
    - bytes 5.74e6 to 5.91e6: xrefs and trailer

    Due to the inept content-stream structure I demonstrated
    above, if the page content streams were uncompressed, they
    would be about 3.3|u as large, going from about 3.12
    megabytes to 10.3 megabytes. This would inflate the Derctuo
    PDF from 5.9 megabytes to 13.1 megabytes, which works out to
    about 13KiB per page.

    This is about three times bigger than plain ASCII text, but
    thatrCOs only because of how badly I screwed the pooch in
    building the content streams.

    I was mostly using Edward TufterCOs rCLET BookrCY TrueType version
    of Bembo, falling back to DejaVu Serif fonts for non-ASCII
    characters, and using Latin Modern Mono Light Condensed (a
    modified Computer Modern Typewriter) for typewriter text,
    falling back to FreeMono and DejaVu Sans Mono fonts.

    Embedding eight typefaces thus cost me 270K. If you want a
    PDF document to be much under 100K, you more or less have to
    restrict yourself to the 14 core PDF fonts instead of
    embedding your own, or hope that the fonts you want to use
    happen to be installed on the reader's system (prohibited in
    PDF/A and, I believe, PDF 2.0).

    Nearly half of the bytes in the Derctuo PDF are hyperlinks,
    which mostly look like this (IrCOve elided the CRs ReportLab
    inserted before LFs):

    % 'Annot.NUMBER2006': class LinkAnnotation
    2343 0 obj
    << /Border [ 0
    0
    .1 ]
    /C [ .6
    .6
    1 ]
    /Contents (notes/lithium-fuel.html)
    /Dest [ 2559 0 R
    /XYZ
    null
    null
    null ]
    /Rect [ 4.8
    595.44
    33.62578
    609.84 ]
    /Subtype /Link
    /Type /Annot >>
    endobj

    2559 0 obj is the /Page object for page 371, where the note
    on lithium fuel begins. I think there are a lot of
    opportunities for optimization here, including unnecessary
    whitespace, and I donrCOt think the PDF spec *requires* an
    /Annot to be a top-level object (I think you can embed it
    inside the /Page object), but honestly most of those would
    go away if you just used a PDF 1.5 deflated object stream.

    Kragen
    --- Synchronet 3.22a-Linux NewsLink 1.2