From Newsgroup: comp.lang.tcl
ANNOUNCE: tclpdf 1.2 released
=============================
tclpdf is a pure Tcl extension for creating PDF documents: pages and
graphics, embedded fonts in three formats - TrueType subset to the
glyphs actually used, OpenType with CFF outlines and Type 1 - plus Type
3 fonts drawn by the caller, JPEG, PNG and TIFF images, SVG drawn as
vectors, tables, layers, tagged and accessible documents up to PDF/UA-2, encryption, digital signatures, and electronic invoices as ZUGFeRD /
Factur-X in PDF/A-3B. It also reads: a page of an existing PDF is taken
over as a form, a finished file is continued by an incremental update,
and what a foreign file says about itself is answered as a dictionary. Developed by Alexander Schoepe, it requires Tcl 8.6.11+ and runs under
Tcl 9 as well; zlib is a built-in command rather than a package, tdom is needed only for the XMP metadata of a PDF/A, PDF/UA or ZUGFeRD document,
and Tk is not used.
Standards it is built on, and measured against: ISO 32000-1 (PDF 1.7)
and ISO 32000-2 (PDF 2.0) for the file itself; ISO 19005-2 and 19005-3
for PDF/A parts 2 and 3, levels B, U and A; ISO 14289-1 and 14289-2 for PDF/UA, with WTPDF for the well-tagged profile; EN 16931 with ZUGFeRD / Factur-X 2.3 and Order-X for electronic invoices and orders; ETSI EN 319
142-1 (PAdES) and RFC 5652 (CMS) for signatures, RFC 3161 for the
timestamps inside them; ISO/IEC 14496-22 (OpenType, the sfnt tables,
GSUB and GPOS) for fonts, beside it Adobe's own specifications for what OpenType does not cover - the Type 1 Font Format, Technical Note #5176
(CFF), #5902 (PostScript name generation), the Adobe Font Metrics of the fourteen standard faces and the Adobe Glyph List - with UAX #9 for right-to-left text; TIFF 6.0 with Technote 2, JFIF and Exif for
pictures. What claims a profile is checked against it on every run:
veraPDF for PDF/A and PDF/UA, Mustangproject for the invoices, qpdf and
pdfsig over every document.
How much of it is tested: 2598 tests across 59 test files, green under
Tcl 8.6.18 and Tcl 9.0.4 alike, and 63 example scripts writing 70
documents - every command, option and value the manual names has a test
of its own, and every example is run, validated and looked at.
1.2 follows 1.1 and closes the two halves the package did not have:
reading and protecting. A page of a foreign PDF is now imported as a
form, a finished file is continued rather than rewritten, a second
signature goes onto a signed document, the file is encrypted with the
revision PDF 2.0 prescribes - all of it in pure Tcl, with no
cryptographic library to build. Beside that: layers, Type 3 and colour
fonts, GPOS mark attachment, hyphenation from Liang patterns, TIFF, the resolution a picture states about itself, Lab colour, text rendering
modes and image masks. The API is unchanged and nothing was withdrawn;
two behaviours changed and are listed under "Changed since 1.1".
Two tech notes go with this release:
What it does, and what it deliberately does not:
https://fossil.sowaswie.de/tclpdf/technote/071ce10bf8
Why tclpdf exists, and what it shows about working with AI:
https://fossil.sowaswie.de/tclpdf/technote/d647175afa
Download / Repository:
https://fossil.sowaswie.de/tclpdf
Author: Alexander Schoepe <
alx.tcl@sowaswie.de> - bug reports, patches
and questions are welcome by mail as well as through the repository.
English
-------
New since 1.1:
* Importing a page: "pdf import" takes one page of an existing PDF over
as a form, and "form place" puts it down as often as wanted, scaled,
rotated, faded - a letterhead is taken over instead of being rebuilt.
Both cross-reference flavours are read, the classic table and
cross-reference streams with object streams, hybrid files included and incremental updates followed over /Prev; inheritable page attributes
come from the page tree, CropBox and /Rotate become the form's box and
matrix, and everything the page's resources reach is copied byte for
byte with its own filters. Only the content streams are decoded, so an
exotic image filter is no obstacle; LZW is among the filters read, /EarlyChange with it.
* Continuing a file: "::tclpdf::update open" appends to a finished PDF
as an incremental update (ISO 32000-2, 7.5.6) - every byte already in
the file stays where it is. It writes new objects, replaces existing
ones, sets trailer entries and pins the file identifier, and it speaks
the same vocabulary as the writer, so a builder written against the one
works on the other. /Info and the metadata stream are never rewritten,
so a PDF/A or ZUGFeRD claim survives an update untouched.
* Reading a file: "::tclpdf::pdf info" answers what a finished PDF says
about itself as a dictionary - version, cross-reference flavour,
sections and revisions, identifier, encryption, page count and size, the information dictionary, the PDF/A and PDF/UA claims, tagging, form type, output intents, attachments and signatures; "pages", "fonts" and
"metadata" are the three answers that cost more, and they are separate
calls for that reason. It is the same parser the import uses, so nothing foreign has to be installed. An encrypted file is answered by "info"
rather than refused, which is what an inventory needs.
* Encryption: "encrypt" writes the standard security handler in revision
6 - AES-256 for strings, streams and the file as a whole (/V 5 /R 6,
crypt filter /AESV3) - in pure Tcl, with no library to build and no
rebuilt TclTLS. The deprecated revisions and RC4 are deliberately not
offered. Permissions are named rather than numbered (print, modify,
copy, annotate, fill, assemble, highres), "-metadata 0" leaves the XMP
packet in the clear for a cataloguing system, and the combination with
PDF/A or ZUGFeRD is refused in whichever order the calls are made.
* Signatures: "sign" prepares an approval signature over the whole file
(ISO 32000-2, 12.8), invisible or with a place on a page, and signs it
as the document is written where "-signer" says what to sign with. The two-stage way is there for signing that happens elsewhere and later - "::tclpdf::sign digest" hands out the bytes, "::tclpdf::sign embed" puts
the CMS object into the space that was reserved - and "::tclpdf::sign
add" puts a second signature onto an already signed file as an
incremental update, so the first signature keeps the bytes it covers. A signature that claims PAdES is checked against ETSI EN 319 142-1 for the signing-time attribute that profile forbids, and a timestamp token
inside the object is told apart from it by position rather than being
refused with it.
* Layers: "layer create" declares an optional content group and "layer
draw -script" puts what the script draws into it, brackets nesting
cleanly with the structure marks of a tagged document. "layer state"
switches a group, "layer radio" makes a set of them mutually exclusive,
"layer configure" names the configuration and says what a reader's panel lists. /OCProperties is written with /OCGs, /Order, /ON and /OFF; /AS
and /BaseState never are, which is what keeps a document with layers conforming under PDF/A. Layers that came in with "pdf import" are merged
into the same entry.
* Type 3 fonts: "font define" makes a font whose glyphs are drawn rather
than read out of a file, and "font glyph -script" draws one from the
same calls that draw a page - a filled shape, a gradient, a picture.
From then on the alias is a family like any other: font -family,
textWidth, a paragraph and a table cell all take it, and the marks are
copied and searched as the characters they stand for. "-color text"
writes d1 so the glyph takes the colour of the text like a letter.
* Colour fonts: "colorFont" builds a Type 3 font out of the colour
glyphs of a COLR version 0 face, one glyph per character asked for, and
hands back the alias. PDF has no colour font - the text operators paint
an outline in one colour, which is why such a face embeds blank - so the layers are read out of the face and drawn instead. The palette index
0xFFFF is a sentinel and takes the colour of the text, so a two-colour
mark follows font -color with one half and keeps its brand colour with
the other; a translucent palette entry costs an ExtGState, an opaque one
costs nothing.
* GPOS mark attachment: lookup types 4, 5 and 6 of ISO/IEC 14496-22 are
read, so a combining mark is set from the anchors of the face rather
than at the pen position, on the baseline as on a path, and marks attach
to marks as well (mkmk). Text that arrives decomposed now comes out
where the composed spelling comes out; the old advice to normalise to
NFC first is gone. With it fall the refusals for nikud, harakat, Thaana
and Samaritan - a face that needs mark attachment for its dots is set
rather than refused.
* Hyphenation: "::tclpdf::hyphenate load" reads a libhyphen .dic file -
the format LibreOffice, Hunspell and the hyphen library install - and "-hyphenate" then breaks the words the text's own soft hyphens do not
cover, in a paragraph, in "textLines" and in "textHeight" alike, so a
block is measured the way it will be set. "hyphenate word" shows what
the patterns make of a word, "-left", "-right", "-min" and "-exceptions"
tune it. tclpdf ships no patterns and that is a licence decision: every published set carries terms of its own, and this package is MIT.
* TIFF as a third image format: a TIFF is embedded and placed like a PNG
or a JPEG. Every compression that occurs has a PDF filter that undoes it
- Deflate, PackBits, CCITT Group 3 and 4, TIFF 6.0 Technote 2 JPEG - so
the strips are passed through, Group 4 fax data included; LZW is
unpacked and written out again as Flate so PDF/A holds. Where a
compression carries state from one row to the next the picture becomes
one image XObject per strip, and nothing about the calls changes.
BigTIFF, tiled files, separate planes and an alpha channel are refused
by name with what to re-save as.
* The resolution a picture states about itself: "-dpi auto" reads a PNG
pHYs chunk, a JFIF density, an Exif resolution or a TIFF XResolution,
and places the picture at the size it was scanned from. Without it a 1728-pixel fax line read as 72 dpi comes out 610 mm wide, off the sheet
in any format. "image info" answers xResolution, yResolution, resolution
- which of the four the number came from - and pixelAspect for a file
that gives a ratio without a measure.
* Lab colour: "{lab L a b}" is a colour as it was measured rather than
as a device would make it, with "-whitePoint" and "-range" as options on
the colour itself. Its point is the spot colour - a separation whose
alternate is a measurement instead of a CMYK guess - and, being device independent, it is admitted under every output intent where the same
colour as {rgb ...} is not.
* Text rendering modes: "-render" takes a word rather than a number -
fill, stroke, fillStroke, invisible - with "-stroke" and "-strokeWidth"
for the two that stroke. "invisible" is the mode a scanned page uses to
carry its recognised text under the picture of itself. The four clipping
modes are refused by name and with the reason.
* Image masks: "image embed -stencil 1" turns a one-bit PNG into a
stencil mask, so one embedding paints the fill colour in force through
the bits of the file - a logo in any number of colours from one object.
"-mask alias" names a second embedded picture as this one's mask, a
stencil becoming the /Mask of explicit masking and any other picture the /SMask of a soft mask. "-invert" and "-interpolate" go with them.
* A fallback chain: "font -fallback {notoJP ...}" names further faces,
and every character the family has no glyph for is set from the first
face in the list that has one. The line falls into segments, each with
its own font resource and its own ToUnicode map, so the text extracts
whole and in order. It is font state like -size, and it is how a colour
font stands beside a real one.
* Error codes: the machine-readable -errorcode of a refusal is a
contract, and the manual now lists every code it promises - the font and colour-font classes, TCLPDF COLR for what the COLR/CPAL reader refuses,
TCLPDF TIFF for the TIFF reader, the two hyphenation codes and the two signature ones.
Changed since 1.1:
* A picture is placed at the resolution it states, because "-dpi auto"
is the default. Before, a pixel was a point unless -dpi said otherwise.
A file that states nothing still falls back on 72, so nothing changes
for a picture without a resolution; a scan changes size, which is the point.
* Text that arrives decomposed sets differently, because GPOS mark
attachment is now read: the mark lands on its anchor instead of at the
pen position. Composed text is unaffected. A face that needs mark
attachment for its dots is no longer refused, so calls that used to fail
now draw.
Fixed since 1.1:
* A fourth norm review before this release closed twelve violations that
no validator sees, among them: an imported page inherited nothing from
its page tree, so a page whose MediaBox sat on a parent node came out at
the wrong size; a CropBox was not intersected with the MediaBox; an
escaped resource name was looked up under its raw spelling; a paragraph reported the position of a missing glyph relative to the chunk it was
breaking rather than to the string handed in; a version lowered after
the fact slipped past the 2.0 guards; and a dozen checks that came after
the first byte was written now come before it.
* A fifth review followed the same way, over the code built since -
import, encryption, signatures, marks, hyphenation, Type 3 and colour
fonts, TIFF - and its findings are built. What it looked for is what
"make check" cannot see: values at the edge of what the standard
permits, state left behind by a refused call, and every sentence of the
manual read back against the code.
Notes:
* Tested against Tcl 8.6.18 and Tcl 9.0.4 on macOS: 2598 tests green
under both over 59 test files, 63 examples writing 70 documents, every
one of them accepted by "qpdf --check" and, where it claims PDF/A or
PDF/UA, by veraPDF against the profile it claims; a document with a
hybrid attachment is read by Mustangproject as well. Every command,
option and value the manual names has a test of its own; "make check"
runs the whole acceptance - suite under every interpreter, examples,
qpdf, veraPDF, Mustangproject, pdfinfo, and whether the manual is newer
than its source.
* The package is 84 modules and loads only what a script actually uses: "package require tclpdf" registers all 84 and loads one, creating a
document loads eight, a document with a line of text fifteen, and one
with shapes and a table twenty - a script that never touches an image,
an SVG, an invoice, a signature or a foreign file never reads the other sixty-four.
* Installing: the release archives tclpdf1.2.zip and tclpdf1.2.tar.gz
hold one directory tclpdf1.2 with the modules, pkgIndex.tcl and the ICC profiles - nothing to build; extract it into a directory on your
auto_path and "package require tclpdf" finds it. The source archive
builds the TEA way in the source directory: ./configure, make test, make install. A separate build directory configures cleanly but cannot run
the tests, because the generated pkgIndex.tcl locates the modules
through $dir while the sources stay where they are.
* tdom is required for the XMP metadata packet, so every document that declares PDF/A, PDF/UA or ZUGFeRD needs it; a document with no such
claim runs without it. The SVG module uses it as well when present: it
parses 27 to 48 times faster than the parser tclpdf brings along and
refuses an entity expansion bomb, and without it SVG still works through
the built-in element tree parser. tzint is optional and only for
barcodes. Nothing else is used - the encryption and the signatures are
pure Tcl.
* No font is part of the package. The faces under examples/assets/fonts
exist so that the tests and examples have something to embed and travel
in the source archive alone; "make install" installs none. Whether you
may embed a font into a PDF that tclpdf writes is a question for that
font's licence, not for this one.
* Some things are refused with a message naming the reason and the way
out rather than written wrong: a progressive JPEG and an interlaced
(Adam7) PNG, a BigTIFF and a tiled or planar TIFF and one with an alpha channel, an encrypted file handed to "pdf import" or "update open" -
this package carries no decryption -, and a line that mixes writing
directions in one call.
* What tclpdf covers is roughly a third of ISO 32000-1, weighted by the
page count of its chapters. Of the rest, what a writer would need is
small: interactive form fields beyond the signature field, multimedia
and 3D, JBIG2 and JPX, linearisation, and object and cross-reference
streams to write - reading them has been there since the import.
Planned:
* Drawing a Tk canvas into a PDF - every item type, as vectors rather
than as a screenshot. This is the one planned feature that involves Tk
at all, and it stays outside the core: a caller who has a canvas has
loaded Tk already, while the package itself neither requires nor loads it.
* The small gaps, each an afternoon and none of them blocked: shading
types 1 and 4 to 7 beside the axial and radial ones that are there,
further annotation types beside the link and the signature widget,
inline images, vertical writing (Identity-V), CCITT streams on import,
DeviceN beside the separation, and a bare CFF table without its sfnt
wrapper.
* Deliberately not coming, each for a measured reason: mixed writing directions in one call, PDF/X-3 and X-4 (no tool here could check the
claim), document timestamps (the same reason), RC4 and the older
encryption revisions, EPS as a template, GIF and BMP.
None of this is promised for a particular release, and nothing on the
list is needed for what 1.2 does.
Thanks to everyone who reported findings on the mailing list, and to
those who wrote to me directly - several of the corrections in this
release came out of private mail rather than a public thread.
Deutsch
-------
Normen, auf denen das Paket aufsetzt und gegen die es geprueft wird: ISO 32000-1 (PDF 1.7) und ISO 32000-2 (PDF 2.0) fuer die Datei selbst; ISO
19005-2 und 19005-3 fuer PDF/A Teil 2 und 3 in den Stufen B, U und A;
ISO 14289-1 und 14289-2 fuer PDF/UA, dazu WTPDF fuer das gut getaggte
Profil; EN 16931 mit ZUGFeRD / Factur-X 2.3 und Order-X fuer
elektronische Rechnungen und Bestellungen; ETSI EN 319 142-1 (PAdES) und
RFC 5652 (CMS) fuer Signaturen, RFC 3161 fuer die Zeitstempel darin;
ISO/IEC 14496-22 (OpenType, die sfnt-Tabellen, GSUB und GPOS) fuer
Schriften, daneben Adobes eigene Spezifikationen fuer das, was OpenType
nicht abdeckt - das Type 1 Font Format, Technical Note #5176 (CFF),
#5902 (PostScript-Namensbildung), die Adobe Font Metrics der vierzehn Standardschriften und die Adobe Glyph List - mit UAX #9 fuer
rechtslaeufigen Text; TIFF 6.0 samt Technote 2, JFIF und Exif fuer
Bilder. Was ein Profil beansprucht, wird bei jedem Lauf dagegen
geprueft: veraPDF fuer PDF/A und PDF/UA, Mustangproject fuer die
Rechnungen, qpdf und pdfsig ueber jedes Dokument.
Wie viel davon geprueft wird: 2598 Tests in 59 Testdateien, gruen unter
Tcl 8.6.18 und Tcl 9.0.4, und 63 Beispielskripte, die 70 Dokumente
schreiben - jeder Befehl, jede Option und jeder Wert, den das Handbuch
nennt, hat einen eigenen Test, und jedes Beispiel wird ausgefuehrt,
validiert und angesehen.
Zwei Technotes gehoeren zu dieser Ausgabe:
Was es kann und was bewusst fehlt:
https://fossil.sowaswie.de/tclpdf/technote/071ce10bf8
Warum es tclpdf gibt und was das ueber die Arbeit mit KI zeigt:
https://fossil.sowaswie.de/tclpdf/technote/d647175afa
Bezug / Repository:
https://fossil.sowaswie.de/tclpdf
Autor: Alexander Schoepe <
alx.tcl@sowaswie.de> - Fehlermeldungen,
Patches und Fragen gern auch per Mail, nicht nur ueber das Repository.
Neu seit 1.1:
* Eine Seite uebernehmen: "pdf import" nimmt eine Seite eines
vorhandenen PDF als Form, und "form place" setzt sie beliebig oft,
skaliert, gedreht, transparent - ein Briefbogen wird uebernommen statt nachgebaut. Beide Verzeichnis-Bauweisen werden gelesen, die klassische
Tabelle und Querverweisstroeme mit Objektstroemen, Hybriddateien eingeschlossen und Fortschreibungen ueber /Prev verfolgt; vererbte Seitenattribute kommen aus dem Seitenbaum, CropBox und /Rotate werden
BBox und Matrix der Form, und alles, was die Ressourcen der Seite
erreichen, wird byteweise mit den eigenen Filtern kopiert. Nur die Inhaltsstroeme werden dekodiert, ein exotischer Bildfilter ist also kein Hindernis; LZW gehoert zu den gelesenen Filtern, /EarlyChange mit ihm.
* Eine Datei fortschreiben: "::tclpdf::update open" haengt an ein
fertiges PDF als Fortschreibung an (ISO 32000-2, 7.5.6) - jedes Byte,
das schon in der Datei steht, bleibt, wo es ist. Der Griff schreibt neue Objekte, ersetzt bestehende, setzt Trailer-Eintraege und legt den Dateibezeichner fest, und er spricht dieselbe Sprache wie der Writer,
sodass ein Erzeuger, der gegen den einen geschrieben ist, auf dem
anderen laeuft. /Info und der Metadatenstrom werden nie neu geschrieben,
eine PDF/A- oder ZUGFeRD-Zusage ueberlebt eine Fortschreibung also
unberuehrt.
* Eine Datei auslesen: "::tclpdf::pdf info" sagt als Woerterbuch, was
ein fertiges PDF ueber sich selbst sagt - Version, Verzeichnisbauweise, Abschnitte und Revisionen, Bezeichner, Verschluesselung, Seitenzahl und Groesse, das Info-Woerterbuch, die PDF/A- und PDF/UA-Zusagen, Tagging, Formulartyp, Output-Intents, Anhaenge und Signaturen; "pages", "fonts"
und "metadata" sind die drei teureren Antworten und darum eigene
Aufrufe. Es ist derselbe Parser, den der Import benutzt, es muss also
nichts Fremdes installiert sein. Eine verschluesselte Datei wird von
"info" beantwortet statt abgelehnt - das ist es, was eine
Bestandsaufnahme braucht.
* Verschluesselung: "encrypt" schreibt den Standard-Sicherheitshandler
in Revision 6 - AES-256 fuer Strings, Stroeme und die Datei als Ganzes
(/V 5 /R 6, Crypt-Filter /AESV3) - in reinem Tcl, ohne eine Bibliothek
zu bauen und ohne neu gebautes TclTLS. Die veralteten Revisionen und RC4 werden bewusst nicht angeboten. Rechte werden benannt statt nummeriert
(print, modify, copy, annotate, fill, assemble, highres), "-metadata 0"
laesst das XMP-Paket im Klartext, damit ein Katalogsystem indizieren
kann, und die Kombination mit PDF/A oder ZUGFeRD wird abgelehnt, in
welcher Reihenfolge die Aufrufe auch kommen.
* Signaturen: "sign" bereitet eine Freigabesignatur ueber die ganze
Datei vor (ISO 32000-2, 12.8), unsichtbar oder mit einem Platz auf einer Seite, und unterschreibt beim Schreiben, wo "-signer" sagt, womit. Der zweistufige Weg ist fuer das Unterschreiben anderswo und spaeter da - "::tclpdf::sign digest" gibt die Bytes heraus, "::tclpdf::sign embed"
legt das CMS-Objekt in den reservierten Platz -, und "::tclpdf::sign
add" setzt eine zweite Signatur als Fortschreibung auf eine bereits unterschriebene Datei, sodass die erste die Bytes behaelt, die sie
deckt. Eine Signatur, die PAdES beansprucht, wird gegen ETSI EN 319
142-1 auf das signing-time-Attribut geprueft, das dieses Profil
verbietet, und ein Zeitstempel-Token im Objekt wird an seiner Lage davon unterschieden statt mit ihm abgelehnt.
* Ebenen: "layer create" erklaert eine Optional-Content-Gruppe, und
"layer draw -script" legt hinein, was das Skript zeichnet; die Klammern verschachteln sich sauber mit den Strukturmarken eines getaggten
Dokuments. "layer state" schaltet eine Gruppe, "layer radio" macht eine
Menge davon gegenseitig ausschliessend, "layer configure" benennt die Konfiguration und sagt, was die Ebenenliste eines Lesers zeigt.
/OCProperties wird mit /OCGs, /Order, /ON und /OFF geschrieben; /AS und /BaseState nie, und genau das haelt ein Dokument mit Ebenen unter PDF/A konform. Ebenen, die mit "pdf import" hereinkamen, gehen in denselben
Eintrag ein.
* Type-3-Schriften: "font define" macht eine Schrift, deren Glyphen
gezeichnet statt aus einer Datei gelesen werden, und "font glyph
-script" zeichnet eine aus denselben Aufrufen, die eine Seite zeichnen -
eine gefuellte Form, ein Verlauf, ein Bild. Von da an ist der Alias eine Familie wie jede andere: font -family, textWidth, ein Absatz und eine Tabellenzelle nehmen ihn, und die Zeichen werden kopiert und gesucht als
die Zeichen, fuer die sie stehen. "-color text" schreibt d1, sodass die
Glyphe die Textfarbe annimmt wie ein Buchstabe.
* Farbige Schriften: "colorFont" baut aus den Farbglyphen einer COLR-Version-0-Schrift eine Type-3-Schrift, eine Glyphe je verlangtem
Zeichen, und gibt den Alias zurueck. PDF kennt keine Farbschrift - die Textoperatoren malen einen Umriss in einer Farbe, weswegen eine solche
Schrift leer einbettet -, also werden die Ebenen aus der Schrift gelesen
und gezeichnet. Der Palettenindex 0xFFFF ist ein Kennwert und nimmt die Textfarbe, sodass eine zweifarbige Marke mit der einen Haelfte font
-color folgt und mit der anderen ihre Hausfarbe behaelt; ein
durchscheinender Paletteneintrag kostet ein ExtGState, ein deckender nichts.
* GPOS-Markenpositionierung: die Lookuptypen 4, 5 und 6 von ISO/IEC
14496-22 werden gelesen, eine kombinierende Marke wird also ueber die
Anker der Schrift gesetzt statt an der Schreibposition, auf der
Grundlinie wie auf einem Pfad, und Marken haengen sich auch an Marken
(mkmk). Zerlegt ankommender Text kommt jetzt dort heraus, wo die
komponierte Schreibweise herauskommt; der alte Rat, vorher nach NFC zu normalisieren, ist weg. Damit fallen die Absagen fuer Nikud, Harakat,
Thaana und Samaritanisch - eine Schrift, die fuer ihre Punkte Markenpositionierung braucht, wird gesetzt statt abgelehnt.
* Silbentrennung: "::tclpdf::hyphenate load" liest eine
libhyphen-.dic-Datei - das Format, das LibreOffice, Hunspell und die hyphen-Bibliothek installieren -, und "-hyphenate" trennt dann die
Woerter, die die eigenen weichen Trennstriche des Textes nicht abdecken,
in einem Absatz, in "textLines" und in "textHeight" gleichermassen,
sodass ein Block so gemessen wird, wie er gesetzt wird. "hyphenate word" zeigt, was die Muster aus einem Wort machen, "-left", "-right", "-min"
und "-exceptions" stellen es ein. tclpdf liefert keine Muster mit, und
das ist eine Lizenzentscheidung: jede veroeffentlichte Sammlung traegt
eigene Bedingungen, und dieses Paket ist MIT.
* TIFF als drittes Bildformat: ein TIFF wird eingebettet und platziert
wie ein PNG oder ein JPEG. Jede vorkommende Kompression hat einen
PDF-Filter, der sie aufloest - Deflate, PackBits, CCITT Gruppe 3 und 4, TIFF-6.0-Technote-2-JPEG -, die Strips werden also durchgereicht, Gruppe-4-Faxdaten eingeschlossen; LZW wird entpackt und als Flate
geschrieben, damit PDF/A traegt. Wo eine Kompression Zustand von Zeile
zu Zeile mitnimmt, wird das Bild ein Image-XObject je Strip, und an den Aufrufen aendert sich nichts. BigTIFF, Kacheldateien, getrennte Ebenen
und ein Alphakanal werden mit Namen abgelehnt und sagen, als was neu zu speichern ist.
* Die Aufloesung, die ein Bild ueber sich selbst angibt: "-dpi auto"
liest einen PNG-pHYs-Block, eine JFIF-Dichte, eine Exif-Aufloesung oder
ein TIFF-XResolution und platziert das Bild in der Groesse, in der es
gescannt wurde. Ohne das kommt eine Faxzeile von 1728 Pixeln, als 72 dpi gelesen, 610 mm breit heraus, in jedem Format ueber das Blatt hinaus.
"image info" antwortet xResolution, yResolution, resolution - welche der
vier Quellen die Zahl lieferte - und pixelAspect fuer eine Datei, die
ein Verhaeltnis ohne Mass angibt.
* Lab-Farbe: "{lab L a b}" ist eine Farbe, wie sie gemessen wurde, und
nicht, wie ein Geraet sie machen wuerde, mit "-whitePoint" und "-range"
als Optionen an der Farbe selbst. Ihr Zweck ist die Sonderfarbe - eine Separation, deren Alternate eine Messung ist statt einer CMYK-Schaetzung
- und sie ist als geraeteunabhaengige Farbe unter jedem Output-Intent zugelassen, wo dieselbe Farbe als {rgb ...} es nicht ist.
* Textrendering-Modi: "-render" nimmt ein Wort statt einer Zahl - fill, stroke, fillStroke, invisible - mit "-stroke" und "-strokeWidth" fuer
die beiden, die stricheln. "invisible" ist der Modus, mit dem eine
gescannte Seite ihren erkannten Text unter ihrem eigenen Bild traegt.
Die vier Clip-Modi werden mit Namen und mit dem Grund abgelehnt.
* Bildmasken: "image embed -stencil 1" macht aus einem Ein-Bit-PNG eine Schablonenmaske, sodass eine Einbettung die gerade geltende Fuellfarbe
durch die Bits der Datei malt - ein Logo in beliebig vielen Farben aus
einem Objekt. "-mask alias" benennt ein zweites eingebettetes Bild als
Maske dieses Bildes, wobei eine Schablone zur /Mask der expliziten
Maskierung wird und jedes andere Bild zur /SMask einer weichen Maske. "-invert" und "-interpolate" gehoeren dazu.
* Eine Ersatzschrift-Kette: "font -fallback {notoJP ...}" nennt weitere Schnitte, und jedes Zeichen, fuer das die Familie keine Glyphe hat, wird
aus dem ersten Schnitt der Liste gesetzt, der eine hat. Die Zeile faellt
in Abschnitte, jeder mit eigener Schriftressource und eigener ToUnicode-Zuordnung, sodass der Text ganz und in der Reihenfolge
extrahiert. Sie ist Schriftzustand wie -size, und sie ist der Weg, auf
dem eine Farbschrift neben einer echten steht.
* Fehlercodes: der maschinenlesbare -errorcode einer Ablehnung ist eine Zusage, und das Handbuch listet jetzt jeden Code, den es zusagt - die
Schrift- und Farbschrift-Klassen, TCLPDF COLR fuer das, was der COLR/CPAL-Leser ablehnt, TCLPDF TIFF fuer den TIFF-Leser, die beiden Silbentrenn-Codes und die beiden Signatur-Codes.
Geaendert seit 1.1:
* Ein Bild wird in der Aufloesung platziert, die es angibt, denn "-dpi
auto" ist die Vorgabe. Vorher war ein Pixel ein Punkt, solange -dpi
nichts anderes sagte. Eine Datei, die nichts angibt, faellt weiterhin
auf 72 zurueck, fuer ein Bild ohne Aufloesung aendert sich also nichts;
ein Scan aendert seine Groesse, und darum geht es.
* Zerlegt ankommender Text setzt anders, weil GPOS-Markenpositionierung
jetzt gelesen wird: die Marke landet auf ihrem Anker statt an der Schreibposition. Komponierter Text ist unberuehrt. Eine Schrift, die
fuer ihre Punkte Markenpositionierung braucht, wird nicht mehr
abgelehnt, Aufrufe, die frueher scheiterten, zeichnen also jetzt.
Behoben seit 1.1:
* Eine vierte Normpruefung vor dieser Fassung schloss zwoelf Verstoesse,
die kein Validator sieht, darunter: eine importierte Seite erbte nichts
aus ihrem Seitenbaum, sodass eine Seite, deren MediaBox an einem
Elternknoten stand, in der falschen Groesse herauskam; eine CropBox
wurde nicht mit der MediaBox geschnitten; ein escapeter Ressourcenname
wurde unter seiner rohen Schreibweise nachgeschlagen; ein Absatz meldete
die Position einer fehlenden Glyphe relativ zu dem Stueck, das er gerade umbrach, statt zur uebergebenen Zeichenkette; eine nachtraeglich
abgesenkte Version rutschte an den 2.0-Waechtern vorbei; und ein Dutzend Pruefungen, die nach dem ersten geschriebenen Byte kamen, kommen jetzt
davor.
* Eine fuenfte Pruefung folgte auf dieselbe Weise, ueber den seither
gebauten Code - Import, Verschluesselung, Signaturen, Marken,
Silbentrennung, Type 3 und Farbschriften, TIFF -, und ihre Befunde sind gebaut. Gesucht war, was "make check" nicht sehen kann: Werte am Rand
dessen, was die Norm erlaubt, Zustand, den ein abgelehnter Aufruf zuruecklaesst, und jeder Satz des Handbuchs gegen den Code zurueckgelesen.
Hinweise:
* Geprueft gegen Tcl 8.6.18 und Tcl 9.0.4 unter macOS: 2598 Tests gruen
unter beiden ueber 59 Testdateien, 63 Beispiele mit 70 Dokumenten, jedes
davon von "qpdf --check" angenommen und, wo es PDF/A oder PDF/UA
beansprucht, von veraPDF gegen das beanspruchte Profil; ein Dokument mit Hybridanhang wird ausserdem von Mustangproject gelesen. Jeder Befehl,
jede Option und jeder Wert, den das Handbuch nennt, hat einen eigenen
Test; "make check" faehrt die ganze Abnahme - Suite unter jedem
Interpreter, Beispiele, qpdf, veraPDF, Mustangproject, pdfinfo, und ob
das Handbuch neuer ist als seine Quelle.
* Das Paket besteht aus 84 Modulen und laedt nur, was ein Skript
wirklich benutzt: "package require tclpdf" meldet alle 84 an und laedt
eines, ein Dokument anzulegen laedt acht, ein Dokument mit einer
Textzeile fuenfzehn und eines mit Formen und einer Tabelle zwanzig - wer
nie ein Bild, kein SVG, keine Rechnung, keine Signatur und keine fremde
Datei anfasst, liest die anderen vierundsechzig nie.
* Installation: die Release-Archive tclpdf1.2.zip und tclpdf1.2.tar.gz enthalten ein Verzeichnis tclpdf1.2 mit den Modulen, pkgIndex.tcl und
den ICC-Profilen - nichts zu bauen; in ein Verzeichnis auf dem auto_path entpacken, und "package require tclpdf" findet es. Das Quellarchiv baut
auf TEA-Art im Quellverzeichnis: ./configure, make test, make install.
Ein eigenes Bauverzeichnis konfiguriert sauber, kann die Tests aber
nicht ausfuehren, weil die erzeugte pkgIndex.tcl die Module ueber $dir
sucht, waehrend die Quellen liegen bleiben.
* tdom wird fuer das XMP-Metadatenpaket gebraucht, also von jedem
Dokument, das PDF/A, PDF/UA oder ZUGFeRD erklaert; ein Dokument ohne
solche Erklaerung laeuft ohne. Das SVG-Modul benutzt es ausserdem, wenn
es vorhanden ist: es parst 27- bis 48-mal schneller als der
mitgelieferte Parser und wehrt eine Entity-Bombe ab, und ohne tdom
laeuft SVG weiter, ueber den eingebauten Elementbaum-Parser. tzint ist optional und nur fuer Barcodes. Sonst wird nichts benutzt - die Verschluesselung und die Signaturen sind reines Tcl.
* Keine Schrift gehoert zum Paket. Die Schnitte unter
examples/assets/fonts sind da, damit Tests und Beispiele etwas zum
Einbetten haben, und reisen allein im Quellarchiv mit; "make install" installiert keine. Ob eine Schrift in ein von tclpdf geschriebenes PDF eingebettet werden darf, entscheidet die Lizenz jener Schrift, nicht
diese hier.
* Manches wird mit einer Meldung abgelehnt, die Grund und Ausweg nennt,
statt falsch geschrieben zu werden: ein progressives JPEG und ein verschraenktes (Adam7) PNG, ein BigTIFF und ein gekacheltes oder ebenengetrenntes TIFF und eines mit Alphakanal, eine verschluesselte
Datei an "pdf import" oder "update open" - dieses Paket bringt keine Entschluesselung mit - und eine Zeile, die in einem Aufruf die Schreibrichtungen mischt.
* Was tclpdf abdeckt, ist rund ein Drittel von ISO 32000-1, gewichtet
nach dem Seitenumfang der Kapitel. Vom Rest ist das, was ein Schreiber braeuchte, wenig: interaktive Formularfelder ueber das Signaturfeld
hinaus, Multimedia und 3D, JBIG2 und JPX, Linearisierung sowie Objekt-
und Querverweisstroeme zum Schreiben - gelesen werden sie seit dem Import.
Geplant:
* Einen Tk-Canvas ins PDF zeichnen - jeder Elementtyp, als Vektoren
statt als Bildschirmfoto. Das ist das einzige geplante Merkmal, in dem
Tk ueberhaupt vorkommt, und es bleibt ausserhalb des Kerns: wer einen
Canvas hat, hat Tk ohnehin geladen, waehrend das Paket selbst es weder verlangt noch laedt.
* Die kleinen Luecken, jede ein Nachmittag und keine blockiert:
Shading-Typen 1 und 4 bis 7 neben den axialen und radialen, die da sind, weitere Annotationstypen neben dem Link und dem Signatur-Widget, Inline-Bilder, vertikale Schreibrichtung (Identity-V), CCITT-Stroeme
beim Import, DeviceN neben der Separation und eine nackte CFF-Tabelle
ohne ihre sfnt-Huelle.
* Ausdruecklich nicht geplant, jedes aus einem gemessenen Grund:
gemischte Schreibrichtungen in einem Aufruf, PDF/X-3 und X-4 (kein
Werkzeug hier koennte die Zusage pruefen), Dokument-Zeitstempel
(derselbe Grund), RC4 und die aelteren Verschluesselungsrevisionen, EPS
als Vorlage, GIF und BMP.
Nichts davon ist fuer eine bestimmte Fassung zugesagt, und nichts auf
dieser Liste wird fuer das gebraucht, was 1.2 leistet.
Dank an alle, die auf der Mailingliste Befunde gemeldet haben, und an
die, die mir direkt geschrieben haben - mehrere Korrekturen dieser
Ausgabe kamen aus persoenlicher Post und nicht aus einem oeffentlichen
Thread.
--- Synchronet 3.22a-Linux NewsLink 1.2