From Newsgroup: comp.lang.tcl
ANNOUNCE: tclpdf 1.1 released
=============================
tclpdf is a pure Tcl extension for creating PDF documents: pages and
graphics, embedded fonts in three formats - TrueType subset to the
glyphs actually used, OpenType with CFF outlines and Type 1 - JPEG and
PNG images, SVG drawn as vectors, tables, tagged and accessible
documents up to PDF/UA-2, and electronic invoices as ZUGFeRD / Factur-X
in PDF/A-3B. Developed by Alexander Schoepe, it requires Tcl 8.6.11+ and
runs under Tcl 9 as well; zlib is a built-in command rather than a
package, tdom is needed only for the XMP metadata of a PDF/A, PDF/UA or ZUGFeRD document, and Tk is not used.
1.1 is the first release after 1.0 of 11 August 2026, and it is a large
one: kerning and ligatures from the font's own tables, tagged PDF with
PDF/A level A and PDF/UA-1/-2, three font formats and variable font
instances, right-to-left text with Arabic contextual forms, text that paginates itself in columns inside a type area, barcodes through tzint,
PDF/A colours held against the output intent, and a test suite that now exercises every option and value the manual names. The API is unchanged;
four behaviours are, and they are listed under "Changed since 1.0". With
this release the second big block of work is done. :-)
Download / Repository:
https://fossil.sowaswie.de/tclpdf
English
-------
New since 1.0:
* Kerning and ligatures: pair kerning from GPOS or the kern table, the standard ligatures (liga) from GSUB, both on by default and switchable
per call (-kerning 0, -ligatures 0); lookup flags and GDEF classes are honoured, so a combining accent between two letters no longer breaks a
pair. A ligature extracts as the letters it was made from. -spacing
counts glyphs, not characters, which is what the PDF operator does.
* Tagged PDF: a structure tree with the 49 standard types (1.7) and the
2.0 additions, nesting checked against Annex L, artifacts kept out and
named by kind (Pagination/Layout/Page/Background with Header/Footer/Watermark), abbreviations with their expansion, structure attributes (-lang, -title, -actualText, -scope, -numbering, -bbox,
-colSpan, -rowSpan), element identifiers (-id, written as /ID and the
IDTree; a Note gets one on its own), and structure destinations that
links and bookmarks point at with -structure. text, table and image tag themselves; -tag takes anything else. A paragraph that paginates stays
one element with a mark on every page.
* PDF/A level A on top of B and U, and PDF/UA-1 and PDF/UA-2 with the Well-Tagged PDF declaration ($doc ua): the claim is refused when the
document does not keep it - title, language, every font embedded,
headings without a gap, a description on every link and picture, lists
whose numbering and labels agree - and every failure is reported at
once, naming the call to change. Every example passes veraPDF against
the profile it claims.
* Right-to-left text: -direction rtl for lines, paragraphs, leaders,
text on a path, page numbers and table cells; Hebrew by order alone,
Arabic with its contextual forms (isol/init/medi/fina, ccmp, rlig) from
the face's own GSUB tables and checked glyph by glyph against HarfBuzz;
digits keep their own order inside such a line, paired brackets are
mirrored and extract as written. A face that needs GPOS mark attachment
for its dots is refused with a message naming it, a script that needs
shaping this package does not have (Devanagari, Thai and the like) is
refused rather than drawn wrong, and -unshaped 1 turns the refusal off
for the caller who decides isolated glyphs are better than nothing. A mixed-direction line is refused - the bidi algorithm is planned work.
* Three font formats: TrueType subset (as before), OpenType with CFF
outlines whole (/FontFile3), Type 1 whole (.pfb, .pfa, .t1 with its
AFM); a variable face is embedded as any instance on its axes (-axes, -instance) with outlines, advances, bounding boxes and hhea summary
computed for the instance; font info answers what the file says about
itself. Type 1 and the fourteen standard faces are addressed through
WinAnsi with the no-break space and the soft hyphen where Annex D puts
them. A character the face has no glyph for is an error, and the manual
now says why a font viewer is no proof of coverage.
* Text flow: -height hands back what did not fit, -height max measures
down to the type area the document was given (-typeArea {top bottom
?left right?}, page typeArea reads it), -paginate 1 runs the loop itself
over as many pages as it takes - in as many columns as asked (-columns
n) with -gutter and -balance - while pageAdded fires per page; text
flows around rectangles and circles (-avoid), soft hyphens (U+00AD)
break where they stand and never reach a font, a break hyphen is written
where a word was divided, leader rows (leader), text along a path
(textPath), page numbers drawn at write time (pageNumbers), first-line
and right indents, paragraph spacing.
* Graphics: sixteen blend modes as state (blend) and per shape (-blend),
style holds every option including the two colours - a shape paints the
sides style coloured plus the ones it names - and the state belongs to
the stream, so save/restore, forms and pages keep it apart; transform
composes several parts like the same calls in sequence, so a
displacement is no longer scaled or turned; the type area for tables and
text; PDF/A colour spaces are held against the output intent at write
time and a CMYK document validates with basICColor's ISO Coated v2, a
grey one with the grey intent, both shipped.
* SVG: text in a drawing uses the same faces as the rest of the
document, embedded ones included, with kerning and ligatures; svg -data
takes the markup from a variable; a drawing that fails leaves neither an
open q nor a font resource behind. Barcodes through tzint (Code 128,
EAN, QR, DataMatrix, PDF417 and 100 more), whose SVG tclpdf draws as
vectors - the clear text line under a barcode can be real OCR-B, as text.
* Images: image embed and image draw take -data with the bytes
themselves; a PNG whose transparency is one colour keeps it as a
colour-key mask on the pass-through path, a 16-bit key becomes a mask
because readers ignore it otherwise; image info says how the
transparency reaches the file (none, colourKey, softMask), and image
size answers what a placement would measure.
* Tables: -top says where a breaking table resumes and -bottom how far
it may run, both defaulting to the type area; a group of rows moves to
the next page as a whole; -border outer frames the table per page;
themes plain/striped/grid; four hooks per cell; valign, decimal
alignment measured over head, body and foot together.
* Documents and files: file attachments with relationship, description
and modification date as a core feature; links to a URL, a page or a
structure element; bookmarks; page labels; viewer preferences with their version rules; document language checked against RFC 3066; the six text
info entries and the two dates paired with the XMP packet as ISO 32000-1
Table 317 pairs them (Creator is xmp:CreatorTool, Keywords is
pdf:Keywords, CreationDate is xmp:CreateDate), while Trapped stays in
the Info dictionary alone; the version is never raised silently, and
every feature is checked against it before it is written; writing is idempotent, so a second write of an unchanged document is byte-identical.
* ZUGFeRD / Factur-X: -type and -profile are validated (INVOICE/ORDER; MINIMUM, BASIC WL, BASIC, EN 16931, EXTENDED, XRECHNUNG), zugferd state carries the relationship written, and pdfa -conformance U or A may be
claimed on top of the invoice's 3B.
* Order-X (Order-X 1.0, 4.1.1 and 4.1.2) rides on the same zugferd call:
the BT-24 identifiers urn:order-x.eu:1p0:basic, :comfort and :extended
are read as BASIC, COMFORT and EXTENDED, the document becomes an ORDER
(or ORDER_RESPONSE, ORDER_CHANGE by -type), the XML is embedded as
order-x.xml with /AFRelationship Data by default, and the extension
schema carries the Order-X namespace URI; levels, types and names are
bound to their family, so an invoice cannot go out labelled an order or
the other way round.
* Extension points: doc/PLUGINS.md describes the event bus (beforeWrite, afterWrite, pageAdded, resources, catalog, info), how a topical module attaches with oo::define, and what is and is not promised.
* Examples: 41 of them writing 43 documents, among them a type-area and pagination example, writing systems in ten Noto faces, barcodes,
variable fonts, Type 1 and OpenType, tagged documents, PDF/UA-2 with
WTPDF, and a ZUGFeRD invoice trying every option; the fonts they need
travel in the source archive, sorted by origin with their licences.
Changed since 1.0:
* transform with several parts in one call now composes them like the
same calls in sequence (translate, rotate, skew, scale): the
displacement is no longer scaled or turned. A call with -at and -rotate
alone is unaffected.
* style -fill and -stroke are in force until changed, like the other
options: a bare shape after "style -fill red -stroke blue" is painted;
before, it was constructed and discarded. line and curve draw black only
when no style stroke stands.
* A two-number -format is turned when an orientation is asked for on
tclpdf new or configure as well, not only on page add; the default
portrait is no request, so a pair without any orientation is still taken
as it stands.
* Kerning and ligatures are on by default, so every document with an
embedded face sets differently than under 1.0 - measured, nine of the
then twenty-five examples changed, exactly the ones with embedding;
PDF/A and ZUGFeRD output stays conformant.
Fixed since 1.0:
* Justified, rotated and anchored text: the alignment shift ran along
the rotated baseline, and the ascender was lost in the return value of a paragraph; word spacing in justified lines is applied through TJ so the
last glyph lands on the edge.
* A font whose PostScript name sits under platform 0 in the name table
(Times New Roman and Arial Narrow among the 20 of 210 faces on this
machine) produced a /BaseFont with null bytes that qpdf rejected;
platform 0 is decoded as UTF-16 like platform 3.
* A breaking table resumed on the next page where it had started on the
first; -top and -bottom are derived from the page, not from -at. valign
top sat on the rule when the AFM said Ascender 0.
* Text on a path measured with kerning and drew without, and measured a -spacing it did not draw; an empty string with -wordSpacing had a
negative width; a coverage index was a running counter instead of the
table's.
* An SVG that failed left its q open, so everything after it stood
upside down and enlarged, and left a font resource behind that made pdfa refuse the document.
* SVG gradient patterns took the SVG id as their name, so two drawings
with the same id shared one gradient.
* The subset prefix follows the glyphs used, the name and the axis
point, so two subsets of one file in one document carry different tags
and two writes the same.
* The XMP packet is rebuilt on every write, so a title or language
changed between two writes reaches the file; Creator no longer
masquerades as the producer in xmp:CreatorTool.
* checkSumAdjustment of an embedded subset was 0 in every file - the
code comment said it was recomputed and it was not; forms, patterns and page-number forms carry their own /Resources as PDF/A 6.2.2 requires; a
16-bit PNG colour key is a mask, not a /Mask array readers ignore.
* Release archives are readable for everyone: 37 modules used to travel
as 0600 because the checkout's modes were copied through.
* A third norm review before 1.1 closed what no validator sees: the
xpacket header carried a double-encoded byte order mark; writeChannel
into a pipe or socket wrote xref offsets of -1; a palette with
transparent indices in more than one run got a four-number /Mask array;
style in a tagged document opened an artifact it never closed; a
paragraph that paginates from a page with no room was refused instead of moved; two font aliases differing by a space shared one font resource;
the IDTree of a structure tree was not in byte order; ua -part 2 could
be lowered again and ua 0 broke the write; a refused zugferd call left a
PDF/A claim behind; /Trapped was a string; and dozens of checks that
used to come after the first byte was written now come before it, so a
refused call leaves stream, tree and resources as they were.
* Behaviour that changed with it: Arabic-Indic numbers follow UAX #9
W2/W4/W5 (a European digit after an Arabic letter is Arabic; separators
join Arabic digits only between two of them), the type-area margins
follow -unit, an axis value outside its range is refused, -numCopies
takes 2 to 5, and unknown events, unknown -style words and unknown -rule values are refused; font info postScript names the instance a variable
face is embedded as, and the byte order mark U+FEFF is dropped by the
standard faces as by an embedded one.
Notes:
* Tested against Tcl 8.6.18 and Tcl 9.0.4 on macOS: 1492 tests green
under both, 41 examples writing 43 documents, every one of them accepted
by "qpdf --check" and, where it claims PDF/A or PDF/UA, by veraPDF
against the profile it claims. Every command, option and value the
manual names has a test of its own; "make check" runs the whole
acceptance - suite under every interpreter, examples, qpdf, veraPDF,
pdfinfo, and whether the manual is newer than its source.
* Installing: the release archives tclpdf1.1.zip and tclpdf1.1.tar.gz
hold one directory tclpdf1.1 with the modules, pkgIndex.tcl and the ICC profiles - nothing to build; extract it into a directory on your
auto_path and "package require tclpdf" finds it. The source archive
builds the TEA way in the source directory: ./configure, make test, make install. A separate build directory configures cleanly but cannot run
the tests, because the generated pkgIndex.tcl locates the modules
through $dir while the sources stay where they are.
* tdom is required for the XMP metadata packet, so every document that declares PDF/A, PDF/UA or ZUGFeRD needs it; a document with no such
claim runs without it. The SVG module uses it as well when present: it
parses 27 to 48 times faster than the parser tclpdf brings along and
refuses an entity expansion bomb, and without it SVG still works through
the built-in element tree parser. tzint is optional and only for
barcodes. Nothing else is used.
* No font is part of the package. The faces under examples/assets/fonts
exist so that the tests and examples have something to embed and travel
in the source archive alone; "make install" installs none. Whether you
may embed a font into a PDF that tclpdf writes is a question for that
font's licence, not for this one.
* Progressive JPEG and interlaced (Adam7) PNG are refused with a message naming the reason and the way out; save the file baseline or
non-interlaced. GPOS mark attachment is not read: text that arrives
composed (NFC) is unaffected, decomposed accents come out displaced - normalise to NFC first.
* What tclpdf covers is roughly a third of ISO 32000-1, weighted by the
page count of its chapters. The remainder is almost entirely what a
reader of foreign PDFs needs rather than a writer; encryption, form
fields (AcroForm) and multimedia are not implemented.
Planned:
* Reading foreign PDFs, so that an existing letterhead can be taken over
as a background instead of being rebuilt. That needs the parser side of
the format - cross-reference and object streams above all, plus the
filters a writer never emits (LZW, CCITT, RunLength) - and it is the
largest single piece still ahead.
* The bidirectional algorithm (UAX #9) for lines that mix directions -
an Arabic sentence with a Latin word in it - which are refused today and
have to be set as separate calls.
* Encryption. The standard security handler needs MD5 and either RC4 or
AES, and none of them is available here without either a rebuilt TclTLS
or a pure-Tcl implementation shipped along; which of the two it becomes
is still open. Note that PDF/A-3 and therefore ZUGFeRD forbid encryption anyway.
* Drawing a Tk canvas into a PDF - every item type, as vectors rather
than as a screenshot. This is the one planned feature that involves Tk
at all, and it stays outside the core: a caller who has a canvas has
loaded Tk already, while the package itself neither requires nor loads it.
None of this is promised for a particular release, and nothing on the
list is needed for what 1.1 does.
Thanks to everyone who reported findings on the mailing list - the
missing glyph that a font viewer showed all the same, the "unknown
subcommand" that came from an older installation shadowing a newer tree,
and the unreadable modules in the tarball all became either code, a
manual paragraph or a note above.
Deutsch
-------
Neu seit 1.0:
* Kerning und Ligaturen: Paarkerning aus GPOS oder der kern-Tabelle, die Standardligaturen (liga) aus GSUB, beide standardmaessig an und je
Aufruf abschaltbar (-kerning 0, -ligatures 0); Lookup-Flags und
GDEF-Klassen werden beachtet, sodass ein kombinierender Akzent zwischen
zwei Buchstaben kein Paar mehr unterbricht. Eine Ligatur extrahiert als
die Buchstaben, aus denen sie entstand. -spacing zaehlt Glyphen, nicht
Zeichen - so wie der PDF-Operator.
* Tagged PDF: ein Strukturbaum mit den 49 Standardtypen (1.7) und den 2.0-Ergaenzungen, die Verschachtelung gegen Anhang L geprueft, Artefakte draussen gehalten und nach Art benannt
(Pagination/Layout/Page/Background mit Header/Footer/Watermark),
Abkuerzungen mit ihrer Langform, Strukturattribute (-lang, -title, -actualText, -scope, -numbering, -bbox, -colSpan, -rowSpan),
Elementkennungen (-id, geschrieben als /ID und IDTree; eine Note bekommt
eine von selbst) und Strukturziele, auf die Links und Lesezeichen mit -structure zeigen. text, table und image taggen sich selbst; -tag nimmt
alles andere. Ein Absatz, der ueber Seiten umbricht, bleibt ein Element
mit einer Marke je Seite.
* PDF/A Stufe A neben B und U, und PDF/UA-1 und PDF/UA-2 mit der Well-Tagged-PDF-Erklaerung ($doc ua): die Zusage wird verweigert, wenn
das Dokument sie nicht haelt - Titel, Sprache, jede Schrift eingebettet, Ueberschriften ohne Luecke, eine Beschreibung an jedem Link und Bild,
Listen, deren Nummerierung und Marken zusammenpassen -, und jeder
Verstoss wird auf einmal gemeldet und nennt den Aufruf, der zu aendern
ist. Jedes Beispiel besteht veraPDF gegen das Profil, das es beansprucht.
* Text von rechts nach links: -direction rtl fuer Zeilen, Absaetze, Fuehrungszeilen, Text auf Pfad, Seitenzahlen und Tabellenzellen;
Hebraeisch allein durch die Reihenfolge, Arabisch mit seinen kontextabhaengigen Formen (isol/init/medi/fina, ccmp, rlig) aus den GSUB-Tabellen der Schrift, glyphenweise gegen HarfBuzz geprueft; Ziffern behalten in einer solchen Zeile ihre Reihenfolge, paarige Klammern
werden gespiegelt und extrahieren wie geschrieben. Eine Schrift, die
fuer ihre Punkte GPOS-Markenpositionierung braucht, wird mit Nennung abgelehnt, eine Schrift, die Shaping braucht, das dieses Paket nicht hat (Devanagari, Thai und Verwandte), wird abgelehnt statt falsch
gezeichnet, und -unshaped 1 schaltet die Ablehnung ab fuer den, der
isolierte Glyphen besser findet als nichts. Eine Zeile mit gemischter
Richtung wird abgelehnt - der Bidi-Algorithmus ist geplant.
* Drei Schriftformate: TrueType als Subset (wie bisher), OpenType mit CFF-Umrissen ganz (/FontFile3), Type 1 ganz (.pfb, .pfa, .t1 mit seiner
AFM); eine variable Schrift wird als beliebige Instanz ihrer Achsen eingebettet (-axes, -instance) mit fuer die Instanz berechneten
Umrissen, Vorschueben, Begrenzungsboxen und hhea-Werten; font info sagt,
was die Datei ueber sich sagt. Type 1 und die vierzehn Standardschnitte
werden ueber WinAnsi angesprochen, mit dem geschuetzten Leerzeichen und
dem weichen Trennstrich dort, wo Anhang D sie hinlegt. Ein Zeichen, fuer
das die Schrift keine Glyphe hat, ist ein Fehler, und das Handbuch sagt
jetzt, warum ein Schriftbetrachter kein Beleg fuer den Zeichenvorrat ist.
* Textfluss: -height gibt zurueck, was nicht passte, -height max misst
bis zum Satzspiegel, den das Dokument bekommen hat (-typeArea {oben
unten ?links rechts?}, page typeArea liest ihn), -paginate 1 fuehrt die Schleife selbst ueber so viele Seiten wie noetig - in so vielen Spalten
wie verlangt (-columns n) mit -gutter und -balance -, waehrend pageAdded
je Seite feuert; Text fliesst um Rechtecke und Kreise (-avoid), weiche Trennstriche (U+00AD) brechen, wo sie stehen, und erreichen nie eine
Schrift, ein Trennstrich wird geschrieben, wo ein Wort geteilt wurde, Fuehrungszeilen (leader), Text entlang eines Pfads (textPath),
Seitenzahlen beim Schreiben gezeichnet (pageNumbers), Erstzeilen- und Rechtseinzug, Absatzabstand.
* Grafik: sechzehn Mischmodi als Zustand (blend) und je Form (-blend),
style haelt jede Option einschliesslich der beiden Farben - eine Form
malt die Seiten, die style gefaerbt hat, plus die, die sie selbst nennt
-, und der Zustand gehoert dem Strom, sodass save/restore, Formulare und Seiten ihn auseinanderhalten; transform setzt mehrere Teile zusammen wie dieselben Aufrufe nacheinander, sodass eine Verschiebung nicht mehr mitskaliert oder mitgedreht wird; der Satzspiegel fuer Tabellen und
Text; PDF/A-Farbraeume werden beim Schreiben gegen den Output-Intent
gehalten, und ein CMYK-Dokument besteht mit basICColors ISO Coated v2,
ein Graudokument mit dem Grau-Intent, beide mitgeliefert.
* SVG: Text in einer Zeichnung benutzt dieselben Schriften wie der Rest
des Dokuments, eingebettete eingeschlossen, mit Kerning und Ligaturen;
svg -data nimmt das Markup aus einer Variablen; eine Zeichnung, die
scheitert, laesst weder ein offenes q noch eine Schriftressource
zurueck. Barcodes ueber tzint (Code 128, EAN, QR, DataMatrix, PDF417 und
100 weitere), dessen SVG tclpdf als Vektoren zeichnet - die
Klartextzeile unter einem Barcode kann echtes OCR-B sein, als Text.
* Bilder: image embed und image draw nehmen -data mit den Bytes selbst;
ein PNG mit einer transparenten Farbe behaelt sie als
Farbschluessel-Maske auf dem Durchreichweg, ein 16-Bit-Schluessel wird
zur Maske, weil Leser ihn sonst ignorieren; image info sagt, wie die Transparenz in die Datei kommt (none, colourKey, softMask), und image
size, was eine Platzierung messen wuerde.
* Tabellen: -top sagt, wo eine umbrechende Tabelle fortsetzt, und
-bottom, wie weit sie laufen darf, beide mit dem Satzspiegel als
Vorgabe; eine Zeilengruppe wandert geschlossen auf die naechste Seite;
-border outer rahmt die Tabelle je Seite; Themen plain/striped/grid;
vier Haken je Zelle; valign, Dezimalausrichtung ueber Kopf, Rumpf und
Fuss gemeinsam gemessen.
* Dokumente und Dateien: Dateianhaenge mit Beziehung, Beschreibung und Aenderungszeit als Kernfunktion; Links auf eine URL, eine Seite oder ein Strukturelement; Lesezeichen; Seitenbeschriftungen; Anzeigeeinstellungen
mit ihren Versionsregeln; Dokumentsprache gegen RFC 3066 geprueft; die
sechs Text-Info-Eintraege und die beiden Daten dem XMP-Paket so
zugeordnet, wie ISO 32000-1 Tabelle 317 es tut (Creator ist
xmp:CreatorTool, Keywords ist pdf:Keywords, CreationDate ist
xmp:CreateDate), waehrend Trapped allein im Info-Woerterbuch bleibt; die Version wird nie stillschweigend erhoeht, und jedes Merkmal wird vor dem Schreiben gegen sie geprueft; das Schreiben ist idempotent, ein zweites Schreiben eines unveraenderten Dokuments ist byteweise gleich.
* ZUGFeRD / Factur-X: -type und -profile werden geprueft (INVOICE/ORDER; MINIMUM, BASIC WL, BASIC, EN 16931, EXTENDED, XRECHNUNG), zugferd state
traegt die geschriebene Beziehung, und pdfa -conformance U oder A darf
ueber dem 3B der Rechnung beansprucht werden.
* Order-X (Order-X 1.0, 4.1.1 und 4.1.2) laeuft ueber denselben zugferd-Aufruf: die BT-24-Kennungen urn:order-x.eu:1p0:basic, :comfort
und :extended werden als BASIC, COMFORT und EXTENDED gelesen, das
Dokument wird ein ORDER (oder ORDER_RESPONSE, ORDER_CHANGE ueber -type),
die XML wird als order-x.xml mit /AFRelationship Data als Vorgabe
eingebettet, und das Erweiterungsschema traegt den Namespace-URI von
Order-X; Stufen, Typen und Namen sind an ihre Familie gebunden, sodass
eine Rechnung nicht als Bestellung hinausgeht oder umgekehrt.
* Erweiterungspunkte: doc/PLUGINS.md beschreibt den Ereignisbus
(beforeWrite, afterWrite, pageAdded, resources, catalog, info), wie sich
ein Themenmodul mit oo::define anhaengt, und was zugesagt ist und was nicht.
* Beispiele: 41, die 43 Dokumente schreiben, darunter ein Satzspiegel-
und Umbruchbeispiel, Schriftsysteme in zehn Noto-Schnitten, Barcodes,
variable Schriften, Type 1 und OpenType, getaggte Dokumente, PDF/UA-2
mit WTPDF und eine ZUGFeRD-Rechnung, die jede Option durchprobiert; die Schriften dafuer reisen im Quellarchiv mit, nach Herkunft sortiert und
mit ihren Lizenzen.
Geaendert seit 1.0:
* transform mit mehreren Teilen in einem Aufruf setzt sie jetzt zusammen
wie dieselben Aufrufe nacheinander (translate, rotate, skew, scale): die Verschiebung wird nicht mehr mitskaliert oder mitgedreht. Ein Aufruf mit
-at und -rotate allein ist unberuehrt.
* style -fill und -stroke gelten bis zur naechsten Aenderung wie die
uebrigen Optionen: eine nackte Form nach "style -fill red -stroke blue"
wird gemalt; vorher wurde sie gebaut und verworfen. line und curve
zeichnen nur dann schwarz, wenn kein style-Strich steht.
* Ein Zahlenpaar als -format wird auch dann gedreht, wenn die
Orientierung auf tclpdf new oder configure verlangt wurde, nicht nur auf
page add; die Vorgabe portrait ist keine Anforderung, ein Paar ohne jede Orientierung bleibt also, wie es ist.
* Kerning und Ligaturen sind standardmaessig an, sodass jedes Dokument
mit eingebetteter Schrift anders setzt als unter 1.0 - gemessen
aenderten sich neun der damals fuenfundzwanzig Beispiele, genau die mit Einbettung; PDF/A- und ZUGFeRD-Ausgabe bleibt konform.
Behoben seit 1.0:
* Blocksatz, gedrehter und verankerter Text: die
Ausrichtungsverschiebung lief entlang der gedrehten Grundlinie, und der Ascender ging im Rueckgabewert eines Absatzes verloren; der Wortabstand
im Blocksatz wird ueber TJ gesetzt, sodass die letzte Glyphe an der
Kante landet.
* Eine Schrift, deren PostScript-Name in der name-Tabelle unter
Plattform 0 steht (Times New Roman und Arial Narrow unter den 20 von 210 Schriften dieses Rechners), erzeugte einen /BaseFont mit Nullbytes, den
qpdf ablehnte; Plattform 0 wird wie Plattform 3 als UTF-16 dekodiert.
* Eine umbrechende Tabelle setzte auf der Folgeseite dort fort, wo sie
auf der ersten begonnen hatte; -top und -bottom leiten sich von der
Seite ab, nicht von -at. valign top sass auf der Linie, wenn die AFM
Ascender 0 sagte.
* Text auf Pfad mass mit Kerning und zeichnete ohne, und mass ein
-spacing, das er nicht zeichnete; ein leerer String mit -wordSpacing
hatte eine negative Breite; ein Coverage-Index war ein laufender Zaehler
statt der der Tabelle.
* Ein SVG, das scheiterte, liess sein q offen, sodass alles danach
kopfueber und vergroessert stand, und liess eine Schriftressource
zurueck, wegen der pdfa das Dokument verweigerte.
* SVG-Verlaufsmuster trugen die SVG-id als Namen, sodass zwei
Zeichnungen mit derselben id einen Verlauf teilten.
* Der Subset-Praefix folgt den benutzten Glyphen, dem Namen und dem Achsenpunkt, sodass zwei Subsets einer Datei in einem Dokument
verschiedene Kennungen tragen und zwei Schreibvorgaenge dieselbe.
* Das XMP-Paket wird bei jedem Schreiben neu gebaut, sodass ein zwischen
zwei Schreibvorgaengen geaenderter Titel oder eine geaenderte Sprache
die Datei erreicht; Creator gibt sich in xmp:CreatorTool nicht mehr als Producer aus.
* checkSumAdjustment eines eingebetteten Subsets war in jeder Datei 0 -
der Kommentar sagte, es werde neu berechnet, und es wurde nicht;
Formulare, Muster und Seitenzahl-Formulare tragen eigene /Resources, wie
PDF/A 6.2.2 verlangt; ein 16-Bit-PNG-Farbschluessel ist eine Maske, kein /Mask-Array, das Leser ignorieren.
* Release-Archive sind fuer jeden lesbar: 37 Module reisten als 0600,
weil die Modi des Checkouts durchgereicht wurden.
* Eine dritte Normpruefung vor 1.1 schloss, was kein Validator sieht:
der xpacket-Kopf trug eine doppelt kodierte Byte-Order-Marke;
writeChannel in eine Pipe oder einen Socket schrieb xref-Offsets von -1;
eine Palette mit transparenten Indizes in mehr als einem Lauf bekam ein /Mask-Array aus vier Zahlen; style in einem getaggten Dokument oeffnete
ein Artefakt, das es nie schloss; ein Absatz, der von einer Seite ohne
Platz aus umbricht, wurde abgelehnt statt verschoben; zwei
Schriftaliase, die sich um ein Leerzeichen unterscheiden, teilten eine Schriftressource; der IDTree eines Strukturbaums war nicht in
Bytereihenfolge; ua -part 2 liess sich wieder senken, und ua 0 brach das Schreiben; ein abgelehnter zugferd-Aufruf liess eine PDF/A-Zusage
zurueck; /Trapped war ein String; und Dutzende Pruefungen, die frueher
nach dem ersten geschriebenen Byte kamen, kommen jetzt davor, sodass ein abgelehnter Aufruf Strom, Baum und Ressourcen laesst, wie sie waren.
* Verhalten, das sich damit geaendert hat: arabisch-indische Zahlen
folgen UAX #9 W2/W4/W5 (eine europaeische Ziffer nach einem arabischen Buchstaben ist arabisch; Trennzeichen verbinden arabische Ziffern nur
zwischen zweien), die Satzspiegelraender folgen -unit, ein Achsenwert ausserhalb seines Bereichs wird abgelehnt, -numCopies nimmt 2 bis 5, und unbekannte Ereignisse, unbekannte -style-Woerter und unbekannte
-rule-Werte werden abgelehnt; font info postScript nennt die Instanz,
als die eine variable Schrift eingebettet wird, und das Byte-Order-Mark
U+FEFF lassen die Standardschriften fallen wie eine eingebettete.
Hinweise:
* Geprueft gegen Tcl 8.6.18 und Tcl 9.0.4 unter macOS: 1492 Tests gruen
unter beiden, 41 Beispiele mit 43 Dokumenten, jedes davon von "qpdf
--check" angenommen und, wo es PDF/A oder PDF/UA beansprucht, von
veraPDF gegen das beanspruchte Profil. Jeder Befehl, jede Option und
jeder Wert, den das Handbuch nennt, hat einen eigenen Test; "make check" faehrt die ganze Abnahme - Suite unter jedem Interpreter, Beispiele,
qpdf, veraPDF, pdfinfo, und ob das Handbuch neuer ist als seine Quelle.
* Installation: die Release-Archive tclpdf1.1.zip und tclpdf1.1.tar.gz enthalten ein Verzeichnis tclpdf1.1 mit den Modulen, pkgIndex.tcl und
den ICC-Profilen - nichts zu bauen; in ein Verzeichnis auf dem auto_path entpacken, und "package require tclpdf" findet es. Das Quellarchiv baut
auf TEA-Art im Quellverzeichnis: ./configure, make test, make install.
Ein eigenes Bauverzeichnis konfiguriert sauber, kann die Tests aber
nicht ausfuehren, weil die erzeugte pkgIndex.tcl die Module ueber $dir
sucht, waehrend die Quellen liegen bleiben.
* tdom wird fuer das XMP-Metadatenpaket gebraucht, also von jedem
Dokument, das PDF/A, PDF/UA oder ZUGFeRD erklaert; ein Dokument ohne
solche Erklaerung laeuft ohne. Das SVG-Modul benutzt es ausserdem, wenn
es vorhanden ist: es parst 27- bis 48-mal schneller als der
mitgelieferte Parser und wehrt eine Entity-Bombe ab, und ohne tdom
laeuft SVG weiter, ueber den eingebauten Elementbaum-Parser. tzint ist optional und nur fuer Barcodes. Sonst wird nichts benutzt.
* Keine Schrift gehoert zum Paket. Die Schnitte unter
examples/assets/fonts sind da, damit Tests und Beispiele etwas zum
Einbetten haben, und reisen allein im Quellarchiv mit; "make install" installiert keine. Ob eine Schrift in ein von tclpdf geschriebenes PDF eingebettet werden darf, entscheidet die Lizenz jener Schrift, nicht
diese hier.
* Progressive JPEG und verschraenktes (Adam7) PNG werden mit einer
Meldung abgelehnt, die Grund und Ausweg nennt; die Datei baseline beziehungsweise nicht verschraenkt speichern. GPOS-Markenpositionierung
wird nicht gelesen: komponiert ankommender Text (NFC) ist unberuehrt,
zerlegte Akzente kommen versetzt heraus - vorher nach NFC normalisieren.
* Was tclpdf abdeckt, ist rund ein Drittel von ISO 32000-1, gewichtet
nach dem Seitenumfang der Kapitel. Der Rest ist fast durchweg das, was
ein Leser fremder PDFs braucht und kein Schreiber; Verschluesselung, Formularfelder (AcroForm) und Multimedia sind nicht umgesetzt.
Geplant:
* Fremde PDFs lesen, damit ein vorhandener Briefbogen als Hintergrund uebernommen statt nachgebaut wird. Dafuer fehlt die Leseseite des
Formats - vor allem Querverweis- und Objektstroeme, dazu die Filter, die
ein Schreiber nie erzeugt (LZW, CCITT, RunLength) - und es ist der
groesste noch offene Einzelposten.
* Der bidirektionale Algorithmus (UAX #9) fuer Zeilen mit gemischter
Richtung - ein arabischer Satz mit einem lateinischen Wort darin -, die
heute abgelehnt werden und als getrennte Aufrufe zu setzen sind.
* Verschluesselung. Der Standard-Sicherheitshandler braucht MD5 und dazu
RC4 oder AES, und keines davon ist hier zu haben ohne ein neu gebautes
TclTLS oder eine mitgelieferte Umsetzung in reinem Tcl; welcher der
beiden Wege es wird, ist offen. Zu beachten: PDF/A-3 und damit ZUGFeRD verbieten Verschluesselung ohnehin.
* Ein Tk-Canvas ins PDF zeichnen - jeder Elementtyp, als Vektoren statt
als Bildschirmfoto. Das ist das einzige geplante Merkmal, in dem Tk
ueberhaupt vorkommt, und es bleibt ausserhalb des Kerns: wer einen
Canvas hat, hat Tk ohnehin geladen, waehrend das Paket selbst es weder verlangt noch laedt.
Nichts davon ist fuer eine bestimmte Fassung zugesagt, und nichts auf
dieser Liste wird fuer das gebraucht, was 1.1 leistet.
Dank an alle, die auf der Mailingliste Befunde gemeldet haben - die
fehlende Glyphe, die ein Schriftbetrachter trotzdem zeigte, das "unknown subcommand" aus einer aelteren Installation, die einen neueren Baum
verdeckte, und die unlesbaren Module im Tarball wurden alle entweder
Code, ein Handbuchabsatz oder ein Hinweis oben.
--- Synchronet 3.22a-Linux NewsLink 1.2