This is background reading for anyone working on the letters subsystem's finalization
pipeline (HTML → PDF) or the email-attachment path. It's conceptual, not a how-to —
see overview.md for the subsystem's architecture and
the codenforce repo's docs/subsystems/letters+emailing/ for phase-by-phase implementation
history. Nothing here is CodeNForce-specific trivia you can find by reading the code once —
it's the "why does this keep breaking" context that isn't obvious from the source alone.
| Library | Version | Role |
|---|---|---|
| openhtmltopdf-pdfbox | 1.0.10 | Takes well-formed XHTML + CSS and lays it out as a PDF content stream — the actual HTML→PDF renderer. |
| Apache PDFBox | 2.0.22 | Low-level PDF object model: reads/writes the actual PDF file structure (objects, xref table, streams, document-info dictionary). openhtmltopdf uses it as its output backend; we also use it directly in BlobCoordinator to strip metadata from uploaded PDFs. |
| jsoup | 1.18.1 | HTML tidying/parsing. Rich-text-editor output (from LetterHtmlRenderer-generated content) is rarely well-formed XML. jsoup parses it leniently, then re-emits it as strict XHTML so openhtmltopdf's XML parser won't choke on it. |
| resend-java | 4.13.0 | Email transport (Resend API client). Not a PDF library, but it's the other place binary document bytes get encoded/transported in this subsystem — see §6. |
A note so you don't go down a rabbit hole: pom.xml still defines an <itext.version>RELEASE</itext.version>
property, but there is no actual iText dependency anywhere in the project. It's vestigial —
grep for com.itextpdf before assuming iText is involved in anything; it isn't.
It's a pure-JVM library — no native binary, no headless-browser subprocess to manage on
WildFly, no sandboxing/process-lifecycle concerns. The tradeoff (see §4) is that it only
implements a subset of CSS and does not execute JavaScript or fetch network resources during
render, which shapes several of our design decisions.
Code entry points, in pipeline order:
LetterCoordinator.letter_generateFinalizationPDF(Letter, UserAuthorized) — the publicLetterFlowBB after finalization).LetterCoordinator.letter_buildPdfReadyHtml(...) — builds the full XHTML document stringLetterCoordinator.letter_inlinePhotoImages(...) — regex-scans for /tcvce/photo/{id} srcdata:<mime>;base64,<bytes> URI, pulling the actual bytesBlobCoordinator.getBlob(id).LetterCoordinator.letter_renderPdfBytes(...) — the actual PdfRendererBuilder call.LetterCoordinator.letter_loadPrintCss() — loads /css/letter-print.css from the classpath.Why inline photos as base64 data URIs instead of just pointing at the URL? openhtmltopdf's
render happens synchronously inside the app server, with no live HTTP request/response cycle
backing it and no default resolver that can authenticate against and hit our own running
servlet. Handing it a data: URI means the image bytes are already in the document it's
parsing — no I/O, no auth, no network dependency at render time.
A PDF is not a picture of a page — it's a program, in a very restricted sense. Conceptually:
xref) at the end of the file maps object numbers to byte offsets soPages/Page objects, each Page referencing itsPDF was Adobe's proprietary format from 1993 until 2008, when Adobe submitted the PDF 1.7
specification to ISO, which published it as ISO 32000-1:2008. That was the moment PDF
stopped being "whatever Adobe's spec document says" and became an independently-governed
standard. In 2017, ISO published ISO 32000-2:2017 ("PDF 2.0") — the first version
developed directly by the ISO committee (not adapted from an existing Adobe spec), adding
things like better accessibility/tagging support, more robust encryption, and cleanup of
ambiguous corners of the 1.7 spec. PDFBox and openhtmltopdf both target the older,
vastly-more-common 1.x object model; we are not doing anything PDF-2.0-specific here.
openhtmltopdf implements a fixed, print-oriented CSS subset — treat it like a well-behaved
CSS 2.1 engine with a handful of CSS3 additions (@page, some @font-face, basic
box-shadow/border-radius), not like a real browser:
letter-print.css is deliberately simple (percentage-width inline-block<table> for the violations table) rather than reusing any of the app'sletter-print.css@page controls page geometry, not a wrapper <div>. Page size/margins/orientation live@page at-rule (letter-print.css's @page { size: letter portrait; margin: 1in; }),.replace("size: letter portrait;", "size: letter landscape;") inLetterCoordinator.letter_buildPdfReadyHtml() — fragile if the CSS text ever gets<p>&, mismatched nesting, etc.; openhtmltopdf's parser (by default) does not.letter_buildPdfReadyHtml() — it'spt, in) over viewport-relative units (vw, % of anrem tied to a browser default) — there's no real "viewport" onceletter-print.css intentionally sticks to the generic families serif/sans-serif rather
than a specific webfont, so PDFBox can satisfy every glyph request from its built-in
base-14 fonts — no font files to embed, no licensing to track, smaller output files.
The catch: the base-14 fonts only guarantee WinAnsiEncoding/StandardEncoding coverage —
essentially Latin-1/Windows-1252. That covers standard English text and most Word-pasted
"smart quotes"/em-dashes fine, but:
PDType0Font (composite/font-family value, since it means shippingThere are two completely separate places byte[] gets turned into base64 text in this
subsystem, and they solve different problems:
LetterCoordinator.letter_inlinePhotoImages() — base64-encodes image bytes into adata: URI so they can be embedded inside the HTML that openhtmltopdf renders. ThisCommunicationCoordinator's Resend attachment mapping (~line 1236) — base64-encodesEmailAttachment.content) because Resend's APIAttachment.Builder ab = Attachment.builder()
.filename(att.getFilename())
.content(Base64.getEncoder().encodeToString(att.getContent()));
This base64 text is transport-only — Resend decodes it back to bytes before actuallyIf you're debugging an attachment/image issue, figure out which of these two encodings
you're actually looking at before changing anything — they have different lifetimes, different
consumers, and fixing one by copying a pattern from the other is a common way to introduce a
new bug.
Every PDF carries a small metadata dictionary (Author/Title/Subject/Keywords/Creator/Producer/
CreationDate/ModificationDate) alongside the actual page content — separate from the visible
text, invisible in any normal viewer's page display, but trivially readable by anyone who opens
the file's properties panel (or just great-search greps the raw bytes; older PDFs often store
this as plain, uncompressed text).
This matters for uploaded PDFs (not ones we generate) because that metadata was written by
whatever software the original author used, and may leak information nobody intended to
share (real names, internal file paths in the Creator/Producer fields, original creation
timestamps that predate what the case record implies, etc.). BlobCoordinator scrubs this on
ingest:
PDDocument doc = PDDocument.load(input.getBytes());
PDDocumentInformation docInfo = doc.getDocumentInformation();
blobMeta.setProperty(new MetadataKey("Author"), docInfo.getAuthor());
docInfo.setAuthor("");
// ...same pattern for Title/Subject/Keywords/Creator/Producer/CreationDate/ModificationDate...
doc.save(output); // re-saved with the info dictionary cleared
input.setBytes(output.toByteArray());
Note it extracts the original values into our own Metadata/MetadataKey model before
blanking them in the PDF itself — so the information isn't lost, just moved somewhere we
control rather than left embedded in a file we're about to store and potentially hand back out.
Two things worth knowing if you touch this code:
PDDocumentInformationdoc.save() produces a clean, fully-rewritten file here (PDFBox's save() isn'tdata: URIs. There's no live request context during/tcvce/photo/{id} as a plain URL will not resolve.pom.xml) — bumping PDFBoxitext.version pom property is vestigial — no iText dependency actually exists in| File | Purpose |
|---|---|
LetterCoordinator.java (letter_generateFinalizationPDF and the private letter_*Pdf*/letter_inlinePhotoImages/letter_loadPrintCss methods) |
Orchestrates the whole HTML→PDF pipeline. |
/css/letter-print.css (classpath resource, src/main/resources/css/) |
The print stylesheet injected before rendering — the CSS subset actually exercised by this pipeline. |
BlobCoordinator.java (PDF metadata scrub, near PDDocument.load(...)) |
Strips Document Information dictionary fields from uploaded PDFs on ingest. |
CommunicationCoordinator.java (Resend attachment mapping, ~line 1236) |
Base64-encodes finished document bytes for email transport — unrelated to the data-URI encoding used during rendering. |
EmailAttachment.java / EmailMessage.java |
Plain POJOs carrying attachment bytes/metadata between the coordinator and the Resend client. |
pdfbox.apache.org) — the object-model API we useBlobCoordinator.github.com/openhtmltopdf/openhtmltopdf) — its README/wiki isjsoup.org) — for the HTML-tidying API used in letter_buildPdfReadyHtml.