Rodney Dela Cruz

Mapping placeholders into .docx templates

Rodney Dela Cruz

Technical Co-founder, eSerbisyo

  • document-automation
  • templates

Generating a document from a template and some data looks like a solved problem: put {{name}} in the template, replace the placeholder, ship the file. It stops looking solved the first time a template has a header on page two, a table whose row count depends on the data, or a field someone typed by hand that your replacement logic now overwrites.

The replacement is the easy part. Everything around it is where document automation earns or loses its reputation.

I co-founded eSerbisyo, where generating DOCX outputs from verified records is one of the jobs the system has to do for barangay documents. What follows is the set of constraints we had to respect to keep those outputs usable.

What a .docx file actually is

A .docx is a zip archive of XML. The content you see is one part of it, and the parts interact:

  • document.xml holds the body flow.
  • header*.xml and footer*.xml are separate parts, and their contents are not in the body.
  • Numbering, styles, and relationships live in their own parts.

The consequence that bites first: a text replacement over the body does not touch the header. A template with the applicant name in the header keeps the old name, and it looks correct until someone compares it to the record.

Word splits text into runs (<w:r> elements). A run is the unit with its own formatting (bold, font, colour). When you edit a document, the word processor creates new runs rather than rewriting characters in one block. What that means for templates is that there is no guarantee the string you type as {{name}} exists as a single contiguous sequence of characters in any single run. The placeholder can be cut across two runs, and the tooling has to account for it.

Why naive string replacement breaks documents

Replacing text at the character level assumes the placeholder is intact and uniquely identifiable. In practice:

  • Word splits a run of text across several XML runs. Typing normally, editing afterwards, or a spell-checker pass can break {{name}} into {{na, me}}.
  • The same placeholder may legitimately appear several times, and “first occurrence” is not a rule anyone wants to explain to an auditor.
  • Replacing across a run boundary can produce a file that opens but has lost formatting on the substituted text.
  • Nested formatting (a placeholder inside bold inside a table cell) can create a chain of runs rather than one.

The robust approach is to parse the XML, recognise placeholders that span run boundaries, and merge or replace runs as needed. It is more code, and it is necessary. Any implementation that treats the document as a big string will occasionally corrupt formatting.

Finding placeholders across runs

The workable strategy is not to regex over raw XML, but to walk the text elements and accumulate characters. When you see {{, start collecting until you see }}. If the start and end are in different runs, the collector must glue their w:t (text) content together and can then remove or rewrite those runs in one go. Only once the whole token is formed do you substitute it.

That preserves styles. If the placeholder was bold, the substituted value should inherit that style unless the substitution has its own rule. For a system generating official documents, the template’s formatting must dominate — the value fills the slot, the typography does not change arbitrarily.

Tables are where templates usually break

A table whose row count depends on the data is the most common source of “the generator broke my template”. Two workable approaches:

  • Template rows as a pattern. One marked row is the prototype; the generator clones it per record, clears the placeholders, and appends. This preserves all the cell formatting, borders, and column widths the designer set up.
  • Build the table programmatically. More control, and you lose the designer’s formatting unless you rebuild it deliberately.

For templates that non-technical staff edit, the first is almost always right. The designer stays in charge of the appearance. Also, table styles in Word are not trivial to reconstruct correctly: borders can be on cells, rows, or the table, and copying the prototype row copies the right thing if you do it in the XML, not by retyping.

Another table edge case: a row that should not appear if there is no data. Cloning an empty row produces an empty row in the output, which looks wrong. So the generator must be able to drop the prototype row if the source list is empty, or insert a “none” placeholder as specified by the template.

Headers, footers, and fields you might miss

Body-only replacement misses headers and footers. A document can have different headers on odd and even pages, a first-page header, and section-specific footers. Each is a separate part in the ZIP, so the generator must scan all of them, not just document.xml.

Other places to check:

  • Footnotes and endnotes. Rare, but when a template puts a case number in a footnote, body replacement will not touch it.
  • Text boxes and shapes. Some templates use DrawingML or VML to put text in shapes. If that is true for a template, it needs explicit handling. For the templates we use, we avoid this and prefer tables and paragraphs.
  • Bookmarks and content controls. Content controls (SDT, structured document tags) are useful because they name the field and can be used for find/replace, but they come with their own XML and are not always present. They are not a substitute for proper placeholder parsing.

The checks that prevent silent corruption

A document that opens is not a document that is correct. Worth asserting on every generation:

  • No unresolved placeholders remain. Scan the output for {{ and fail the job rather than shipping a file with {{applicant_name}} printed on it.
  • No empty required fields. A missing value should fail generation, not produce a blank line in an official document.
  • Header and footer values were substituted. Explicitly assert against the header and footer parts, because the body-only pass will not catch it.
  • The document opens. Round-trip it through a reader, not just a ZIP check. A ZIP can be structurally valid and still unreadable by Word.
  • No unexpected run splits. A cheap smoke test: the output is smaller in a sane way and the key fields are present. This catches the “replaced and broke all runs” class of bugs.

What this means for templates

Templates should be authored for the generator, not against it. A few rules save time:

  1. Put placeholders in normal text runs where possible, not split across edits.
  2. If a value can be long, use a table cell or a paragraph that allows wrapping.
  3. Name placeholders unambiguously — {{applicant.last_name}} rather than {{name}} that could mean three different things.
  4. Avoid content that spans very specific run-splitting behaviour. For example, editing the same token repeatedly in Word can create many tiny runs.
  5. Keep the prototype row as a single row, and mark it clearly (a column with a flag, or a comment) so it cannot be confused with data rows.

None of these require programming knowledge. They are editing habits, and they remove the most frequent failures.

The part that matters more than the generator

The document is not the deliverable. The record is. A generated PDF that cannot be traced back to the request that produced it is a piece of paper, and paper is what the process was trying to get away from.

So the ID of the request goes into the document, the same document is reachable from the record, and the verification path is a lookup rather than a re-read of the paper. Everything else in the generator is a means to that. The goal is fidelity: the document says what the record says, and a person can prove that without trusting the file itself.