Skip to content

Escape XML metacharacters in EPUB metadata and page titles - #2270

Merged
josevalim merged 1 commit into
elixir-lang:mainfrom
dajiaohuang:epub-escape-metadata
Sep 17, 2026
Merged

josevalim merged 1 commit into
elixir-lang:mainfrom
dajiaohuang:epub-escape-metadata

Conversation

@dajiaohuang

Copy link
Copy Markdown
Contributor

Summary

OEBPS/content.opf is written verbatim, without the numeric-entity pass that is applied to generated pages, and the EPUB templates interpolated project, version, language, author names and page titles into it unescaped. The same unescaped values were also written into the <title> element of every page header.

Any of those values containing &, < or > therefore produced a bare metacharacter in an XML document. With project: "A & B", content.opf contains:

<dc:title>A & B - 1.0.1</dc:title>

That is a fatal well-formedness error — :xmerl_scan reports {:error_scanning_entity_ref, ...} and readers reject the file outright, so the whole EPUB fails to open. authors: ["AT&T"] fails the same way, as does any extra page whose heading contains &, since the heading is used as the page <title>.

Changes

  • lib/ex_doc/formatter/epub/templates/content_template.eex — escape project, version, language, author names and static-file hrefs with h/1.
  • lib/ex_doc/formatter/epub/templates/title_template.eex — escape the cover page project and version.
  • lib/ex_doc/formatter/epub/templates/head_template.eex — escape the page title, project, version and language.
  • test/ex_doc/formatter/epub_test.exs — new test covering the package document, the cover page and a generated extra page; each document is parsed with :xmerl_scan, so malformed XML fails the test.
  • test/fixtures/ExtraPageWithAmpersand.md — fixture with & in its heading.

This follows the pattern of #2269, which escaped the manifest/nav ids and hrefs but left the metadata and page-title values untouched.

Validation

  • mix compile --warnings-as-errors — clean
  • mix format --check-formatted — clean
  • mix test — 1 doctest, 457 tests, 0 failures (456 tests before this change)

Reverting only the three templates makes the new test fail with {:fatal, {:error_scanning_entity_ref, ..., {:line, 3}, {:col, 46}}}, confirming the test exercises the defect rather than merely passing alongside it.

The EPUB templates interpolated project, version, language, author and
page-title values straight into XHTML and into the OPF package document,
which is written verbatim without the numeric-entity pass applied to
generated pages. A value such as "A & B" or "AT&T" therefore produced a
bare ampersand in content.opf and in every page header, making the EPUB
unparseable by any XHTML/XML reader.

Escape those values with h/1, matching what generated page content and
the nav/manifest ids already do.
@github-actions

Copy link
Copy Markdown

@josevalim
josevalim merged commit 9d4d337 into elixir-lang:main Sep 17, 2026
6 checks passed
@josevalim

Copy link
Copy Markdown
Member

💚 💙 💜 💛 ❤️

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants