CJK fonts in PDFs: Chinese, Japanese, and Korean rendering pitfalls

Generating PDFs with Chinese, Japanese, or Korean text looks easy until tofu boxes show up in production. Here's what actually breaks and how to fix it.

pdffontsinternationalizationcjknode.jspython

CJK fonts in PDFs: Chinese, Japanese, and Korean rendering pitfalls

If you've ever shipped a PDF generator to users in Asia, you've probably seen the dreaded tofu — those little empty rectangles where Chinese, Japanese, or Korean characters should be. CJK rendering is one of those problems that seems trivial ("just embed a font") until you hit the actual scale of the writing systems involved. Let's walk through why CJK PDFs break and what to do about it.

Why CJK is different from Latin scripts

A typical Latin font covers a few hundred glyphs. CJK is a different universe:

  • Chinese (Simplified): ~7,000 commonly used characters, 20,000+ for full coverage
  • Chinese (Traditional): similar range, with regional variants (Taiwan vs Hong Kong)
  • Japanese: Hiragana, Katakana, and ~2,000–10,000 Kanji (which overlap with but are not identical to Chinese Hanzi)
  • Korean: 11,172 precomposed Hangul syllables, plus Hanja for formal contexts

A complete CJK font file is often 5–20 MB. That single fact drives most of the problems you'll run into.

Pitfall 1: The default font has no CJK glyphs

Most PDF libraries default to one of the 14 PDF "standard fonts" (Helvetica, Times, Courier, etc.). None of them contain CJK glyphs. If you pass Chinese text to a default-configured PDFKit, ReportLab, or wkhtmltopdf, you'll get tofu — or worse, silent character dropping.

javascript
// PDFKit — this will produce tofu for CJK const doc = new PDFDocument(); doc.text('你好世界'); // ❌ default Helvetica has no Chinese // You must register a CJK font first doc.registerFont('NotoSC', 'fonts/NotoSansSC-Regular.ttf'); doc.font('NotoSC').text('你好世界'); // ✅

Pitfall 2: Subsetting that drops characters

To keep file sizes reasonable, most PDF generators subset fonts — embedding only the glyphs actually used. This is fine until your subsetting logic doesn't fully understand the input.

Common failure modes:

  • Late-bound text: Headers and footers rendered after subsetting locks in.
  • Form fields: Interactive PDF forms need glyphs for characters the user might type, not just what's there at generation time. For CJK forms, you generally need to embed the full font (or a much larger subset).
  • Mixed scripts in one run: A string like "注文 #A-1024" requires glyphs from two different fonts if you're switching fonts manually.

If you need form-fillable CJK PDFs, embed the full font and accept the file size hit.

Pitfall 3: Han unification ambiguity

Unicode unifies many Chinese, Japanese, and Korean ideographs into single code points even when the preferred glyph shape differs by region. The character U+8FD4 (返) renders differently in mainland Chinese, Taiwanese, Japanese, and Korean typography.

If you render Japanese text with a Simplified Chinese font, the characters will display but look wrong to a native reader — and they often won't notice it's "wrong," just that it looks off.

The fix is to pick a region-appropriate font:

  • Simplified Chinese: Noto Sans SC, Source Han Sans SC
  • Traditional Chinese: Noto Sans TC, Source Han Sans TC
  • Japanese: Noto Sans JP, Source Han Sans JP
  • Korean: Noto Sans KR, Source Han Sans KR

Detect the locale from your data (user profile, document language tag) and route to the correct font. Don't just pick one CJK font and call it done.

Pitfall 4: Vertical text and line breaking

Japanese and Traditional Chinese documents are sometimes typeset vertically (top-to-bottom, right-to-left columns). Most PDF generation stacks have no idea how to do this. If you need vertical text, you generally need a layout engine that supports CSS writing-mode: vertical-rl and a font with vertical metrics.

Line breaking is also different. CJK text can break between almost any two characters — there are no spaces. But certain characters (closing brackets, punctuation) must not start a new line. This is called kinsoku shori in Japanese typography. Naive word-break algorithms will produce ugly line breaks like leaving 。 (full stop) alone at the start of a line.

If your PDF library uses text-align: justify with naive whitespace-based breaking, you'll get bad CJK output. Look for libraries that implement Unicode line-breaking (UAX #14).

Pitfall 5: PDF/A and font embedding rules

If you're generating archival PDFs (PDF/A), every font must be fully embedded — no references, no system font fallbacks. With CJK fonts, this means each PDF can balloon by ~10 MB. Strategies:

  • Use font subsetting combined with PDF/A — most modern tools support this.
  • Pre-build a minimal subset for known-fixed text (e.g., template labels) and a fuller subset for user-generated content.
  • Reuse font streams across documents in the same archive when possible.

A working example

Here's a Python ReportLab snippet that handles a CJK invoice correctly:

python
from reportlab.pdfbase import pdfmetrics from reportlab.pdfbase.ttfonts import TTFont from reportlab.pdfgen import canvas pdfmetrics.registerFont(TTFont('NotoJ