How to extract embedded font subsets and text encoding from PDF files
Understanding Embedded Fonts and Text Encoding in PDFs
Extracting embedded font subsets and understanding text encoding from PDF files is critical for ensuring text fidelity, accurate data extraction, and proper document rendering across different systems. Without correctly identifying the fonts and their encoding, text can appear garbled, unsearchable, or fail to extract altogether.
What are Embedded Font Subsets
PDF files can embed fonts to ensure consistent rendering, regardless of whether the fonts are installed on the viewing system. Font subsets are partial versions of a font, containing only the characters actually used within the document. This significantly reduces file size compared to embedding the entire font family.
- Identification: Subset fonts are typically identified by a 6-uppercase-letter tag followed by a plus sign and the font name (e.g.,
AAAAAA+Calibri). - Purpose: They preserve the document's visual appearance and allow for text selection and copying.
- Limitations: Full font metrics and glyphs for unused characters are not included, meaning you cannot reconstruct the complete font from a subset alone.
Grasping Text Encoding in PDFs
Text encoding defines how character codes map to visual glyphs. PDFs use various encoding schemes, which are crucial for interpreting extracted text correctly. Misinterpreting the encoding can lead to "mojibake" (garbled text).
- Standard Encodings: Many PDFs use standard encodings like
WinAnsiEncoding(for Windows) orMacRomanEncoding(for macOS). - Custom Encodings: Fonts can have custom encodings defined directly within their font dictionaries.
- ToUnicode CMaps: For complex scripts, symbol fonts, or non-standard encodings, PDFs often include a
ToUnicodeCMap (Character Map). This map explicitly translates character codes from the font's internal encoding to Unicode values, ensuring universal interpretation.
Method One Command Line Tools
For a quick overview of fonts embedded in a PDF, command-line utilities are often the fastest approach. Tools from the Poppler-utils package are widely used for this.
Using pdffonts
pdffonts is a utility included in Poppler-utils that lists all embedded fonts, their types, encoding, and whether they are subsets.
Installation (Linux example):
sudo apt-get install poppler-utils
Command:
pdffonts your_document.pdf
Output Interpretation:
| Column | Description |
| name | The font name (e.g., AAAAAA+ArialMT). |
| type | Font type (e.g., TrueType, Type1, CIDType2). |
| emb | "yes" if embedded, "no" if not. |
| sub | "yes" if a subset, "no" if full. |
| enc | Encoding used (e.g., MacRoman, Ansi, Custom). |
While pdffonts provides essential metadata, it does not extract the actual font files or the detailed ToUnicode CMaps.
Method Two Programmatic Extraction with Python
For extracting font files, CMap data, or performing detailed text analysis based on encoding, a programmatic approach using libraries like pdfminer.six or PyPDF2 (for lower-level parsing) is necessary.
Identifying Font Objects and Resources
Every page in a PDF has a /Resources dictionary, which references various resources, including /Font dictionaries. Each font dictionary contains details like its type, name, encoding, and crucially, a reference to its embedded font stream.
General Steps:
- Parse PDF Structure: Open the PDF file and parse its objects.
- Iterate Pages: Access each page object within the document.
- Access Resources: Retrieve the
/Resourcesdictionary for the current page. - Locate Font Dictionary: Inside
/Resources, find the/Fontentry, which lists all fonts used on that page. - Extract Font Details: For each font, extract its
/BaseFontname,/Encoding, and look for a/ToUnicodeCMap stream. - Extract Font File: If the font is embedded, its dictionary will contain a
/FontDescriptor, which in turn points to a/FontFile,/FontFile2, or/FontFile3stream object. These stream objects contain the raw font data (e.g., PFB for Type1, TTF for TrueType, OTF for OpenType).
Working with ToUnicode CMaps
The ToUnicode CMap is a stream object containing PostScript-like syntax that maps byte sequences from the font's encoding to Unicode code points. Extracting and parsing this stream allows for accurate text conversion, especially for fonts with custom encodings or symbol sets.
You can identify the ToUnicode CMap as a direct stream object within the font dictionary. Parsing it requires understanding its syntax (e.g., beginbfchar, endbfchar, beginbfrange, endbfrange blocks).
Conceptual Python with pdfminer.six for Font and Encoding Insights
pdfminer.six provides a more high-level way to extract text while often handling encoding automatically. To dive into font specifics, you'd access the underlying PDF objects.
from pdfminer.pdfpage import PDFPage
from pdfminer.pdfparser import PDFParser
from pdfminer.pdfdocument import PDFDocument
from pdfminer.pdftypes import resolve1
def extract_font_info(pdf_path):
with open(pdf_path, 'rb') as fp:
parser = PDFParser(fp)
document = PDFDocument(parser)
# Iterate through pages to find font resources
for page_num, page in enumerate(PDFPage.create_pages(document)):
print(f"--- Page {page_num + 1} ---")
if 'Resources' in page.attrs and 'Font' in page.attrs['Resources']:
font_dict = resolve1(page.attrs['Resources']['Font'])
for font_name, font_obj_ref in font_dict.items():
font_obj = resolve1(font_obj_ref)
print(f" Font Name: {font_name}")
print(f" BaseFont: {font_obj.get('BaseFont', 'N/A')}")
print(f" Subtype: {font_obj.get('Subtype', 'N/A')}")
print(f" Encoding: {font_obj.get('Encoding', 'N/A')}")
# Check for ToUnicode CMap
if 'ToUnicode' in font_obj:
to_unicode_stream = resolve1(font_obj['ToUnicode']).get_data().decode('utf-8', errors='ignore')
print(" ToUnicode CMap found (first 200 chars):")
print(f" {to_unicode_stream[:200]}...")
# Check for FontFile (actual font data)
if 'FontDescriptor' in font_obj:
font_descriptor = resolve1(font_obj['FontDescriptor'])
if 'FontFile' in font_descriptor:
font_file_stream = resolve1(font_descriptor['FontFile'])
# You can now extract font_file_stream.get_data()
print(f" FontFile stream found (Type1, {len(font_file_stream.get_data())} bytes)")
elif 'FontFile2' in font_descriptor:
font_file_stream = resolve1(font_descriptor['FontFile2'])
print(f" FontFile2 stream found (TrueType, {len(font_file_stream.get_data())} bytes)")
elif 'FontFile3' in font_descriptor:
font_file_stream = resolve1(font_descriptor['FontFile3'])
print(f" FontFile3 stream found (OpenType/CID, {len(font_file_stream.get_data())} bytes)")
# Example usage:
# extract_font_info('your_document.pdf')
This conceptual code provides a starting point for iterating through pages and identifying relevant font objects and their properties. Extracting the raw font file data requires writing the stream data to a file with the appropriate extension (.ttf, .otf, .pfb).
For more advanced data handling, particularly when dealing with complex structures and making sense of raw data, tools designed for AI PDF extraction can significantly streamline the process, automating the parsing and interpretation of embedded information.
Key Takeaways for Font and Encoding Extraction
Understanding these elements is fundamental for any advanced PDF processing task, from accurate text search to content migration.
| Aspect | Description | Tool Relevance |
| Font Subsets | Partial fonts, only containing used glyphs. Identified by AAAAAA+ prefix. Essential for rendering but not for full font reconstruction. |
pdffonts, Programmatic Parsing |
| Text Encoding | Maps character codes to glyphs. Can be standard (e.g., WinAnsiEncoding) or custom. | pdffonts, Programmatic Parsing |
| ToUnicode CMaps | Explicit mappings from font's internal encoding to Unicode. Crucial for non-standard or complex scripts. | Programmatic Parsing |
| Embedded Font Files | Raw font data (Type1, TrueType, OpenType) stored in streams, referenced by FontDescriptor. | Programmatic Parsing |
Whether you're debugging rendering issues or ensuring data integrity, grasping font embedding and text encoding is crucial. When faced with complex PDF analysis or content transformation, consider leveraging PDFjin's robust set of online tools to handle your document needs efficiently.