Files
Jonas Jenwald a8057e037b Convert the CMap and ToUnicodeMap classes to store its data in Maps
The CMap-data and the ToUnicode-data is often sparse[1], which means that Arrays are not ideal data-structures for this purpose.
Note how there are separate paths, in the existing code, depending on the size of the data and that we're forced to either iterate over non-existing keys or use `for...in` iteration which "unnecessarily" stringify the keys.
By using Maps instead both of these issues can be avoided, and given that performance of Maps have been improved recently (in Firefox) this shouldn't be an issue.

Given how intertwined all of this functionality is, it unfortunately wasn't really possible to easily split this into several patches.
However, all of this code should (famous last words) be well covered by existing test-cases.

One notable difference is that iterating through CMap-data and ToUnicode-data now happens in insertion order, but given how this data is being used that's likely not an issue.

Also, copy the data returned by `CMap.prototype.getMap` since it's used as input to the `ToUnicodeMap` class. Note that we may amend the ToUnicode-data at the end of font parsing, hence we should not modify the underlying CMap-data.
Given how/where the CMap-data is accessed it's unlikely that this pre-existing "bug" has caused any issues, but it nonetheless seems like something that should be fixed.

*Note:* We purposely keep the `forEach` methods, since making the classes iterable seemed to be approximately an order or magnitude slower (based on very quick `console.{time, timeEnd}` benchmarking).

---

[1] In some cases even *extremely* sparse, see e.g. `issue8372.pdf`.
2026-09-22 08:58:23 +02:00
..

Font tests

The font tests check if PDF.js can read font data correctly. For validation the ttx tool (from the Python fonttools library) is used that can convert font data to an XML format that we can easily use for assertions in the tests. In the font tests we let PDF.js read font data and pass the PDF.js-interpreted font data through ttx to check its correctness. The font tests are successful if PDF.js can successfully read the font data and ttx can successfully read the PDF.js-interpreted font data back, proving that PDF.js does not apply any transformations that break the font data.

Running the font tests

The font tests are run on GitHub Actions using the workflow defined in .github/workflows/font_tests.yml, but it is also possible to run the font tests locally. The current stable versions of the following dependencies are required to be installed on the system:

The recommended way of installing fonttools is using pip in a virtual environment because it avoids having to do a system-wide installation and therefore improves isolation, but any other way of installing fonttools that makes ttx available in the PATH environment variable also works.

Using the virtual environment approach the font tests can be run locally by creating and sourcing a virtual environment with fonttools installed in it before running the font tests:

python3 -m venv venv
source venv/bin/activate
pip install fonttools
npx gulp fonttest