The licence file at the top of a repository describes that repository. It does not describe the compression engine downloaded during the build, or the WebAssembly module fetched at runtime.
We found that out while picking a PDF library. People upload PDFs to us, and the obvious next thing they want is to fix one without downloading it, editing it elsewhere, and uploading it again. So we read four toolkits closely: BentoPDF, PDFCraft, Stirling PDF and EmbedPDF.
All four are better documented than the average dependency. What follows is what we found, project by project, and the questions we wish we had asked on day one.
Worth stating, because it explains the weighting. We wanted the work to happen in the customer's browser, partly for cost and partly because "your document never leaves your browser" is a claim we would like to make honestly. The feature is for paying customers, our product is closed-source, and we are small enough that a five-figure annual licence is not an option.
A team with a server fleet and an enterprise budget would rank these differently, and reasonably so.
Licence: AGPL-3.0, or a commercial licence at $79 lifetime. Runs: in the browser. Shape: a Vite multi-page app, over a hundred tools, one page each.
This was the most capable browser-side option we found. Merge, split, organise, crop, watermark, forms, redaction, annotation, text editing, OCR through Tesseract, Office conversion through a LibreOffice WebAssembly build. It has a workflow builder and ships in eighteen languages.
The part to understand before choosing it is where the heaviest features come from. The project's licensing page is direct about this:
BentoPDF does not bundle AGPL-licensed processing libraries. The following components must be configured separately via Advanced Settings if you wish to use their features: PyMuPDF (AGPL-3.0), Ghostscript (AGPL-3.0), CoherentPDF / CPDF (AGPL-3.0)
and
The commercial license covers BentoPDF's own code only. It does not bypass the AGPL licensing of these components. Users must comply with the AGPL v3 terms for these components.
That is more disclosure than most projects offer, and it shaped our whole evaluation. The three modules load at runtime from a CDN by default, which the documentation describes as the compliance boundary: the CDN distributes them, you do not.
We counted what depends on them, because the answer is not in the documentation. Of 125 tool pages, 42 call one of the three.
| Module | Tools | What it powers |
|---|---|---|
| PyMuPDF | 31 | Compression, extraction, most conversions |
| CoherentPDF | 10 | Merge, split, attachments, bookmarks |
| Ghostscript | 3 | PDF/A and font outlining |
PyMuPDF carries the longest tail: every compress variant, extract images and tables, deskew, rasterize, conversion out to text, Word, Excel, CSV, Markdown and SVG, and the inbound conversions from image, JPG, TXT, email, EPUB, MOBI, XPS, CBZ and PSD.
Two details matter more than the count. There is no graceful degradation: several tools check whether their module is available, but the check gates an explanatory dialog rather than an alternative implementation. And the split is uneven, because CoherentPDF accounts for only ten tools and two of them are merge and split, which are among the most-used PDF operations there are.
One asymmetry surprised us. PNG, WebP, SVG and HEIC to PDF use pdf-lib and work without any module. JPG to PDF and the generic image-to-PDF tool go through PyMuPDF and do not.
One more thing sits outside both lists: opening a password-protected PDF routes through a decrypt helper that tries CoherentPDF then PyMuPDF. Nearly every tool imports the password prompt, so an encrypted input can reach an AGPL module from a tool that otherwise never touches one.
If you deploy to a platform with per-file asset limits, check the payload before you design around it. The LibreOffice WebAssembly build is about 74 MB in two files of 46.5 MB and 27.3 MB, and the PyMuPDF bundle is about 38 MB.
Licence: AGPL-3.0. Runs: in the browser. Shape: Next.js with static export, 90+ tools.
PDFCraft covers similar ground and credits its lineage openly in its acknowledgements:
PDFCraft stands on the shoulders of giants. We gratefully acknowledge BentoPDF for their pioneering work in privacy-first, client-side PDF tools.
Two differences decided it for us. It is AGPL-3.0 with no commercial option, so there is no path to using it inside a closed-source product. And it vendors PyMuPDF directly in the repository along with a full Pyodide runtime, around 61 MB of it, rather than loading it from elsewhere.
Both are reasonable engineering decisions. For an open-source project it is one less thing to configure. For a proprietary one it closes a door rather than leaving it ajar.
Licence: MIT, with named directories under a separate paid licence. Runs: on your server, as a Java service. Shape: roughly 94 REST endpoints plus its own React frontend.
Stirling has the most careful licence engineering of the four, which we did not expect and came to appreciate. Its Java dependency allow-list permits MIT, BSD, Apache, MPL and similar while omitting plain GPL and AGPL, which is deliberate hygiene you rarely see documented.
Its root licence carves out eleven directories, and the shape of them is worth reading carefully because we misread it at first. Three are on the backend and seven are on the frontend, and what they have in common is that they are editions and delivery channels rather than features. frontend/editor/src/core, which holds roughly fifty tool components including the text editor, annotation, form filling and redaction, is not carved out and is MIT. So it is not a free backend with a paid frontend. Both halves are an MIT core with paid edges around it.
One configuration note: the published Docker image builds the flavour that includes the carved-out directories, so a commercial deployment wants the core flavour build flag.
The copyleft that does arrive comes through the operating system layer of the image: Ghostscript built from Artifex source, plus Calibre, unpaper and pngquant. Here is the part we liked. The application has a first-class option for disabling dependency groups, and endpoints with an alternative implementation keep working through it:
endpoints:
groupsToRemove: ['Ghostscript', 'Calibre', 'OCRmyPDF']
We traced the resolution logic against the endpoint registry to see what that costs. Five endpoints out of 94: two vector conversions, a colour invert, a scanner effect, and EPUB output. Compression, repair and crop all survive because they resolve to alternative implementations, and PDF/A survives because Stirling routes it through LibreOffice rather than Ghostscript.
OCRmyPDF belongs in that list even though it is not a licensing question itself. It depends on Ghostscript, and the OCR controller uses it when available with no fallback, routing to Tesseract only when the OCRmyPDF group is disabled. Disable Ghostscript alone and OCR stops working. Disable both and OCR runs on Tesseract.
Licence: MIT for the stable line, Apache-2.0 for the next major. Runs: in the browser. Shape: a plugin architecture over a WebAssembly engine, with React, Vue, Preact and Svelte bindings.
We checked this one against the published packages rather than the repository, since those are what you install. Every package ships a LICENSE file and declares MIT, the engine package carries both its own MIT licence and PDFium's BSD notice, and the strings "Affero" and "General Public License" appear nowhere in any published artifact. One loose end: @embedpdf/plugin-form ships with no license field in its manifest, so that is worth pinning before you depend on it.
The repository also contains a server product under the Fair Core License, which requires a licence key and restricts competing use. The project states the dividing line clearly:
The rule behind the line: libraries are Apache-2.0; the deployable server product is Fair Source. Client SDKs, contracts, engines, viewers, plugins, anything you link into your own software, carry no strings.
Capability surprised us in the useful direction. We had it filed as a viewer with annotations, and the engine's type definitions describe 236 members including a good deal more: saveAsCopy returning the modified document, merge, mergePages, extractPages, importPages, deletePage, applyRedaction, redactTextInRects, form field setters, text extraction, search, document encryption, bookmarks and attachments.
Merge and split being native matters if you had budgeted a second library for them.
What is genuinely absent: text content editing, OCR, compression and format conversion. The text APIs are all read-side, so you can extract and locate text but not retype it.
Interactive editing usually lives in a frontend. Across several of these projects, the tools that feel like editing, meaning clicking text and changing it or drawing an annotation, are implemented in the client application rather than exposed as an API. In Stirling there is no controller for the text editor, annotation, image placement or compare, and the text editor's only backend is a small font glyph encoding helper. If your plan is "call their API from our own UI", check that the specific features you care about exist as endpoints before you commit.
Per-copy licensing and browser delivery do not fit together. Artifex, which owns Ghostscript and PyMuPDF, publishes its commercial models without prices. Both options are described as a per-copy cost with a quarterly minimum fee, plus annual reporting on volume of distribution, and the page says that each licence is crafted around your individual use case. That is a clean model for shipping an application to customers you can count. When the same library is compiled to WebAssembly and sent to every visitor's browser, the unit of counting stops being obvious.
CoherentPDF, by contrast, publishes a price list. Its per-machine licences run from $599 for the command line tools to $899 for the API, and the list says plainly that cloud usage needs a quotation instead. The unlimited-user Developer Licence is quote-only, and its wording covers building into the back end of your own software, which is a different thing from shipping a module to a browser.
Incremental saving can undo a redaction. PDFs can be saved by appending changes rather than rewriting the file, which keeps the earlier revision inside the document. Redact something, save incrementally, and the original text is still there and recoverable. This is a property of the format rather than a flaw in any library, and it was the most important thing we learned. We wrote it up separately in the text you blacked out is probably still in the PDF, because it affects anyone who redacts a PDF, not only people building with these libraries.
Twenty minutes, and it is the same twenty minutes whether you spend it now or after a customer's security questionnaire arrives.
The four are not ranked, because the ranking depends entirely on constraints you have and we do not.
If your own product is open-source, or you are willing to comply with AGPL across it, the calculation changes completely. BentoPDF with all three modules configured is the largest browser-side feature set of the four, and PDFCraft covers similar ground with the same obligation.
If you already run servers and need volume, Stirling is hard to argue with. The core flavour is MIT at both ends, the dependency groups you cannot accept come out in configuration rather than in a fork, and its MIT frontend carries the interactive editing that no headless API in this space provides.
If you want browser-side processing with no licence question to defend, EmbedPDF is the cleanest chain we read. You give up text editing, OCR, compression and format conversion, which is a real cost and an easy trade only if those are not what your customers ask for.
If you need supported text editing with someone to call, none of these four is the answer and a commercial SDK is. Ask early what the licence is counted by, because per-domain, per-developer and per-document produce very different bills from the same product.
A repository licence describes a repository. What you ship is a build, and a build pulls in things the repository never contained: a compression engine downloaded during the build, packages installed by a Dockerfile, a WebAssembly module fetched at runtime from somebody else's CDN.
None of that is hidden. All four projects document it, two of them better than most commercial vendors do. It is just not in the file everyone checks, which is why the twenty minutes above is worth spending before the architecture is decided rather than after.
Upload a PDF and share the link
All licence terms, file sizes and behaviours described here were checked against the source repositories and vendor licensing pages in September 2026. Projects change. Check them yourself rather than trusting an article, and nothing here is legal advice.
