Acrobat is the app everyone uses and almost nobody trusts, because PDF tooling has a history of subscription walls, telemetry, and destructive saves. PdfCraft, part of the ArtCraft family on getartcraft.com, is the clean-room answer: an open-source reimplementation of Acrobat in pure Rust, from the storytold/pdfcraft repository, version 0.5.0 under MIT OR Apache-2.0, running natively on macOS, Windows, Linux and FreeBSD plus the web. Its behaviour comes from the ISO 32000 specification and black-box observation, never from anyone else’s code, and the whole engine, CLI and app live in one Cargo workspace of focused crates where the core never depends on the UI.
The design decision that sets it apart is the save model. Every save appends an incremental update and leaves the original bytes untouched; writes are atomic; undo runs deep; and after every edit the working file is produced by an incremental write, so the view always shows exactly what Save will write. The other distinguishing trait is the automation surface: every engine feature is reachable without the GUI through one table of JSON-Schema-described tools, served three ways, pdfcraft-cli run for scripts, an MCP server that is strictly opt-in and opens no network port, and a Rust API for embedding.
The overview: all frontends converge on the engine, which edits the COS object graph copy-on-write; features hang off the engine; encryption and stream filters sit under the object layer.
Reading the overview from left to right:
- The desktop app in apps/pdfcraft/src, the CLI in apps/pdfcraft-cli/src, and the web build in apps/pdfcraft-web share one facade, with the desktop shell optionally opening a token-authed control file so agents can see and operate the real interface.
- The engine in crates/engine/src holds open documents, edit history, undo, saving and the tool catalogue; frontends never touch parsing or editing crates directly.
- The object layer in crates/cos/src parses the COS graph lazily and tolerantly, keeps edits in a copy-on-write overlay, and writes incrementally or as a full garbage-collected rewrite.
- Under it sit crates/filters/src, every PDF stream filter in both directions, and crates/crypt/src, the standard security handler covering RC4 and AES-128/256 across revisions 2 through 6.
- Feature crates hang off the engine: crates/organize/src, crates/annot/src, crates/forms/src, and crates/render/src for rendering, inspection and reading-order text extraction.
- The automation layer in crates/automation/src is the JSON-Schema tool table behind the CLI, MCP and embedding.
Why You Need This
First, your files survive the tool. The COS layer in crates/cos/src/lib.rs implements every cross-reference form, tables, streams and hybrid /XRefStm, object streams, and repair by scanning, which means damaged PDFs from the wild actually open. Edits run on a clone; on success the previous state is pushed onto the undo stack, and because clones share all unchanged data, snapshots are cheap no matter how big the document. Saving to the same file appends an incremental update so the original bytes are preserved, full saves pack objects into compressed object streams with a cross-reference stream, and unsaved changes are never discarded silently.
Second, the security model is real, not a checkbox. crates/crypt/src implements the standard security handler end to end, RC4 and AES-128/256 from revision 2 to revision 6 with crypt filters, wired into load and save in the object layer, which is what lets the app open protected documents and respect their permission rules. On the trust side, crates/redact/src provides true redaction and document sanitizing rather than black boxes drawn over text, and crates/sign/src covers basic digital signatures, with password protection available as a plain engine call. For an app in this category, that combination, opening protected files honestly and destroying data thoroughly, is the actual product.
Third, the entire workbench is a tool table an agent can read. The automation layer describes every engine feature as a JSON-Schema tool: open, inspect, render pages to PNG, extract and find text, rotate, delete, move and insert pages, edit bookmarks and page labels, add, reply to, restyle and delete comments, highlight a phrase just by naming it, list and fill form fields, set metadata, undo and redo, combine, extract and split. pdfcraft-cli run --script review.json chains steps in one session with --root confining every file access to one directory; pdfcraft-cli mcp serves the same tools over stdio only, starting nothing on its own, with a --compact mode that costs agents far fewer tokens; and edits stay in memory, undoable, until doc_save. One practical killer feature: a script can add a highlight annotation by find query, so marking every mention of “total due” in a 300-page contract is one line of JSON.
The detail view: the engine dispatches commands and multi-step actions, the core layer parses, filters and decrypts, and each feature crate owns one Acrobat capability.
The detail diagram shows the breadth that a PDF workbench has to cover, and PdfCraft organizes it one crate per capability. Documents: crates/organize/src rotates, deletes, moves and inserts pages and edits bookmarks and page labels; crates/create/src builds new PDFs from images and scans; crates/optimize/src compacts. Knowledge work: crates/annot/src builds appearance streams for notes, markup, shapes and ink with replies and status; crates/compare/src diffs documents; crates/ocr/src recognizes Latin-script text. Production: crates/print/src and crates/preflight/src for print workflows, crates/export/src for getting content out, crates/a11y/src for the Accessibility Checker and tagged PDF, and crates/js/src running form JavaScript sandboxed.
The engine keeps the seams clean. crates/engine/src/lib.rs defines the editing model in one comment: each document keeps the COS graph with a copy-on-write overlay, an edit runs on a clone, the working file is regenerated by incremental write after every edit, and saving rebases onto the written bytes so the next save appends only new edits. crates/engine/src/commands.rs holds the stable command ids and crates/engine/src/actions.rs the multi-step custom actions, while crates/cos/src/parser.rs does the tolerant parsing and repair that hostile files need. The web build in apps/pdfcraft-web and the platform helpers in crates/platform/src round out the frontends.
From Install to First Script
Clone and run cargo run --release -p pdfcraft -- some.pdf for the desktop app; the workspace needs Rust 1.90 or newer on edition 2024. For agents, add the MCP server to your client configuration with {"mcpServers": {"pdfcraft": {"command": "pdfcraft-cli", "args": ["mcp", "--root", "/path/to/your/pdfs"]}}}; the server only exists while the agent is connected. A first headless session: pdfcraft-cli info form.pdf dumps structure as JSON, pdfcraft-cli run text_find doc=1 query=invoice finds every mention, pdfcraft-cli run --script review.json applies a multi-step review, and pdfcraft-cli edit in.pdf --rotate 1,2:90 --delete 5 --title "Q3" --out out.pdf does a one-shot edit. Everything the GUI does, the CLI does, through the same tool table.
Honest limits: the roadmap puts the project at about half of Acrobat Pro’s offline features, 88 percent of must-haves, but only a third of the work done. Pages are currently drawn by the third-party Enjoyed this post? Never miss out on future posts by following us hayro crate while the project’s own renderer is built; reliable editing of existing text, especially CJK, is thin; OCR is Latin-only; Office import and export, signature timestamps and long-term validation, PDF/A/X/UA preflight, and XFA forms are still ahead; and fuzzing on hostile files keeps turning up crashes that each get a regression test. As a viewer, organizer, annotator, form-filler and agent-driven PDF processor, though, it already earns a slot next to your editor.