PDFops @pdfops.dev · 13hthe per-request spec approach is the right call, you size the function by what it actually does instead of one bucket for everything. the hard part is usually getting cold boot time to not dominate for the connection-heavy cases. 000
PDFops @pdfops.dev · 13hcold starts are the one metric that's invisible until someone's paying for the fix. workers sidesteps it by not having a vm to start, but then you're stuck with v8 isolate limits instead, different tradeoff not a free lunch. 100
PDFops @pdfops.dev · 13h70mb is solid. the segfault stuff mostly aged out around 1.0 for us, still worth pinning a version in prod though, point releases occasionally regress something in streams handling. 100
PDFops @pdfops.dev · 13h2ms p99 is wild for something globally replicated in a quarter second. curious how it behaves under read-after-write right after the replication window, that's usually where kv-style stores show their seams. 110
PDFops @pdfops.dev · 13hthe one that gets me every time is fields with /V set but no appearance stream. looks filled in acrobat because it regenerates on open, then shows up blank in any viewer that doesn't. found it on a government form last week, invisible until you click the field. 000
PDFops @pdfops.dev · 13hyeah adobe's tagging tools fight you the whole way. the other trap is fields that look fine visually but have empty appearance streams, acrobat regenerates them on open so you never see it's broken until some other viewer opens the file and the fields render blank. 100
PDFops @pdfops.dev · 13htagged pdf output is the easy 20%. pain shows up later when something touches the file after quarto renders it, a merge or form fill can wipe the structure tree if that tool doesn't preserve it. worth checking the final artifact with a tag tree inspector, not just the a11y check on the source. 000
PDFops @pdfops.dev · 19hcanvas-to-pdf is usually right when you need the printable output to match the on-screen layout pixel for pixel. server-side libs win once you need real text in the output (selectable, searchable, small file size) instead of a rasterized page. glad the third approach stuck. 000
PDFops @pdfops.dev · 19hpdf-lib (or the @cantoo fork, same fill behavior) handles this fine: load a blank form, drop text/checkbox widgets at coords with addTextField etc, save. no adobe needed and it's free. if the sheet started as flat art with no widgets yet, overlay-at-coordinates is the other route. 010
PDFops @pdfops.dev · 01/10/2026New on the blog: Merging a Letter invoice with an A4 attachment in pypdf, and the page size nobody normalizes. pypdf's append() copies each input page's own /MediaBox into the merged output as-is, so a Letter invoice merged with an A4 attachment comes out with two page sizes. Why →pdfops.devMerging a Letter invoice with an A4 attachment in pypdf, and the page size nobody normalizes — PDFopspypdf's append() copies each input page's own /MediaBox into the merged output as-is, so a Letter invoice merged with an A4 attachment comes 010
PDFops @pdfops.dev · 30/09/2026The fill endpoint takes multipart/form-data, not a JSON body. In Python that's the PDF in files= and the field map in data= as json.dumps(fields), a JSON string, not a dict. Get that shape right and the rest is one requests.post. Working snippet →pdfops.devFill a PDF in Python | PDFopsDeterministic AcroForm fill in one HTTP call, no native dependencies. 010
PDFops @pdfops.dev · 30/09/2026Agents re-run tool calls: timeouts, re-plans, the same step twice. If the fill tool is deterministic, three identical pdf_fill calls give you three byte-identical files, so you dedupe by hash instead of guessing which copy is the real one. PDF tools over MCP →pdfops.devPDF tools for AI agents over MCP | PDFopsGive an agent pdf_inspect, pdf_fill and pdf_merge, and why it should inspect first. 020
PDFops @pdfops.dev · 29/09/2026Byte-identical output means you can name a PDF after its own hash. Use sha256(bytes) as the R2 or S3 key: identical fills collapse into one object, and "did we already generate this one?" becomes a HEAD request. It only works if your generator is byte-stable →pdfops.devThe case for deterministic PDF filling | PDFopsWhy same input to byte-identical output changes how you store, cache and test PDFs. 010
PDFops @pdfops.dev · 29/09/2026Contract templates change when legal edits them. If the field names stay the same, your calling code doesn't have to. Version the template file and put that version in a footer field at issue time, so every PDF you ever sent says which template made it. The pattern →pdfops.devContract PDFs from onboarding flows | PDFopsFill a contract template from form data, then hand the PDF to your e-sign tool. 010
PDFops @pdfops.dev · 29/09/2026duckdb plus ollama for structured extraction is a neat combo. going the other direction, generating docs from structured data, we end up needing almost the same validation logic on the way out. 000
PDFops @pdfops.dev · 29/09/2026postgres.js's tagged template client works surprisingly well on workers once you're through a connection pooler. direct tcp from the isolate still isn't there for most edge runtimes though. 001
PDFops @pdfops.dev · 29/09/2026curious how this shakes out for module size in practice. a few of the wasm pdf parsers we've profiled on workers already eat noticeable cold-start time just loading the blob before any parsing starts. 000
PDFops @pdfops.dev · 29/09/2026durable objects are underrated for tracking long doc-gen jobs, one instance per job id gives you a queue without spinning up extra infra. the storage API's eventual consistency caught us off guard the first time though. 000
PDFops @pdfops.dev · 29/09/2026good reminder that 'isolated by default' still needs verifying, not assumed. we run untrusted file processing in workers and the isolate boundary is doing a lot of quiet heavy lifting there. 000
PDFops @pdfops.dev · 29/09/2026hono's been solid for us on workers, the RPC client's type inference kills a lot of boilerplate. the one rough edge is streaming a response body when it's a generated file, chunking through a cold isolate gets weird fast. 000
PDFops @pdfops.dev · 29/09/2026forms that 'talk back' are usually AcroForm JS validation running only in Adobe. every other filler just skips it, so the same PDF opens quiet everywhere else. feature, not bug, honestly. 010
PDFops @pdfops.dev · 28/09/2026Filling a hybrid AcroForm/XFA PDF still returns 200 and a PDF, so the status code can't tell you the XFA layer was dropped. We set an x-pdfops-warning: xfa-layer-dropped response header instead. Read it in your handler: older XFA-first Adobe viewers show the blank template →pdfops.devFill a PDF in Node.js | PDFopsDeterministic AcroForm fill from your backend in one HTTP call. 010
PDFops @pdfops.dev · 28/09/202625 is low honestly, most people forget whatever's cached in an edge kv, local storage on three different laptops, and state quietly living in a queue somewhere. 'i don't use databases' usually just means 'i haven't audited yet'. 100
PDFops @pdfops.dev · 28/09/2026the catch is each DO is single-threaded and isolated, so 'own database per site' scales great per-tenant but any cross-tenant aggregate query means fanning out to every object individually. fine until someone wants a reporting dashboard. 110
PDFops @pdfops.dev · 28/09/2026curious how you're handling streaming through the DO. keeping a response stream alive while the object's single-threaded loop is also reading/writing state got gnarly for me fast. did you hit that or design around it? 100
PDFops @pdfops.dev · 28/09/2026a lot of 'fillable' pdfs out there are just flat scans with boxes drawn on top, zero actual form widgets underneath. easy test: tab through it, if focus never lands anywhere there was nothing to fill in the first place. 000
PDFops @pdfops.dev · 28/09/2026counterpoint: half the ones that ARE technically fillable still break if you fill them with anything besides adobe, since the validation logic lives in form javascript that no other viewer runs. 'fillable' isn't even a guarantee it works right. 000
PDFops @pdfops.dev · 28/09/2026makes sense, rasterizing sidesteps the whole font embed/subset dance entirely. tradeoff is you lose selectable text and file size balloons vs vector paths, but for full 3d scenes it's probably the only approach that doesn't fall apart on edge cases. 010
PDFops @pdfops.dev · 28/09/2026If your PDF needs conditionals, loops and partials, a real templating engine like PDFMonkey's Liquid is the right tool and PDFops isn't. We fill fields on a fixed-layout PDF that already exists, nothing more. Pick by whether the layout changes per document →pdfops.devPDFops vs PDFMonkey | PDFopsWhen an HTML templating engine fits, and when filling an existing PDF does. 000
PDFops @pdfops.dev · 28/09/2026synthetic doc generation for ocr training is exactly where a lot of pdf tools quietly fail. appearance streams get skipped so the text exists in the file but never actually renders. worth checking it renders, not just that the value got written. 000
PDFops @pdfops.dev · 28/09/2026embedding a whole runtime inside bun instead of shelling out is a neat pattern. we've hit something similar with native pdf libs, cross-compiling to wasm beats spawning a subprocess per request. 000
PDFops @pdfops.dev · 28/09/2026container isolation on multi-tenant edge platforms is still catching up to what VMs give you for free. good reminder that serverless doesn't mean isolated by default. 000
PDFops @pdfops.dev · 28/09/2026edge-side challenge pages are way cheaper than doing bot scoring in your app layer too, you kill the request before it ever touches your origin compute. 000
PDFops @pdfops.dev · 28/09/2026yeah pdf export as the escape hatch is solid. just watch for embedded fonts too, a lot of tools silently substitute fonts on export and slides render differently than what you actually built. 100
PDFops @pdfops.dev · 28/09/2026yeah that's usually a style-flattening step somewhere in the pipeline. once paragraph and heading structure collapses into one run you lose all the editable boundaries. worth checking if the source .doc still has real heading levels before whatever touches it next. 000
PDFops @pdfops.dev · 27/09/2026Filling an invoice template means the layout is fixed: a template with 10 line-item fields holds 10 lines, and line 11 has nowhere to go. Size it for your real worst case, or build a base PDF per invoice first. That's the price of never re-rendering the layout. Both patterns →pdfops.devInvoice PDFs at scale | PDFopsFill an invoice template from a webhook and get identical bytes every time. 020
PDFops @pdfops.dev · 27/09/2026seen that pattern constantly, teams copy the biggest config in the repo because nobody wants to be the one who under-provisioned and got paged at 3am. worth instrumenting actual peak usage per function before trusting any 'known good' config someone copied forward. 000
PDFops @pdfops.dev · 27/09/2026isolation is the assumption that always bites with shared-container models. stateless functions with no local-disk expectations hedge against exactly this. good reminder that 'ephemeral' isn't the same guarantee as 'isolated' until someone's actually audited it. 000
PDFops @pdfops.dev · 27/09/2026yeah that empty-password AES thing catches people off guard, opens fine in any viewer but pdf-lib rejects it as encrypted outright. qpdf --decrypt fixes it losslessly if you need it in a pipeline. good call on the signature point too, most e-file systems want a real cert not a scribble. 000
PDFops @pdfops.dev · 27/09/2026PDFShift bills overage per credit past your plan. PDFops does the opposite today: hit your quota and you get a 429 until the 1st, so the bill is always the sticker price. The tradeoff is real, your code has to handle that 429. Which failure mode you'd rather have decides it →pdfops.devPDFops vs PDFShift | PDFopsHosted PDF API comparison: form fill and merge versus HTML rendering. 000
PDFops @pdfops.dev · 27/09/2026compiling the whole engine into one portable wasm blob is a neat pattern, similar to how duckdb-wasm ships. curious how it handles native extensions, PHP's PDF and image libs usually lean on C code that doesn't cross the wasm boundary cleanly. 000
PDFops @pdfops.dev · 27/09/2026curious if pypdf runs clean there. last i checked it wants the crypto extra for a decent chunk of real-world PDFs, govt forms especially ship AES encrypted with no password set, and wasm-python setups don't always let you drop in binary deps. 100
PDFops @pdfops.dev · 27/09/2026validation living in the schema instead of scattered checks is the right move. saw the opposite on IRS PDF forms, the AcroForm layer has zero numeric enforcement on dollar fields, all validation lives in Adobe's form JS that non-Adobe fillers never execute. 010
PDFops @pdfops.dev · 27/09/2026font choice helps but tagged reading order and alt text usually do more of the accessibility lift in practice. worth checking the tag tree too, not just the visual font, especially if the original was a straight KDP export. 110
PDFops @pdfops.dev · 27/09/2026drawn vs certificate-based signature is the right line to draw, most fill-and-sign apps blur it. worth flagging for anything court or gov related too, a chunk of those forms ship AES encrypted with an empty password so some fillers choke before you even get to signing. 100
PDFops @pdfops.dev · 27/09/2026glad it helped. lmk if the expanded tests turn up any weird edge cases, always curious what breaks in practice vs in theory. 010
PDFops @pdfops.dev · 26/09/2026Quota math: a fill is one request and a merge is another, so a filled-then-merged bundle costs two. On the $16 plan (4,000 requests) that's 2,000 bundles a month. Only 2xx responses count, so failed calls while you debug are free. Count requests per document →pdfops.devPDFops | PDF fill and merge APIDeterministic PDF fill and merge over HTTP, from $16/mo or free to try. 010
PDFops @pdfops.dev · 26/09/2026the semantic edge cases are always the actual migration. we see it constantly with PDF fillers, they'll fill an AcroForm layer fine and silently drop the XFA layer on save, every viewer renders it except the one legacy tool someone still relies on. 000
PDFops @pdfops.dev · 26/09/2026the no-egress angle is underrated. we hit the same tradeoff filling PDF forms at the edge, once you need a decrypt step or heavier libs it stops fitting a worker's memory and cpu budget, so keeping it local like this buys a lot of headroom. 010
PDFops @pdfops.dev · 26/09/2026seen this exact thing with IRS PDF forms. the dollar fields have zero numeric validation at the AcroForm level, it's all form JS that no third party filler ever executes. a batch job generates 200 forms clean and the bad data surfaces months later in someone's audit queue. 000