Live Browser Demo

Generate real synthetic documents, right now.

Pick a document type, dial in OCR degradation, and hit generate — every run produces unique synthetic PDFs from fresh randomized data. Runs 100% in your browser: nothing is uploaded, no signup, no server. Demo output carries a watermark; the full product does not.

Document Type
Count
3 docs
OCR Degradation
clean

Demo limits: 3 of 33 types, max 5 docs per run, watermarked output. Every run includes a manifest.jsonl with per-document seeds, corruption settings & ground-truth text — just like the full product, which adds corrupt text-layer mode, Bates numbering & batch packaging.

Full version only — 30 more types
Unlock all 33 types →
Generated Output
No documents yet Choose a type and hit generate

Like what you see? This is the watered-down version.

The desktop product generates 33 document types — court filings, patents, financial statements, transcripts, checks, faxes & more — with the same OCR degradation engine plus corrupt text-layer mode, seeded reproducible runs, Bates numbering, and a full ground-truth manifest recording corruption settings for every file. No watermarks. Runs entirely offline on Windows & Linux.

Purchase — $199 →

A free synthetic document generator, in your browser.

This demo is a working slice of DocSet Generator, a desktop tool that produces synthetic documents for OCR training data, document AI evaluation, and extraction-pipeline testing. Everything here is generated client-side with JavaScript — pick a type, set a count, dial in degradation, and download real PDFs plus a manifest.jsonl with the seed and corruption settings for every file.

Because every name, company, address, amount, and date is fabricated, the output is safe to share, commit to test fixtures, or feed to third-party OCR services without any PII or privacy risk — a drop-in substitute for real paperwork in demos, QA suites, and benchmarks.

Want more volume without the watermark? Grab the free sample pack — 50 documents across 10+ types — or see the full feature list.

What people use it for

  • OCR stress-testing — generate degraded invoices and contracts and run them through Tesseract, AWS Textract, Azure Document Intelligence, or Google Document AI to see where extraction breaks.
  • Fake invoices & emails for software testing — realistic sample PDFs for document-upload features, parsers, and viewers, with internally consistent totals.
  • Document AI evaluation sets — seeded, reproducible PDFs with known corruption levels for before/after benchmarking.
  • RAG & LLM pipeline tests — noisy business documents for testing chunking, extraction, and retrieval against imperfect scans.
  • Training & demos — placeholder paperwork for screenshots, tutorials, and e-discovery walkthroughs with zero real data.

Demo FAQ

Is the online demo really free?

Yes. The browser demo is completely free with no signup, no email, and no trial period. It generates watermarked synthetic emails, invoices, and agreements. The full desktop product ($199, one-time) removes the watermark and adds 30 more document types.

Is anything uploaded to a server?

No. Every document is generated locally in your browser with JavaScript (pdf-lib). No document data, analytics events, or files are sent to any server.

What is the OCR degradation slider?

It controls how corrupted the generated text is, from 0% (clean) to 100% (heavily degraded), using character substitutions modeled on real OCR failure modes — the same concept as the desktop product's quality slider. Use it to stress-test OCR engines and extraction pipelines against realistic noisy input.

Can I use the demo documents to test my OCR pipeline?

Yes. Download the generated PDFs and manifest.jsonl as a zip and feed them to any OCR or document AI tool — Tesseract, AWS Textract, Azure Document Intelligence, Google Document AI, or your own models. Demo output carries a DEMO watermark; every value in it is synthetic, so there is no PII risk.

Is the generated data real?

No. All names, companies, addresses, amounts, and dates are synthetically generated. Nothing refers to real people or organizations.

How is the demo different from the full product?

The demo offers 3 of 33 document types, a maximum of 5 documents per run, and watermarked output. The full desktop product adds corrupt text-layer mode, image-only PDFs, Bates numbering, batch packaging, full per-document ground truth, and a CLI — with no watermarks, running entirely offline on Windows.

Privacy: this demo generates everything locally in your browser with JavaScript. No document data, analytics events, or files are sent to any server.