Free Download · No Signup

Free Sample Documents for Testing

50 realistic synthetic documents — invoices, letters, forms, emails, receipts, agreements and more — ready to feed into any OCR, extraction, or document-AI pipeline. Every entity is fabricated, so there's nothing to redact and nothing to leak.

Download the sample pack

50 documents across 10+ types: clean and OCR-degraded versions of the same documents, corrupt-text-layer examples, and native formats (.docx, .xlsx, .eml, .csv, .txt).

What's inside

What people use them for

Need more than 50?

The full desktop app generates unlimited documents across 33 types with a 0–100% degradation slider, seeded reproducible output, ground-truth manifests, and ML-ready COCO / LayoutLM / FUNSD / DocVQA annotation exports — $199 one-time, runs fully offline on Windows and Linux. See pricing.

Frequently asked questions

Are these documents free to use commercially?

Yes. The samples are synthetic and free to use for testing, development, demos, and evaluation. No signup or attribution required.

Is any of the data real?

No. Every name, address, company, SSN, and dollar figure is fabricated — enforced by the generator's test suite.

Can I upload these to cloud OCR services?

Yes. Because no real entities appear anywhere, they're safe to send to any third-party API with zero PII exposure.