Guide · eDiscovery

Synthetic Test Data for eDiscovery and Litigation-Support Platforms

Testing eDiscovery platforms, review workflows, and processing pipelines requires document populations that behave like real matters: mixed types, native formats, Bates ranges, load files, custodians. Using actual matter documents for testing is a confidentiality breach waiting to happen. Here's the synthetic alternative.

What eDiscovery testing actually needs

A pile of PDFs isn't a test corpus. Litigation-support workflows consume productions: documents plus the metadata scaffolding that review platforms ingest.

Why synthetic beats sample matters

Real matter documents are privileged, confidential, or both — using them to test software (or worse, to demo it) is a professional-responsibility problem. Public document dumps (Enron) are stale, structurally monotonous, and every platform has been tuned on them. Synthetic batches give you fresh, structurally varied populations with zero confidentiality exposure, regenerable on demand at any size.

Degradation for processing QC

Scanned productions arrive with terrible OCR. Generating degraded documents — and documents whose embedded text layer is corrupted while the page looks clean — lets you test how your processing pipeline handles low-confidence text: does search still work, does deduplication break, do privilege screens miss terms? These are exactly the failure modes that surface mid-review at the worst possible time.

One-command productions

DocSet Generator's batch packager emits mixed-type batches as a single zip with page-accurate Bates ranges in metadata.csv, custodian/author/date fields, watermarks, and native-format files — with CLI flags (--bates, --bates-prefix, --bates-start) for scripting nightly test-data refreshes.

Frequently asked questions

Can the load file be imported into Relativity / Everlaw / DISCO-style platforms?

The metadata.csv is a standard flat load file with Bates ranges, custodian, author, and date fields — mappable in any mainstream platform's import wizard. Field mapping specifics vary by platform.

Are the email documents real .eml files?

Yes — native .eml files (plus .docx, .xlsx, .csv, and .txt for other types), with all addresses on reserved .example domains and no real entities anywhere.

Generate this data yourself, locally

DocSet Generator is a $199 one-time desktop app (Windows & Linux) that produces synthetic documents across 33 types with controllable OCR degradation, ground truth, and ML-ready COCO / LayoutLM / FUNSD / DocVQA exports — all offline, with zero real PII.