Two layers of truth
Every file in a build carries two independent truth layers:- the printed layer — what the page says (item name, value, unit, reference range, flag; the five MedRepBench fields)?
- the semantic layer — what it means (indicator key, LOINC code, UCUM unit, observation date).
unit.missing, a faithful transcription legitimately leaves the
printed unit empty while the semantic truth still says glucose, mmol/L — and an invariance check
must therefore live at the semantic layer.
Named difficulty and minimal pairs
Every file declares which hazard classes it carries, from a taxonomy of 62 classes distilled from a study of 627 real documents (unit.glued_to_value appears on 22% of real documents;
unit.in_header_or_reference_only on 9%). For each generatable class, the generator renders the
same person’s same visit twice — clean and hazard-bearing — so “how much does this phenomenon
cost” becomes a paired, causal estimate rather than an observational correlation.
Values that obey physiology
No value is sampled from a distribution fitted to patient data. Reference intervals cite China’s WS/T 404 (biochemistry) and WS/T 405 (haematology) series; within- and between-subject variation follows the public EFLM / Westgard tables; derived quantities (BMI, LDL, MCH/MCHC, eGFR) are computed by their defining identities. A cross-generator audit the project runs on its own check-ups is telling: Synthea and PySynthea complete blood counts mostly violate the red-cell identity MCH = MCHC × MCV and carry zero reference intervals. ESL-Doc carries both, by construction.Privacy by construction
Nothing in a build is derived from a real person. The document-shape statistics came from a private de-identified reference set as format tokens and aggregate counts only, each carrying a provenance tag (source), behind an allow-list .gitignore, a pre-commit privacy hook, and an
n-gram replay gate. The threat model is written down in the repository’s docs/PRIVACY.md,
including what to do if you believe a file is not synthetic.
Using it
--lang-mix, the same --paraphrase) are
byte-identical, so out/ is disposable and the generator plus the seed is the artefact.
Status. Mirobody reads a build directory through an environment variable; as of engine
release 1.5.3 this is not yet wired (feat/1.5.4 does not yet read a mirobody-gen build), and
today’s consumers are the scoring CLIs plus the vendor-payload shapes pinned against the decoder
test suite. The working article composing this corpus with
ESL-Bench — provisionally the ESL-Doc benchmark — lives in
the repository’s docs/zh-CN/paper.md.