# Backlog for the 12-week demo The demo needs 16 features, about 48 engineer-weeks and four engineers. Feature IDs and acceptance tests come from `docs/PRD.md` section 5. Effort comes from `docs/demo-architecture.md`, "Build order". With the thin golden path built, the sprint plan commits about 35 of the 48 engineer-weeks and keeps the rest as buffer. Replay and calibration are protected; the design critique and model-based theming are the first cuts. The sprint-by-sprint plan is in `docs/sprint-plan.md`. ## Status after the first thin build (1 Oct 2026) The golden path runs end to end on the synthetic package. See `docs/build-plan.md`, "Results". Done or started: - 0.1 skeleton with the two zones and a boundary test: done (SQLite per zone, no CI yet) - 0.2 synthetic package generator: done, as Markdown and CSV. Word and scanned PDF versions are not generated yet. - 0.3 behaviour parameter schema: moved to draft 0.2 in `shared/schema/` (adds `within_study_interval`). The rows are generated, not hand-filled; the team still needs to agree the fields. - 0.6 simulator skeleton: done. Power matches the analytic value within Monte Carlo error. - Every epic has a thin version on the path. None meets its PRD acceptance test on real documents yet. ## Week 1 to 2: before any feature work | # | Item | Done when | | --- | --- | --- | | 0.1 | Repository skeleton with the two zones as separate deployable units, continuous integration, and a check that fails if a platform-zone module imports from the hospital zone | A trivial end-to-end run passes in both zones | | 0.2 | **Synthetic study package generator.** A made-up two-arm diabetes trial with 3 sites and about 300 participants. It produces a protocol and an analysis plan as Word and scanned PDF in Vietnamese and English, a patient-level dataset with dated visits, enrolment, deviation and query logs, two ethics review letters, a monitoring report and a final report. The true parameters are known, so every estimator can be checked. | The package regenerates from a seed, and its true dropout, enrolment and adherence parameters are written to a file | | 0.3 | **Behaviour parameter schema.** Hand-fill ten rows for the synthetic study against `schema/behaviour-parameter.schema.json` and change the schema where it does not fit | Ten rows validate, and the team agrees on the fields | | 0.4 | Spike: extraction accuracy on Vietnamese protocol documents. Born-digital files feed the demo; scans are measured for F1.10 only | Character and field accuracy measured on at least 5 surrogate documents | | 0.5 | Spike: mapping a raw dataset to the small clinical schema | 50 variables mapped, with the time taken recorded | | 0.6 | Spike: simulator skeleton | One design simulated 1,000 times, with power within Monte Carlo error of the analytic value | | 0.7 | Study inventory and permission check on the real studies (not an engineering task; it gates real data) | Go or no-go on real data by the end of week 2 | | 0.8 | Request to place a server inside the hospital network | Request filed in week 1 | ## Epic 1. Environment, intake, de-identification, clinical store (6 engineer-weeks) | Feature | Work | Acceptance test | | --- | --- | --- | | F1.1 | Upload, file typing, and a checklist against the 39-item package list in `docs/data-request.md`. Scanned pages listed as not read | File type correct for at least 90% of a hand-labelled set of 100 files | | F9.1 | De-identification with rule checks, a human review step and a log; the clinical store lives in the hospital zone only | No patient-level record is reachable from the platform zone; an automated test tries and fails | | F9.2 | Consent scope recorded on each dataset | A dataset without a consent scope is refused by the estimator | ## Epic 2. Document pipeline and study record (7 engineer-weeks) | Feature | Work | Acceptance test | | --- | --- | --- | | F1.2 | Schema-constrained extraction from born-digital files of arms, visits, eligibility criteria, endpoints and estimand into the USDM subset, each field linked to its source passage; a review screen for low-confidence fields; dataset mapping to the clinical schema | At least 90% correct on a hand-labelled set of 200 protocol fields and 50 dataset variables | ## Epic 3. Retrospective and feedback digest (6 engineer-weeks) | Feature | Work | Acceptance test | | --- | --- | --- | | F1.3 | Planned-against-actual table: enrolment, screen failure by criterion, dropout, variance, effect, timeline, amendments. Planned values come from the study record, and actual values arrive as aggregates from the hospital zone. | Matches a statistician's hand extraction on at least 90% of 30 items | | F1.4 | Split review letters, monitoring reports and debrief notes into comments, group them into themes, and link each to its source text; human review of themes | A reviewer agrees with at least 85% of theme assignments | ## Epic 4. Behaviour estimator and profile store (8 engineer-weeks) | Feature | Work | Acceptance test | | --- | --- | --- | | F1.5 | Time-to-event models for dropout, count models for enrolment and deviations, adherence and visit attendance; bootstrap intervals; export of aggregates only, above the minimum cell size, with an approval record | 30 parameters re-computed by hand agree; on the synthetic package the true values fall inside the stated intervals at the stated rate | | F8.2 | Label, interval type, tags and provenance stored on every parameter; the minimum-evidence rule enforced in the store | No parameter can be saved as `learned` below 30 patients or 3 sites; no screen or report shows a number without its label | ## Epic 5. Clinician review and web app shell (6 engineer-weeks) | Feature | Work | Acceptance test | | --- | --- | --- | | F1.6 | Review queue in which an investigator confirms, corrects or rejects each parameter; a structured debrief form that fills gaps as `elicited` values; four roles; an audit log | At least 80% of parameters reviewed by a named investigator; every decision appears in the audit log | | Shell | Navigation and the six screens in `design/screens/` | Each screen renders real data from the golden path | ## Epic 6. Simulator, forecast, replay and calibration (8 engineer-weeks) | Feature | Work | Acceptance test | | --- | --- | --- | | F3.4 | Adherence, missed visits and dropout drawn from the profile | Re-running the ingested study's own design reproduces its observed dropout curve. This checks the code, not the model. | | F3.6 | Discrete-event engine on a job queue; at least 1,000 runs per design and 10,000 for null scenarios; power, type I error, size and duration with intervals | Identical output on re-run from the same seed, profile version and code version | | F3.7 | The same design under global defaults and under the local profile, side by side; every global default cited or flagged | Re-run of 1,000 trials in at most 1 hour; no default shown without its citation or flag | | F5.4 | Enrolment forecast by site type with an interval | Back-tested on the ingested study's enrolment log | | F3.8 | Replay: re-fit the profile on the first two-thirds of the study, predict the last third, then reveal it. The cut must be by calendar date, and nothing from the last third may reach the fit. Calibration: coverage of every stated interval on at least 100 synthetic studies. Protected. | A test proves no record dated after the cut enters the fit; calibration coverage reported | ## Epic 7. Synopsis and report builder (4 engineer-weeks) | Feature | Work | Acceptance test | | --- | --- | --- | | F2.1 | A study brief becomes a protocol synopsis as USDM subset data | Each element of the synopsis traces to a structured field | | F2.8 | Feasibility and design report: recommended design, sample size, eligibility, enrolment forecast, retention, timeline, risks and an evidence table; sign-off; Word and PDF export | A reviewer can trace every number to its source label; export is blocked without a named signer | ## Epic 8. Gold sets, tests and dry runs (3 engineer-weeks) - Hand-labelled sets: 100 files for typing, 200 protocol fields, 50 dataset variables, 30 retrospective items, theme assignments - The golden path as one automated test on the synthetic package - Two dry runs: one with a biostatistician and one with an investigator, with every unexplained number logged as a defect - A "Limits" screen or closing section stating what one study can and cannot show ## Schedule | Weeks | Epics | Demo acts ready | | --- | --- | --- | | 1–2 | Week-1 items | Go or no-go on real data | | 3–5 | 1, 2, 3 | Acts 1 and 2: ingest and understand | | 6–8 | 4, 5, forecast from 6 | Act 3 and the enrolment forecast | | 9–10 | 6, 7 | Act 4: simulate and report | | 11–12 | Replay, 8 | Act 5 and fixes | ## Known disputes to keep in mind These are recorded in `docs/prd-review.md`. They are not resolved. - The review's critic holds that 15 features in 48 engineer-weeks is still too much, and that operations records such as deviation and query logs often sit with sponsors and research organisations, not hospitals. If the real package has none, the site layer comes from the debrief and is labelled `elicited`. - Intervals from one study understate uncertainty. The between-study spread is elicited and must be shown as such. - The demo ends on the signed design report (decided 1 Oct 2026, PRD §10).