Sprint 2 · 19–30 Oct 2026 (run early, 1 Oct)

Sprint 2 board

What was planned, what is done, what is blocked and on whom, and whether the checks are green.

Sample values. Synthetic study, not real patients.
Demo role: Viewer: reads only. Switch. Roles are a demo switch, not sign-in.

Sprint 1 · Sprint 2

Sprint goal

A real Word or PDF protocol goes in and comes out as structured fields with page-level sources, and the web app shows acts 1 and 2. Added mid-sprint by the owner: synthetic studies for several diseases, and a protocol tool that starts from a disease or a published trial.

33 of 45 engineer-days done · gate: Sprint 2 review: extraction on Word and PDF due 2026-10-30, open Word and PDF extraction works on synthetic protocols in both languages. It has not met a real protocol, and the model has not been called live: no API key in this environment.

Engineering items

IDItemOwnerEstimateStatusEvidence
S2-1Hosted model behind the gateway; prompts and schemas in version controlML13dpartialplatform_zone/gateway/claude.py: Claude Opus 5.5, structured output, effort medium, cached system prompt, server-side fallback, refusal and error handling. The gateway also checks the study's data-use terms. Prompt and schema in platform_zone/documents/prompts/. Tested against a fake client (tests/test_model_extraction.py). Never called live: no ANTHROPIC_API_KEY here.
S2-2Model field extraction v1, compared with the rule baselineML15dpartialplatform_zone/documents/model_extract.py. Each value needs a passage id and an exact quote; the quote must be in the passage and a number must be in its quote, or the value is marked unsupported. scripts/extraction_eval.py scores rules and model against known truth. Rules 100% on 18 synthetic protocols. The model is not yet measured, and there is no hand-labelled gold set.
S2-3Package, retrospective and feedback screens in the web app; package page redesign (design #4)FS5ddoneFastAPI app serves every act-1 and act-2 screen. The package page now leads with the checklist summary, a 'What this study can teach' panel by layer and source label, and a data-quality card; the file list sits underneath. New: Read a protocol (/upload), with datasets refused.
E-T7Two command-line entry points; golden path as two processesDE3ddonepython -m hospital_zone prepare|approve|console and python -m platform_zone documents|trust|import|design. tests/test_two_process.py runs the stroke study as separate processes and checks neither zone imports the other.
E-T8Site types read from the dataML23ddoneEstimator, simulator, forecast, synopsis and report take site types from the data and the brief. A test with a 'district health centre' type found the forecast silently dropping unknown types; fixed.
S2-4Estimator hardening: visit window from the protocol, missing dates, data qualityML24ddoneThe protocol's window travels to the hospital in protocol_params.json and sets the missed-visit rule and the outside-window count. An undated attended visit is placed on its planned date and counted; entries dated before their visit leave the lag estimate. Data-quality counts travel as cells. tests/test_estimator_hardening.py.
D-T2Hospital-zone UI: de-identification review and export approvalFS3ddonehospital_zone/console.py on port 8820, separate from the platform app and importing no platform code. A run can stop at the hospital's approval step; the data steward approves by name; the platform imports after the signature check.
E-T11Observed effect gated by data-use termsML21ddoneWhen results_may_leave_hospital is false, observed effects are held back with the reason. Test in tests/test_estimator_hardening.py.
S2-5PDF reading with page-level sourcesML12ddoneBorn-digital PDFs read line by line, anchored page-N-line-M. A page without a text layer is refused by name.
S2-6Synthetic studies for three diseases and three endpoint kinds (owner request)DE4ddonesynthetic/diseases.py: type 2 diabetes (continuous), minor stroke or TIA (binary, 90 days), EGFR-mutated NSCLC (progression-free survival). Each comes with Word, PDF and Markdown protocols in two languages, patient-level tables and letters. Injected data-quality problems use their own random stream.
S2-7Endpoint-generic simulator, synopsis and reportBS4ddoneBinary (risk difference) and time-to-event (hazard ratio) simulated and checked against analytic power, with unequal allocation. The report adds power at the weak end of the planning range. Endpoint and trial names with digits go in a per-report glossary.
S2-8Protocol tool: start from an indication or a published trial (owner request)FS4ddone/design in the demo app. The library holds FLAURA, KEYNOTE-024, KEYNOTE-189 and CHANCE, with each fact marked checked or not yet checked against the paper; no protocol text is copied. Outputs: a synopsis in Word and Markdown in both languages, simulations and a design report that can be signed. The local profile is used only for the same indication.
StretchReal dataset mapping, if the gate said goDE4dnot startedNo real data. The go or no-go on real data (16 Oct) has not been taken.

Measured this sprint

WhatResultNote
Rule extraction against known truth100% of fields, 18 protocolsSynthetic text written beside the rules: an upper bound, not a real number.
Model extraction against known truthnot measuredNo API key. scripts/extraction_eval.py runs it when one is set.
Simulated against analytic powercontinuous 0.861 against 0.861; binary 0.815 against 0.805; time to event 0.797 against 0.783; time to event at 2:1, 0.752 against 0.752Measured once, at the start of the sprint. The tests allow 3 points. Type I error is near 5% for every kind.
Interval calibration, 150 synthetic studiesparameters 90.3% (stated 90%); replay 86.0% (stated 80%); back-test 75.3% (stated 80%)The back-test interval is slightly overconfident. Recorded, not tuned.
Test suite147 tests passingscripts/ci.py also runs the golden path for all three diseases.

Pulled forward from later sprints

FromItemStatus
design #4Package page leads with checklist and layersdone
Sprint 3De-identification review and export approval screens in the hospital zonedone

Checks

No CI run yet.

Decisions taken this sprint

DecisionTakenWhy
Hosted model for documentsProposed: Claude Opus 5.5 through the gateway, structured output, documents only, gated by the study's data-use termsNeeds the owner's confirmation. Until a key is set, the rule baseline runs alone and screens say so.
When the local profile is usedOnly when the ingested study's indication matches the brief'sDiabetes clinic behaviour is not evidence about an oncology trial. Otherwise the design runs on global defaults and published figures, and says so.
Published trials in the libraryFacts only, each with a citation and a checked flag; no protocol text copiedCopyright, and a figure nobody has checked must not look checked.
Data-quality problems in synthetic dataInjected from their own random stream; the de-identification regression fixture regeneratedOnly the two data-entry lag estimates moved, because bad rows are now left out. The de-identification code did not change.
Uploads to the platformDocuments only; datasets refused by file typeHard rule 1: patient-level data is loaded in the hospital zone.

Not engineering, still on this sprint

ItemOwnerStatus
Label the gold sets: 100 files for typing, 200 protocol fieldsPL, RA, CLnot started
No update recorded. Model extraction cannot be measured honestly without it.
Schedule the investigator debriefsPLnot started
No update recorded.
Confirm the hosted model provider and supply a keyPLnot started
Built for Claude Opus 5.5 behind the gateway; needs the owner's yes and an API key.
Check the library figures marked not yet checked against their papersBSnot started
FLAURA enrolment and registry, KEYNOTE-189 registry, CHANCE event proportions.

Documents

PRD · CEO review · Engineering review · Design review · Sprint plan · Build plan · Design system · Backlog