Sprint 2 board
What was planned, what is done, what is blocked and on whom, and whether the checks are green.
Sprint 1 · Sprint 2
Sprint goal
A real Word or PDF protocol goes in and comes out as structured fields with page-level sources, and the web app shows acts 1 and 2. Added mid-sprint by the owner: synthetic studies for several diseases, and a protocol tool that starts from a disease or a published trial.
33 of 45 engineer-days done · gate: Sprint 2 review: extraction on Word and PDF due 2026-10-30, open Word and PDF extraction works on synthetic protocols in both languages. It has not met a real protocol, and the model has not been called live: no API key in this environment.
Engineering items
| ID | Item | Owner | Estimate | Status | Evidence |
|---|---|---|---|---|---|
| S2-1 | Hosted model behind the gateway; prompts and schemas in version control | ML1 | 3d | partial | platform_zone/gateway/claude.py: Claude Opus 5.5, structured output, effort medium, cached system prompt, server-side fallback, refusal and error handling. The gateway also checks the study's data-use terms. Prompt and schema in platform_zone/documents/prompts/. Tested against a fake client (tests/test_model_extraction.py). Never called live: no ANTHROPIC_API_KEY here. |
| S2-2 | Model field extraction v1, compared with the rule baseline | ML1 | 5d | partial | platform_zone/documents/model_extract.py. Each value needs a passage id and an exact quote; the quote must be in the passage and a number must be in its quote, or the value is marked unsupported. scripts/extraction_eval.py scores rules and model against known truth. Rules 100% on 18 synthetic protocols. The model is not yet measured, and there is no hand-labelled gold set. |
| S2-3 | Package, retrospective and feedback screens in the web app; package page redesign (design #4) | FS | 5d | done | FastAPI app serves every act-1 and act-2 screen. The package page now leads with the checklist summary, a 'What this study can teach' panel by layer and source label, and a data-quality card; the file list sits underneath. New: Read a protocol (/upload), with datasets refused. |
| E-T7 | Two command-line entry points; golden path as two processes | DE | 3d | done | python -m hospital_zone prepare|approve|console and python -m platform_zone documents|trust|import|design. tests/test_two_process.py runs the stroke study as separate processes and checks neither zone imports the other. |
| E-T8 | Site types read from the data | ML2 | 3d | done | Estimator, simulator, forecast, synopsis and report take site types from the data and the brief. A test with a 'district health centre' type found the forecast silently dropping unknown types; fixed. |
| S2-4 | Estimator hardening: visit window from the protocol, missing dates, data quality | ML2 | 4d | done | The protocol's window travels to the hospital in protocol_params.json and sets the missed-visit rule and the outside-window count. An undated attended visit is placed on its planned date and counted; entries dated before their visit leave the lag estimate. Data-quality counts travel as cells. tests/test_estimator_hardening.py. |
| D-T2 | Hospital-zone UI: de-identification review and export approval | FS | 3d | done | hospital_zone/console.py on port 8820, separate from the platform app and importing no platform code. A run can stop at the hospital's approval step; the data steward approves by name; the platform imports after the signature check. |
| E-T11 | Observed effect gated by data-use terms | ML2 | 1d | done | When results_may_leave_hospital is false, observed effects are held back with the reason. Test in tests/test_estimator_hardening.py. |
| S2-5 | PDF reading with page-level sources | ML1 | 2d | done | Born-digital PDFs read line by line, anchored page-N-line-M. A page without a text layer is refused by name. |
| S2-6 | Synthetic studies for three diseases and three endpoint kinds (owner request) | DE | 4d | done | synthetic/diseases.py: type 2 diabetes (continuous), minor stroke or TIA (binary, 90 days), EGFR-mutated NSCLC (progression-free survival). Each comes with Word, PDF and Markdown protocols in two languages, patient-level tables and letters. Injected data-quality problems use their own random stream. |
| S2-7 | Endpoint-generic simulator, synopsis and report | BS | 4d | done | Binary (risk difference) and time-to-event (hazard ratio) simulated and checked against analytic power, with unequal allocation. The report adds power at the weak end of the planning range. Endpoint and trial names with digits go in a per-report glossary. |
| S2-8 | Protocol tool: start from an indication or a published trial (owner request) | FS | 4d | done | /design in the demo app. The library holds FLAURA, KEYNOTE-024, KEYNOTE-189 and CHANCE, with each fact marked checked or not yet checked against the paper; no protocol text is copied. Outputs: a synopsis in Word and Markdown in both languages, simulations and a design report that can be signed. The local profile is used only for the same indication. |
| Stretch | Real dataset mapping, if the gate said go | DE | 4d | not started | No real data. The go or no-go on real data (16 Oct) has not been taken. |
Measured this sprint
| What | Result | Note |
|---|---|---|
| Rule extraction against known truth | 100% of fields, 18 protocols | Synthetic text written beside the rules: an upper bound, not a real number. |
| Model extraction against known truth | not measured | No API key. scripts/extraction_eval.py runs it when one is set. |
| Simulated against analytic power | continuous 0.861 against 0.861; binary 0.815 against 0.805; time to event 0.797 against 0.783; time to event at 2:1, 0.752 against 0.752 | Measured once, at the start of the sprint. The tests allow 3 points. Type I error is near 5% for every kind. |
| Interval calibration, 150 synthetic studies | parameters 90.3% (stated 90%); replay 86.0% (stated 80%); back-test 75.3% (stated 80%) | The back-test interval is slightly overconfident. Recorded, not tuned. |
| Test suite | 147 tests passing | scripts/ci.py also runs the golden path for all three diseases. |
Pulled forward from later sprints
| From | Item | Status |
|---|---|---|
| design #4 | Package page leads with checklist and layers | done |
| Sprint 3 | De-identification review and export approval screens in the hospital zone | done |
Checks
No CI run yet.
Decisions taken this sprint
| Decision | Taken | Why |
|---|---|---|
| Hosted model for documents | Proposed: Claude Opus 5.5 through the gateway, structured output, documents only, gated by the study's data-use terms | Needs the owner's confirmation. Until a key is set, the rule baseline runs alone and screens say so. |
| When the local profile is used | Only when the ingested study's indication matches the brief's | Diabetes clinic behaviour is not evidence about an oncology trial. Otherwise the design runs on global defaults and published figures, and says so. |
| Published trials in the library | Facts only, each with a citation and a checked flag; no protocol text copied | Copyright, and a figure nobody has checked must not look checked. |
| Data-quality problems in synthetic data | Injected from their own random stream; the de-identification regression fixture regenerated | Only the two data-entry lag estimates moved, because bad rows are now left out. The de-identification code did not change. |
| Uploads to the platform | Documents only; datasets refused by file type | Hard rule 1: patient-level data is loaded in the hospital zone. |
Not engineering, still on this sprint
| Item | Owner | Status |
|---|---|---|
| Label the gold sets: 100 files for typing, 200 protocol fields | PL, RA, CL | not started No update recorded. Model extraction cannot be measured honestly without it. |
| Schedule the investigator debriefs | PL | not started No update recorded. |
| Confirm the hosted model provider and supply a key | PL | not started Built for Claude Opus 5.5 behind the gateway; needs the owner's yes and an API key. |
| Check the library figures marked not yet checked against their papers | BS | not started FLAURA enrolment and registry, KEYNOTE-189 registry, CHANCE event proportions. |
Documents
PRD · CEO review · Engineering review · Design review · Sprint plan · Build plan · Design system · Backlog