Frozen training and validation studies¶
Use an explicit prediction-study contract when you already have liquid-biopsy feature tables and a reviewed training/validation split. This complements the assay-table workflows; it does not infer clinical labels, automatically design a valid clinical cohort, or turn exploratory classification into external validation.
Prepare the source¶
Keep a JSON contract and its four input tables in the source folder. Tables may
be CSV, TSV or Parquet, with one row per sample and a shared sample_id column.
Feature tables contain numeric columns with matching definitions. Metadata needs
binary label values (0 and 1) and patient_id. Each partition must contain both
classes, with at least five training observations per class. Samples and patients
must not overlap between partitions. Repeated patient measurements require a
separate reviewed longitudinal design and are rejected by this workflow.
For example, study.prediction.json:
{
"version": 1,
"assay": "small-rna",
"value_kind": "processed_features",
"train_features": "train.csv",
"validation_features": "validation.csv",
"train_metadata": "train_metadata.csv",
"validation_metadata": "validation_metadata.csv",
"independent_patients_confirmed": false,
"split_description": "A fixed internal holdout; patient identity requires source review.",
"preprocessing_provenance": "Published processed array intensities; upstream normalization was not reproduced.",
"limitations": ["Internal holdout only; batch balance and patient independence require review."],
"models": ["logistic", "rbf_svm", "random_forest"],
"max_features": 200,
"specificity_target": 0.95,
"seed": 42
}
Do not replace unknown patient identities with invented identities and claim
independence. independent_patients_confirmed: false preserves this limitation;
it does not bypass detected overlaps. All paths must stay within the source.
Raw counts must be complete nonnegative integers; beta values must be within
0–1. Other declared numeric inputs can contain missing values, handled inside
training pipelines. Published upstream processing remains a separate limitation.
Use it in Web or CLI¶
- Attach the folder and ask: “Inspect the prediction contract and explain the split, measurement type and prerequisites. Do not fit yet.”
- After review, ask: “Run the declared study, retain training-only model and threshold selection, and publish ROC, PR, calibration and performance results.”
- Ask: “Reopen the saved summary. Which model was selected on training, and what specificity did it actually achieve on the holdout?”
- For different feature sets with the same cohort, ask to compare their completed studies using each contract and its registered summary. The comparison checks frozen predictions and inputs, then produces paired AUROC differences and confidence intervals without refitting the parent models.
Both interfaces use the same scientific tools. Prediction studies currently run
through conversational tools (inspect_prediction_study, run_prediction_study,
compare_prediction_studies); they do not create standalone prediction entries
for the Plan panel's Run next step button. Existing explicit assay entries
remain discoverable and runnable when their contracts share the folder.
Model choice uses training CV. Preprocessing and feature selection are fitted within training folds. The selected pipeline's OOF predictions determine a training specificity target; this target is not guaranteed on validation. The stored choice and validation probabilities are frozen before scoring. OOF threshold estimates reuse the training-selected hyperparameters and are not an independent, fully nested estimate of model-selection performance.
Prediction fitting and comparison share the existing six-scientific-jobs-per-turn
limit. They run locally and check cancellation between phases; Stop may wait for
an in-progress scikit-learn fit to return. This differs from the isolated worker
used by run_analysis.
Results and boundaries¶
Outputs include the input fingerprints, frozen training choice, model files, validation probabilities, per-model metrics, bootstrap AUROC intervals, ROC/PR/ calibration figures and a local report. They follow normal task ownership and trash/restore/purge rules. Large reports can appear in Results; brief discussion can stay in the conversation. Saved summary reads verify the probability-file hash and identify the selected model rather than average scores across models. ROC, precision–recall and calibration figures expose their types with their registered handles. Reports use locally bound captions for these figures so a model-supplied caption cannot swap their meanings.
Comparisons require identical training metadata and matching validation samples, labels and patient groups; current training-metadata comparison is byte-for-byte, so even reordered equivalent metadata may need reconciliation before fitting. Intervals are exploratory and not multiplicity adjusted. A confidence interval crossing zero does not establish superiority or equivalence. Paired differences explicitly mean study A minus study B.
Existing source data, frozen parent results and accepted personal skills are not changed by inspection or comparison. Starting a study is a new computation. Training-only processing cannot remove leakage already introduced in published inputs. High internal AUROC, including 1.0, is not clinical screening validation.