Metadata and Label-Aware Planning¶
Liquid Agent treats metadata and labels as part of dataset understanding, not only as optional plot-colouring inputs. During scan, the assistant profiles candidate metadata files, estimates whether labels are usable, and then chooses an unsupervised, grouped, or exploratory supervised analysis route.
This is still exploratory research analysis. It does not produce diagnostic claims.
What Is Detected¶
The scanner looks for user-provided metadata and label tables in common formats:
- CSV and TSV
- Excel workbooks (
.xlsx,.xls) - Parquet
- JSON arrays
- JSON objects containing
samples,metadata,records,data, orrows - JSONL / NDJSON records
Generated output folders such as analysis/, visualisation/,
visualization/, assistant/, preprocessed/, preprocessing/, features/,
models/, and run-output folders are excluded from metadata detection so result
tables are not accidentally treated as user labels.
Metadata Profile¶
Each scanned project receives a metadata_profile attached to the project
profile and exposed to Web scan/plan responses.
The profile records:
- candidate metadata tables
- selected metadata table
- sample identifier column
- selected label column
- label classes and counts
- labeled and matched sample counts
- sample coverage
- confidence
- supervision mode
- recommended modeling backend
- warnings and ambiguity notes
The selection priority is:
- explicit user target or manual override
- high sample coverage
- common label names such as
condition,status,response,label,group,class,cancer, orcontrol - reasonable categorical class distribution
Planning Modes¶
| Mode | Trigger | Planner behavior |
|---|---|---|
| Unsupervised | no reliable classification label, one class after sample matching, or ignored metadata | QC, PCA/UMAP where applicable, outlier summaries, clustering-style summaries, raw-signal and matrix review |
| Grouped | labels exist but sample size or class balance is too small for model training | grouped summaries, group-coloured plots, feature/effect summaries, no classifier training |
| Supervised | labels have sufficient matched samples and at least two usable classes | grouped outputs plus exploratory supervised modeling |
Current automatic supervised threshold is intentionally conservative: at least 20 matched labeled samples, 2 to 5 classes, and at least 5 samples in the smallest class.
Modeling Backends¶
The default supervised backend is lightweight and deterministic:
sklearn_logistic_regressionwith balanced class weights and stratified cross-validation when scikit-learn is availablepytorch_linear_probewhen PyTorch is available and the labeled sample count is at least 100
If neither backend is available, planning falls back to grouped analysis and the report states why supervised modeling was skipped.
For high-dimensional raw matrices, the analysis layer prefers existing feature stores or embeddings when available. Otherwise it uses conservative feature selection/PCA-style preparation before training. Output reports mark all model metrics as exploratory.
CLI Control¶
Inside liquid-agent:
/metadata
/metadata use <table> <sample_col> <label_col>
/metadata ignore
/metadata rescan
Use /metadata to inspect candidates and the current mode. Use
/metadata use ... when the automatic selection chose the wrong table or label.
Use /metadata ignore to force an unsupervised plan for the current session.
After changing metadata state, run /plan again.
Natural-language equivalents also work:
Use metadata.csv and group samples by condition.
Ignore labels for this run and focus on unsupervised QC.
Use the response column for grouped plots and exploratory modeling.
Web Client¶
The Web client shows a compact Metadata card under Sources:
- selected label or
no active label - coverage
- mode
- backend
- confidence
- Change / Ignore controls
The full table is not shown by default. The card is meant to explain the planning decision without turning the sidebar into a spreadsheet viewer.
Reports¶
Run reports include a Metadata and Label Decision section when metadata was available or explicitly ignored. It states:
- which table and label column were used
- how many samples matched
- why the planner chose unsupervised, grouped, or supervised mode
- which modeling backend was selected or skipped
- which samples or labels were insufficient when relevant
This section is designed to make the analysis route auditable for collaborators who were not present during the interactive session.