Skip to content
Back to Liquid Agent DOCUMENTATION

Metadata and Label-Aware Planning

Liquid Agent treats metadata and labels as part of dataset understanding, not only as optional plot-colouring inputs. During scan, the assistant profiles candidate metadata files, estimates whether labels are usable, and then chooses an unsupervised, grouped, or exploratory supervised analysis route.

This is still exploratory research analysis. It does not produce diagnostic claims.

What Is Detected

The scanner looks for user-provided metadata and label tables in common formats:

  • CSV and TSV
  • Excel workbooks (.xlsx, .xls)
  • Parquet
  • JSON arrays
  • JSON objects containing samples, metadata, records, data, or rows
  • JSONL / NDJSON records

Generated output folders such as analysis/, visualisation/, visualization/, assistant/, preprocessed/, preprocessing/, features/, models/, and run-output folders are excluded from metadata detection so result tables are not accidentally treated as user labels.

Metadata Profile

Each scanned project receives a metadata_profile attached to the project profile and exposed to Web scan/plan responses.

The profile records:

  • candidate metadata tables
  • selected metadata table
  • sample identifier column
  • selected label column
  • label classes and counts
  • labeled and matched sample counts
  • sample coverage
  • confidence
  • supervision mode
  • recommended modeling backend
  • warnings and ambiguity notes

The selection priority is:

  1. explicit user target or manual override
  2. high sample coverage
  3. common label names such as condition, status, response, label, group, class, cancer, or control
  4. reasonable categorical class distribution

Planning Modes

Mode Trigger Planner behavior
Unsupervised no reliable classification label, one class after sample matching, or ignored metadata QC, PCA/UMAP where applicable, outlier summaries, clustering-style summaries, raw-signal and matrix review
Grouped labels exist but sample size or class balance is too small for model training grouped summaries, group-coloured plots, feature/effect summaries, no classifier training
Supervised labels have sufficient matched samples and at least two usable classes grouped outputs plus exploratory supervised modeling

Current automatic supervised threshold is intentionally conservative: at least 20 matched labeled samples, 2 to 5 classes, and at least 5 samples in the smallest class.

Modeling Backends

The default supervised backend is lightweight and deterministic:

  • sklearn_logistic_regression with balanced class weights and stratified cross-validation when scikit-learn is available
  • pytorch_linear_probe when PyTorch is available and the labeled sample count is at least 100

If neither backend is available, planning falls back to grouped analysis and the report states why supervised modeling was skipped.

For high-dimensional raw matrices, the analysis layer prefers existing feature stores or embeddings when available. Otherwise it uses conservative feature selection/PCA-style preparation before training. Output reports mark all model metrics as exploratory.

CLI Control

Inside liquid-agent:

/metadata
/metadata use <table> <sample_col> <label_col>
/metadata ignore
/metadata rescan

Use /metadata to inspect candidates and the current mode. Use /metadata use ... when the automatic selection chose the wrong table or label. Use /metadata ignore to force an unsupervised plan for the current session. After changing metadata state, run /plan again.

Natural-language equivalents also work:

Use metadata.csv and group samples by condition.
Ignore labels for this run and focus on unsupervised QC.
Use the response column for grouped plots and exploratory modeling.

Web Client

The Web client shows a compact Metadata card under Sources:

  • selected label or no active label
  • coverage
  • mode
  • backend
  • confidence
  • Change / Ignore controls

The full table is not shown by default. The card is meant to explain the planning decision without turning the sidebar into a spreadsheet viewer.

Reports

Run reports include a Metadata and Label Decision section when metadata was available or explicitly ignored. It states:

  • which table and label column were used
  • how many samples matched
  • why the planner chose unsupervised, grouped, or supervised mode
  • which modeling backend was selected or skipped
  • which samples or labels were insufficient when relevant

This section is designed to make the analysis route auditable for collaborators who were not present during the interactive session.