Skip to content
Back to Liquid Agent DOCUMENTATION

Explicit Assay Tables

An assay contract makes units and measurement semantics explicit before a local analysis. It does not turn the agent into a fixed workflow. QC and differential expression are separate capabilities: the user can inspect, discuss, run, stop, redirect or decline either one. Web and CLI use the same analysis implementation.

Supported Operations

Declared assay Value kinds Computation
Digital PCR partitions Poisson-occupancy concentration, exact-binomial transformed 95% intervals, zero/saturation QC
Circulating tumour cells cell_counts Counts per mL, exact Poisson 95% intervals
Cell-free RNA, small RNA raw_counts, abundance, log_abundance Sample/feature QC, missingness, distributions, variable-feature profiles, exploratory PCA
Plasma proteins, metabolites abundance, log_abundance The same processed-matrix QC with declared units; no count-model substitution
EV cargo raw_counts, abundance, log_abundance Molecular-table QC, not proof of EV purity or origin
CTC molecular profiles raw_counts Count-matrix QC, distinct from enumeration
Methylation arrays beta Beta-table QC; no raw IDAT normalization or DMR claim
Declared independent-group RNA count contrast raw_counts only PyDESeq2 negative-binomial fitting, Wald tests and adjusted p-values

Install optional scientific dependencies in the environment used to launch the agent:

python -m pip install -e '.[assay-tables]'

The supported PyDESeq2 0.5.4 requires NumPy 2 and is not compatible with the legacy optional blood-models dependency set's NumPy constraint. Do not force an incompatible combined environment; those legacy encoder runtimes need separate dependency resolution. The normal agent can inspect prerequisites without installing all optional model runtimes.

A Source Contract

Put a small JSON file ending in .assay.json beside its input tables. For example:

{
  "version": 1,
  "assay": "cell-free-rna",
  "table": "counts.csv",
  "value_kind": "raw_counts",
  "units": "reads",
  "metadata": "metadata.csv",
  "group_column": "condition",
  "contrast": ["case", "control"],
  "independent_samples": true,
  "synthetic": false
}

Matrix tables are CSV or TSV with a unique feature_id first column followed by unique sample columns. Metadata has unique sample_id values and the declared group column. No groups are inferred from filenames or sample names.

feature_id,S01,S02,S03,S04,S05,S06
GENE_A,12,18,14,42,38,45
GENE_B,31,29,34,27,30,32

This tiny schema example is not large enough for dispersion fitting. The differential wrapper requires at least ten features with total count >= 10, exactly two declared groups, at least three independent samples per group and exact metadata/sample matching. Paired, patient, batch or time covariates require a reviewed multifactor workflow and are rejected by this single-factor wrapper. Those checks are minimum technical gates, not a power calculation.

Omit contrast, metadata and independent_samples for an unlabelled QC-only source. Missing values in abundance/beta tables stay missing. Missing raw counts, negative counts, fractional counts, invalid beta values and duplicate IDs are rejected rather than silently repaired. Relative input paths must stay inside the manifest folder; symlink escapes are rejected.

Digital PCR and CTC Schemas

Digital PCR uses assay: digital-pcr and value_kind: partitions. Its table needs:

sample_id,target,positive_partitions,total_partitions,partition_volume_nl,dilution_factor
S01,TARGET_A,120,18000,0.85,1
S02,TARGET_A,0,18000,0.85,2

Counts refer to the accepted partitions after upstream gating. Concentration is -ln(1 - positive/total) * 1000 / partition_volume_nl * dilution_factor, in copies per microlitre of the input before that dilution. Plasma-volume conversion requires further explicit extraction/recovery information and is not inferred. Zero positives retains a positive upper uncertainty bound. Fully positive wells are saturated; the report marks their point estimate and upper bound unbounded, not zero. Replicate aggregation, blank/detection thresholds and variant duplex cross-talk fitting are not supplied by this count wrapper.

CTC enumeration uses assay: circulating-tumour-cells, value_kind: cell_counts:

sample_id,cell_count,volume_ml
S01,10,7.5
S02,0,7.5

Cell identity and accepted counting criteria must already be established. The engine does not estimate enrichment efficiency or assign a clinical threshold.

Use From Web or CLI

Attach the source folder normally. Ask the LLM to inspect its assay contract, then request a focused plan or authorize the desired operation. inspect_workspace exposes distinct QC and contrast task IDs; a question alone runs neither.

If no contract exists, you can supply the same details in natural language:

This CSV contains accepted digital-PCR partitions. Its columns are sample_id, target, positive_partitions, total_partitions, partition_volume_nl and dilution_factor. Configure it for occupancy quantification in copies/uL input. Do not run the analysis yet.

The configure_assay tool validates a structured declaration and copies only the declared inputs into this conversation's workspace. It does not write into the original source or run the analysis. The model must ask about unknown assay, units or experimental-design details rather than guess. Configuration is not evidence that upstream gating, sample identity or clinical validation is correct.

For a deterministic, reproducible CLI operation without an LLM:

liquid-agent assay /path/to/source/input.assay.json --output /path/to/new/qc-run
liquid-agent assay /path/to/source/input.assay.json --output /path/to/new/contrast-run --differential

The second command is an explicit contrast request. The first never runs the contrast merely because it is present in the manifest.

Outputs and Interpretation

Each completed run publishes a Markdown report, real PNG figures, full CSV measurements/QC tables and input hashes. The report embeds figure links and a bounded table preview. Original tables are never overwritten. Existing output directories are rejected; cancelled or failed runs do not publish a completed report. Agent runs live in conversation-owned workspaces and follow the same Trash/restore/purge rules as other results.

Distribution plots show at most 30 samples; heatmaps show at most 40 variable features. Full CSVs preserve all supplied measurements. PCA uses up to 2,000 complete, variable features across all samples and reports that bound; incomplete and constant features are excluded rather than replaced by invented zeros. Count QC uses log2(CPM+1) for display; PyDESeq2 receives the unnormalized integer counts, not those plots. Model-fitting warnings and non-estimable tests are retained in the scientific limitations and a diagnostic record. Differential reports show effects and adjusted p-values, not only sample QC. Technical provenance stays in audit files, not in the report's default scientific table section.

Reports must distinguish exploratory findings from evidence of clinical utility. Processed-table support is not a claim of end-to-end alignment, protein identification, metabolite identification or real-cohort validation.

Reused Scientific Methods

The implementation calls maintained libraries rather than recreating a count model: PyDESeq2 and its documented workflow, statsmodels binomial intervals, and SciPy's inverse chi-square distribution for exact Poisson intervals. Input contracts, ownership, cancellation and scientific reporting remain Liquid Agent's code.