Skip to content
Back to Liquid Agent DOCUMENTATION

Blood-Signal Encoding

The project keeps one internal blood-signal encoding core, with type-specific user entry scripts on top.

For normal guided use, start from liquid-agent or liquid-agent web and ask to encode the active dataset. The scripts below remain direct reproducible entrypoints for advanced users.

Type-Specific Entrypoints

  • scripts/encode_cfdna_foundation_features.py
  • scripts/encode_epigenomic_signal_features.py
  • scripts/encode_lpwgs_features.py
  • scripts/encode_variant_features.py

Support Matrix

Signal family Default encoder Optional encoders Notes
cfchip_seq ntv2 dnabert2, hyenadna, caduceus, epibert, epcot, enformer, coverage_profile Use foundation encoders for interval/alignment inputs; coverage_profile for continuous tracks
cfmedip_seq epibert ntv2, dnabert2, hyenadna, caduceus, epcot, enformer, coverage_profile Default leans methylation-aware
medip_seq epibert same as above Same routing policy as cfmedip_seq
lpwgs / ulpwgs lpwgs_cnv_profile coverage_profile, DNA foundation encoders Foundation encoders remain exploratory here
ctdna_variant / variant vcf_signature variant_effect_profile variant_effect_profile expects pre-annotated effect scores

Accepted Inputs

Depending on signal family, the encoding layer supports:

  • bed, bed.gz
  • narrowPeak, broadPeak, gappedPeak
  • bam, cram
  • bedGraph, WIG, bigWig
  • vcf, vcf.gz
  • maf, maf.gz, maf.tsv, maf.tsv.gz
  • cnv_parquet

Design Rule

Use a foundation model when there is a biologically reasonable sequence-window interpretation. Otherwise, keep the signal in a deterministic profile encoder that preserves its native semantics.

Why These Defaults

The current defaults are not simply the newest available models.

  • ntv2 remains the default interval encoder because it is mature, public, and still benchmark-competitive for general genomic representation learning.
  • epibert remains the methylation-enrichment default because it is more modality-aligned than a generic sequence model, but users should treat its packaged checkpoint provenance more carefully than ntv2 or dnabert2.
  • lpwgs_cnv_profile stays the LPWGS default because LPWGS is still best treated as a copy-number / coverage problem, not a generic sequence-embedding problem.
  • vcf_signature stays the variant default because raw VCF/MAF tables are sparse event tables, and optional effect-model aggregation should only be used when the upstream annotations are genuinely present.

For the full keep / optional / watchlist reasoning, see Components and Methods.

Core References

Examples

cfDNA foundation:

python scripts/encode_cfdna_foundation_features.py \
  --input_dir <dataset_or_subdir> \
  --fasta <fasta_path> \
  --model ntv2

Epigenomic:

python scripts/encode_epigenomic_signal_features.py \
  --signal cfchip_seq \
  --input_format bed.gz \
  --input_dir <input_dir> \
  --encoder ntv2

LPWGS:

python scripts/encode_lpwgs_features.py \
  --signal lpwgs \
  --input_format bed.gz \
  --input_dir <input_dir> \
  --encoder lpwgs_cnv_profile

Variant:

python scripts/encode_variant_features.py \
  --signal variant \
  --input_format vcf.gz \
  --input_dir <input_dir> \
  --encoder vcf_signature

Metadata-Aware Follow-Up

After encoding, use metadata-aware cfDNA analysis or visualization to inspect sample groups. The assistant can fuzzy-match approximate label names, such as her2, HER2, Her2, or response, against available metadata columns and ask for clarification only when multiple plausible choices remain.