Blood-Signal Encoding¶
The project keeps one internal blood-signal encoding core, with type-specific user entry scripts on top.
For normal guided use, start from liquid-agent or liquid-agent web and ask
to encode the active dataset. The scripts below remain direct reproducible
entrypoints for advanced users.
Type-Specific Entrypoints¶
scripts/encode_cfdna_foundation_features.pyscripts/encode_epigenomic_signal_features.pyscripts/encode_lpwgs_features.pyscripts/encode_variant_features.py
Support Matrix¶
| Signal family | Default encoder | Optional encoders | Notes |
|---|---|---|---|
cfchip_seq |
ntv2 |
dnabert2, hyenadna, caduceus, epibert, epcot, enformer, coverage_profile |
Use foundation encoders for interval/alignment inputs; coverage_profile for continuous tracks |
cfmedip_seq |
epibert |
ntv2, dnabert2, hyenadna, caduceus, epcot, enformer, coverage_profile |
Default leans methylation-aware |
medip_seq |
epibert |
same as above | Same routing policy as cfmedip_seq |
lpwgs / ulpwgs |
lpwgs_cnv_profile |
coverage_profile, DNA foundation encoders |
Foundation encoders remain exploratory here |
ctdna_variant / variant |
vcf_signature |
variant_effect_profile |
variant_effect_profile expects pre-annotated effect scores |
Accepted Inputs¶
Depending on signal family, the encoding layer supports:
bed,bed.gznarrowPeak,broadPeak,gappedPeakbam,crambedGraph,WIG,bigWigvcf,vcf.gzmaf,maf.gz,maf.tsv,maf.tsv.gzcnv_parquet
Design Rule¶
Use a foundation model when there is a biologically reasonable sequence-window interpretation. Otherwise, keep the signal in a deterministic profile encoder that preserves its native semantics.
Why These Defaults¶
The current defaults are not simply the newest available models.
ntv2remains the default interval encoder because it is mature, public, and still benchmark-competitive for general genomic representation learning.epibertremains the methylation-enrichment default because it is more modality-aligned than a generic sequence model, but users should treat its packaged checkpoint provenance more carefully thanntv2ordnabert2.lpwgs_cnv_profilestays the LPWGS default because LPWGS is still best treated as a copy-number / coverage problem, not a generic sequence-embedding problem.vcf_signaturestays the variant default because raw VCF/MAF tables are sparse event tables, and optional effect-model aggregation should only be used when the upstream annotations are genuinely present.
For the full keep / optional / watchlist reasoning, see Components and Methods.
Core References¶
- Nucleotide Transformer
- DNABERT-2
- HyenaDNA
- Caduceus
- EpiBERT
- EPCOT / EPCOTv2
- Enformer
- 2025 DNA foundation-model benchmark
Examples¶
cfDNA foundation:
python scripts/encode_cfdna_foundation_features.py \
--input_dir <dataset_or_subdir> \
--fasta <fasta_path> \
--model ntv2
Epigenomic:
python scripts/encode_epigenomic_signal_features.py \
--signal cfchip_seq \
--input_format bed.gz \
--input_dir <input_dir> \
--encoder ntv2
LPWGS:
python scripts/encode_lpwgs_features.py \
--signal lpwgs \
--input_format bed.gz \
--input_dir <input_dir> \
--encoder lpwgs_cnv_profile
Variant:
python scripts/encode_variant_features.py \
--signal variant \
--input_format vcf.gz \
--input_dir <input_dir> \
--encoder vcf_signature
Metadata-Aware Follow-Up¶
After encoding, use metadata-aware cfDNA analysis or visualization to inspect
sample groups. The assistant can fuzzy-match approximate label names, such as
her2, HER2, Her2, or response, against available metadata columns and
ask for clarification only when multiple plausible choices remain.