Components And Methods¶
This page explains the main liquid-biopsy components used by the current product layer, why they are used, and which newer methods are worth watching. The selection principle is conservative: prefer mature, inspectable, reproducible tools over newer components that are impressive but hard to operate safely.
For a table-first inventory of all callable workflows, encoders, external runtimes, method guidance, LLM engines, and supported data types, see Capability Matrix.
Decision Principles¶
- Keep raw biological semantics visible whenever possible.
- Use foundation encoders only when the input representation matches the biology.
- Prefer deterministic profiles for sparse event tables and copy-number style signals.
- Keep preprocessing explicit and auditable.
- Treat exploratory plots as hypothesis-generating, not clinically validated evidence.
- Track newer models, but do not make them defaults until access, weights, runtime, and validation are practical.
Blood-Signal Preprocessing¶
Epigenomic Interval Preprocessing¶
Used for cfChIP, cfMeDIP/MeDIP, peak-like intervals, and enrichment-style tables.
Why it is kept:
- interval cleanup and region summarization are mature and easy to audit
- background-aware summaries reduce misleading signal shifts
- downstream feature stores can consume region-level tables directly
LPWGS / ULPWGS Preprocessing¶
Used for shallow whole-genome cfDNA copy-number style analysis.
Why it is kept:
- LPWGS is primarily a coverage and copy-number problem
- GC-aware bin preparation, segmentation, and arm-burden summaries are more defensible than generic sequence embeddings for this signal family
- outputs connect directly to CNV-focused analysis and visualization
Variant Preprocessing¶
Used for VCF, MAF, and variant-style tables.
Why it is kept:
- variant data are sparse event tables, not continuous sequence windows
- matched-normal awareness, strict somatic-style filtering, and VAF summaries are central to credible interpretation
- effect-model aggregation should only be used when upstream annotations are genuinely present
Blood-Signal Encoders¶
| Encoder | Current role | Why used |
|---|---|---|
ntv2 |
default cfChIP interval encoder | mature public genomic foundation model; strong general-purpose DNA representations |
epibert |
default cfMeDIP / MeDIP route | better aligned with methylation-enrichment style questions than generic sequence models |
lpwgs_cnv_profile |
default LPWGS / ULPWGS encoder | preserves coverage and copy-number semantics |
coverage_profile |
continuous track encoder | deterministic, inspectable profile for bedGraph/WIG/bigWig style inputs |
vcf_signature |
default variant encoder | robust sparse-event summary for VCF/MAF tables |
variant_effect_profile |
optional variant encoder | useful only when effect scores such as CADD, SpliceAI, or DeepSEA already exist |
dnabert2 |
optional sequence encoder | mature DNA language model baseline |
hyenadna |
optional long-context sequence encoder | useful watchlist model for long genomic contexts |
caduceus |
optional sequence encoder | bidirectional DNA state-space architecture worth tracking |
epcot |
optional regulatory sequence encoder | strong regulatory modeling family; use only when input assumptions fit |
enformer |
optional regulatory sequence encoder | established enhancer/promoter-style sequence model baseline |
Default Encoder Rationale¶
Nucleotide Transformer¶
Nucleotide Transformer remains the most practical default for general DNA interval representation because it is public, mature, and broadly benchmarked.
Watchlist: NT v3 should be evaluated when a stable public workflow and local resource profile are clear enough for routine users.
Reference: Nucleotide Transformer
EpiBERT¶
EpiBERT is kept as the methylation-enrichment default because it is closer to the assay semantics than generic DNA sequence embedding. It should still be documented carefully because checkpoint provenance and operating assumptions matter.
Reference: EpiBERT
LPWGS CNV Profile¶
For LPWGS/ULPWGS, the correctness-first route is coverage and copy-number profiling. A generic sequence model can obscure the signal that analysts actually need to inspect.
VCF Signature¶
Variant tables should start with event-count, VAF, gene, recurrence, and filtering summaries. Effect-model features are optional and annotation-dependent.
cfDNA Analysis Methods¶
Standard cfDNA analysis covers:
- feature-space summaries
- sample similarity and correlation
- outlier detection
- group comparisons
- region-signal summaries
- CNV summaries
- arm-burden summaries
- segment-aware cohort outputs when available
Visualization covers:
- UMAP, t-SNE, and PCA projections
- metadata-coloured sample scatter plots
- heatmaps and cohort summaries
- region-level plots
- CNV cohort plots
UMAP should be the preferred default projection when installed, t-SNE is the next nonlinear option, and PCA is the stable fallback.
Raw-Signal Methods¶
Raw-signal suites are included because many liquid-biopsy workflows first inspect the measured signal before model-based analysis.
Covered outputs:
- fragment-length distributions
- genome-wide signal profiles
- sample-by-bin heatmaps
- region metaprofiles
- VAF views
- arm-level burden plots and tables
- motif summaries
- browser-track inventories
- longitudinal summaries when sample-time metadata exists
Method Advisor¶
The method advisor adds a structured bridge between internal analysis routes, mature external tools, and explicitly marked research/watchlist liquid-biopsy methods. It is used when the user asks which method, tool, or algorithm should be considered for a dataset or assay question.
It currently tracks mature or commonly used routes for:
- fragmentomics: FinaleToolkit, cfDNAPro, DELFI-style features, Griffin, LIQUORICE, LBFextract, cfDNAFE, cfDNAanalyzer, EMIT, DeepFRAG
- broad cfDNA WGS/WGBS workflows: cfDNApipe and cfDNA UniFlow as optional external routes
- copy-number analysis: ichorCNA, QDNAseq, WisecondorX, HMMcopy readcount correction, CNVkit, Control-FREEC, CopywriteR, FACETS/facetsSuite, PureCN, BayesCNV
- methylation enrichment: QSEA and MEDIPS
- bisulfite, EM-seq, or nanopore methylation: Bismark, MethylDackel, nf-core/methylseq, Dorado, modkit
- methylation follow-up and advanced models: FinaleMe, cfTools/cfSort, MethylBERT, cfDecon, CelFiE-ISH, CelFEER, UXM, MethAtlas, cfNOMe, MetDecode, CpGPT, MethylGPT, MethFormer, cfMethylPre
- ctDNA variants: fgbio UMI consensus and low-VAF callers such as Mutect2, LoFreq, and VarDict
- cfRNA, EV-miRNA, CTC tables, and plasma proteomics as guidance-first routes
The advisor checks local dependency status and writes JSON/Markdown reports. If an external tool is not installed or needs assay-specific reference resources, the agent should report that clearly and continue with safe internal summaries when possible.
Usage reference: Liquid-Biopsy Method Advisor