Blood-Signal Preprocessing¶
The preprocessing layer sits between raw or near-raw blood-biopsy inputs and downstream encoding or analysis.
Main module:
src/liquidbiopsy_agent/preprocessing/
Main entry scripts:
scripts/preprocess_epigenomic_signal.pyscripts/preprocess_lpwgs_signal.pyscripts/preprocess_variant_signal.py
For normal guided use, start from liquid-agent or liquid-agent web and ask
for preprocessing in natural language. The direct scripts are still the
reproducible low-level route.
What It Does¶
Current executable subset:
- interval / peak cleanup for BED-like inputs
- BAM / CRAM fragment materialization
- optional focus-region selection
- optional blacklist filtering
- optional fragment-length filtering
- epigenomic region-panel aggregation
- cfChIP background-aware summaries when a background BED is supplied
- cfMeDIP / MeDIP scale-factor normalization
- LPWGS bin-count export, lightweight correction, segmentation, and arm-burden summaries
- variant-table normalization with optional QUAL / VAF / CHIP filtering
- matched-normal / PBMC overlap filtering for variant tables
Still future-facing:
- true UMI consensus generation
- richer CNV models
- assay-specific cfChIP background models beyond the implemented supplied-background route
Implemented Profiles¶
cfchip_seq¶
cfchip_interval_cleanupcfchip_panel_summarycfchip_background_aware
cfmedip_seq / medip_seq¶
- cleanup-only profiles
- panel-summary profiles
- scale-normalized panel profiles
lpwgs / ulpwgs¶
lpwgs_interval_cleanup/ulpwgs_interval_cleanuplpwgs_cleanup_only/ulpwgs_cleanup_onlylpwgs_gc_corrected/ulpwgs_gc_corrected
ctdna_variant / variant¶
variant_table_qcvariant_strict_somaticvariant_matched_normal
Natural-Language Steering¶
When preprocessing is auto-inserted from a downstream task, the assistant now passes the downstream goal into profile selection. That means cues such as:
promoter panelbackground-awareGC-correctedstrict somaticmatched normal
still influence the selected implemented preprocessing route.
Assistant Behavior¶
The assistant can:
- reuse compatible preprocessing outputs
- rerun preprocessing with a better-matched implemented profile
- skip preprocessing only when the downstream task is already directly runnable
This is the one place where LLM input participates in the decision but remains constrained by explicit workflow legality checks.
Method Provenance¶
The preprocessing layer is literature-grounded, but some implementations are intentionally lighter than the canonical assay pipelines:
- LPWGS correction and segmentation are inspired by HMMcopy / ichorCNA-style workflows, but the current implementation is a pragmatic lightweight variant rather than a full reproduction.
- cfChIP preprocessing supports practical supplied-background normalization, not a full assay-native background model.
- cfMeDIP preprocessing includes useful normalization hooks, but not the full spike-in-centric quantitative workflow.
- variant preprocessing covers table-level QC, somatic-style filtering, and matched-normal overlap, but not true UMI-family consensus.
For references and rationale, see Components and Methods.
Key references:
CLI Examples¶
Epigenomic:
python scripts/preprocess_epigenomic_signal.py \
--signal cfchip_seq \
--input_format bed.gz \
--input_dir <raw_interval_dir> \
--region_set_bed <region_panel_bed>
LPWGS:
python scripts/preprocess_lpwgs_signal.py \
--signal lpwgs \
--input_format bed.gz \
--input_dir <lpwgs_interval_dir> \
--bin_annotation_table <bin_annotation_table> \
--arm_annotation_table <arm_annotation_table>
Variant:
python scripts/preprocess_variant_signal.py \
--signal ctdna_variant \
--input_format csv \
--input_dir <variant_table_dir> \
--exclude_chip_genes \
--matched_normal_dir <matched_normal_variant_dir>
Output Convention¶
Default root:
<dataset>/preprocessed/...
Common outputs:
intervals/regions/bins/corrected_bins/segments/arm_burden/variants/preprocessing_summary.csvpreprocessing_summary.json