Python API¶
The package keeps Python-callable entrypoints alongside the npm user interface. Python is the analysis kernel and direct API; regular users should start from liquid-agent, liquid-agent web, or liquid-agent client.
For direct module use, execution is available as:
python -m liquidbiopsy_agent.cli <command>
The project no longer installs Python console scripts from pyproject.toml. Those names are owned by the npm package in frontend/.
Top-Level Package Exports¶
From src/liquidbiopsy_agent/__init__.py, currently exported convenience functions include:
scan_project_profile(...)plan_assistant_tasks(...)summarize_project_outputs(...)run_liquidbiopsy_assistant(...)run_liquid_agent_shell(...)list_supported_blood_preprocessing_specs(...)plan_blood_preprocessing(...)preprocess_blood_signal_dataset(...)list_liquid_biopsy_methods(...)recommend_liquid_biopsy_methods(...)write_method_advice_report(...)list_feature_specs(...)compile_user_analysis_idea(...)
Preprocessing API¶
from liquidbiopsy_agent.preprocessing import preprocess_blood_signal_dataset
Use this for epigenomic, LPWGS/ULPWGS, and variant preprocessing when you want to call the same preprocessing layer directly from Python.
Assistant API¶
from liquidbiopsy_agent.agent.assistant import (
scan_project_profile,
plan_assistant_tasks,
execute_assistant_task,
summarize_project_outputs,
)
Typical direct API flow:
profile = scan_project_profile("<dataset_or_subdir>")
plan = plan_assistant_tasks(profile, goal="Run a first-pass liquid-biopsy analysis")
Plans include backend-only data_state and FeatureBook context when applicable.
The data state records visible signal families, input/output counts, metadata
coverage, blockers, and safe next actions. FeatureBook tracks which
liquid-biopsy signal contracts are relevant. Both are internal planning context;
regular users do not need to choose extra buttons or modes.
scan_project_profile(...) also attaches profile.metadata_profile. It records
candidate metadata tables, selected sample and label columns, label class
counts, matched sample coverage, confidence, supervision mode, backend, and
warnings. plan_assistant_tasks(...) consumes that profile to populate
labels_table, labels_sample_col, labels_label_col, supervised_modeling,
and supervised_backend task parameters when analysis inputs support them.
from liquidbiopsy_agent import build_liquid_biopsy_data_state, compile_user_analysis_idea, list_feature_specs
data_state = build_liquid_biopsy_data_state(profile)
feature_specs = list_feature_specs()
idea = compile_user_analysis_idea(
"Compare HER2 positive and negative methylation signals and make a figure.",
profile,
)
FeatureBook entries describe expected artifacts, QC checks, interpretation limits, and task sequences for fragmentomics, methylation, copy-number, variant, signal-matrix, archive, and metadata/grouped-comparison contexts.
Closed-loop result evaluation can be called directly for custom evaluation or report workflows:
from liquidbiopsy_agent.agent.result_evaluator import evaluate_project_results
from liquidbiopsy_agent.agent.result_signals import extract_result_signals
evaluation = evaluate_project_results(profile)
data_state = evaluation["data_state"]
signals = evaluation["result_signals"]
concepts = evaluation["analysis_concepts"]
pending = evaluation["pending_analysis_concepts"]
actioned = evaluation["actioned_analysis_concepts"]
blocked = evaluation["blocked_analysis_concepts"]
from liquidbiopsy_agent.agent.ledger import PlanLedger
concept_memory = PlanLedger(profile.dataset_root).concept_memory()
data_state is the same backend summary persisted into plan ledger records and
autopilot reports. analysis_concepts are backend audit records that connect
parsed result signals to follow-up questions, candidate task families, QC
checks, interpretation limits, stable novelty_key values, and backend
priority_score values. They also include backend lifecycle fields such as
lifecycle_stage, refined_task_sequence, required_outputs, and
verification_standard. They are consumed by the planner and reports without
adding user-facing modes. actioned_analysis_concepts records which concepts
were covered by completed task families in the current run and carries
ToolCard-derived verification status when available, so downstream reports and
replans can avoid treating already-actioned follow-up questions as new work
without evidence. pending_analysis_concepts is the concept subset still
eligible for automatic task promotion and is marked as
pending_execution. blocked_analysis_concepts records concepts whose
candidate task families failed or produced incomplete ToolCard verification, so
the planner can avoid blind repeats and move to alternative ready work, method
advice, or an explicit blocker. The planner keeps actioned and blocked records
for audit but does not use them as pending concept-driven task promotions unless
new inputs, fresh result signals, verification gaps, or explicit user intent
change the evidence.
Result signals that would recreate the same actioned concept use the same
novelty key, so they cannot bypass concept-level dedupe and silently requeue
the same automatic follow-up. The same guard applies to blocked concepts: a raw
result signal cannot bypass a blocked concept novelty key and silently requeue a
failed task family.
When evaluate_project_results(profile) is called without an in-memory
executed_runs list, it also inspects the dataset's persisted
assistant/ledger/run_*.json records. This preserves concept action state after
restarting the shell or Web backend and prevents a new process from forgetting
that a matching task family already covered a follow-up concept.
PlanLedger.concept_memory() rolls up recent result_evaluation_*.json
records by concept novelty key and returns compact pending, actioned, and
blocked counts plus the latest concept rows. The planner can use this backend
memory when the newest evaluation is incomplete, without adding user-facing
controls or modes. The merge is safety-biased:
generated < pending < blocked < actioned, so a generic pending concept does
not erase a previous blocker, while a later verified action can supersede it.
result_signals are conservative follow-up clues extracted from generated
effect tables, grouped summaries, outlier tables, and summary JSON files. They
help the next plan and report reference actual outputs instead of repeating a
generic scan-only recommendation.
Source Management API¶
from liquidbiopsy_agent.agent.sources import (
discover_dataset_sources,
source_inventory_row,
write_joint_source_inventory,
)
Use discover_dataset_sources("<parent_or_dataset_folder>") when a user-selected folder may contain multiple liquid-biopsy datasets. The returned source objects are session associations only; removing one from an agent task should not delete files. Joint autopilot writes source inventories with source_id, source path, likely signal, available assay summaries, labels, and generated output counts.
Analysis API¶
Standard cfDNA:
from liquidbiopsy_agent.analysis import run_cfdna_analysis_suite
summary = run_cfdna_analysis_suite(
output_dir="<analysis_output_dir>",
cfdna_features_dir="<feature_store_dir>",
)
Supplied CNV, methylation, EPIC-like, or generic liquid-biopsy signal matrices:
from liquidbiopsy_agent.analysis import analyze_cfdna_signal_matrix, run_cfdna_analysis_suite
summary = run_cfdna_analysis_suite(
output_dir="<analysis_output_dir>",
cnv_matrix_table="<cnv_matrix.tsv.gz>",
methylation_matrix_table="<methylation_matrix.tsv.gz>",
matrix_max_features=1000,
)
matrix_summary = analyze_cfdna_signal_matrix(
matrix_table="<matrix.tsv.gz>",
output_dir="<analysis_output_dir>/matrix",
signal_kind="methylation_matrix",
)
Raw-signal numeric:
from liquidbiopsy_agent.analysis import run_cfdna_raw_signal_analysis_suite
Visualization API¶
Standard cfDNA:
from liquidbiopsy_agent.visualization import run_cfdna_plot_suite
summary = run_cfdna_plot_suite(
output_dir="<visualization_output_dir>",
cfdna_features_dir="<feature_store_dir>",
projection="auto",
)
run_cfdna_plot_suite(...) accepts projection="auto" | "umap" | "tsne" | "pca". The same visualization API accepts cnv_matrix_table, methylation_matrix_table, or signal_matrix_table and writes matrix heatmaps, projection CSVs, PNG figures, and optional Plotly HTML files when Plotly is installed. The current Web Results collector lists reports, tables, JSON, text, and static figures by default and filters HTML artifacts from the general result list.
Raw-signal visualization:
from liquidbiopsy_agent.visualization import run_cfdna_raw_signal_suite
run_cfdna_raw_signal_suite(...) writes PNG and CSV artifacts and, when Plotly
is installed, may also generate HTML files for genome-wide profiles, sample/bin
heatmaps, and VAF distributions. The Web Results panel currently surfaces the
PNG/CSV/JSON/report outputs by default.
Internal compatibility proxy for legacy CopywriteR-like off-target/bin-count CNV screening:
from liquidbiopsy_agent.analysis import run_copywriter_like_cnv_proxy
summary = run_copywriter_like_cnv_proxy(
input_path="<interval_or_bin_dir>",
output_dir="<output_dir>",
exclude_regions="<targets_or_peaks.bed>",
)
Method Advisor API¶
from liquidbiopsy_agent.methods import (
bootstrap_external_tools,
external_tool_status,
install_external_tool,
list_liquid_biopsy_methods,
recommend_liquid_biopsy_methods,
run_external_tool_command,
smoke_external_tool,
write_method_advice_report,
write_external_tool_status_report,
)
Use this layer to compare liquid-biopsy methods and tools against a dataset path or a natural-language question:
advice = recommend_liquid_biopsy_methods(
input_path="<dataset_or_subdir>",
query="fragmentomics CNV methylation",
)
Focused method queries work the same way:
recommend_liquid_biopsy_methods(query="cfDNAPro FinaleToolkit LBFextract fragmentomics")
recommend_liquid_biopsy_methods(query="WisecondorX HMMcopy low-pass CNV")
recommend_liquid_biopsy_methods(query="FinaleMe cfTools cfSort methylation tissue of origin")
recommend_liquid_biopsy_methods(query="MethylBERT CelFEER UXM MethAtlas cfNOMe MetDecode methylation deconvolution")
recommend_liquid_biopsy_methods(query="CpGPT MethylGPT MethFormer methylation foundation model")
recommend_liquid_biopsy_methods(query="PureCN FACETS BayesCNV CopywriteR targeted ctDNA CNV")
recommend_liquid_biopsy_methods(query="cfDNAFE cfDNAanalyzer EMIT DeepFRAG fragmentomics")
Write JSON and Markdown reports:
summary = write_method_advice_report(
output_dir="<output_dir>",
input_path="<dataset_or_subdir>",
query="fragmentomics CNV methylation",
)
External tool runtime helpers:
status = external_tool_status("purecn")
bootstrap_plan = bootstrap_external_tools(profile="core", execute=False)
install_plan = install_external_tool("purecn", execute=False)
smoke = smoke_external_tool("purecn")
result = run_external_tool_command("cnvkit", ("cnvkit.py", "--help"))
report = write_external_tool_status_report("<output_dir>")
Skill API¶
from liquidbiopsy_agent.agent.skills import (
list_skill_documents,
ingest_skill_source,
remember_professional_observation,
remember_user_preference,
get_skill_context,
refresh_skill_cache,
)
Use this layer for paper ingestion, expert notes, private user preference memory, local skill cache refresh, and skill-context retrieval.
Compatibility aliases list_skills and remember_expert_note remain available
for older local scripts, but new code should use list_skill_documents,
remember_professional_observation, and remember_user_preference.
Web Backend API¶
The local service lives under src/liquidbiopsy_agent/web/ and exposes session,
chat, source, planning, task, artifact, skill, method, LLM, upload, and job
endpoints. It calls the same assistant, autopilot, method-advisor, result
browser, and skill logic used by the terminal shell.
The closed-loop agent layer also exposes durable plan and result state under
each conversation-owned source workspace's assistant/ledger/ folder (standalone
analysis commands use their explicit output root). The main Python helpers are:
from liquidbiopsy_agent.agent.ledger import PlanLedger
from liquidbiopsy_agent.agent.result_evaluator import evaluate_project_results
profile = scan_project_profile("<dataset_or_subdir>")
plan = plan_assistant_tasks(profile)
ledger = PlanLedger(profile.dataset_root)
plan_record = ledger.write_plan(plan, profile)
evaluation = ledger.write_evaluation(evaluate_project_results(profile))
These helpers are used by the terminal shell and Web autopilot. They are safe to
read directly when building dashboards or external audit tools.
PlanLedger.concept_memory() returns the compact pending/actioned/blocked
follow-up memory used by the planner, reports, and final immediate-next-action
generation.
PlanLedger.write_evaluation(...) refreshes that book automatically after each
saved evaluation; PlanLedger.write_concept_book() can also refresh it
explicitly for backend audit tools and result-driven replanning.
Plan records written after a result evaluation include the concept book path and
a compact concept-memory snapshot, so external audit tools can explain why a
next plan continued, paused, or avoided a repeated follow-up.
The plan_novelty.concept_memory_delta field compares the previous and current
snapshots for count changes and concept lifecycle transitions.
PlanLedger.summary() carries this compact novelty block in recent plan
history so reports and external audit tools can show the explanation without
loading full plan JSON files.
Main endpoint groups:
- Session and LLM:
POST /api/session,GET/PATCH /api/session/{session_id},GET/POST /api/llm/config,DELETE /api/llm/config/{provider} - Reviewed local-model library:
GET /api/llm/local/catalog,POST /api/llm/local/catalog/{identifier}/install|pause|resume|acknowledge,DELETE /api/llm/local/catalog/{identifier}/download; profile CRUD/probes remain under/api/llm/local - Dataset and source management:
POST /api/session/{session_id}/dataset,GET/POST /api/session/{session_id}/sources,DELETE /api/session/{session_id}/sources/{selector},DELETE /api/session/{session_id}/sources - Scan and planning:
POST /api/session/{session_id}/scan?include_plan=0|1,POST /api/session/{session_id}/plan,POST /api/session/{session_id}/message - Method and tools:
GET /api/methods,POST /api/session/{session_id}/methods/advice,GET /api/tools,GET /api/tools/{tool_key} - Skills and uploads:
GET /api/skills,GET /api/skills/{skill_id},POST /api/skills/refresh,POST /api/session/{session_id}/skills/ingest,POST /api/session/{session_id}/skills/remember,POST /api/session/{session_id}/uploads - Runs and jobs:
POST /api/session/{session_id}/autopilot,POST /api/session/{session_id}/tasks/{task_key}/run,GET /api/jobs/{job_id},GET /api/jobs/{job_id}/events,POST /api/jobs/{job_id}/cancel - Results and artifacts:
GET/DELETE /api/session/{session_id}/results,GET /api/session/{session_id}/file,GET /api/session/{session_id}/artifact - Local page lifecycle:
/api/system/health,/api/system/page-open,/api/system/heartbeat,/api/system/page-close,/api/system/page-close-mode,/api/system/page-watch,/api/system/force-shutdown
include_plan=0 lets the Web source drawer attach and scan a folder quickly
without immediately building a runnable plan. The Plan button, /plan, or a
planning-oriented message can request the full plan after the source is known.
Plan responses include the current plan_id, ledger_path, matched
skills_used, and a compact ledger summary when a runnable plan is generated.
Autopilot job events can include run_recorded, result_evaluated,
workflow_replanned, and workflow_completed payloads.
Workspace continuity and current contracts¶
The running service exposes its complete generated schema at /api/docs and /openapi.json; /docs/ is the bilingual user guide. Use that schema for request/response models rather than copying historical payloads.
- Task lifecycle:
POST /api/conversations/status,GET /api/conversations/trash,POST /api/conversations/trash,POST /api/conversations/{thread_id}/restore,DELETE /api/conversations/{thread_id}/purge, andGET /api/activity. - Scientific jobs and multi-turn requests retain the shared session/job contracts above. Linking and memory use registered conversation tools and task-owned storage; they do not introduce arbitrary filesystem access.
- For model profile CRUD/probes, images, skill reviews and linked preparation, inspect the current OpenAPI paths and the shared tool schemas in
agent/llm.py. Context snapshots and result references remain subject to privacy and ownership checks.
See current capabilities, task memory and local model profiles for user-facing contracts.