Local data and the remote model¶
Liquid Agent separates its local scientific workspace from the external GPT API. The LLM remains the conversational decision-maker; local tools do the data work.
User request -> LLM tool request -> local validated tool
|
local data / analysis
|
local summary gateway
|
aggregate evidence -> LLM reply
Detailed tables and figures -> local report renderer -> user's browser / CLI
What crosses the boundary¶
The model receives the user's natural-language request, maintained professional
guidance, tool availability, declared assay/value semantics, aggregate dataset
and metadata counts, QC/statistical summaries, execution status, and opaque
resource IDs. A remote model also receives a bounded relevant set of explicitly
remembered preferences marked remote_allowed; local_only memories stay on
the device. Public literature explicitly retrieved for knowledge work is not
treated as a private dataset.
It does not receive file previews, rows of a result table, sequence records, full matrices, image payloads, arbitrary Markdown/JSON report contents, local execution logs or metadata identity values from the scientific tools. A CSV summary is computed locally, not produced by sending CSV text to a second LLM. Unrecognised fields are omitted rather than forwarded by default.
read_artifact now means summarize this artifact locally. For bounded CSV/TSV
outputs it reports table dimensions and aggregates of recognised scientific
metric columns. It does not return row lists or arbitrary column names. Mean and
standard deviation require at least two finite observations; missingness and
fixed QC-flag counts remain available. Adjusted-p-value tables can report the
number of features below the declared summary threshold (currently 0.05).
Large/unsupported artifacts return an explicit limitation, not a raw-text fallback.
What stays fully functional¶
The model can still choose tools, ask questions, revise a plan, perform an
authorised off-plan analysis, interpret measured aggregate results and propose
progressive follow-ups. It can select artifact: IDs for a report without seeing
the underlying records. The local report renderer embeds the selected tables and
figures for the user; large tables use bounded previews with full files retained
locally. The privacy boundary does not remove them.
The boundary introduces no extra approval dialogue or fixed scientific workflow. If a task needs additional evidence, extend the local scientific summarizer rather than disable the task or send the dataset to another model. Useful measurement semantics and statistics should remain available for decisions.
Both Web and CLI use the same gateway. Their result-follow-up controls no longer append file excerpts to chat. The server also removes excerpts from messages created by older clients. Task errors are available locally; only a safe error category and recovery guidance go back to the model.
Skills and saved conversations¶
The shared Local Data Inspection and Model Privacy skill
explains what a local inspection should establish and how to use its summaries.
It supplements enforced Python contracts; a prompt or skill cannot override them.
Maintained project skills may be loaded as guidance. Unreviewed personal source
documents and automatically learned observations remain local rather than being
silently uploaded for distillation. Explicit user workflow preferences can still
be phrased conversationally for GPT. Structured memory retains a disclosure flag:
only relevant remote_allowed values are included in the workspace snapshot;
their evidence, source task and revision history are not sent.
Personal workflow preferences and skills distilled from explicitly supplied
public URLs remain loadable. For a locally imported professional document, review
and remove any dataset records or identifying information first; setting
metadata.model_safe: true in its SKILL.md records that local review and makes
its guidance loadable. This flag applies to the skill's guidance/references, not
to dataset files or result tables, and never authorises analysis execution.
An older provider transcript may contain previews created before this boundary. On migration, Liquid Agent starts a fresh model context and retains local chat, plans and results. It does not re-upload the old transcript or claim to delete anything already processed by the provider. New safe conversational history is kept separately from the local display history.
Practical limits¶
This is useful data minimisation, not formal anonymisation, a compliance certification, or a guarantee that all free-text identifiers can be recognised. Aggregate statistics can still be sensitive. Avoid typing identifiers into chat; attach local files instead. Pasted structured record blocks are withheld and common paths, credentials, emails and long nucleotide strings receive basic redaction. User-entered ordinary prose is still sent to the selected API.
An unsupported summary should prompt a new local summarizer or an honest limitation, not either a silent upload or a claim that analysis is impossible. No cross-provider fallback or extra model is introduced by this boundary.