Local Language Models¶
OpenAI remains the default. Local inference is an explicit option for users who want to run a language model on their own computer or private LAN server. The same scientific tools, LangGraph execution, skills, task memory, linked tasks, image provenance and review controls remain in use. Model reasoning quality and scientific reliability still depends on the selected model. Text-only models are no longer eligible.
One preferred engine, one official compatibility path¶
- Ollama is the preferred managed engine. Use its local model tags or supported
hf.co/...GGUF references. It is not limited to a company or country of origin. - An explicit Hugging Face
organisation/modelselection with a reviewed Ollama mapping uses that counterpart. Otherwise, setup inspects the official model card/configuration, pins the HF commit and routes it to Hugging Face's official Transformers server. This optional runtime is installed only when needed, into the user-selected Python environment; review dependency compatibility before installation. - An existing private OpenAI-compatible Chat Completions server can also be registered. This covers model-card recipes that require vLLM, SGLang or another specialised runtime, without installing every framework into LIQUID-Agent.
Routing is resolved in the installation plan and saved in each profile. A network,
quota, authentication, memory or generation failure does not trigger another
model, engine or cloud provider. Arbitrary Ollama tags cannot be reliably mapped
back to an HF checkpoint: specify its official HF repository to request that path.
To use the original HF weights even when an Ollama counterpart is known, use a
model-list entry with "runtime": "transformers".
There is no promise that every HF repository will run. Models need a compatible architecture, instruct/chat template, adequate context, multi-image input, multi-turn conversation and working function calls. Custom repository code is not executed automatically. Specialised quantization kernels and model-card-specific servers require their own reviewed setup. A model may fit in memory and still perform poorly on scientific planning.
Installation and storage¶
Local models can also be added after either installer mode. In the Web workspace,
open the model menu and choose Manage local models. The reviewed library is
hardware-screened before download and uses the installation's private user-data
directory automatically; model files go to its models/ child. No research-data
folder is reused for weights. Eligible models have an install control; unsupported
models remain visible but disabled with the capacity reason.
Web downloads are independent, so several eligible models can download together. Pause closes that model's active pull stream while its isolated staging service keeps the verified partial files; Continue resumes against the same service. Abandon stops that staging service, removes only that partial model and returns its row to Not installed. Closing the library does not stop a download: the compact model menu keeps its progress and controls. Completion performs a local manifest/capability check without loading the model for scientific inference, then shows a green completion mark. Selecting the model clears that mark and starts the installation-owned Ollama service when needed. Cloud keys and the previous cloud selection are preserved throughout.
The base installer has an optional local-model choice. Select an absolute user
data directory outside the project source; weights go in its models/ child. On macOS/Linux:
./scripts/install_liquid_agent_cli.sh --user-data-dir "$HOME/Liquid Agent Data" --with-local-models
On Windows, use an activated Python 3.12 and Node.js environment:
.\scripts\install_liquid_agent_cli.ps1 -UserDataDir "D:\Liquid Agent Data" -WithLocalModels
For an existing installation, first review a download-free capacity plan:
liquid-agent local-models hardware
liquid-agent local-models setup --models-dir "$HOME/Liquid Agent Data/models" --dry-run
liquid-agent local-models setup --models-dir "$HOME/Liquid Agent Data/models" --install-runtime
The second command does not create model directories or change provider defaults.
HF selections may fetch metadata/model cards, but not weights, during planning.
Installation prompts for confirmation after displaying eligible and skipped models.
Use --yes for an already-reviewed unattended installation. Gated HF downloads
require the appropriate license/access and an HF_TOKEN in the installing process;
it is not stored in the local inference profile.
Setup conservatively reserves RAM/VRAM for the OS and analysis. The reviewed initial recommendations at 32K context are estimates: 16GB Apple Silicon → Qwen3.5 4B, 32GB Apple Silicon → Qwen3.5 9B, and a sufficiently free 32GB NVIDIA GPU → Qwen3.8 27B. These are capacity candidates, not completed hardware benchmarks or scientific-quality endorsements. Active GPU use changes selection. Large MoE models need their full weights available; active parameter count alone is not a memory estimate. Longer context consumes additional memory.
Managed Ollama requires macOS 14+ or Windows 10 22H2+ on the supported
architectures (Linux is also supported). Its release is pinned with per-platform SHA256 checks. Setup downloads
eligible models, binds the planned context to an owned liquid-…:managed alias,
and runs the existing small synthetic tool/result round trip. It records successes/failures
in liquid-local-install-report.json in the model directory. Failed probes do not
activate a profile. Partial/cached weights remain for diagnosis or retry; reports
and existing source datasets are never removed. Existing profile IDs and default
model pins are not overwritten.
Meta Muse Glimmer 30B¶
Reviewed 23 September 2026: the official repository is
meta-models/Muse-Glimmer-30B.
Meta's Ollama recipe
supports the existing local Chat Completions interface. An explicit selection of
that HF repository now resolves to muse-glimmer:30b-q4_K_M, the text-and-image
quantization in Ollama's model library.
This adds an optional model; it does not change automatic recommendations, the
active provider, saved keys or an existing model profile.
From a repository checkout, use the supplied configs/local-models-muse.json
(profile ID muse), replacing the model storage path:
liquid-agent local-models setup --models-dir "$HOME/Liquid Agent Data/models" --models-list configs/local-models-muse.json --dry-run
liquid-agent local-models setup --models-dir "$HOME/Liquid Agent Data/models" --models-list configs/local-models-muse.json --install-runtime
liquid-agent local-models serve
Keep the server terminal open. In another terminal, test and explicitly select it:
liquid-agent local-models probe muse
liquid-agent local-models use muse
liquid-agent
Alternatively pass --model meta-models/Muse-Glimmer-30B or the exact Ollama tag
instead of --models-list; that route generates a profile ID, shown in the plan.
An existing installation can also register a manually served model through
Manage local models. A manually registered endpoint still needs the correct
runtime/model/context configuration and its own tool probe.
| Setting | Managed Muse profile |
|---|---|
| Model files | Approximately 17 GiB, including the vision projector |
| Capacity estimate | 22 GiB working memory at 64K context; includes cache/runtime headroom, not a measured peak |
| Context / output budget | 65,536 / 8,192 tokens; context is bound to the owned Ollama alias |
| Sampling / reasoning | Temperature 1.0, top-p 0.95, high reasoning; top-k 64 inherited from the official Ollama model |
| Request timeout | 900 seconds |
| Runtime | Stable Ollama 0.32.7+; checked before downloading weights. The bundled release may be newer. |
The 64K default leaves room for LIQUID-Agent's full scientific tool catalog,
workspace context, image attachments and the reasoning/output reserve. The capacity
estimate is anchored at 64K (estimate_context_tokens); larger contexts increase
the screening estimate instead of treating the published maximum as free memory.
A 16GB Mac is skipped before download. A 32GB Apple Silicon machine or a 32GB
NVIDIA GPU with sufficient free memory is an estimated candidate, not a
guaranteed fit. Close other memory-heavy work and verify peak use with real images
and scientific tools. Longer contexts undergo additional memory screening; a
published 128K limit does not mean 128K fits these machines. The main HF BF16
checkpoint alone is about 55.5 GiB. "runtime": "transformers" or an explicit HF
revision keeps that original checkpoint and its capacity checks instead of
silently selecting the quantization. The 3B assistant/drafter is not a standalone
replacement. MLX and DFlash variants are not automatically selected by this recipe.
The managed runtime still uses an internal liquid-muse-glimmer-…:managed alias,
but Web/CLI menus show the human name Muse Glimmer 30B. The alias binds the
planned context and remains visible only as a technical server identifier. Text,
tool calls, image attachments, task history and scientific safeguards
use the shared local adapter. A successful installation probe establishes a basic
tool/result round trip only. Confirm image interpretation, scientific accuracy and
long-task reliability on your own hardware before relying on the model.
Multiple models and automatic routing¶
Pass a JSON list using --models-list /path/models.json:
[
"qwen3.5:4b",
"qwen3.8:27b",
{"id": "hf-qwen", "model": "Qwen/Qwen3.5-4B", "runtime": "transformers"},
{"id": "custom-gguf", "model": "hf.co/OWNER/REPOSITORY:Q4_K_M",
"weights_gib": 5, "working_gib": 14, "estimate_context_tokens": 65536,
"context_tokens": 65536, "vision": true}
]
Replace the illustrative OWNER/REPOSITORY with a reviewed multimodal GGUF bundle,
including its vision projector; a text-only GGUF is not supported.
For unknown Ollama tags, provide conservative weights_gib and working_gib
(including weights, context cache and runtime overhead). Unknown sizes are skipped
instead of guessed. HF safetensors file sizes can supply an initial estimate;
quantized/custom checkpoints may still need explicit estimates and separate
runtime dependencies. Oversized models are explained and skipped before download.
Disk requirements accumulate across installed models; memory is assessed per model.
Start the selected engine in a terminal:
liquid-agent local-models serve
# Only for profiles routed to the optional HF engine:
liquid-agent local-models serve --runtime transformers
Keep the appropriate terminal open. Ctrl+C stops the owned server; closing the last Web workspace page also stops a Liquid Agent-managed Ollama server. An independently started Ollama is untouched. Managed ports are 11435 for Ollama and 11436 for HF by default. The HF wrapper serializes generations and releases the previous model on a model switch; its official generation/tool parser is unchanged. Do not run both engines with loaded models on a memory-constrained machine. Setup never takes over an unknown process already using those ports. A downloaded model is not a running model server.
Connect, install, test and select¶
In Web, open the model menu → Manage local models. Each reviewed model has one compact row and an install button. Liquid Agent downloads the runtime and model into the personal storage folder selected at installation, verifies the model manifest, and configures the endpoint, context and display name internally. There is no manual configuration form. Click an installed model to use it.
The preparation phase has a moving indicator; model transfers have a thin, theme-coloured progress line. Pause keeps the partial files; Continue resumes; Abandon removes that model's partial files. Closing the library leaves the model menu open with the same progress and controls. The green completion check clears from both places after successful selection. Downloads continue across page refreshes; restarting the service preserves them in the paused state.
Download eligibility checks total device capacity and available disk space; transient free RAM does not disable an otherwise compatible download. Models that exceed the device capacity remain unavailable with an explanation when clicked. Installing does not load a model or certify its inference quality.
Existing private-server profiles can still be configured through the CLI or local API. They are kept separate from this automatic download library, and existing profiles and conversation selections remain supported.
Saved local profiles appear by their display name in the same model menu. Managed Qwen and Muse profiles receive names such as Qwen 3.5 4B automatically; their hashed server aliases are not used as UI titles. Older reviewed managed profiles are recognized by profile ID, while custom profiles can set or edit a display name. Choosing one is explicit; setup itself leaves the cloud default untouched. CLI equivalents:
liquid-agent local-models list
liquid-agent local-models add --id lab --display-name "Lab multimodal model" --model my-model --runtime ollama --base-url http://127.0.0.1:11434/v1 --context-tokens 65536 --vision
liquid-agent local-models probe lab
liquid-agent local-models use lab
liquid-agent cli
Inside the CLI, /llm use local lab selects the profile; /llm save persists it.
local-models remove lab removes only the profile, not model files or conversations.
An active task with a removed profile reports an error; it never falls back to GPT.
The optional local bearer token can be passed via --api-key-env LOCAL_MODEL_TOKEN;
use a distinct variable, never a cloud API key. Advanced sampling/template options
can be supplied as a JSON file with --request-options.
Local adapters use streaming Chat Completions, validated tools, cancellation, steering and locally persisted history. Complete history groups are summarized with the same local model before context exhaustion. Current prompts, images or tool catalogs that cannot fit fail explicitly; the adapter does not truncate scientific evidence silently. Interrupted tool receipts direct the next turn to check saved task evidence before retrying work. The full tool catalog needs more context than a small standalone chat. The Qwen profiles retain their existing 32K default; Muse uses 64K. When using several images with the complete tool catalog, monitor memory and adjust context/output budgets within the available capacity. Model-card maximums are not automatically allocated.
Data safety and deployment requirements¶
Profiles live in local_models.json under the normal user configuration directory
(LIQUIDBIOPSY_CONFIG_DIR can redirect it). POSIX writes use owner-only files;
Windows deployments must verify the directory's user ACL. API keys are encrypted
with an OS credential-store master key; unavailable stores fail without plaintext
fallback. No keys belong in Git. Advanced Transformers dependencies install into
the Python environment chosen by the user; they do not create another venv.
Managed Ollama uses OLLAMA_NO_CLOUD=1, loopback binding, an explicit model storage
path and one loaded model. Remote/cloud forwarding models are rejected. The owned
HF server uses HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1, telemetry-disabled
settings and no repository custom code. Assets are downloaded during installation,
not fetched on behalf of a scientific inference request. Local request clients
ignore proxy environment variables and reject public endpoints and redirects.
Connections use the checked private IP directly while retaining the original
HTTP Host and TLS server name, closing the gap between DNS validation and connection.
Changing or revoking a profile's token interrupts its active turn. A damaged local
registry produces a repair error in the local manager without disabling cloud menus.
A private endpoint can still be a misconfigured forwarding server; configure and
verify the server you operate.
The local Web API rejects foreign browser origins and rebound Host names before running an operation. The public portal does not probe the local API; Try it opens the installation documentation directly. The managed HF daemon rejects direct browser requests. These checks are not user authentication: keep the application on loopback; a private LAN inference server needs its own access controls.
These settings are not a proof that the entire application is offline. Literature retrieval, dependency/model downloads and explicitly networked tools have separate network needs. For sensitive data, predownload assets, restrict OS/container outbound access, and observe network traffic on the target device. Offline mode in a library is not a firewall.
References: Ollama privacy/offline FAQ, Ollama compatibility, Ollama imports, HF official serving, Qwen model card, DeepSeek local deployment guidance.