Ollama

Run Ollama to manage and serve supported language models locally, with model storage, compute capacity, API access, and security considered.

On this page

Start with workload, not a model label

Ollama manages and runs supported models on infrastructure you control and exposes an API for compatible clients. The target task decides whether a model is suitable: summarisation, extraction, code assistance, chat, and retrieval-grounded answers place different demands on quality, context, response length, and concurrency. Test candidate models against representative prompts before committing shared infrastructure.

Model size and quantisation affect memory, storage, startup time, and response behaviour. A model that works for one developer's short request may fail when several users send long contexts at once. Measure on hardware like the target host, record the model version and source, and keep licensing or permitted-use review with that record. Do not describe a model as approved merely because it downloaded successfully.

A small evaluation set is useful. Include safe examples of the tasks users will perform, expected failure cases, a maximum prompt size, and a response-quality review. Keep the set free of production secrets. It is better to reject a poor model during this stage than to discover its limits after applications depend on its output.

Plan the host and API boundary

An Ollama host holds model artifacts and consumes CPU, memory, and often GPU resources. Treat it as a service host, not a personal laptop that happens to have an API port. Give it persistent storage sized for selected artifacts, monitor free capacity, and plan how model files are introduced, removed, and verified. An uncontrolled shared host becomes difficult to reproduce or patch.

Bind network access deliberately. A local runtime can become a shared service if an application reaches it across a network. Put the API behind the intended ingress and authentication boundary, restrict callers, and avoid exposing it broadly by accident. Observe failed requests, resource exhaustion, and unexpected callers. A firewall rule alone does not identify which application is allowed to use scarce inference capacity.

Keep service credentials and model-management permissions separate. Application clients should receive only the access they need; platform administrators need a documented path for updates and incident response. If the service is used with a chat interface or gateway, map the request path from user to gateway to Ollama so the owner can identify where identity, rate limiting, and audit data are applied.

Measure capacity and data controls

Capacity is shaped by model memory needs, prompt and output length, parallel requests, batching behaviour, accelerator availability, and host contention. Benchmark realistic prompts on the actual operating hardware. Record latency distribution, failure mode under load, memory headroom, and the point at which requests queue or fail. Plan for peaks rather than assuming the average is enough.

A local model endpoint does not remove data-handling choices. Callers can still send sensitive prompts, and logs or tracing may retain request metadata. Define approved data classes, request logging policy, retention, and incident handling. If applications have tenant boundaries, carry those boundaries at the gateway or application layer; a model runtime does not understand business authorisation by itself.

Keep model artifacts and runtime versions under change control. Driver, accelerator toolkit, host operating system, model file, and client-library changes can all shift behaviour. Test the combination that will run in production and retain a simple record of what passed. That makes an unexpected quality or performance change diagnosable later.

Operate an inference host

Monitor host health, GPU or accelerator condition where present, memory pressure, disk use, queue symptoms, API errors, and response time. Tie alerts to an action: stop accepting new traffic, route to a controlled alternative, or involve the host owner. A vague utilisation threshold offers little help if no one knows which workload is consuming memory.

Plan maintenance windows for model and runtime updates. Drain or communicate with callers, preserve the prior known-good configuration, and test a safe prompt set after a change. Make clear whether clients may retry requests and how they should handle timeouts. That prevents a capacity incident from becoming a retry storm.

Recovery may mean rebuilding a host rather than restoring every generated response. Document which model artifacts, configuration, network rules, secret references, and monitoring hooks are required to restore service. Test the process in a separate environment. The accepted recovery boundary should be agreed with the applications that depend on it.

Model-serving decisions

Use observed workload results to settle these items.

AreaDecisionEvidence
Model setWhich named model versions are permitted for each task?Evaluation record, source, and usage review.
CapacityWhat concurrency, context, and response length does the host support?Target-hardware load test and headroom notes.
API accessWhich applications may call the service and through which boundary?Network and caller inventory with a rejected-call test.
RecoveryHow is a host rebuilt with approved artifacts and configuration?Isolated rebuild record.

Questions for platform and application owners

Can a model that fits in memory be treated as production-ready?

No. Memory fit says little about task quality, concurrency, prompt length, response latency, licensing, or failure behaviour. Test the intended workload on target hardware.

Should the API be reachable on the office network?

Only if an explicit access design requires it. Put an authentication and network boundary in place, enumerate callers, and monitor resource use before sharing capacity.

What should be retained after model changes?

Record host and runtime versions, model identifier, test prompts or acceptance criteria, benchmark result, owner, and rollback point. Do not retain sensitive test inputs.

Bring a model service online

Pick tasks, approved model candidates, target hardware, safe prompt set, data policy, and expected concurrency before deployment.

Handover evidence

Keep this information with the service, not only with the original builder.

Approved-model register

The register identifies model versions, owners, intended tasks, evaluation results, and permitted uses.

Caller boundary

The intended applications and network path are documented, and an unapproved call has been rejected in testing.

Host rebuild

A recorded exercise rebuilds an approved runtime with expected artifacts, configuration, and monitoring.

Sources and further reading

Talk to our team.

Tell us what you're working on, whether it's a deployment, an audit, a security test or a cyber range. You'll speak with an engineer who can help you scope it.

  • 30-minute call: free, with no obligation.
  • NDA on request: we can sign before you share details.
  • Clear next steps: a scope and plan after the call.