Outcome
You will run one small language model on a Linux host that you already own and verify a local response. The example uses Ollama because it has a documented Linux installation, CLI, and local API. Model speed, memory needs, quality, and accelerator support depend on your particular host, model, drivers, and settings. Nothing here establishes compatibility or performance for a planned PebbleRack device.
Is local the right choice?
Local execution can give you direct control of a model and work during some internet outages, but you maintain the host, updates, power, backups, and security. It can be slower or less capable than a hosted assistant for a given task. A hosted assistant or API may be simpler for occasional use and offer a wider model selection, while sending requests to that provider under its current data controls. Renting a GPU gives you more capacity without buying hardware, but adds account, network, storage, and ongoing cost decisions. Do not use this local recipe for a task requiring a model too large for your host, a managed service-level commitment, or an unreviewed sensitive-data workflow. Compare the same real task across options, including total cost and data path, before choosing.
Before you start
Use a supported Linux machine or VM with administrator access, available disk space, and enough memory for the model you choose. Check the current Ollama Linux documentation for CPU architecture and optional GPU drivers. Choose a small official-library model; read its card and license. Consider download size on metered connections. Have a terminal and a way to inspect systemd service status. Keep the host on a trusted private network for this first exercise.
Steps
- Record operating system version, architecture (
uname -m), free disk (df -h), memory (free -h), and whether you have a supported accelerator. Save this baseline privately. Select a model small enough to leave headroom for the OS and other workloads. A model's advertised parameter count is not an exact memory requirement; context length and quantization matter too. - Read the official Linux install instructions. For this lab, use the documented manual installation and service instructions after checking the download, architecture and commands against your host. Ollama also offers an installation script, but do not run a fetched script directly without inspecting what it will change. Avoid third-party install commands.
- After installation, run
ollama --versionandsystemctl status ollamaif installed as a service. If your chosen installation did not create a service, useollama servein a separate terminal as the documentation describes. Record which path you used; avoid running a second server on the same port. - Select a current small model from the official library. Run
ollama pull MODEL_NAME, replacing the placeholder with the exact model tag. Then runollama run MODEL_NAMEand ask a simple non-sensitive question, such as “List three ways to check free disk space on Linux.” End the session using the documented CLI interaction. - Check the local API only from the same host. The documented endpoint is
http://localhost:11434/api/generate; send a request withmodel,prompt, andstream:false, for example viacurlusing a JSON file that contains no personal data. Compare the response to the CLI result. This tests that the server works locally; it does not measure correctness or privacy beyond the local service boundary. - Record model name and tag, install date, runtime version, response time, observed memory use, and one incorrect answer if you find one. This small evidence log prevents a later guide from turning a single local response into a benchmark claim.
Check it worked
The service is running, the chosen model is listed locally, the CLI returns a response, and the loopback API returns a response without exposing the port beyond the host. Confirm binding with your host's networking tools. Run one known-answer prompt and assess it yourself; fluent output is not proof of accuracy.
If it fails / rollback
If download or load fails, check disk, memory, architecture, service logs (journalctl -u ollama for a service install), and the exact model tag. Try a smaller model. If the server cannot start, check port conflicts before changing firewall rules. To roll back, stop the service and use the official uninstall steps for your installation method; preserve any model files you need first.
Safety and data notes
Do not expose the Ollama API to the public internet. The local API should be treated as an administrative capability, especially when integrated with tools. Avoid sensitive prompts until you understand model files, logs, client applications, and any external dependencies. The inference request in this specific localhost exercise goes to the local Ollama server; the model download uses the internet, and any later web UI, cloud model, search tool, telemetry, or integration may send data elsewhere. Inspect each component and network path before entering sensitive material. Local execution does not make generated answers reliable or license-free.
Sources
Official sources checked 2026-09-29: Ollama Linux installation; Ollama quickstart; Ollama API introduction; Ollama model library; OpenAI API data controls; AWS GPU instance options.
Next guide
Continue with private remote access.