PRACTICAL AI & AUTOMATION

Ollama on Ubuntu: Run a Small Model and Add Context Knowledge

Install Ollama on Ubuntu, run qwen2.5:3b, add a dated knowledge note, create a Modelfile and test grounded answers before growing into RAG.

Manabi Kōbō··10 min·808 words
AI AutomationOllamaLocal AIUbuntuRAG
Ollama on Ubuntu: Run a Small Model and Add Context Knowledge
ENNATURAL ARTICLE

Install Ollama on Ubuntu, run qwen2.5:3b, add a dated knowledge note, create a Modelfile and test grounded answers before growing into RAG.

Step 1: Start with a small local model

Ollama makes it straightforward to run a language model on your own computer. For this tutorial we use qwen2.5:3b: an established small text model, not a claim that it is the newest or best model. It is suitable for experimenting with short notes and prompts before investing in larger hardware.

Choose a Linux machine with space for the model download and enough free RAM for the weights, context and operating system. GPU acceleration can help, but model size alone does not predict performance. Start with a short context and measure response time on your own hardware. This text-only example does not read images or operate your desktop.

Check the model’s official library entry and license before deciding whether it fits your intended use.

Step 2: Install Ollama on Ubuntu

Download the official Linux installer, read it, then run it. Check the installed version and service. If your installation does not start a background server, run ollama serve in a separate terminal instead; do not launch a second server when the port is already occupied.

A local daemon normally listens on port 11434. Keep it bound to localhost for this exercise. See Ollama’s Linux installation guide for service and GPU-specific details.

Open the complete copyable example — example-01.txt

# Open full code above.

Step 3: Give the model a compact knowledge note

Context knowledge is information you supply while a model answers. It does not retrain the model’s weights. Start with one small, sanitized note that has a source and date. For example, create notes.txt with: ‘Lab policy, 2026-10-09: backups go to ./backups; disposable test files go to ./scratch; never delete originals during a demo.’ These are facts about this tutorial’s fictional lab.

Pass that note with the question using a short Python script. The example calls Ollama’s local chat endpoint and prints the answer. It includes an explicit unknown-answer rule and keeps the note separate from instructions. That separation helps, but model behavior is still something to test.

Create notes.txt and ask.py in the same folder. Run python3 ask.py. API shape and basic local requests are documented in the Ollama quickstart.

Open the complete copyable example — ask.py

import json
import urllib.request
from pathlib import Path

Step 4: Reuse instructions with a Modelfile

For a stable assistant style, save the block below as Modelfile. Build a named configuration with ollama create lab-helper -f Modelfile and run it with ollama run lab-helper. Keep changing knowledge notes outside this file so you can date, update and cite them independently.

SYSTEM sets a reusable instruction; num_ctx requests a context size; temperature controls sampling. Creating this configuration is not fine-tuning and does not make a note permanent learned knowledge. See the Modelfile reference.

Open the complete copyable example — Modelfile.txt

FROM qwen2.5:3b
PARAMETER temperature 0.2
PARAMETER num_ctx 4096

Step 5: Check what it knows and what it does not

Try three checks: ask where disposable files belong, ask the lab’s administrator password, and supply an updated note that changes the scratch path. Expect ./scratch for the first, an admission that the password is absent for the second, and the updated path for the third. Never insert a real password to make the exercise work.

If the answer invents details, shorten the note, make the question more precise and inspect the actual request. A successful single answer is not proof that all future questions will be grounded. Keep a small set of representative questions and rerun them when changing model or prompt.

Step 6: Grow into retrieval only when needed

When you have many documents, add retrieval: split approved material into useful chunks, preserve filename/date/page, select a few relevant passages, then send those passages with the question. For ten short notes, keyword search may be enough. An embedding index is helpful when wording differs, but it adds maintenance and evaluation work.

A bigger context window uses more memory and is not a substitute for selecting good evidence. Keep sources current and treat retrieved text as untrusted data. RAG is retrieval plus generation; it does not train model weights.

Troubleshooting: connection refused means check the server; slow generation means inspect ollama ps and memory pressure; unrelated answers mean inspect the selected note. A downloaded local model can answer without a cloud API, but a :cloud model or external search service changes that data path.

Reading practice: compare ‘The model reads the note’ with the Japanese roles beside it. Continue with the OpenClaw guide when you want a tool-using agent, rather than only text answers.

Continue the series

JP日本語

UbuntuでOllama:小さなモデルに文脈知識を渡す

SubjectObjectVerbParticleTechnicalConnectorAdjectiveTime / context
CONTINUE IN MANABI KŌBŌ

Put this idea into practice

Reading Nook

Practice bilingual reading with Japanese scaffolding.

Open tool →
Guided Japanese Course

Continue with structured lessons and practice.

Open tool →
Adaptive Recall

Turn useful phrases into spaced practice.

Open tool →
RELATED NOTES
← Back to all articles
Share this note Email