PRACTICAL AI & AUTOMATION
Ollama on Ubuntu: Run a Small Model and Add Context Knowledge
Install Ollama on Ubuntu, run qwen2.5:3b, add a dated knowledge note, create a Modelfile and test grounded answers before growing into RAG.

Install Ollama on Ubuntu, run qwen2.5:3b, add a dated knowledge note, create a Modelfile and test grounded answers before growing into RAG.
Step 1: Start with a small local model
Ollama makes it straightforward to run a language model on your own computer. For this tutorial we use qwen2.5:3b: an established small text model, not a claim that it is the newest or best model. It is suitable for experimenting with short notes and prompts before investing in larger hardware.
Choose a Linux machine with space for the model download and enough free RAM for the weights, context and operating system. GPU acceleration can help, but model size alone does not predict performance. Start with a short context and measure response time on your own hardware. This text-only example does not read images or operate your desktop.
Check the model’s official library entry and license before deciding whether it fits your intended use.
Step 2: Install Ollama on Ubuntu
Download the official Linux installer, read it, then run it. Check the installed version and service. If your installation does not start a background server, run ollama serve in a separate terminal instead; do not launch a second server when the port is already occupied.
A local daemon normally listens on port 11434. Keep it bound to localhost for this exercise. See Ollama’s Linux installation guide for service and GPU-specific details.
Open the complete copyable example — example-01.txt
# Open full code above.Step 3: Give the model a compact knowledge note
Context knowledge is information you supply while a model answers. It does not retrain the model’s weights. Start with one small, sanitized note that has a source and date. For example, create notes.txt with: ‘Lab policy, 2026-10-09: backups go to ./backups; disposable test files go to ./scratch; never delete originals during a demo.’ These are facts about this tutorial’s fictional lab.
Pass that note with the question using a short Python script. The example calls Ollama’s local chat endpoint and prints the answer. It includes an explicit unknown-answer rule and keeps the note separate from instructions. That separation helps, but model behavior is still something to test.
Create notes.txt and ask.py in the same folder. Run python3 ask.py. API shape and basic local requests are documented in the Ollama quickstart.
Open the complete copyable example — ask.py
import json
import urllib.request
from pathlib import PathStep 4: Reuse instructions with a Modelfile
For a stable assistant style, save the block below as Modelfile. Build a named configuration with ollama create lab-helper -f Modelfile and run it with ollama run lab-helper. Keep changing knowledge notes outside this file so you can date, update and cite them independently.
SYSTEM sets a reusable instruction; num_ctx requests a context size; temperature controls sampling. Creating this configuration is not fine-tuning and does not make a note permanent learned knowledge. See the Modelfile reference.
Open the complete copyable example — Modelfile.txt
FROM qwen2.5:3b
PARAMETER temperature 0.2
PARAMETER num_ctx 4096Step 5: Check what it knows and what it does not
Try three checks: ask where disposable files belong, ask the lab’s administrator password, and supply an updated note that changes the scratch path. Expect ./scratch for the first, an admission that the password is absent for the second, and the updated path for the third. Never insert a real password to make the exercise work.
If the answer invents details, shorten the note, make the question more precise and inspect the actual request. A successful single answer is not proof that all future questions will be grounded. Keep a small set of representative questions and rerun them when changing model or prompt.
Step 6: Grow into retrieval only when needed
When you have many documents, add retrieval: split approved material into useful chunks, preserve filename/date/page, select a few relevant passages, then send those passages with the question. For ten short notes, keyword search may be enough. An embedding index is helpful when wording differs, but it adds maintenance and evaluation work.
A bigger context window uses more memory and is not a substitute for selecting good evidence. Keep sources current and treat retrieved text as untrusted data. RAG is retrieval plus generation; it does not train model weights.
Troubleshooting: connection refused means check the server; slow generation means inspect ollama ps and memory pressure; unrelated answers mean inspect the selected note. A downloaded local model can answer without a cloud API, but a :cloud model or external search service changes that data path.
Reading practice: compare ‘The model reads the note’ with the Japanese roles beside it. Continue with the OpenClaw guide when you want a tool-using agent, rather than only text answers.
Continue the series
UbuntuでOllama:小さなモデルに文脈知識を渡す
手順1:小さなローカルモデルから始める
Ollamaを使うと、自分のパソコンで言語モデルを動かせます。この例は小さなテキストモデルqwen2.5:3bを使います。最新版や最良モデルという意味ではなく、大きな機材へ投資する前に短いメモで試すための選択です。
モデル用の保存領域と、重み、文脈、OSに必要な空きRAMを用意します。GPUは高速化に役立ちますが、実際の応答時間を短い文脈から測りましょう。この例はテキスト専用で、画像理解やデスクトップ操作は行いません。
利用目的に合うか、公式モデル情報とライセンスを確認してください。
手順2:UbuntuにOllamaを入れる
公式Linuxインストーラーをダウンロードして確認後に実行します。バージョンとサービスを確認します。バックグラウンドで動いていなければ別端末でollama serveを実行します。すでにポートが使用中なら二重起動しません。
通常のローカルサービスは11434番ポートを使います。この練習ではlocalhostのままにします。サービスやGPUの詳細は公式Linuxガイドを参照してください。
コピー可能な完全なコードを開く — example-01.txt
# Open full code above.手順3:短い知識メモを渡す
文脈知識とは、回答時に渡す情報です。モデルの重みを再学習することではありません。notes.txtに日付付きの短いメモを作ります。例:「2026-10-09の練習ルール:バックアップは./backups、破棄できる試験ファイルは./scratch、元ファイルは削除しない」。これは記事内の架空ラボの設定です。
次のスクリプトでローカルのチャットAPIへ渡します。未知の情報は不明と答えるように指定し、資料と指示を分けています。ただし動作の検証は必要です。
notes.txtとask.pyを同じフォルダーに保存し、python3 ask.pyを実行します。APIは公式クイックスタートを参照してください。
import json
import urllib.request
from pathlib import Path手順4:Modelfileで設定を再利用する
下をModelfileとして保存し、ollama create lab-helper -f Modelfileで設定を作り、ollama run lab-helperで起動します。変化する知識メモは外に置き、日付と出典を別に管理します。
SYSTEMは指示、num_ctxは文脈サイズ、temperatureはサンプリングの設定です。設定作成はファインチューニングではなく、メモが永続的に学習されるわけでもありません。Modelfile仕様を確認してください。
コピー可能な完全なコードを開く — Modelfile.txt
FROM qwen2.5:3b
PARAMETER temperature 0.2
PARAMETER num_ctx 4096手順5:答えられる範囲を確認する
破棄用ファイルの場所、管理者パスワード、scratchの場所を変えた新版メモの3つを試します。最初は./scratch、次は情報がないという回答、最後は新版の場所が期待値です。本物のパスワードは入れません。
情報を作り上げたら、メモを短くし、質問と送信内容を確認します。一度の正解だけでは将来の正確さは証明できません。モデルや指示を変更したら代表的な質問を再実行します。
手順6:必要になってから検索を加える
文書が増えたら検索を追加します。承認済み資料を適切な長さに分割し、ファイル名、日付、ページを残し、関連箇所だけを質問と一緒に渡します。短いメモが10個ならキーワード検索で十分かもしれません。埋め込み検索には保守と評価も必要です。
大きな文脈はメモリーを多く使い、良い資料選択の代わりにはなりません。RAGは検索と生成の組合せで、重みの学習ではありません。資料は信頼できる指示として扱わないでください。
接続拒否ならサーバー、遅いならollama psとメモリー、無関係な回答なら選択資料を確認します。ダウンロード済みローカルモデルはクラウドAPIなしで回答できますが、:cloudや外部検索では経路が変わります。
モデルがメモを読みます。ツールを使うエージェントへ進む場合はOpenClawの記事を読んでください。
技術資料の確認日:2026年10月9日。コマンドとAPIは更新されることがあります。実行時に公式資料と利用中のバージョンを確認してください。
