← Back to Tutorials
TutorialFor: AI Engineers, ML Engineers, Platform Engineers, AI Systems Architects

Install Ollama and Run Llama 3.2 Locally: Python, LangChain

Install Ollama, run Llama 3.2 on your own machine, and call it from the terminal, a browser sidebar, plain HTTP, the official Python client and LangChain - then reach it from another machine without exposing it to the internet.

Updated
#tutorial#beginner#ollama#llama#python#langchain#page-assist

Updated 29 September 2026: this tutorial on how to install Ollama and run Llama 3.2 locally is rewritten for Ollama 0.34. It corrects the Llama 3.2 lineup (the 11B and 90B models are a separate Ollama model, llama3.2-vision), covers the desktop chat app that macOS and Windows installs now include, adds the official ollama Python client and LangChain's ChatOllama, and replaces the old remote-access advice. The earlier version told you to set OLLAMA_HOST=0.0.0.0 and OLLAMA_ORIGINS=*. Do not do that: Ollama has no login, and that setting puts it on every network your machine can reach. Step 8 shows the safe way.

1. What you'll build: Ollama running Llama 3.2 locally, five ways to reach it

text
$ ollama run llama3.2 "State Newton's first law of motion in one sentence."An object at rest will remain at rest, and an object in motion willcontinue to move with a constant velocity, unless acted upon by anexternal force.$ python chat_client.pyQ1: What is Newton's second law of motion?A1: Newton's second law of motion states that the force applied to an object isequal to the mass of the object multiplied by its acceleration. Mathematically,this is represented by the equation F = ma, where F is the net force acting onthe object, m is its mass, and a is its acceleration.Q2: Give one everyday example of it.A2: When you push a heavy box across the floor, the force you apply to the box(F) is equal to its mass (m) multiplied by its acceleration (a), resulting inthe box moving at a certain speed.

That is Llama 3.2 answering from your own laptop, with no API key and no network call after the download. The same model then answers a browser sidebar, a Python script over HTTP, the official Python client, and a LangChain chain, because all of them talk to one local server.

By the end you will have installed Ollama, with Llama 3.2 running on your machine, and you will be able to reach it from the terminal, from your browser, from Python, and from LangChain. If you have a second machine, you will also reach it from there through an SSH tunnel, without opening it to the network. This is a beginner tutorial. You need to be able to open a terminal and run a Python script; everything else is explained as it comes. Budget about 30 minutes, most of it waiting for a 2 GB download.

Verified against Ollama 0.34.4, Python 3.13.9, requests==2.34.2, ollama==0.6.3, langchain-ollama==1.1.0 and langchain-core==1.6.5 on 2026-09-29, on Windows 11. Four things were not run as written: the macOS and Linux install and activation commands (the test machine runs Windows), the Page Assist extension (a browser interface), the SSH tunnel in Step 8 (it needs a remote host), and the Set-ExecutionPolicy fix in When it breaks, which would change the test machine's settings. Step 8's OLLAMA_HOST switch was checked against a second local Ollama server instead.

2. Prerequisites: Ollama 0.34, Python 3.10+, and 4 GB of free RAM

You need:

  • A computer with about 4 GB of free memory and 3 GB of free disk. The 3B model needs roughly 2 GB of disk and more than that in memory while it runs. On a machine with less, use the 1B model instead (Step 2 shows both). No GPU is required; Ollama uses one if it finds one.
  • macOS 14 or later, Windows 10 22H2 or later, or Linux.
  • Python 3.10 or later for Steps 5 to 7. requests and langchain-core both require 3.10.
  • Chrome, Edge or Firefox for Step 4.
  • No accounts, no API keys, and nothing to pay for.
  • Optional, for Step 8 only: a second machine you can log in to with ssh, such as a home server or a cloud server. You can skip Step 8 and still finish everything else.

Check your Python version first with python --version (on macOS and Linux, python3 --version). It must be 3.10 or later. Then create a folder for the Python part of the tutorial and a virtual environment inside it. A virtual environment keeps these packages separate from anything else installed on your machine.

bash
mkdir ollama-quickstartcd ollama-quickstartpython -m venv .venv

On macOS and Linux the command is often python3, not python, so type python3 -m venv .venv there. Once the environment is active, python works on every system. On Ubuntu and Debian, if the command fails with a message that mentions ensurepip, install the venv module first with sudo apt install python3-venv.

Activate it. In Windows PowerShell:

powershell
.venv\Scripts\Activate.ps1

On macOS and Linux:

bash
source .venv/bin/activate

If PowerShell refuses to run Activate.ps1, see the execution-policy entry in When it breaks. Your prompt now starts with (.venv). Every python and pip command below assumes it is active.

Create a file named requirements.txt inside ollama-quickstart with these four lines:

File: requirements.txt

text
requests==2.34.2ollama==0.6.3langchain-ollama==1.1.0langchain-core==1.6.5

Install it:

bash
pip install -r requirements.txt

Check that all four imported cleanly:

bash
python -c "import requests, ollama, langchain_ollama, langchain_core; print('ok')"
text
ok

If you see ok, the Python side is ready. If you see ModuleNotFoundError, the virtual environment is not active; activate it and run pip install again.

3. How does Ollama work? One server, many clients

Every later step uses one idea in a different way, so learn the idea before you type anything.

Ollama is a small server program. It does three jobs. It downloads model files and stores them on your disk. It loads a model into memory when something asks for it, and unloads it after a few idle minutes. And it answers requests over HTTP on port 11434 of your own machine.

Everything you use to talk to the model in this tutorial is a client of that server: the ollama command in your terminal, the desktop chat app, Page Assist in your browser, and your Python scripts. None of them contain the model. They all send a request to http://127.0.0.1:11434 and print what comes back.

The address matters. 127.0.0.1 means "this machine only". By default Ollama listens there and nowhere else, so nothing on your network or the internet can reach it. That default is what keeps a server with no password safe, and Step 8 is about keeping it that way while still reaching it from somewhere else.

4. Steps: install Ollama and call Llama 3.2 from each client

Step 1: Install Ollama on macOS, Windows or Linux

Goal: install Ollama and confirm its server is running.

Why this step: everything else needs the server. On macOS and Windows the installer also sets Ollama to start in the background when you log in, so you do not have to start it by hand.

Download the installer for your system from ollama.com/download and run it. The download page always serves the newest release. This tutorial was checked on 0.34.4. A newer version should behave the same way, but small details of the output can change. On Windows, run the installer. On macOS, open the downloaded file, drag Ollama into Applications, and start it once; the app sets up the ollama command. On Linux, use the install script instead. It asks for your password, because it installs Ollama as a system service:

bash
curl -fsSL https://ollama.com/install.sh | sh

On macOS and Windows, the installer also adds a desktop chat app. Since July 2025 you can pick a model and chat in a window, without the terminal. This tutorial uses the terminal because it shows what is happening underneath, but the app talks to the same server.

Run it: open a new terminal and check the version.

bash
ollama --version

Expected output:

text
ollama version is 0.34.4

On Windows machines without an NVIDIA graphics card, Ollama 0.34.4 also prints a line ending in CHECK failed: mlx_compile_cache_new_ before the version. It is harmless noise from a library Ollama bundles; When it breaks has the details.

What just happened: the ollama command asked the server for its version and got an answer, so both halves are installed and the server is running. If the output starts with Warning: could not connect to a running Ollama instance, the server is not running; see When it breaks.

Stay in this new terminal for the rest of the tutorial. Go back into your project folder with cd ollama-quickstart (or the full path to wherever you created it) and activate the virtual environment again with the command from section 2. Your prompt starts with (.venv) again.

Step 2: Download Llama 3.2 and chat with it in the terminal

Goal: download Llama 3.2 and get an answer from it.

Why this step: Ollama ships with no models. ollama run downloads the model the first time you ask for it, loads it, and sends your prompt, so one command proves the whole chain works.

First, pick a size. Meta released four sizes of Llama 3.2. The "B" stands for billions of parameters. Parameters are the numbers the model learned during training, so their count tells you how big the model is. A bigger model is usually better at reasoning. It is also slower and needs more memory.

SizeOllama nameDownloadInputUse it when
1Bllama3.2:1b1.3 GBtextyour machine has little memory
3Bllama3.22.0 GBtextthe default for this tutorial
11Bllama3.2-vision7.8 GBtext and imagesyou need to ask about pictures
90Bllama3.2-vision:90b55 GBtext and imagesyou have a large GPU server

The 11B and 90B models are a separate Ollama model called llama3.2-vision, not bigger versions of llama3.2. ollama run llama3.2:11b fails, and When it breaks shows the error.

Run it: pass the prompt on the command line, so Ollama answers once and exits.

bash
ollama run llama3.2 "State Newton's first law of motion in one sentence."

Expected output (the download progress appears only the first time):

text
An object at rest will remain at rest, and an object in motion willcontinue to move with a constant velocity, unless acted upon by anexternal force.

The exact wording of the answer can differ on your machine. What matters is that you get a one-sentence answer about inertia.

Run it without a prompt to chat instead. The >>> prompt shows grey hint text, Send a message (/? for help), until you start typing. Type a message, press Enter, and type /bye to leave:

bash
ollama run llama3.2
text
>>> Explain inertia to a ten-year-old in two sentences.Here's an explanation of inertia that a ten-year-old can understand:Inertia is a fancy word for the idea that things don't like to changetheir motion - like a ball just keeps rolling if it's rolling, or stays onthe ground if it's sitting still. When something is moving or sittingstill, it wants to keep doing that, so it will keep moving or stayingstill unless something else pushes or pulls it to change its motion.>>> /bye

The model does not always follow instructions like "two sentences" exactly; small models are loose about length. Your answer will be worded differently.

Check what is on your disk now:

bash
ollama list
text
NAME               ID              SIZE      MODIFIED       llama3.2:latest    a80c4f17acd5    2.0 GB    18 seconds ago    

If you pulled other models before, they appear in this list too.

What just happened: Ollama downloaded Llama 3.2 once and stored it, so the next ollama run llama3.2 starts in seconds. The model stays loaded in memory for about five minutes after the last request, then Ollama frees the memory. If you want the smaller model, run ollama run llama3.2:1b the same way. It is small enough for running the 1B model on a Raspberry Pi.

Step 3: Check the HTTP API that every client uses

Goal: talk to the Ollama server directly, without the ollama command.

Why this step: Steps 4 to 7 all use this API. If you see it answer once by hand, a failure later is easy to place: either the server is not answering, or your code is wrong.

Run it: ask for the server version, then for the list of models. On Windows, type curl.exe, not curl. In Windows PowerShell, curl is a different command with different output.

bash
curl http://127.0.0.1:11434/api/version
text
{"version":"0.34.4"}

The model list comes back as one long line of JSON. Pipe it through Python's built-in formatter to read it. The -s flag stops curl from printing a download progress meter above the JSON, which it does whenever its output goes into a pipe. This uses the Python from your virtual environment, so keep it active. On Windows, the same command is curl.exe -s http://127.0.0.1:11434/api/tags | python -m json.tool.

bash
curl -s http://127.0.0.1:11434/api/tags | python -m json.tool

Expected output (your modified_at time will differ):

text
{    "models": [        {            "name": "llama3.2:latest",            "model": "llama3.2:latest",            "modified_at": "2026-09-29T09:06:01.7225009+05:30",            "size": 2019393189,            "digest": "a80c4f17acd55265feec403c7aef86be0c25983ab279d83f3bcd3abbcb5b8b72",            "details": {                "parent_model": "",                "format": "gguf",                "family": "llama",                "families": [                    "llama"                ],                "parameter_size": "3.2B",                "quantization_level": "Q4_K_M",                "context_length": 131072,                "embedding_length": 3072            },            "capabilities": [                "completion",                "tools"            ]        }    ]}

What just happened: you got JSON back from the same server the terminal used. That confirms the server listens on 127.0.0.1:11434, and that the models you pulled are visible to anything that can reach that address. Every client in the next four steps sends requests like these for you.

Step 4: Chat with Llama 3.2 in your browser with Page Assist

Goal: use the local model from a sidebar in your browser.

Why this step: a terminal is awkward for long answers, and it cannot see the web page in front of you. Page Assist is an open-source browser extension that gives the local model a chat sidebar and a full-page chat view. It can also use the page you are reading as context.

Install it from the Chrome Web Store (this also works in Edge) or from Firefox Add-ons.

Run it: press Ctrl+Shift+L to open the full-page chat, or Ctrl+Shift+Y for the sidebar. If the shortcut does nothing, click the Page Assist icon in the browser toolbar instead. Chrome and Edge hide new extensions behind the puzzle-piece icon, so pin Page Assist there first. Choose llama3.2:latest from the model menu at the top and send a message.

Expected output: Page Assist finds your local Ollama by itself, and the model menu lists the models from Step 2.

To confirm the browser used your local server, run this in a terminal right after you get an answer:

bash
ollama ps
text
NAME               ID              SIZE      PROCESSOR    CONTEXT    UNTIL              llama3.2:latest    a80c4f17acd5    2.6 GB    100% CPU     4096       4 minutes from now    

ollama ps lists the models loaded in memory right now. PROCESSOR shows 100% CPU on a machine without a supported graphics card, and a GPU share if Ollama found one. UNTIL counts down to when Ollama unloads the model, and it restarts from about five minutes with every request. A fresh countdown right after the browser answers means its request reached this server. CONTEXT is the context length: how much text the model can consider at once, prompt and answer together, counted in tokens (pieces of words). CONTEXT and SIZE follow the server's setting: 4096 tokens and 2.6 GB at the server default shown here, and larger if you raise the context length in the desktop app's settings.

What just happened: the extension sent the same kind of HTTP request you sent with curl in Step 3. You did not have to configure anything, because Page Assist looks for Ollama at 127.0.0.1:11434 and handles the browser's cross-origin check itself. That check normally stops a web page or extension from calling a server on another address unless the server allows it. If you see Ollama call failed with status code 403, see When it breaks.

Step 5: Call Llama 3.2 from Python with the HTTP API and requests

Goal: send a prompt from a Python script and read the answer.

Why this step: this is the lowest-level way to use Ollama from code. It needs only the requests library, and it shows you the exact fields the server sends back. The higher-level clients in Steps 6 and 7 hide those fields from you.

Create generate_http.py in the ollama-quickstart folder:

File: generate_http.py

python
import textwrapimport requestsOLLAMA_URL = "http://127.0.0.1:11434/api/generate"payload = {    "model": "llama3.2",    "prompt": "State Newton's first law of motion in one sentence.",    "stream": False,    "options": {"temperature": 0, "seed": 42},}response = requests.post(OLLAMA_URL, json=payload, timeout=120)response.raise_for_status()data = response.json()print("Model:", data["model"])print(textwrap.fill("Answer: " + data["response"], width=80))print("Tokens generated:", data["eval_count"])

Three settings in payload matter. "stream": False asks for the whole answer as one JSON object. Without it, the server streams the answer as many small JSON lines, and response.json() fails. temperature 0 and a fixed seed make the answer repeatable: run the script twice and you get the same text twice. Your text can still differ slightly from the output below, because generation also depends on your hardware and on server settings such as the context length.

The textwrap.fill lines at the end are only for display. They wrap the answer at 80 characters so it reads well in a terminal.

Run it (from the ollama-quickstart folder, with the virtual environment active):

bash
python generate_http.py

Expected output:

text
Model: llama3.2Answer: Newton's first law of motion states that an object at rest will remainat rest, and an object in motion will continue to move with a constant velocity,unless acted upon by an external force.Tokens generated: 40

What just happened: your script sent one HTTP request to /api/generate and printed three fields from the reply. response is the answer. eval_count is how many tokens (pieces of words) the model generated. Watch this number when you care about speed. The reply also carries timing fields, all in nanoseconds.

Step 6: Hold a conversation with the official ollama Python client

Goal: ask a question, then a follow-up that only makes sense with the first answer.

Why this step: /api/generate from Step 5 takes one prompt and forgets it. A conversation needs history. The server's /api/chat endpoint takes a list of messages instead, and the official ollama Python package wraps it, so you write a function call rather than a request body.

Create chat_client.py in the ollama-quickstart folder:

File: chat_client.py

python
import textwrapfrom ollama import chatOPTIONS = {"temperature": 0, "seed": 42}messages = [    {"role": "system", "content": "Answer in one or two short sentences."},    {"role": "user", "content": "What is Newton's second law of motion?"},]first = chat(model="llama3.2", messages=messages, options=OPTIONS)print("Q1:", messages[-1]["content"])print(textwrap.fill("A1: " + first.message.content, width=80))# Keep the answer in the history, then ask a question that depends on it.messages.append({"role": "assistant", "content": first.message.content})messages.append({"role": "user", "content": "Give one everyday example of it."})second = chat(model="llama3.2", messages=messages, options=OPTIONS)print("Q2:", messages[-1]["content"])print(textwrap.fill("A2: " + second.message.content, width=80))

Run it:

bash
python chat_client.py

Expected output:

text
Q1: What is Newton's second law of motion?A1: Newton's second law of motion states that the force applied to an object isequal to the mass of the object multiplied by its acceleration. Mathematically,this is represented by the equation F = ma, where F is the net force acting onthe object, m is its mass, and a is its acceleration.Q2: Give one everyday example of it.A2: When you push a heavy box across the floor, the force you apply to the box(F) is equal to its mass (m) multiplied by its acceleration (a), resulting inthe box moving at a certain speed.

What just happened: the second question says only "it", and the model still knew you meant Newton's second law. That works because the script sent the whole list of messages both times, including the model's own first answer. The server keeps no memory between requests. The history lives in your messages list. The first message has the role system: a standing instruction that applies to the whole conversation, which is why both answers stay short. How the system role gets rendered into the text the model actually reads is its own topic.

The package also reads the OLLAMA_HOST environment variable. Step 8 uses it to point the same script at another machine.

Step 7: Use Llama 3.2 from LangChain with ChatOllama

Goal: call the local model through LangChain, with a reusable prompt template.

Why this step: LangChain is a Python library for building applications out of model calls, and it lets you switch between model providers without rewriting your code. If you build with it, the local model should work wherever a cloud model would. ChatOllama from the langchain-ollama package does that. LangChain also offers OllamaLLM, but that class is for plain text-completion models, which continue a piece of text instead of answering messages. Llama 3.2 is a chat model, so use ChatOllama.

Create chat_langchain.py in the ollama-quickstart folder:

File: chat_langchain.py

python
import textwrapfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_ollama import ChatOllamaprompt = ChatPromptTemplate.from_messages(    [        (            "system",            "Answer the question in one sentence. "            "If you cannot answer it, reply only: I don't know.",        ),        ("human", "{query}"),    ])model = ChatOllama(model="llama3.2", temperature=0, seed=42)chain = prompt | modelfor query in [    "What is Newton's third law of motion?",    "What did I have for breakfast this morning?",]:    reply = chain.invoke({"query": query})    print("Q:", query)    print(textwrap.fill("A: " + reply.content, width=80))

Run it:

bash
python chat_langchain.py

Expected output:

text
Q: What is Newton's third law of motion?A: Newton's third law of motion states that for every action, there is an equaland opposite reaction.Q: What did I have for breakfast this morning?A: I don't know.

What just happened: prompt | model built a chain. The template fills in {query}, and ChatOllama sends the messages to the same /api/chat endpoint Step 6 used. invoke returns a message object, so the text is in .content. In the template, human is LangChain's name for the role that Step 6 called user. The second question tests the system instruction. The model has no way to know your breakfast, and it says so instead of inventing an answer.

Step 8 (optional): Reach Ollama from another machine with an SSH tunnel

Goal: use an Ollama server running on another machine, such as a GPU box or a cloud server, from your laptop.

Why this step: the tempting shortcut is to make the server listen on every network with OLLAMA_HOST=0.0.0.0. Ollama has no login and no password. Anyone who can reach the port can use your models, pull new ones, and delete yours. In January 2026, researchers counted 175,000 Ollama servers open on the public internet. CVE-2026-7482 let an unauthenticated attacker read server memory through /api/create on versions before 0.17.1.

An SSH tunnel avoids all of this. The server keeps listening on 127.0.0.1 only, and you reach it through SSH, which already checks who you are. This step needs a second machine you can log in to with ssh. Set that machine up first. Log in with ssh you@gpu-box, run the Linux install command from Step 1, then run ollama pull llama3.2, and type exit to leave. ollama pull downloads a model without starting a chat; Step 2's ollama run did the same download as a side effect. The tunnel carries requests, not models, so a remote server with no models has nothing to answer with.

Run it: on your laptop, open the tunnel and leave this terminal open. Replace you@gpu-box with your login name on the remote machine and its hostname or IP address.

bash
ssh -N -L 11435:127.0.0.1:11434 you@gpu-box

This forwards port 11435 on your laptop to port 11434 on the remote machine. It uses 11435 so it does not clash with a local Ollama on 11434. -N means "forward only, run no remote command". The first time you connect, SSH asks you to confirm the remote machine's fingerprint; type yes. It may also ask for your password. After that the terminal shows nothing, and that silence means the tunnel is up.

Open a second terminal, go into ollama-quickstart, and activate the virtual environment as before. Then point the ollama command and the Python client at the tunnel. In Windows PowerShell:

powershell
$env:OLLAMA_HOST = "127.0.0.1:11435"

On macOS and Linux:

bash
export OLLAMA_HOST=127.0.0.1:11435
bash
ollama listpython chat_client.py

Expected output: ollama list shows the models on the remote machine, not your laptop's, so you see llama3.2:latest from the pull you ran there. chat_client.py prints its two answers from the remote model. generate_http.py does not read OLLAMA_HOST. To use it through the tunnel, change 11434 to 11435 in OLLAMA_URL.

What just happened: your requests went into SSH on your laptop, came out on the remote machine, and reached Ollama on its own 127.0.0.1. Nothing on the remote machine listens on a public address. Close the first terminal and the path is gone.

Beyond this tutorial, sometimes you do need Ollama to listen on a network address, for example so several people on a trusted office network can use it. In that case, set OLLAMA_HOST to that machine's private address and allow only your network in the firewall. Never do this on a machine with a public IP. The Ollama FAQ shows how to set the variable on each system: launchctl setenv on macOS, systemctl edit ollama.service on Linux, and the user environment variables dialog on Windows. For access from outside, put an authenticating reverse proxy or a VPN in front of it. Do not set OLLAMA_ORIGINS=*; Page Assist does not need it.

5. When it breaks: Ollama errors and their fixes

Every error below was produced on purpose on Ollama 0.34.4 on Windows. The macOS and Linux variants are the shells' standard wording or come from the linked issue reports.

ollama : The term 'ollama' is not recognized as the name of a cmdlet, function, script file, or operable program.

The terminal cannot find the ollama command. On Windows this usually means the terminal was already open when you installed Ollama, so it does not know the new program path. Close it and open a new one. On macOS and Linux the shell says command not found: ollama or ollama: command not found. On macOS, start the Ollama app once, because the app sets up the command. On Linux, run the install script from Step 1 again.

Activate.ps1 cannot be loaded because running scripts is disabled on this system

Windows PowerShell blocks scripts by default on many machines, and the virtual environment's activation file is a script. Allow scripts you create yourself for your user account only, then activate again:

powershell
Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser

Warning: could not connect to a running Ollama instance

The full output is two lines, and the second is Warning: client version is 0.34.4. The ollama command is installed but the server is not running. On macOS and Windows, start Ollama from the Applications folder or the Start menu; it runs in the menu bar or the system tray. On Linux the same problem looks like Error: ollama server not responding - could not connect to ollama server, run 'ollama serve' to start it (ollama#11424). Start the service with sudo systemctl start ollama, or run ollama serve in a separate terminal.

Error: pull model manifest: file does not exist

You asked for a model name that does not exist in the Ollama library. This is exactly what ollama run llama3.2:11b prints, because the 11B model lives under a different name. Check the exact name and tag on ollama.com/library. For the 11B vision model, it is llama3.2-vision.

Error: listen tcp 127.0.0.1:11434: bind: Only one usage of each socket address (protocol/network address/port) is normally permitted.

You ran ollama serve while Ollama was already running. That is the Windows wording; macOS and Linux say bind: address already in use (ollama#707). Nothing is wrong. The desktop app or the Linux service already started the server, so close the extra terminal and use the running one.

CHECK failed: mlx_compile_cache_new_

On Windows machines without an NVIDIA graphics card, Ollama 0.34.4 prints Sep 24 2026 03:47:58 - ERROR - generated.c:2695 - CHECK failed: mlx_compile_cache_new_ before the output of almost every command. It comes from a bundled library that Ollama loads but does not use on that hardware, and it does not affect results. It is reported in ollama#18283, and an open pull request, ollama#18335, stops it being printed.

requests.exceptions.JSONDecodeError: Extra data: line 2 column 1

Your /api/generate request did not set "stream": False. The server then sends one small JSON object per line as the answer is generated, and response.json() can only read one object. Add "stream": False to the payload, as in Step 5. The message ends with a character count, such as (char 98), that changes from run to run.

requests.exceptions.HTTPError: 404 Client Error: Not Found for url: http://127.0.0.1:11434/api/generate

The model in your payload is not on this server. The response body says which one, for example {"error":"model 'llama3.2:1b' not found"}. Run ollama pull with that name, or fix the spelling in the script. Unlike ollama run, the API never downloads a model for you. (The captured run used a second test server on port 11500, so its URL named that port. The message is otherwise identical.)

requests.exceptions.ConnectionError: HTTPConnectionPool(host='127.0.0.1', port=11434): Max retries exceeded

Nothing is listening at that address, so requests could not connect. On Windows the end of the message reads [WinError 10061] No connection could be made because the target machine actively refused it; on Linux it reads [Errno 111] Connection refused (ollama#3200). Start Ollama, or check the port in OLLAMA_URL. (The captured run used a closed port, 11999, to produce the error, so its message named that port.)

ConnectionError: Failed to connect to Ollama. Please check that Ollama is downloaded, running and accessible. https://ollama.com/download

This is the same problem seen from the ollama Python package. LangChain's ChatOllama reports it differently, as httpx.ConnectError: [WinError 10061] No connection could be made because the target machine actively refused it. Either the server is not running, or OLLAMA_HOST points somewhere nothing is listening. After Step 8, the usual cause is a closed SSH tunnel. Open it again, or clear the variable with Remove-Item Env:OLLAMA_HOST (PowerShell) or unset OLLAMA_HOST (macOS and Linux).

Ollama call failed with status code 403

Page Assist reached an Ollama server that refused the browser's request (page-assist#309). With Ollama on 127.0.0.1:11434 this should not happen, because Page Assist fixes up the request itself. It appears when Page Assist talks to Ollama at another address. Turn on the custom origin setting described in Page Assist's connection issues guide, rather than setting OLLAMA_ORIGINS=* on the server.

6. How the pieces connect

Every client on the left sends the same kind of HTTP request to one server, and only the server touches the model files. The dashed line is the one setting that changes for Step 8. Point OLLAMA_HOST at the tunnel, and the same client reaches a remote server that still listens on its own 127.0.0.1 only. Nothing on the public network has a way in.

d2
direction: right

laptop: Your machine {
  style: {fill: "#F4F6F7"; stroke: "#95A5A6"; font-color: "#2C2C2A"}

  cli: "ollama CLI\n(Steps 1-2)" {style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}}
  app: "Desktop app" {style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}}
  pa: "Page Assist\n(Step 4)" {style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}}
  gen: "generate_http.py\n(Step 5)" {style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}}
  chat: "chat_client.py\n(Step 6)" {style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}}
  lc: "chat_langchain.py\n(Step 7)" {style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}}

  server: "Ollama server\n127.0.0.1:11434" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; bold: true}}
  disk: "Model files\nllama3.2, llama3.2:1b" {shape: cylinder; style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}

  cli -> server: "HTTP"
  app -> server: "HTTP"
  pa -> server: "HTTP"
  gen -> server: "/api/generate"
  chat -> server: "/api/chat"
  lc -> server: "/api/chat"
  server -> disk: "loads"
  chat -> tunnel: "OLLAMA_HOST" {style: {stroke-dash: 4}}
  tunnel: "SSH tunnel\n127.0.0.1:11435" {style: {fill: "#7B68EE"; stroke: "#5543C4"; font-color: "#FFFFFF"}}
}

remote: Remote machine (Step 8) {
  style: {fill: "#F4F6F7"; stroke: "#95A5A6"; font-color: "#2C2C2A"}
  rserver: "Ollama server\n127.0.0.1:11434 only" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; bold: true}}
}

internet: "Public network\n(no path in)" {style: {fill: "#E74C3C"; stroke: "#A93226"; font-color: "#FFFFFF"}}

laptop.tunnel -> remote.rserver: "SSH, port 22"
internet -> remote.rserver: "blocked: not listening" {style: {stroke: "#E74C3C"; stroke-dash: 4; font-color: "#A93226"}}

7. The complete project

text
ollama-quickstart/├── .venv/├── requirements.txt├── generate_http.py├── chat_client.py└── chat_langchain.py

File: requirements.txt

text
requests==2.34.2ollama==0.6.3langchain-ollama==1.1.0langchain-core==1.6.5

File: generate_http.py

python
import textwrapimport requestsOLLAMA_URL = "http://127.0.0.1:11434/api/generate"payload = {    "model": "llama3.2",    "prompt": "State Newton's first law of motion in one sentence.",    "stream": False,    "options": {"temperature": 0, "seed": 42},}response = requests.post(OLLAMA_URL, json=payload, timeout=120)response.raise_for_status()data = response.json()print("Model:", data["model"])print(textwrap.fill("Answer: " + data["response"], width=80))print("Tokens generated:", data["eval_count"])

File: chat_client.py

python
import textwrapfrom ollama import chatOPTIONS = {"temperature": 0, "seed": 42}messages = [    {"role": "system", "content": "Answer in one or two short sentences."},    {"role": "user", "content": "What is Newton's second law of motion?"},]first = chat(model="llama3.2", messages=messages, options=OPTIONS)print("Q1:", messages[-1]["content"])print(textwrap.fill("A1: " + first.message.content, width=80))# Keep the answer in the history, then ask a question that depends on it.messages.append({"role": "assistant", "content": first.message.content})messages.append({"role": "user", "content": "Give one everyday example of it."})second = chat(model="llama3.2", messages=messages, options=OPTIONS)print("Q2:", messages[-1]["content"])print(textwrap.fill("A2: " + second.message.content, width=80))

File: chat_langchain.py

python
import textwrapfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_ollama import ChatOllamaprompt = ChatPromptTemplate.from_messages(    [        (            "system",            "Answer the question in one sentence. "            "If you cannot answer it, reply only: I don't know.",        ),        ("human", "{query}"),    ])model = ChatOllama(model="llama3.2", temperature=0, seed=42)chain = prompt | modelfor query in [    "What is Newton's third law of motion?",    "What did I have for breakfast this morning?",]:    reply = chain.invoke({"query": query})    print("Q:", query)    print(textwrap.fill("A: " + reply.content, width=80))

8. Where to go next

  • Stream the answer word by word. Pass stream=True to chat() in chat_client.py and loop over the result, printing chunk.message.content as each piece arrives. This is how chat apps show text appearing as it is generated.
  • Ask about an image. Pull llama3.2-vision (7.8 GB) and pass an image path in a message's images list with the ollama client. It is the same chat() call with one more field.
  • Build an offline voice assistant. This audio bot tutorial answers spoken questions from your own PDFs with a local model.
  • Decide what to run when you outgrow one laptop. This guide to LLM inference frameworks compares Ollama with servers built for many users.
  • Build a small question-answering app on your own documents. This chatbot tutorial uses a local Llama model with Qdrant and Streamlit.

9. References

Genai

Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:

Books by Ranjan Kumar

The 7 GenAI Architectures cover

The 7 GenAI Architectures

Building Real-World Agentic AI Systems with LangGraph cover

Building Real-World Agentic AI Systems

The ChatML Handbook cover

The ChatML Handbook

The Chat Templates Handbook cover

The Chat Templates Handbook

Comments