Running Codex Against a Local Model on an HPC Cluster

5 minute read

Published:

Codex (and most other AI coding tools) assume you’re talking to OpenAI’s API. But if you have access to an HPC cluster with GPUs, you can serve an open-weight model yourself with Ollama and point Codex at it instead — no external API calls, no API costs, full control over the model.

Here’s how I set this up on our cluster.

1. Serve the model with Slurm

The core piece is a Slurm batch script that allocates a GPU node, starts ollama serve, and writes out the connection details once the server is up. Here’s the script (ollama-endpoint.sbatch):

#!/bin/bash
#SBATCH --job-name=ollama-endpoint
#SBATCH --partition=gpu
#SBATCH --gres=gpu:l40s:1
#SBATCH --cpus-per-task=8
#SBATCH --mem=96G
#SBATCH --time=08:00:00
#SBATCH --output=ollama-endpoint-%j.log
#
# Serves an open-weight model from a Slurm allocation as an HTTP endpoint.
# Any OpenAI-compatible client can point at it: LiteLLM, LangChain, the
# openai SDK, Continue, Aider, or your own code.
#
#   sbatch ollama-endpoint.sbatch
#   cat  ollama-endpoint-<jobid>.info     <- connection details
#   tail -f ollama-endpoint-<jobid>.log   <- server log
#
# Overrides:
#   sbatch --export=ALL,MODEL=llama3.1:70b,PORT=11500 ollama-endpoint.sbatch
#
#   MODEL          model to serve; must already be in the store (ollama list)
#   PORT           listen port                      (default 11434)
#   CONTEXT        context window in tokens         (default 65536)
#   OLLAMA_MODELS  model store path

set -euo pipefail
source /etc/profile.d/modules.sh 2>/dev/null || true
module load ollama

# The site module points OLLAMA_MODELS at $HOME/.ollama/models, a per-user
# copy seeded by user_sync.sh. Use the shared store instead.
# Verify contents with:  OLLAMA_MODELS=<path> ollama list
export OLLAMA_MODELS="${OLLAMA_MODELS:-/mnt/nasapps/production/ollama/models}"

MODEL="${MODEL:-qwen3-coder:30b}"
PORT="${PORT:-11434}"
CONTEXT="${CONTEXT:-65536}"

export OLLAMA_HOST="0.0.0.0:${PORT}"
export OLLAMA_CONTEXT_LENGTH="${CONTEXT}"
export OLLAMA_KEEP_ALIVE=-1     # keep the model resident for the whole job
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_NUM_PARALLEL=1    # one client, full context, no KV cache split

NODE=$(hostname -f)
INFO="${SLURM_SUBMIT_DIR:-$PWD}/ollama-endpoint-${SLURM_JOB_ID}.info"

echo "node=${NODE} port=${PORT} gpus=${CUDA_VISIBLE_DEVICES:-none}"
echo "OLLAMA_MODELS=${OLLAMA_MODELS}"

ollama serve &
SERVER_PID=$!
trap 'kill $SERVER_PID 2>/dev/null || true' EXIT

for _ in $(seq 1 60); do
  curl -sf "http://127.0.0.1:${PORT}/api/tags" >/dev/null && break
  sleep 2
done

# Written before warm-up so it exists even if the model load is slow.
cat > "$INFO" <<EOF
JOB=${SLURM_JOB_ID}
NODE=${NODE}
PORT=${PORT}
MODEL=${MODEL}
CONTEXT=${CONTEXT}

# OpenAI-compatible API - most clients, LiteLLM, LangChain, openai SDK
OPENAI_BASE_URL=http://${NODE}:${PORT}/v1
OPENAI_API_KEY=not-needed

# Native Ollama API - the ollama python package, langchain_ollama
# (langchain_ollama ignores its base_url argument; set OLLAMA_HOST instead)
OLLAMA_HOST=http://${NODE}:${PORT}

# Quick check:
#   curl http://${NODE}:${PORT}/v1/chat/completions \
#     -d '{"model":"${MODEL}","messages":[{"role":"user","content":"hi"}]}'
EOF

# No pull. The model must already be in the store; a pull against a read-only
# shared store would kill the job under set -e.
echo "warming ${MODEL}, first load off NFS can take several minutes..."
curl -s --max-time 900 "http://127.0.0.1:${PORT}/api/generate" \
  -d "{\"model\":\"${MODEL}\",\"prompt\":\"ready\",\"stream\":false}" >/dev/null \
  && echo "warm-up ok" \
  || echo "WARNING: warm-up did not complete; first request will be slow"

echo "=============================================================="
cat "$INFO"
echo "=============================================================="
echo "connection details written to: ${INFO}"

wait $SERVER_PID

A few details worth calling out:

  • OLLAMA_KEEP_ALIVE=-1 keeps the model resident in GPU memory for the life of the job, so you’re not paying a reload penalty between requests.
  • OLLAMA_NUM_PARALLEL=1 is a deliberate choice for a single-user session — it gives that one client the full context window instead of splitting the KV cache across multiple concurrent slots.
  • The script deliberately does not run ollama pull. The shared model store is read-only, and a pull failing under set -euo pipefail would otherwise kill the job. The model has to already exist in the store.
  • The .info file is written before the warm-up request completes, so you always get connection details even if the first model load (often slow, especially off NFS) is still in progress.

Launching it

sbatch ollama-endpoint.sbatch

Or to override the model and port:

sbatch --export=ALL,MODEL=qwen3:8b ollama-endpoint.sbatch

Once the job is running, check:

cat ollama-endpoint-<jobid>.info     # connection details
tail -f ollama-endpoint-<jobid>.log  # server log

The .info file gives you the node and port the server landed on — since Slurm picks the node dynamically, this changes every time you submit the job.

2. Point Codex at it

Codex reads provider configuration from ~/.codex/config.toml. Add a custom provider entry pointing at the Ollama endpoint:

[model_providers.frce-ollama]
name = "FRCE Ollama"
base_url = "http://fsitgl-hpc182p.ncifcrf.gov:11434/v1"
wire_api = "responses"
env_key = "OLLAMA_API_KEY"

Then create a profile that uses this provider — I keep mine in a separate file, ~/.codex/frce.config.toml:

model_provider = "frce-ollama"
model = "qwen3-coder:30b"
model_reasoning_effort = "none"

Ollama doesn’t check the API key, but Codex still expects the environment variable to exist, so set a placeholder in your shell profile:

export OLLAMA_API_KEY=not-needed

3. Run Codex against the local model

codex --profile frce

This routes all requests through the frce-ollama provider to whichever node is currently running the Slurm job, instead of hitting OpenAI’s API.

Notes and gotchas

  • The node name in base_url has to match whatever’s in the current .info file — since Slurm reassigns nodes per job, you’ll need to update config.toml (or template it) each time you resubmit.
  • Match model in the Codex profile to whatever MODEL the Slurm job actually launched.
  • The 8-hour --time limit means the endpoint disappears when the job ends — fine for interactive sessions, less fine if you want something persistent. A cron-resubmitted job or a longer time limit are both reasonable fixes depending on your cluster’s policies.
  • CONTEXT=65536 is a reasonable default for coding tasks, but it eats GPU memory — drop it if you’re running a bigger model on the same card.

This setup works for any OpenAI-compatible client, not just Codex — the same .info file and OPENAI_BASE_URL work for LangChain, the openai Python SDK, Aider, Continue, or your own scripts.

Reference

https://ncifrederick.cancer.gov/staff/FRCE/OllamaEndpoint