Running Codex Against a Local Model on an HPC Cluster
Published:
Codex (and most other AI coding tools) assume you’re talking to OpenAI’s API. But if you have access to an HPC cluster with GPUs, you can serve an open-weight model yourself with Ollama and point Codex at it instead — no external API calls, no API costs, full control over the model.
Here’s how I set this up on our cluster.
1. Serve the model with Slurm
The core piece is a Slurm batch script that allocates a GPU node, starts ollama serve, and writes out the connection details once the server is up. Here’s the script (ollama-endpoint.sbatch):
#!/bin/bash
#SBATCH --job-name=ollama-endpoint
#SBATCH --partition=gpu
#SBATCH --gres=gpu:l40s:1
#SBATCH --cpus-per-task=8
#SBATCH --mem=96G
#SBATCH --time=08:00:00
#SBATCH --output=ollama-endpoint-%j.log
#
# Serves an open-weight model from a Slurm allocation as an HTTP endpoint.
# Any OpenAI-compatible client can point at it: LiteLLM, LangChain, the
# openai SDK, Continue, Aider, or your own code.
#
# sbatch ollama-endpoint.sbatch
# cat ollama-endpoint-<jobid>.info <- connection details
# tail -f ollama-endpoint-<jobid>.log <- server log
#
# Overrides:
# sbatch --export=ALL,MODEL=llama3.1:70b,PORT=11500 ollama-endpoint.sbatch
#
# MODEL model to serve; must already be in the store (ollama list)
# PORT listen port (default 11434)
# CONTEXT context window in tokens (default 65536)
# OLLAMA_MODELS model store path
set -euo pipefail
source /etc/profile.d/modules.sh 2>/dev/null || true
module load ollama
# The site module points OLLAMA_MODELS at $HOME/.ollama/models, a per-user
# copy seeded by user_sync.sh. Use the shared store instead.
# Verify contents with: OLLAMA_MODELS=<path> ollama list
export OLLAMA_MODELS="${OLLAMA_MODELS:-/mnt/nasapps/production/ollama/models}"
MODEL="${MODEL:-qwen3-coder:30b}"
PORT="${PORT:-11434}"
CONTEXT="${CONTEXT:-65536}"
export OLLAMA_HOST="0.0.0.0:${PORT}"
export OLLAMA_CONTEXT_LENGTH="${CONTEXT}"
export OLLAMA_KEEP_ALIVE=-1 # keep the model resident for the whole job
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_NUM_PARALLEL=1 # one client, full context, no KV cache split
NODE=$(hostname -f)
INFO="${SLURM_SUBMIT_DIR:-$PWD}/ollama-endpoint-${SLURM_JOB_ID}.info"
echo "node=${NODE} port=${PORT} gpus=${CUDA_VISIBLE_DEVICES:-none}"
echo "OLLAMA_MODELS=${OLLAMA_MODELS}"
ollama serve &
SERVER_PID=$!
trap 'kill $SERVER_PID 2>/dev/null || true' EXIT
for _ in $(seq 1 60); do
curl -sf "http://127.0.0.1:${PORT}/api/tags" >/dev/null && break
sleep 2
done
# Written before warm-up so it exists even if the model load is slow.
cat > "$INFO" <<EOF
JOB=${SLURM_JOB_ID}
NODE=${NODE}
PORT=${PORT}
MODEL=${MODEL}
CONTEXT=${CONTEXT}
# OpenAI-compatible API - most clients, LiteLLM, LangChain, openai SDK
OPENAI_BASE_URL=http://${NODE}:${PORT}/v1
OPENAI_API_KEY=not-needed
# Native Ollama API - the ollama python package, langchain_ollama
# (langchain_ollama ignores its base_url argument; set OLLAMA_HOST instead)
OLLAMA_HOST=http://${NODE}:${PORT}
# Quick check:
# curl http://${NODE}:${PORT}/v1/chat/completions \
# -d '{"model":"${MODEL}","messages":[{"role":"user","content":"hi"}]}'
EOF
# No pull. The model must already be in the store; a pull against a read-only
# shared store would kill the job under set -e.
echo "warming ${MODEL}, first load off NFS can take several minutes..."
curl -s --max-time 900 "http://127.0.0.1:${PORT}/api/generate" \
-d "{\"model\":\"${MODEL}\",\"prompt\":\"ready\",\"stream\":false}" >/dev/null \
&& echo "warm-up ok" \
|| echo "WARNING: warm-up did not complete; first request will be slow"
echo "=============================================================="
cat "$INFO"
echo "=============================================================="
echo "connection details written to: ${INFO}"
wait $SERVER_PID
A few details worth calling out:
OLLAMA_KEEP_ALIVE=-1keeps the model resident in GPU memory for the life of the job, so you’re not paying a reload penalty between requests.OLLAMA_NUM_PARALLEL=1is a deliberate choice for a single-user session — it gives that one client the full context window instead of splitting the KV cache across multiple concurrent slots.- The script deliberately does not run
ollama pull. The shared model store is read-only, and a pull failing underset -euo pipefailwould otherwise kill the job. The model has to already exist in the store. - The
.infofile is written before the warm-up request completes, so you always get connection details even if the first model load (often slow, especially off NFS) is still in progress.
Launching it
sbatch ollama-endpoint.sbatch
Or to override the model and port:
sbatch --export=ALL,MODEL=qwen3:8b ollama-endpoint.sbatch
Once the job is running, check:
cat ollama-endpoint-<jobid>.info # connection details
tail -f ollama-endpoint-<jobid>.log # server log
The .info file gives you the node and port the server landed on — since Slurm picks the node dynamically, this changes every time you submit the job.
2. Point Codex at it
Codex reads provider configuration from ~/.codex/config.toml. Add a custom provider entry pointing at the Ollama endpoint:
[model_providers.frce-ollama]
name = "FRCE Ollama"
base_url = "http://fsitgl-hpc182p.ncifcrf.gov:11434/v1"
wire_api = "responses"
env_key = "OLLAMA_API_KEY"
Then create a profile that uses this provider — I keep mine in a separate file, ~/.codex/frce.config.toml:
model_provider = "frce-ollama"
model = "qwen3-coder:30b"
model_reasoning_effort = "none"
Ollama doesn’t check the API key, but Codex still expects the environment variable to exist, so set a placeholder in your shell profile:
export OLLAMA_API_KEY=not-needed
3. Run Codex against the local model
codex --profile frce
This routes all requests through the frce-ollama provider to whichever node is currently running the Slurm job, instead of hitting OpenAI’s API.
Notes and gotchas
- The node name in
base_urlhas to match whatever’s in the current.infofile — since Slurm reassigns nodes per job, you’ll need to updateconfig.toml(or template it) each time you resubmit. - Match
modelin the Codex profile to whateverMODELthe Slurm job actually launched. - The 8-hour
--timelimit means the endpoint disappears when the job ends — fine for interactive sessions, less fine if you want something persistent. A cron-resubmitted job or a longer time limit are both reasonable fixes depending on your cluster’s policies. CONTEXT=65536is a reasonable default for coding tasks, but it eats GPU memory — drop it if you’re running a bigger model on the same card.
This setup works for any OpenAI-compatible client, not just Codex — the same .info file and OPENAI_BASE_URL work for LangChain, the openai Python SDK, Aider, Continue, or your own scripts.
Reference
https://ncifrederick.cancer.gov/staff/FRCE/OllamaEndpoint
