When Codex Meets NFS: Debugging a Corrupted SQLite Session Database on HPC
Published:
Running developer tools on an HPC system often exposes assumptions that are invisible on a laptop. One of those assumptions is that an application’s local state directory really is local.
I ran into this while using the OpenAI Codex CLI on an HPC login node. Codex had been working normally until I tried to resume an existing session and received this error:
Failed to resume session from ~/.codex/sessions/2026/10/02/rollout-....jsonl:
thread/resume failed during TUI bootstrap:
thread/resume failed:
failed to list thread history:
thread-store internal error:
failed to open thread history database:
failed to open thread history DB at ~/.codex/thread_history_1.sqlite:
error returned from database: (code: 26) file is not a database
The important part was:
(code: 26) file is not a database
That error comes from SQLite. Codex was trying to open its thread-history database, but the file at ~/.codex/thread_history_1.sqlite was no longer a valid SQLite database.
This post walks through how I diagnosed the problem, what I checked along the way, why HPC storage made the issue more likely, and the configuration I ultimately chose to reduce the risk of it happening again.
The environment: Codex on a shared HPC system
The first clue was that this was not a normal single-user workstation. The machine was an HPC login node with several users running Codex-related processes.
I checked with:
ps aux | grep -i codex | grep -v grep
There were Codex processes owned by several accounts, including two processes belonging to my own user. That by itself was not a problem: Unix permissions keep each user’s ~/.codex directory separate. The relevant question was whether my own Codex processes were sharing the same state directory and what kind of filesystem backed that directory.
Before changing anything, I made sure to distinguish my own processes from other users’ processes. On a shared system, never kill a process merely because its command contains codex.
A safer view is:
pgrep -afu "$USER" codex
or:
ps -u "$USER" -f | grep -i codex | grep -v grep
Step 1: Treat the SQLite file as corrupted, but preserve the session data
The failed resume referenced two different kinds of files:
~/.codex/sessions/.../rollout-....jsonl
~/.codex/thread_history_1.sqlite
That distinction matters.
The session rollout file was still present, while the SQLite database used to enumerate or manage thread history was failing to open. My first goal was therefore not to delete the session JSONL.
A useful first inspection is:
file ~/.codex/thread_history_1.sqlite
ls -lh ~/.codex/thread_history_1.sqlite*
head -c 32 ~/.codex/thread_history_1.sqlite | xxd
A normal SQLite database starts with the header:
SQLite format 3
If the file does not contain that signature, or SQLite reports code 26, it is reasonable to treat it as damaged until proven otherwise.
With Codex stopped, another useful check is:
sqlite3 ~/.codex/thread_history_1.sqlite 'PRAGMA quick_check;'
For a healthy database, the expected result is:
ok
Before attempting recovery, I would also copy the relevant files somewhere safe:
ts=$(date +%Y%m%d-%H%M%S)
mkdir -p ~/.codex/recovery-$ts
cp -a ~/.codex/thread_history_1.sqlite* ~/.codex/recovery-$ts/ 2>/dev/null || true
cp -a ~/.codex/sessions/ ~/.codex/recovery-$ts/sessions/
The principle here is simple: preserve the durable session material before modifying the database that indexes or supplements it.
Step 2: Check where ~/.codex actually lives
On a laptop, ~/.codex is usually on a local filesystem. On HPC systems, $HOME is very often a network mount.
I checked:
df -T "$HOME"
df -T ~/.codex
The result was effectively:
Filesystem Type ... Mounted on
server:/shared-home nfs ... /home
So both my home directory and ~/.codex lived on NFS.
That immediately changed the diagnosis.
SQLite is designed primarily around filesystem semantics that are easiest to guarantee on a local filesystem. Network filesystems can introduce additional complexity around locking, caching, atomicity, and failure recovery. SQLite’s own documentation explicitly warns about caveats when using a database directly over a network filesystem, and SQLite’s WAL mode requires all processes using the database to be on the same host.
References:
This does not prove that NFS caused my corrupted file. A corrupted SQLite database can have several causes. But an SQLite runtime database stored in an HPC shared home directory is a strong environmental risk factor, especially when sessions may be opened from multiple login nodes over time.
Step 3: Check whether the cluster’s scratch filesystem is really local
A common HPC reflex is: “Move it to scratch.”
But not all scratch storage is node-local.
I checked the available project scratch filesystem:
df -T /scratch/ccrsf_scratch
It also reported:
Type: nfs
In other words, moving Codex’s SQLite database from $HOME to that scratch path would change the directory but not the underlying storage model.
This is an important HPC troubleshooting lesson:
Never infer that a path is local from its name. Check the filesystem type.
Use:
df -T /path/to/storage
Names such as /scratch, /work, /project, and /data can all be network-backed on a cluster.
Step 4: Find genuinely node-local storage
I then checked the usual candidates:
for p in /tmp /var/tmp /dev/shm; do
echo "=== $p ==="
df -T "$p"
done
The useful result was:
/tmp -> xfs
/var/tmp -> xfs
/dev/shm -> tmpfs
/tmp was particularly attractive because it had plenty of free space and was backed by XFS rather than NFS.
The environment had already defined:
TMPDIR=/tmp/xies4
That gave me a natural user-specific location for Codex’s local runtime database.
At this point the storage picture looked like this:
$HOME/.codex -> NFS, persistent, shared between login nodes
/scratch/... -> NFS, persistent/shared
$TMPDIR -> local XFS on the current node
/dev/shm -> local RAM-backed tmpfs
/dev/shm would also avoid NFS, but it is RAM-backed and highly ephemeral, so I did not choose it for this purpose.
The solution: separate persistent Codex data from SQLite runtime state
Recent Codex versions expose a dedicated sqlite_home setting for the directory that stores the SQLite-backed state database. Codex also supports the CODEX_SQLITE_HOME environment variable; the current Codex source documents sqlite_home as defaulting to $CODEX_SQLITE_HOME when that variable is set, and otherwise to $CODEX_HOME.
The current Codex configuration reference describes sqlite_home as the directory where Codex stores SQLite-backed runtime state used for agent jobs and resumable state:
- OpenAI Codex configuration reference: sqlite_home
That makes it possible to avoid moving the entire Codex home directory.
The arrangement I wanted was:
Persistent/shared data:
~/.codex/config.toml
~/.codex/sessions/
other durable Codex files
|
+--> NFS home directory
SQLite runtime state:
thread/history/state SQLite databases
|
+--> node-local XFS under $TMPDIR
This is preferable to setting all of CODEX_HOME to /tmp, because I still want important session files and configuration to survive logout, cleanup, and node reboot.
One-session test
First create a private local directory:
mkdir -p "$TMPDIR/codex-sqlite"
chmod 700 "$TMPDIR/codex-sqlite"
Then set:
export CODEX_SQLITE_HOME="$TMPDIR/codex-sqlite"
Verify that it is really local:
echo "$CODEX_SQLITE_HOME"
df -T "$CODEX_SQLITE_HOME"
The important part is that the filesystem type should be something like xfs or ext4, not nfs.
Then start Codex normally:
codex
After startup, inspect the local directory:
ls -lah "$CODEX_SQLITE_HOME"
Making it automatic
Because this environment already supplies a node-local $TMPDIR, I prefer an environment-variable setup in ~/.bashrc:
# Keep Codex SQLite runtime state off the NFS home filesystem.
if [ -n "${TMPDIR:-}" ] && [ -d "$TMPDIR" ]; then
export CODEX_SQLITE_HOME="$TMPDIR/codex-sqlite"
mkdir -p "$CODEX_SQLITE_HOME"
chmod 700 "$CODEX_SQLITE_HOME"
fi
Reload the shell:
source ~/.bashrc
and verify:
echo "HOST: $(hostname)"
echo "CODEX_HOME: ${CODEX_HOME:-$HOME/.codex}"
echo "CODEX_SQLITE_HOME: $CODEX_SQLITE_HOME"
df -T "$CODEX_SQLITE_HOME"
Codex also supports the equivalent configuration key in ~/.codex/config.toml:
sqlite_home = "/absolute/local/path/codex-sqlite"
For my HPC environment, however, $TMPDIR is host-specific and already managed by the system, so exporting CODEX_SQLITE_HOME is convenient.
Why this is especially relevant on multi-node HPC systems
The key complication is that $HOME is shared while /tmp is not.
Imagine two login nodes:
login01
/tmp/<user>/codex-sqlite local database A
login02
/tmp/<user>/codex-sqlite local database B
shared NFS
/home/<user>/.codex/sessions
This is intentional.
With the default layout, both nodes can potentially see the same SQLite file through NFS:
login01 ----\
>---- NFS ---- ~/.codex/thread_history_1.sqlite
login02 ----/
With a node-local SQLite directory, each host gets its own database:
login01 ----> local XFS SQLite A
login02 ----> local XFS SQLite B
both -------> shared persistent session files
That avoids asking multiple hosts to coordinate SQLite runtime state through the shared network filesystem.
This does introduce a trade-off: the local SQLite state is no longer inherently shared between nodes. If Codex stores information in SQLite that has not yet been reconstructed from durable session files, switching hosts may initially produce different local state. In practice, this is still a much better failure mode than corrupting the shared database, but it is worth understanding the distinction.
The /tmp trade-off
/tmp is local, fast, and compatible with SQLite’s normal locking assumptions, but it is also temporary.
Administrators may clean it periodically, and its contents may disappear after a reboot.
That means this design should be understood as:
NFS: durable Codex data
/tmp: disposable/rebuildable SQLite runtime state
I would prefer a persistent node-local filesystem if the cluster offered one. For example, a site-specific local SSD path would provide both local filesystem semantics and persistence across reboots.
If /tmp is the only local option, preserve the session JSONL files under the shared home directory and treat the SQLite directory as runtime state that may occasionally need to be recreated.
Do not blindly move the entire CODEX_HOME to /tmp unless you are comfortable losing everything stored there when /tmp is purged.
Recovery after the database is already corrupted
Changing CODEX_SQLITE_HOME prevents future Codex processes from using the bad NFS database, but it does not magically repair an existing corrupt file.
My recovery procedure is:
- Stop my own Codex processes.
- Preserve the existing session JSONLs.
- Back up the damaged SQLite database and any
-wal/-shmcompanions. - Move the corrupt database out of the active path rather than immediately deleting it.
- Start Codex with a clean local SQLite directory.
- Attempt to resume or reconstruct the session from the preserved durable data.
For example:
pgrep -afu "$USER" codex
After confirming that my own Codex processes are stopped:
ts=$(date +%Y%m%d-%H%M%S)
mkdir -p ~/.codex/recovery-$ts
cp -a ~/.codex/sessions/ ~/.codex/recovery-$ts/sessions/
cp -a ~/.codex/thread_history_1.sqlite* ~/.codex/recovery-$ts/ 2>/dev/null || true
Then configure the new local SQLite home:
mkdir -p "$TMPDIR/codex-sqlite"
chmod 700 "$TMPDIR/codex-sqlite"
export CODEX_SQLITE_HOME="$TMPDIR/codex-sqlite"
I intentionally avoid deleting the original rollout/session JSONL simply because the history database is damaged.
Additional precautions
Exit Codex cleanly when possible
Avoid kill -9 unless the process is truly stuck. A clean shutdown gives applications and databases the best chance to flush pending state.
Do not manually modify SQLite sidecar files while Codex is running
Files such as:
*.sqlite
*.sqlite-wal
*.sqlite-shm
can be part of one coordinated database state. Removing one while the process is active is not a safe recovery strategy.
Check database health periodically
With Codex stopped:
for db in "$CODEX_SQLITE_HOME"/*.sqlite; do
[ -e "$db" ] || continue
echo "=== $db ==="
sqlite3 "$db" 'PRAGMA quick_check;'
done
Healthy databases should report:
ok
Check the filesystem, not the directory name
Before using any HPC path for SQLite:
df -T /candidate/path
The useful distinction is not “home versus scratch.” It is network filesystem versus genuinely local filesystem.
A compact diagnostic checklist
If Codex on an HPC system reports an error such as:
file is not a database
for a file under ~/.codex, I would run:
# 1. Identify my Codex processes only.
pgrep -afu "$USER" codex
# 2. Check where Codex state is stored.
df -T "$HOME"
df -T ~/.codex
# 3. Inspect the suspect DB.
file ~/.codex/thread_history_1.sqlite
head -c 32 ~/.codex/thread_history_1.sqlite | xxd
# 4. Check candidate local filesystems.
df -T /tmp /var/tmp /dev/shm
# 5. Check whether the site already assigned a local temp directory.
echo "TMPDIR=${TMPDIR:-<unset>}"
# 6. Move SQLite state to local storage.
mkdir -p "$TMPDIR/codex-sqlite"
chmod 700 "$TMPDIR/codex-sqlite"
export CODEX_SQLITE_HOME="$TMPDIR/codex-sqlite"
# 7. Verify the final filesystem type.
df -T "$CODEX_SQLITE_HOME"
The desired final line should show a local filesystem such as xfs or ext4.
What I learned
The most useful lesson was not specific to Codex.
Many modern CLI tools quietly embed SQLite, LevelDB, RocksDB, lock files, Unix sockets, or other local-state mechanisms under $HOME. On an HPC system, $HOME may be a shared NFS mount visible from many nodes. The application thinks it is writing “local state,” but the infrastructure is providing a distributed filesystem with different failure and locking characteristics.
The troubleshooting pattern is therefore broadly reusable:
Application state error
|
v
Identify the actual state files
|
v
Check the backing filesystem with df -T
|
+--> local filesystem: investigate ordinary corruption/process issues
|
+--> network filesystem: look for a supported way to relocate runtime state
to node-local storage
For Codex specifically, the availability of sqlite_home / CODEX_SQLITE_HOME makes that separation possible without abandoning the persistent shared home directory.
My final arrangement is conceptually:
~/.codex/sessions and configuration
-> shared NFS home
-> persistent
Codex SQLite runtime state
-> $TMPDIR/codex-sqlite
-> node-local XFS
-> temporary/rebuildable
It is a small configuration change, but on an HPC system it aligns each kind of data with the storage system best suited to it.
References
- OpenAI, Codex Configuration Reference —
sqlite_home: https://developers.openai.com/codex/config-reference - SQLite, SQLite Over a Network, Caveats and Considerations: https://www.sqlite.org/useovernet.html
- SQLite, Write-Ahead Logging: https://www.sqlite.org/wal.html
- OpenAI Codex source,
config_toml.rs(sqlite_home): https://github.com/openai/codex/blob/main/codex-rs/config/src/config_toml.rs
