When Codex Meets NFS: Debugging a Corrupted SQLite Session Database on HPC

13 minute read

Published:

Running developer tools on an HPC system often exposes assumptions that are invisible on a laptop. One of those assumptions is that an application’s local state directory really is local.

I ran into this while using the OpenAI Codex CLI on an HPC login node. Codex had been working normally until I tried to resume an existing session and received this error:

Failed to resume session from ~/.codex/sessions/2026/10/02/rollout-....jsonl:
thread/resume failed during TUI bootstrap:
thread/resume failed:
failed to list thread history:
thread-store internal error:
failed to open thread history database:
failed to open thread history DB at ~/.codex/thread_history_1.sqlite:
error returned from database: (code: 26) file is not a database

The important part was:

(code: 26) file is not a database

That error comes from SQLite. Codex was trying to open its thread-history database, but the file at ~/.codex/thread_history_1.sqlite was no longer a valid SQLite database.

This post walks through how I diagnosed the problem, what I checked along the way, why HPC storage made the issue more likely, and the configuration I ultimately chose to reduce the risk of it happening again.


The environment: Codex on a shared HPC system

The first clue was that this was not a normal single-user workstation. The machine was an HPC login node with several users running Codex-related processes.

I checked with:

ps aux | grep -i codex | grep -v grep

There were Codex processes owned by several accounts, including two processes belonging to my own user. That by itself was not a problem: Unix permissions keep each user’s ~/.codex directory separate. The relevant question was whether my own Codex processes were sharing the same state directory and what kind of filesystem backed that directory.

Before changing anything, I made sure to distinguish my own processes from other users’ processes. On a shared system, never kill a process merely because its command contains codex.

A safer view is:

pgrep -afu "$USER" codex

or:

ps -u "$USER" -f | grep -i codex | grep -v grep

Step 1: Treat the SQLite file as corrupted, but preserve the session data

The failed resume referenced two different kinds of files:

~/.codex/sessions/.../rollout-....jsonl
~/.codex/thread_history_1.sqlite

That distinction matters.

The session rollout file was still present, while the SQLite database used to enumerate or manage thread history was failing to open. My first goal was therefore not to delete the session JSONL.

A useful first inspection is:

file ~/.codex/thread_history_1.sqlite
ls -lh ~/.codex/thread_history_1.sqlite*
head -c 32 ~/.codex/thread_history_1.sqlite | xxd

A normal SQLite database starts with the header:

SQLite format 3

If the file does not contain that signature, or SQLite reports code 26, it is reasonable to treat it as damaged until proven otherwise.

With Codex stopped, another useful check is:

sqlite3 ~/.codex/thread_history_1.sqlite 'PRAGMA quick_check;'

For a healthy database, the expected result is:

ok

Before attempting recovery, I would also copy the relevant files somewhere safe:

ts=$(date +%Y%m%d-%H%M%S)
mkdir -p ~/.codex/recovery-$ts

cp -a ~/.codex/thread_history_1.sqlite* ~/.codex/recovery-$ts/ 2>/dev/null || true
cp -a ~/.codex/sessions/ ~/.codex/recovery-$ts/sessions/

The principle here is simple: preserve the durable session material before modifying the database that indexes or supplements it.


Step 2: Check where ~/.codex actually lives

On a laptop, ~/.codex is usually on a local filesystem. On HPC systems, $HOME is very often a network mount.

I checked:

df -T "$HOME"
df -T ~/.codex

The result was effectively:

Filesystem                         Type  ...  Mounted on
server:/shared-home                nfs   ...  /home

So both my home directory and ~/.codex lived on NFS.

That immediately changed the diagnosis.

SQLite is designed primarily around filesystem semantics that are easiest to guarantee on a local filesystem. Network filesystems can introduce additional complexity around locking, caching, atomicity, and failure recovery. SQLite’s own documentation explicitly warns about caveats when using a database directly over a network filesystem, and SQLite’s WAL mode requires all processes using the database to be on the same host.

References:

This does not prove that NFS caused my corrupted file. A corrupted SQLite database can have several causes. But an SQLite runtime database stored in an HPC shared home directory is a strong environmental risk factor, especially when sessions may be opened from multiple login nodes over time.


Step 3: Check whether the cluster’s scratch filesystem is really local

A common HPC reflex is: “Move it to scratch.”

But not all scratch storage is node-local.

I checked the available project scratch filesystem:

df -T /scratch/ccrsf_scratch

It also reported:

Type: nfs

In other words, moving Codex’s SQLite database from $HOME to that scratch path would change the directory but not the underlying storage model.

This is an important HPC troubleshooting lesson:

Never infer that a path is local from its name. Check the filesystem type.

Use:

df -T /path/to/storage

Names such as /scratch, /work, /project, and /data can all be network-backed on a cluster.


Step 4: Find genuinely node-local storage

I then checked the usual candidates:

for p in /tmp /var/tmp /dev/shm; do
    echo "=== $p ==="
    df -T "$p"
done

The useful result was:

/tmp      -> xfs
/var/tmp  -> xfs
/dev/shm  -> tmpfs

/tmp was particularly attractive because it had plenty of free space and was backed by XFS rather than NFS.

The environment had already defined:

TMPDIR=/tmp/xies4

That gave me a natural user-specific location for Codex’s local runtime database.

At this point the storage picture looked like this:

$HOME/.codex              -> NFS, persistent, shared between login nodes
/scratch/...              -> NFS, persistent/shared
$TMPDIR                    -> local XFS on the current node
/dev/shm                   -> local RAM-backed tmpfs

/dev/shm would also avoid NFS, but it is RAM-backed and highly ephemeral, so I did not choose it for this purpose.


The solution: separate persistent Codex data from SQLite runtime state

Recent Codex versions expose a dedicated sqlite_home setting for the directory that stores the SQLite-backed state database. Codex also supports the CODEX_SQLITE_HOME environment variable; the current Codex source documents sqlite_home as defaulting to $CODEX_SQLITE_HOME when that variable is set, and otherwise to $CODEX_HOME.

The current Codex configuration reference describes sqlite_home as the directory where Codex stores SQLite-backed runtime state used for agent jobs and resumable state:

That makes it possible to avoid moving the entire Codex home directory.

The arrangement I wanted was:

Persistent/shared data:
    ~/.codex/config.toml
    ~/.codex/sessions/
    other durable Codex files
            |
            +--> NFS home directory

SQLite runtime state:
    thread/history/state SQLite databases
            |
            +--> node-local XFS under $TMPDIR

This is preferable to setting all of CODEX_HOME to /tmp, because I still want important session files and configuration to survive logout, cleanup, and node reboot.

One-session test

First create a private local directory:

mkdir -p "$TMPDIR/codex-sqlite"
chmod 700 "$TMPDIR/codex-sqlite"

Then set:

export CODEX_SQLITE_HOME="$TMPDIR/codex-sqlite"

Verify that it is really local:

echo "$CODEX_SQLITE_HOME"
df -T "$CODEX_SQLITE_HOME"

The important part is that the filesystem type should be something like xfs or ext4, not nfs.

Then start Codex normally:

codex

After startup, inspect the local directory:

ls -lah "$CODEX_SQLITE_HOME"

Making it automatic

Because this environment already supplies a node-local $TMPDIR, I prefer an environment-variable setup in ~/.bashrc:

# Keep Codex SQLite runtime state off the NFS home filesystem.
if [ -n "${TMPDIR:-}" ] && [ -d "$TMPDIR" ]; then
    export CODEX_SQLITE_HOME="$TMPDIR/codex-sqlite"
    mkdir -p "$CODEX_SQLITE_HOME"
    chmod 700 "$CODEX_SQLITE_HOME"
fi

Reload the shell:

source ~/.bashrc

and verify:

echo "HOST: $(hostname)"
echo "CODEX_HOME: ${CODEX_HOME:-$HOME/.codex}"
echo "CODEX_SQLITE_HOME: $CODEX_SQLITE_HOME"
df -T "$CODEX_SQLITE_HOME"

Codex also supports the equivalent configuration key in ~/.codex/config.toml:

sqlite_home = "/absolute/local/path/codex-sqlite"

For my HPC environment, however, $TMPDIR is host-specific and already managed by the system, so exporting CODEX_SQLITE_HOME is convenient.


Why this is especially relevant on multi-node HPC systems

The key complication is that $HOME is shared while /tmp is not.

Imagine two login nodes:

login01
    /tmp/<user>/codex-sqlite     local database A

login02
    /tmp/<user>/codex-sqlite     local database B

shared NFS
    /home/<user>/.codex/sessions

This is intentional.

With the default layout, both nodes can potentially see the same SQLite file through NFS:

login01 ----\
             >---- NFS ---- ~/.codex/thread_history_1.sqlite
login02 ----/

With a node-local SQLite directory, each host gets its own database:

login01 ----> local XFS SQLite A
login02 ----> local XFS SQLite B

both -------> shared persistent session files

That avoids asking multiple hosts to coordinate SQLite runtime state through the shared network filesystem.

This does introduce a trade-off: the local SQLite state is no longer inherently shared between nodes. If Codex stores information in SQLite that has not yet been reconstructed from durable session files, switching hosts may initially produce different local state. In practice, this is still a much better failure mode than corrupting the shared database, but it is worth understanding the distinction.


The /tmp trade-off

/tmp is local, fast, and compatible with SQLite’s normal locking assumptions, but it is also temporary.

Administrators may clean it periodically, and its contents may disappear after a reboot.

That means this design should be understood as:

NFS:     durable Codex data
/tmp:    disposable/rebuildable SQLite runtime state

I would prefer a persistent node-local filesystem if the cluster offered one. For example, a site-specific local SSD path would provide both local filesystem semantics and persistence across reboots.

If /tmp is the only local option, preserve the session JSONL files under the shared home directory and treat the SQLite directory as runtime state that may occasionally need to be recreated.

Do not blindly move the entire CODEX_HOME to /tmp unless you are comfortable losing everything stored there when /tmp is purged.


Recovery after the database is already corrupted

Changing CODEX_SQLITE_HOME prevents future Codex processes from using the bad NFS database, but it does not magically repair an existing corrupt file.

My recovery procedure is:

  1. Stop my own Codex processes.
  2. Preserve the existing session JSONLs.
  3. Back up the damaged SQLite database and any -wal/-shm companions.
  4. Move the corrupt database out of the active path rather than immediately deleting it.
  5. Start Codex with a clean local SQLite directory.
  6. Attempt to resume or reconstruct the session from the preserved durable data.

For example:

pgrep -afu "$USER" codex

After confirming that my own Codex processes are stopped:

ts=$(date +%Y%m%d-%H%M%S)
mkdir -p ~/.codex/recovery-$ts

cp -a ~/.codex/sessions/ ~/.codex/recovery-$ts/sessions/
cp -a ~/.codex/thread_history_1.sqlite* ~/.codex/recovery-$ts/ 2>/dev/null || true

Then configure the new local SQLite home:

mkdir -p "$TMPDIR/codex-sqlite"
chmod 700 "$TMPDIR/codex-sqlite"
export CODEX_SQLITE_HOME="$TMPDIR/codex-sqlite"

I intentionally avoid deleting the original rollout/session JSONL simply because the history database is damaged.


Additional precautions

Exit Codex cleanly when possible

Avoid kill -9 unless the process is truly stuck. A clean shutdown gives applications and databases the best chance to flush pending state.

Do not manually modify SQLite sidecar files while Codex is running

Files such as:

*.sqlite
*.sqlite-wal
*.sqlite-shm

can be part of one coordinated database state. Removing one while the process is active is not a safe recovery strategy.

Check database health periodically

With Codex stopped:

for db in "$CODEX_SQLITE_HOME"/*.sqlite; do
    [ -e "$db" ] || continue
    echo "=== $db ==="
    sqlite3 "$db" 'PRAGMA quick_check;'
done

Healthy databases should report:

ok

Check the filesystem, not the directory name

Before using any HPC path for SQLite:

df -T /candidate/path

The useful distinction is not “home versus scratch.” It is network filesystem versus genuinely local filesystem.


A compact diagnostic checklist

If Codex on an HPC system reports an error such as:

file is not a database

for a file under ~/.codex, I would run:

# 1. Identify my Codex processes only.
pgrep -afu "$USER" codex

# 2. Check where Codex state is stored.
df -T "$HOME"
df -T ~/.codex

# 3. Inspect the suspect DB.
file ~/.codex/thread_history_1.sqlite
head -c 32 ~/.codex/thread_history_1.sqlite | xxd

# 4. Check candidate local filesystems.
df -T /tmp /var/tmp /dev/shm

# 5. Check whether the site already assigned a local temp directory.
echo "TMPDIR=${TMPDIR:-<unset>}"

# 6. Move SQLite state to local storage.
mkdir -p "$TMPDIR/codex-sqlite"
chmod 700 "$TMPDIR/codex-sqlite"
export CODEX_SQLITE_HOME="$TMPDIR/codex-sqlite"

# 7. Verify the final filesystem type.
df -T "$CODEX_SQLITE_HOME"

The desired final line should show a local filesystem such as xfs or ext4.


What I learned

The most useful lesson was not specific to Codex.

Many modern CLI tools quietly embed SQLite, LevelDB, RocksDB, lock files, Unix sockets, or other local-state mechanisms under $HOME. On an HPC system, $HOME may be a shared NFS mount visible from many nodes. The application thinks it is writing “local state,” but the infrastructure is providing a distributed filesystem with different failure and locking characteristics.

The troubleshooting pattern is therefore broadly reusable:

Application state error
        |
        v
Identify the actual state files
        |
        v
Check the backing filesystem with df -T
        |
        +--> local filesystem: investigate ordinary corruption/process issues
        |
        +--> network filesystem: look for a supported way to relocate runtime state
                                  to node-local storage

For Codex specifically, the availability of sqlite_home / CODEX_SQLITE_HOME makes that separation possible without abandoning the persistent shared home directory.

My final arrangement is conceptually:

~/.codex/sessions and configuration
        -> shared NFS home
        -> persistent

Codex SQLite runtime state
        -> $TMPDIR/codex-sqlite
        -> node-local XFS
        -> temporary/rebuildable

It is a small configuration change, but on an HPC system it aligns each kind of data with the storage system best suited to it.


References

  1. OpenAI, Codex Configuration Reference — sqlite_home: https://developers.openai.com/codex/config-reference
  2. SQLite, SQLite Over a Network, Caveats and Considerations: https://www.sqlite.org/useovernet.html
  3. SQLite, Write-Ahead Logging: https://www.sqlite.org/wal.html
  4. OpenAI Codex source, config_toml.rs (sqlite_home): https://github.com/openai/codex/blob/main/codex-rs/config/src/config_toml.rs