Applied, thanks!

Samuel

Milos Nikic, le lun. 14 sept. 2026 10:28:21 -0700, a ecrit:
> Hey Samuel,
> 
> Thanks for that. Yes, I managed to reproduce it reliably.
> 
> Here is what caused the deadlock: when the journal gets full, it needs to 
> force
> a checkpoint. It did this by calling journal_sync_everything, which eventually
> calls write_all_disknodes() and iterates the node cache. This creates a lock
> inversion. If the filesystem is under heavy load and iterating the node cache
> while the journal simultaneously tries to force a checkpoint, they deadlock.
> 
> Solution: Decouple the journal's force-checkpoint from the filesystem locks.
> The journal already has all the data it needs in its shadow buffers, so it
> doesn't need to iterate the VFS cache. I implemented an "active checkpoint"
> mechanism where the journal streams directly to disk as needed, completely
> eliminating the deadlock surface. A nice side effect is that the deferred 
> queue
> is no longer needed, so that code got ripped out.
> 
> Furthermore, I fixed a couple of cache coherency bugs related to how pager
> write hazards were handled (the Lifeboat mechanism), which were the culprit 
> for
> the data corruption. I also added a new hook, diskfs_journal_shutdown(), which
> can be called from libdiskfs to safely quiesce the journal to Head 0.
> 
> Now, during testing, I tracked down the Deleted inode X has zero dtime fsck
> warnings on nodes like /dev/log and similar files. This is actually a
> pre-existing Hurd timing issue. Mach delivers the final asynchronous 
> port-death
> notifications after the system is put into read-only mode during shutdown, so
> the dtime update is abandoned in RAM. The journal doesn't cause this; it just
> faithfully records the power-cut state.
> 
> I added a manual pager teardown loop in ext2fs/pager.c (diskfs_shutdown_pager)
> that forces these unlinked nodes to drop synchronously. This gets us a
> pristine, clean fsck when a drive is cleanly unmounted as a translator.
> However, this workaround doesn't (and cannot) apply to the root filesystem
> (which bypasses this and uses startup_dosync), so the root system still
> observes the native dtime drop behavior.
> 
> The architecturally correct way to fix this OS-wide is to utilize the
> pre-existing s_last_orphan Ext3 mechanism that is already present on the ext2
> headers and fsck "speaks it". 
> If we implement the Orphan Inode List, fsck will silently clean up these late
> drops on boot and make the warnings go away completely. That is something i 
> can
> do next, if you agree.
> Please take a look at the patch when you get the time.
> 
> Milos
> 
> On Sun, Sep 13, 2026 at 4:07 PM Samuel Thibault <[1][email protected]>
> wrote:
> 
>     Hello,
> 
>     Milos Nikic, le jeu. 03 sept. 2026 20:35:44 -0700, a ecrit:
>     > I run it with
>     > qemu-system-x86_64 \
>     >                                                  -m 512
>     > I give it only 512mb of ram hoping it would stress test it a bit.
> 
>     I am running with 8G, that can on the contrary put stress by caching a
>     lot of data before writing it all in a row.
> 
>     > -smp 4
> 
>     That's not needed, we don't enable smp by default for now :)
> 
>     > I am doing upgrades daily (two days now).
> 
>     Maybe wait for some time, in order to have larger upgrades to perform.
> 
>     > I am also recompiling hurd and gnumach.
> 
>     That is actually not very i/o intensive since the compilation cpu time
>     is large.
> 
>     > linux-7.2.3.tar.xz              
> 100%[====================================
>     ======
>     > ===========>] 152.64M  7.66MB/s    in 53s
>     >
>     > 2026-09-04 04:17:04 (2.90 MB/s) - 'linux-7.2.3.tar.xz' saved [160060344/
>     > 160060344]
>     >
>     > loshmi@debian:/mnt/stress/test$ tar -xf linux-7.2.3.tar.xz
> 
>     That is more heavy indeed. But large upgrades, such as texlive-full, can
>     take way more room with a myriad of files.
> 
>     Samuel
> 
> 
> References:
> 
> [1] mailto:[email protected]

> From b600ef3218b931eca9a10e9ad2849a9776a617e7 Mon Sep 17 00:00:00 2001
> From: Milos Nikic <[email protected]>
> Date: Tue, 8 Sep 2026 08:56:39 -0700
> Subject: [PATCH] ext2fs: Rework JBD2 checkpointing, fix cache coherency, and
>  eliminate deadlocks
> 
> This patch completely overhauls the ext2fs journal checkpointing architecture
> and fixes several severe race conditions with the Mach VM pager during system
> shutdown and heavy I/O load.
> 
> 1. Lock-Safe Active Checkpointer
> Previously, the journal attempted to reclaim space by recursively calling back
> into the VFS via `journal_sync_everything`. Under heavy load (e.g., 
> compiling),
> this caused lock inversions and thread exhaustion. Checkpointing is now a 
> fully
> self-sufficient process that streams in-memory WAL shadow buffers directly to
> the metal, completely bypassing the VFS node cache.
> 
> 2. Lifeboat Cache Coherency (Page Boundary Corruption)
> Fixed a critical bug where file data intercepted by the Lifeboat mechanism was
> silently overwritten by stale shadow metadata from older checkpoints. Lifeboat
> flushes now utilize the global notification system 
> (`journal_notify_blocks_written_locked`)
> to mark blocks as written across all active transactions, fixing file 
> corruption
> at 4KB page boundaries.
> 
> 3. Shutdown Consistency & Ghost Inodes
> Unlinked files (nlink=0) were previously left with unset `dtime` because their
> weak pager references were not dropped before the final journal commit.
> `diskfs_shutdown_pager` now explicitly drops these references to trigger
> synchronous `diskfs_drop_node` calls.
> 
> 4. `startup_dosync` Integration
> Injected `diskfs_journal_shutdown` into `libdiskfs/init-startup.c` to ensure
> the root filesystem journal is properly quiesced (Head 0) before the 
> microkernel
> forces the disk into read-only mode during `sudo halt`.
> ---
>  ext2fs/ext2fs.h          |   7 -
>  ext2fs/journal.c         | 686 ++++++++++++++++++++++++++++-----------
>  ext2fs/journal.h         |   7 -
>  ext2fs/pager.c           | 104 ++++--
>  libdiskfs/diskfs.h       |   8 +
>  libdiskfs/init-startup.c |   1 +
>  libdiskfs/journal.c      |   9 +-
>  libdiskfs/shutdown.c     |   1 +
>  8 files changed, 580 insertions(+), 243 deletions(-)
> 
> diff --git a/ext2fs/ext2fs.h b/ext2fs/ext2fs.h
> index 6f9d184d2..0975457d1 100644
> --- a/ext2fs/ext2fs.h
> +++ b/ext2fs/ext2fs.h
> @@ -345,13 +345,6 @@ extern struct journal *ext2_journal;
>  error_t
>  journal_dirty_block (diskfs_transaction_t * txn, block_t fs_blocknr);
>  
> -/**
> - * This function exists to sync all AND avoid a deadlock with commit.
> - * It doesn't call journal_commit back yet it syncs everything.
> - **/
> -void
> -journal_sync_everything (void);
> -
>  void journal_notify_block_changed (block_t block);
>  
>  /* ---------------------------------------------------------------- */
> diff --git a/ext2fs/journal.c b/ext2fs/journal.c
> index b25bf9de7..5c4cc711a 100644
> --- a/ext2fs/journal.c
> +++ b/ext2fs/journal.c
> @@ -122,32 +122,6 @@
>  
>  #define JRNL_LIFEBOAT_ALLOC_MASK_LEN 8
>  
> -/* Thread-Local Deferred Block Queue (The Checkpoint Circuit Breaker)
> - *
> - * Problem (The Recursion Deadlock):
> - * When the journal fills up, a VFS thread must force a checkpoint.
> - * Forcing a checkpoint calls write_all_disknodes(), which acquires the
> - * global, non-recursive libdiskfs node-cache mutex. Flushing those inodes
> - * modifies memory, triggering journal_notify_block_changed(), which attempts
> - * to start a transaction. If the journal is still full, it recursively calls
> - * journal_force_checkpoint_locked(), attempts to re-acquire the libdiskfs
> - * mutex, and permanently deadlocks against itself.
> - *
> - * The Lockless Sweep:
> - * We use Thread-Local Storage (__thread) to detect if the CURRENT thread
> - * is actively flushing a checkpoint. If it is, we break the recursion by
> - * intercepting the block notifications and saving them in a private array.
> - * Once the thread finishes the flush and safely drops the libdiskfs locks,
> - * it "sweeps" these deferred blocks into a new transaction. This safely
> - * bypasses the lock inversion while maintaining strict Write-Ahead Log
> - * (WAL) crash consistency.
> - */
> -#define MAX_DEFERRED_BLOCKS 128
> -
> -__thread int thread_is_checkpointing = 0;
> -__thread block_t deferred_blocks[MAX_DEFERRED_BLOCKS];
> -__thread int deferred_count = 0;
> -
>  /* Temporary storage for blocks rushed by the Mach VM pager.
>   * Because we cannot block or delay the pager when it needs to flush a page
>   * belonging to an active (RUNNING/COMMITTING) transaction, this cache
> @@ -183,6 +157,7 @@ typedef struct journal_buffer
>    /* -1 if normal, 0-127 if holding a spoofed payload in the lifeboat */
>    int16_t lifeboat_index;
>    uint8_t jb_is_flushing;    /* 1 if commit thread is actively flushing it. 
> */
> +  uint8_t jb_escaped;
>  } journal_buffer_t;
>  
>  /**
> @@ -210,6 +185,7 @@ typedef enum
>    T_LOCKED,                  /* Locked, no new handles, waiting for updates
>                                  to finish */
>    T_FLUSHING,                        /* Writing to the journal ring buffer */
> +  T_COMMITTED,                       /* WAL Commit Record is on disk. Safe 
> to write directly to main FS. */
>    T_FINISHED                 /* Done, waiting to be checkpointed */
>  } transaction_state_t;
>  
> @@ -282,6 +258,7 @@ typedef struct journal
>  
>    pthread_mutex_t j_state_lock;      /* Protects the pointers below */
>    pthread_cond_t j_commit_wait;      /* Cond. var. while waiting for the tx 
> to be ready. */
> +  pthread_cond_t j_flush_wait;       /* Cond. var for safely waiting on 
> physical flushes */
>    /* The Transactions */
>    diskfs_transaction_t *j_running_transaction;       /* Currently filling */
>    diskfs_transaction_t *j_committing_transaction;    /* Transaction that is
> @@ -587,6 +564,7 @@ journal_alloc_buffer (journal_t *journal)
>        jb->jb_next = NULL;
>        jb->jb_is_written = 0;
>        jb->jb_is_flushing = 0;
> +      jb->jb_escaped = 0;
>        goto out;
>      }
>    jb = calloc (1, sizeof (journal_buffer_t));
> @@ -1086,6 +1064,7 @@ journal_try_advance_tail_locked (journal_t *journal)
>    return advanced;
>  }
>  
> +
>  static void
>  journal_stop_transaction_locked (journal_t *journal,
>                                diskfs_transaction_t *txn)
> @@ -1111,22 +1090,14 @@ journal_stop_transaction_locked (journal_t *journal,
>       {
>         if (jb_exp->needs_copy)
>           {
> -           if (jb_exp->lifeboat_index >= 0)
> -             {
> -               memcpy (jb_exp->jb_shadow_data,
> -                       &(ext2_lifeboat.payloads)[jb_exp->lifeboat_index *
> -                                                 block_size], block_size);
> -               jb_exp->needs_copy = 0;
> -             }
> +           /* ALWAYS hydrate from the live VM cache. The lifeboat is for
> +              delayed physical I/O, not for sourcing WAL shadow data! */
> +           jb_exp->jb_next = NULL;
> +           if (!copy_list_head)
> +             copy_list_head = jb_exp;
>             else
> -             {
> -               jb_exp->jb_next = NULL;
> -               if (!copy_list_head)
> -                 copy_list_head = jb_exp;
> -               else
> -                 copy_list_tail->jb_next = jb_exp;
> -               copy_list_tail = jb_exp;
> -             }
> +             copy_list_tail->jb_next = jb_exp;
> +           copy_list_tail = jb_exp;
>           }
>       }
>  
> @@ -1181,35 +1152,6 @@ journal_stop_transaction_locked (journal_t *journal,
>      }
>  }
>  
> -/**
> - * Drains the thread-local deferred block queue into a new transaction.
> - * When a thread is forced to execute a synchronous checkpoint (which locks 
> the
> - * global libdiskfs node-cache), any memory mutations triggered by the VFS 
> flush
> - * are intercepted and stored in a thread-local queue to prevent a recursive
> - * deadlock against the journal lock.
> - * This function "sweeps" those intercepted blocks by explicitly starting a
> - * new transaction. The act of starting the transaction automatically injects
> - * the deferred blocks into the new transaction's map (via the internal
> - * diskfs_journal_start_transaction_locked logic). We then immediately stop
> - * the transaction to allow the normal journal commit pipeline to process 
> them.
> - *
> - * Must strictly be called OUTSIDE the journal lock.
> - */
> -static void
> -journal_drain_deferred_blocks (void)
> -{
> -  if (deferred_count > 0)
> -    {
> -      diskfs_transaction_t *drain_txn = diskfs_journal_start_transaction ();
> -      if (drain_txn)
> -     {
> -       JOURNAL_LOCK (ext2_journal);
> -       journal_stop_transaction_locked (ext2_journal, drain_txn);
> -       JOURNAL_UNLOCK (ext2_journal);
> -     }
> -    }
> -}
> -
>  /**
>   * Checks a range of written blocks against a single transaction's map.
>   * Marks any matching buffers as written and decrements the outstanding I/O
> @@ -1220,17 +1162,31 @@ static int
>  journal_notify_txn_locked (diskfs_transaction_t *txn,
>                          block_t start_block, size_t n_blocks)
>  {
> -  for (size_t i = 0; i < n_blocks && txn->t_outstanding_io > 0; i++)
> +  /* Process all blocks to ensure jb_is_flushing is safely cleared 
> everywhere */
> +  for (size_t i = 0; i < n_blocks; i++)
>      {
>        block_t b = start_block + i;
>        journal_buffer_t *jb = journal_map_lookup (&txn->t_buffer_map, b);
> -      if (jb && !jb->jb_is_written)
> +
> +      if (jb)
>       {
> -       jb->jb_is_written = 1;
> -       txn->t_outstanding_io--;
> -       JRNL_LOG_DEBUG
> -         ("[NOTIFY] Block %u written for TID %u (outstanding: %d)", b,
> -          txn->t_tid, txn->t_outstanding_io);
> +       if (!jb->jb_is_written)
> +         {
> +           jb->jb_is_written = 1;
> +           if (txn->t_outstanding_io > 0)
> +             txn->t_outstanding_io--;
> +
> +           JRNL_LOG_DEBUG
> +             ("[NOTIFY] Block %u written for TID %u (outstanding: %d)", b,
> +              txn->t_tid, txn->t_outstanding_io);
> +         }
> +
> +       /* Pager finished the I/O. Unblock Checkpoint thread! */
> +       if (jb->jb_is_flushing)
> +         {
> +           jb->jb_is_flushing = 0;
> +           pthread_cond_broadcast (&ext2_journal->j_flush_wait);
> +         }
>       }
>      }
>    return txn->t_outstanding_io == 0;
> @@ -1244,16 +1200,32 @@ journal_notify_blocks_written_locked (block_t 
> start_block, size_t n_blocks)
>  {
>    int sb_changed = 0;
>    error_t err = 0;
> -  if (!ext2_journal || n_blocks == 0)
> +  if (!ext2_journal || n_blocks == 0 || ext2_journal->j_must_exit)
>      return 0;
>  
>    JRNL_LOG_DEBUG ("Got notification for %zu blocks starting at %u",
>                 n_blocks, start_block);
>  
> -  /* Check Running Transaction */
> -  diskfs_transaction_t *run = ext2_journal->j_running_transaction;
> -  if (run)
> -    journal_notify_txn_locked (run, start_block, n_blocks);
> +  /* Do NOT notify the running transaction here.
> +   *
> +   * The running transaction has not yet crossed the WAL barrier — its shadow
> +   * buffers may still be updated by journal_stop_transaction_locked as VFS
> +   * threads continue to modify blocks.  If we mark a block jb_is_written=1 
> in
> +   * the running transaction, a later checkpoint will skip writing the 
> (correct)
> +   * shadow data to the main disk, permanently leaving stale metadata behind.
> +   *
> +   * This can happen due to a race: the pager checks hazards (block not in 
> any
> +   * transaction), unlocks for physical I/O, then a VFS thread adds the 
> block to
> +   * the running transaction.  When the pager re-locks and calls us, the 
> block
> +   * IS in the running transaction, but the data we just wrote to disk may be
> +   * stale relative to the final shadow copy.
> +   *
> +   * Blocks in the running transaction that legitimately need to be marked as
> +   * written are handled by journal_flush_lifeboat_payloads(), which runs 
> AFTER
> +   * the WAL barrier is crossed (T_COMMITTED) and does its own marking.
> +   *
> +   * The committing transaction and checkpoint list are safe to notify: their
> +   * shadow data is frozen and the WAL has been (or is being) committed. */
>  
>    /* Check Committing Transaction */
>    diskfs_transaction_t *commit = ext2_journal->j_committing_transaction;
> @@ -1359,6 +1331,8 @@ journal_create (struct node *journal_inode)
>    pthread_mutex_init (&j->j_state_lock, NULL);
>    pthread_cond_init (&j->j_commit_wait, NULL);
>    pthread_cond_init (&j->j_flusher_wakeup, NULL);
> +  pthread_cond_init (&j->j_flush_wait, NULL);
> +
>    j->j_must_exit = 0;
>    if (pthread_create (&kjournald_tid, NULL, kjournald_thread, j) != 0)
>      JRNL_LOG_WARN ("Failed to create a flusher thread.");
> @@ -1427,76 +1401,193 @@ journal_clear_checkpoint_list_locked (journal_t 
> *journal)
>  }
>  
>  /**
> - * Safely marks the journal as clean on disk.
> - * MUST only be called after sync_global(1) ensures no pager I/O is in 
> flight,
> - * otherwise asynchronous pager notifications will cause a Use-After-Free!
> + * Checks if a block exists in any transaction newer than 'txn'.
> + * This prevents older checkpoints from overwriting fresh data on the disk.
>   */
> -void
> -journal_quiesce_checkpoints (void)
> +static int
> +journal_is_block_in_newer_transaction_locked (journal_t *journal,
> +                                           diskfs_transaction_t *txn,
> +                                           block_t b)
> +{
> +  /* Check all transactions in the checkpoint list newer than 'txn'.
> +     These are fully committed to the WAL, so relying on them is safe. */
> +  diskfs_transaction_t *t = txn->t_checkpoint_next;
> +  while (t)
> +    {
> +      if (journal_map_lookup (&t->t_buffer_map, b))
> +     return 1;
> +      t = t->t_checkpoint_next;
> +    }
> +
> +  /* Check the committing transaction ONLY if it has safely crossed the WAL
> +     barrier. If it is still T_FLUSHING or T_LOCKED, a crash would lose it,
> +     so we cannot rely on it to skip physical I/O! */
> +  t = journal->j_committing_transaction;
> +  if (t && t->t_state == T_COMMITTED
> +      && journal_map_lookup (&t->t_buffer_map, b))
> +    return 1;
> +
> +  /* NEVER check the running transaction. It is not on disk yet.
> +     Relying on it would permanently delete the older safely committed WAL 
> backup
> +     before the new one is written, causing unrecoverable data loss on 
> crash! */
> +
> +  return 0;
> +}
> +
> +/**
> + * Internal helper to flush checkpoint transactions to the main filesystem.
> + * If target_free is UINT32_MAX, it flushes ALL transactions in the list.
> + * Otherwise, it flushes until j_free >= target_free.
> + * Returns 1 if a hardware I/O error occurred, 0 on success.
> + * MUST be called with JOURNAL_LOCK held.
> + */
> +static int
> +journal_flush_checkpoints_locked (journal_t *journal, uint32_t target_free)
>  {
> -  if (!ext2_journal)
> -    return;
> +  int sb_changed = 0;
> +  int global_io_error = 0;
>  
> -  JOURNAL_LOCK (ext2_journal);
> +  while (journal->j_free < target_free && journal->j_checkpoint_list)
> +    {
> +      diskfs_transaction_t *txn = journal->j_checkpoint_list;
> +      size_t iter = 0;
> +      journal_buffer_t *jb;
> +      int io_error = 0;
>  
> -  /* Set a 10-second deadline for the active commit to finish. */
> -  struct timespec ts;
> -  clock_gettime (CLOCK_MONOTONIC, &ts);
> -  ts.tv_sec += 10;
> +      /* Stream the fully committed WAL shadow blocks directly to the metal 
> */
> +      while ((jb = journal_map_iterate (&txn->t_buffer_map, &iter)) != NULL)
> +     {
> +     retry_block:
> +       if (!jb->jb_is_written)
> +         {
> +           if (jb->jb_is_flushing)
> +             {
> +               pthread_cond_wait (&journal->j_flush_wait,
> +                                  &journal->j_state_lock);
> +               goto retry_block;
> +             }
>  
> -  int err = 0;
> +           jb->jb_is_flushing = 1;
> +           block_t b = jb->jb_blocknr;
> +           char *data = jb->jb_shadow_data;
> +           size_t amount = 0;
> +           error_t err = 0;
>  
> -  /* Wait for any active commit to finish writing to the log */
> -  while (ext2_journal->j_committing_transaction != NULL && err == 0)
> -    err = pthread_cond_clockwait (&ext2_journal->j_commit_done,
> -                               &ext2_journal->j_state_lock, CLOCK_MONOTONIC, 
> &ts);
> -  if (err)
> +           if (journal_is_block_in_newer_transaction_locked
> +               (journal, txn, b))
> +             {
> +               /* A newer transaction already captured this block.
> +                  Skip physical I/O to protect the fresh data on disk. */
> +               amount = block_size;
> +             }
> +           else
> +             {
> +               /* Unlock to perform physical I/O without stalling the 
> journal */
> +               JOURNAL_UNLOCK (journal);
> +               store_offset_t dev_block =
> +                 (store_offset_t) b << log2_dev_blocks_per_fs_block;
> +               err =
> +                 store_write (store, dev_block, data, block_size, &amount);
> +               JOURNAL_LOCK (journal);
> +             }
> +
> +           /* ONLY mark as written if the hardware actually accepted the 
> full block! */
> +           if (!err && amount == block_size)
> +             {
> +               if (!jb->jb_is_written)
> +                 {
> +                   jb->jb_is_written = 1;
> +                   if (txn->t_outstanding_io > 0)
> +                     txn->t_outstanding_io--;
> +                 }
> +             }
> +           else
> +             {
> +               JRNL_LOG_WARN
> +                 ("Checkpoint I/O failed for block %u! err=%d, wrote=%zu",
> +                  b, err, amount);
> +               io_error = 1;
> +             }
> +
> +           jb->jb_is_flushing = 0;
> +           pthread_cond_broadcast (&journal->j_flush_wait);
> +
> +           if (io_error)
> +             break;
> +         }
> +     }
> +
> +      /* All blocks are physically on disk. Reclaim the space! */
> +      if (!io_error)
> +     {
> +       if (journal_try_advance_tail_locked (journal))
> +         {
> +           sb_changed = 1;
> +         }
> +       else
> +         {
> +           JRNL_LOG_WARN
> +             ("Logic bug: Failed to advance tail after active checkpoint!");
> +           break;
> +         }
> +     }
> +      else
> +     {
> +       global_io_error = 1;
> +       break;                /* Stop checkpointing on device failure */
> +     }
> +    }
> +
> +  if (sb_changed)
>      {
> -      /* If we hit ETIMEDOUT, a VFS thread likely leaked a t_updates refcount
> -         due to a signal interruption or crash. We MUST bail out without
> -         clearing the checkpoint list so the WAL replays on next boot! */
> -      JRNL_LOG_WARN
> -     ("Quiesce timed out! Transaction deadlocked. Leaving journal dirty.");
> -      JOURNAL_UNLOCK (ext2_journal);
> -      return;
> +      JOURNAL_UNLOCK (journal);
> +      flush_to_disk ();              /* Ensure all checkpoint data is on the 
> platter first */
> +      JOURNAL_LOCK (journal);
> +
> +      uint32_t tail_seq;
> +      diskfs_transaction_t *oldest =
> +     journal_get_oldest_transaction_locked (journal);
> +
> +      if (oldest)
> +     tail_seq = oldest->t_tid;
> +      else
> +     tail_seq = journal->j_transaction_sequence;
> +
> +      error_t err = journal_update_superblock (journal, tail_seq);
> +      if (err)
> +     JRNL_LOG_WARN ("Failed to update superblock during checkpoint. %s",
> +                    strerror (err));
> +      else
> +     {
> +       JOURNAL_UNLOCK (journal);
> +       flush_to_disk ();     /* Ensure the SB update itself hits the platter 
> */
> +       JOURNAL_LOCK (journal);
> +     }
>      }
>  
> -  /* Clear the list and write s_start = 0 to the JBD2 superblock */
> -  journal_clear_checkpoint_list_locked (ext2_journal);
> -  JOURNAL_UNLOCK (ext2_journal);
> +  return global_io_error;
>  }
>  
>  /**
> - * Called when we are running out of space.
> - * Since we do a version of sync() on every commit, we can safely declare all
> - * previous transactions "checkpointed" and reset the log.
> - * Must be called with a journal lock held, and that state will remain such
> - * after returning.
> + * Actively flushes the oldest checkpointed transactions to the main 
> filesystem.
> + * This guarantees space is freed instantly without relying on the lazy Mach 
> pager,
> + * and executes entirely without acquiring VFS node locks.
>   */
>  static void
>  journal_force_checkpoint_locked (journal_t *journal)
>  {
> -  JRNL_LOG_DEBUG ("[CHECKPOINT] Journal Full (Free: %u). Squeezing disk...",
> -               journal->j_free);
> -  JOURNAL_UNLOCK (journal);
> -
> -  /* Arm the circuit breaker and reset the queue */
> -  thread_is_checkpointing = 1;
> -  deferred_count = 0;
> +  JRNL_LOG_DEBUG
> +    ("[CHECKPOINT] Journal Full (Free: %u). Actively checkpointing...",
> +     journal->j_free);
>  
> -  journal_sync_everything ();
> +  /* Add a 1/8th safety runway to prevent thrashing */
> +  uint32_t runway = (journal->j_last - journal->j_first) / 8;
> +  uint32_t target_free = journal->j_min_free + runway;
>  
> -  /* Disarm the circuit breaker */
> -  thread_is_checkpointing = 0;
> +  journal_flush_checkpoints_locked (journal, target_free);
>  
> -  JOURNAL_LOCK (journal);
> -  journal_clear_checkpoint_list_locked (journal);
> -  JOURNAL_UNLOCK (journal);
> -  flush_to_disk ();
> -  JOURNAL_LOCK (journal);
> -
> -  JRNL_LOG_DEBUG ("[CHECKPOINT] Space reclaimed. Free: %u. Tail: %u",
> -               journal->j_free, journal->j_tail);
> +  JRNL_LOG_DEBUG ("[CHECKPOINT] Done checkpointing (Free: %u).",
> +               journal->j_free);
>  }
>  
>  /**
> @@ -1552,7 +1643,12 @@ journal_dirty_block_locked (diskfs_transaction_t *txn, 
> block_t fs_blocknr)
>    journal_buffer_t *new_jb;
>    error_t err = 0;
>  
> -  assert_backtrace (txn);
> +  if (!txn)
> +    {
> +      JRNL_LOG_DEBUG ("[TRX] Transaction null but block dirty.");
> +      goto out;
> +    }
> +
>    assert_backtrace (txn->t_state == T_RUNNING || txn->t_state == T_LOCKED);
>    jb = journal_map_lookup (&txn->t_buffer_map, fs_blocknr);
>  
> @@ -1600,6 +1696,9 @@ out:
>  static diskfs_transaction_t *
>  diskfs_journal_start_transaction_locked (journal_t *journal)
>  {
> +  if (journal->j_must_exit)
> +    return NULL;
> +
>    diskfs_transaction_t *txn;
>    if (ext2_journal->j_free < ext2_journal->j_min_free)
>      {
> @@ -1640,17 +1739,6 @@ diskfs_journal_start_transaction_locked (journal_t 
> *journal)
>        journal->j_running_transaction = txn;
>        JRNL_LOG_DEBUG ("[TRX] Created NEW TID %u", txn->t_tid);
>      }
> -  /* THE SWEEP: Safely inject deferred blocks into our brand new transaction 
> */
> -  if (deferred_count > 0)
> -    {
> -      /* Copy to local var and reset count immediately to prevent any
> -         impossible recursion loops during dirty_block */
> -      int count = deferred_count;
> -      deferred_count = 0;
> -
> -      for (int i = 0; i < count; i++)
> -     journal_dirty_block_locked (txn, deferred_blocks[i]);
> -    }
>  
>    return txn;
>  }
> @@ -1733,6 +1821,24 @@ journal_write_batch (journal_t *journal, const 
> diskfs_transaction_t *txn,
>    return 0;
>  }
>  
> +static void
> +restore_escaped_magic (const diskfs_transaction_t *txn,
> +                    size_t batch_start_iter, uint32_t batch_count)
> +{
> +  size_t iter = batch_start_iter;
> +  uint32_t magic_const = htobe32 (JBD2_MAGIC_NUMBER);
> +
> +  for (uint32_t i = 0; i < batch_count; i++)
> +    {
> +      journal_buffer_t *jb = journal_map_iterate (&txn->t_buffer_map, &iter);
> +      if (jb->jb_escaped)
> +     {
> +       memcpy (jb->jb_shadow_data, &magic_const, sizeof (magic_const));
> +       jb->jb_escaped = 0;
> +     }
> +    }
> +}
> +
>  /* Writes the Descriptor Block + All Data Blocks (Escaped) */
>  static error_t
>  journal_write_payload (journal_t *journal, const diskfs_transaction_t *txn)
> @@ -1775,12 +1881,14 @@ journal_write_payload (journal_t *journal, const 
> diskfs_transaction_t *txn)
>         if (err)
>           return err;
>  
> +       restore_escaped_magic (txn, batch_start_iter, batch_count);
> +
>         /* Prepare for the next batch */
>         descriptor_loc = journal_next_log_block_safe (journal);
>         memset (descriptor_buf, 0, block_size);
>         setup_header (descriptor_buf, txn, JBD2_DESCRIPTOR_BLOCK);
>         tag_offset = sizeof (journal_header_t);
> -       batch_start_iter = iter - 1;
> +       batch_start_iter = iter - 1;  /* Point to the item that caused the 
> flush */
>         batch_count = 0;
>       }
>  
> @@ -1802,6 +1910,11 @@ journal_write_payload (journal_t *journal, const 
> diskfs_transaction_t *txn)
>       {
>         flags |= JBD2_FLAG_ESCAPE;
>         memset (jb->jb_shadow_data, 0, sizeof (data_head));
> +       jb->jb_escaped = 1;
> +     }
> +      else
> +     {
> +       jb->jb_escaped = 0;
>       }
>  
>        tag->t_flags = htobe32 (flags);
> @@ -1817,6 +1930,7 @@ journal_write_payload (journal_t *journal, const 
> diskfs_transaction_t *txn)
>        err =
>       journal_write_batch (journal, txn, descriptor_buf, descriptor_loc,
>                            batch_start_iter, batch_count);
> +      restore_escaped_magic (txn, batch_start_iter, batch_count);
>      }
>  
>    return err;
> @@ -1890,12 +2004,14 @@ journal_flush_lifeboat_payloads (journal_t *journal,
>         /* We always free the raw slot we just finished using */
>         lifeboat_free_slot (lb_idx);
>  
> -       /* Mark it as written so checkpointing can advance! */
> -       if (!err && !jb_lb->jb_is_written)
> +       /* Mark it as written so checkpointing can advance!
> +          Use the global notification system so ALL transactions
> +          that contain this block are marked as written, preventing
> +          the Active Checkpointer from overwriting fresh data with
> +          stale shadow metadata from older checkpoint transactions. */
> +       if (!err)
>           {
> -           jb_lb->jb_is_written = 1;
> -           if (txn->t_outstanding_io > 0)
> -             txn->t_outstanding_io--;
> +           journal_notify_blocks_written_locked (jb_lb->jb_blocknr, 1);
>           }
>         JOURNAL_UNLOCK (journal);
>       }
> @@ -1979,6 +2095,8 @@ journal_commit_running_transaction_locked (journal_t 
> *journal)
>    /* Ensure Commit is persistent */
>    flush_to_disk ();
>  
> +  /* The WAL barrier is crossed! Tell the Pager it can write safely! */
> +  txn->t_state = T_COMMITTED;
>    /* Flush any intercepted VM pager blocks to the primary disk */
>    journal_flush_lifeboat_payloads (journal, txn);
>  
> @@ -2016,11 +2134,9 @@ journal_commit_running_transaction_locked (journal_t 
> *journal)
>      flush_to_disk ();
>  
>    journal_forget_freed_blocks (journal, freed_extents);
> -  journal_drain_deferred_blocks ();
>    JOURNAL_LOCK (journal);
>    goto out;
>  abort_commit:
> -  journal_drain_deferred_blocks ();
>    /* We hit a physical I/O error. We must clear the pipeline slot and wake
>       up any sleeping threads so they don't deadlock, before we free the txn. 
> */
>    JOURNAL_LOCK (journal);
> @@ -2032,6 +2148,58 @@ out:
>    return err;
>  }
>  
> +/**
> + * Safely marks the journal as clean on disk.
> + * MUST only be called after sync_global(1) ensures no pager I/O is in 
> flight,
> + * otherwise asynchronous pager notifications will cause a Use-After-Free!
> + */
> +static void
> +journal_quiesce_checkpoints (void)
> +{
> +  JRNL_LOG_DEBUG ("Journal in quiesce checkpoints.");
> +  JOURNAL_LOCK (ext2_journal);
> +
> +  journal_commit_running_transaction_locked (ext2_journal);
> +  /* Set a 10-second deadline for the active commit to finish. */
> +  struct timespec ts;
> +  clock_gettime (CLOCK_MONOTONIC, &ts);
> +  ts.tv_sec += 10;
> +
> +  int err = 0;
> +
> +  /* Wait for any active commit to finish writing to the log */
> +  while (ext2_journal->j_committing_transaction != NULL && err == 0)
> +    err = pthread_cond_clockwait (&ext2_journal->j_commit_done,
> +                               &ext2_journal->j_state_lock, CLOCK_MONOTONIC, 
> &ts);
> +  if (err)
> +    {
> +      /* If we hit ETIMEDOUT, a VFS thread likely leaked a t_updates refcount
> +         due to a signal interruption or crash. We MUST bail out without
> +         clearing the checkpoint list so the WAL replays on next boot! */
> +      JRNL_LOG_WARN
> +     ("Quiesce timed out! Transaction deadlocked. Leaving journal dirty.");
> +      JOURNAL_UNLOCK (ext2_journal);
> +      return;
> +    }
> +
> +  /* Write ALL shadow data from every checkpoint transaction to the main 
> disk. */
> +  /* Pass UINT32_MAX to guarantee we drain the entire checkpoint list */
> +  int io_error = journal_flush_checkpoints_locked (ext2_journal, UINT32_MAX);
> +
> +  if (io_error)
> +    {
> +      JRNL_LOG_WARN ("[QUIESCE] I/O error during final checkpoint. "
> +                  "Leaving journal dirty for WAL replay.");
> +      JOURNAL_UNLOCK (ext2_journal);
> +      return;
> +    }
> +
> +  /* Now safe to clear the list and write s_start = 0 to the JBD2 superblock 
> */
> +  journal_clear_checkpoint_list_locked (ext2_journal);
> +
> +  JOURNAL_UNLOCK (ext2_journal);
> +}
> +
>  static void
>  diskfs_journal_stop_transaction_locked (journal_t *journal,
>                                       diskfs_transaction_t *txn)
> @@ -2079,7 +2247,6 @@ diskfs_journal_stop_transaction (diskfs_transaction_t 
> *txn)
>    JOURNAL_LOCK (ext2_journal);
>    diskfs_journal_stop_transaction_locked (ext2_journal, txn);
>    JOURNAL_UNLOCK (ext2_journal);
> -  journal_drain_deferred_blocks ();
>  }
>  
>  /* Forces the currently running transaction (if any) to safely commit to the
> @@ -2178,7 +2345,6 @@ diskfs_journal_commit_transaction (diskfs_transaction_t 
> *opaque_txn)
>    journal_wait_on_tid_locked (ext2_journal, tid);
>  out:
>    JOURNAL_UNLOCK (ext2_journal);
> -  journal_drain_deferred_blocks ();
>  }
>  
>  /**
> @@ -2191,17 +2357,27 @@ static int
>  journal_handle_write_hazard_locked (block_t b, char *b_data)
>  {
>    int intercepted = 0;
> -
> -  diskfs_transaction_t *commit = ext2_journal->j_committing_transaction;
> -  diskfs_transaction_t *run = ext2_journal->j_running_transaction;
> -
> -  journal_buffer_t *jb_run =
> -    run ? journal_map_lookup (&run->t_buffer_map, b) : NULL;
> -  journal_buffer_t *jb_commit =
> -    commit ? journal_map_lookup (&commit->t_buffer_map, b) : NULL;
> -
> -  /* Deadlock Hazard Check */
> -  if ((jb_run && (run->t_updates > 0 || commit != NULL)) || jb_commit)
> +  diskfs_transaction_t *commit;
> +  diskfs_transaction_t *run;
> +  journal_buffer_t *jb_run;
> +  journal_buffer_t *jb_commit;
> +  diskfs_transaction_t *chk;
> +
> +retry:
> +  commit = ext2_journal->j_committing_transaction;
> +  run = ext2_journal->j_running_transaction;
> +
> +  jb_run = run ? journal_map_lookup (&run->t_buffer_map, b) : NULL;
> +  jb_commit = commit ? journal_map_lookup (&commit->t_buffer_map, b) : NULL;
> +
> +  /* If it's in the committing transaction BUT the WAL barrier is crossed,
> +     it is safe to write to disk. Remove it from the intercept hazard list. 
> */
> +  if (commit && commit->t_state == T_COMMITTED && jb_commit)
> +    jb_commit = NULL;
> +
> +  /* Hazard Interception: If it's trapped in an active transaction,
> +     send it straight to the lifeboat and RETURN EARLY. Do not touch 
> checkpoints. */
> +  if (jb_run || jb_commit)
>      {
>        int lb_idx_run = jb_run ? lifeboat_alloc_slot () : -1;
>        int lb_idx_commit = jb_commit ? lifeboat_alloc_slot () : -1;
> @@ -2224,6 +2400,8 @@ journal_handle_write_hazard_locked (block_t b, char 
> *b_data)
>           {
>             memcpy (&(ext2_lifeboat.payloads)[lb_idx_run * block_size],
>                     b_data, block_size);
> +           /* If the old slot is NOT being flushed, we must free it to avoid 
> a leak.
> +              If it IS being flushed, the commit thread owns it and will 
> free it. */
>             if (jb_run->lifeboat_index >= 0)
>               lifeboat_free_slot (jb_run->lifeboat_index);
>             jb_run->lifeboat_index = (int16_t) lb_idx_run;
> @@ -2232,8 +2410,6 @@ journal_handle_write_hazard_locked (block_t b, char 
> *b_data)
>           {
>             memcpy (&(ext2_lifeboat.payloads)[lb_idx_commit * block_size],
>                     b_data, block_size);
> -           /* If the old slot is NOT being flushed, we must free it to avoid 
> a leak.
> -              If it IS being flushed, the commit thread owns it and will 
> free it. */
>             if (jb_commit->lifeboat_index >= 0
>                 && !jb_commit->jb_is_flushing)
>               lifeboat_free_slot (jb_commit->lifeboat_index);
> @@ -2245,32 +2421,62 @@ journal_handle_write_hazard_locked (block_t b, char 
> *b_data)
>           ("Intercepted rushed pager write for block %u into Lifeboat slots 
> (run:%d, commit:%d)",
>            b, lb_idx_run, lb_idx_commit);
>       }
> +      /* Do not claim any flushing flags! */
> +      return intercepted;
>      }
> -  else if (jb_run)
> +
> +  /* We are definitively going to physical disk.
> +     Now safely check for hardware races against Checkpoint/Committing 
> threads. */
> +  if (commit && commit->t_state == T_COMMITTED)
>      {
> -      /* No hazard, but block is in the running transaction.
> -         Force a synchronous commit to satisfy the WAL barrier. */
> -      JRNL_LOG_DEBUG ("Pager forcing synchronous commit for TID %u",
> -                   run->t_tid);
> -      error_t err = journal_commit_running_transaction_locked (ext2_journal);
> -      if (err)
> -     JRNL_LOG_WARN ("Synchronous commit failed for TID %u: %s",
> -                    run->t_tid, strerror (err));
> +      journal_buffer_t *jb_com_real =
> +     journal_map_lookup (&commit->t_buffer_map, b);
> +      if (jb_com_real)
> +     {
> +       if (jb_com_real->jb_is_flushing)
> +         {
> +           pthread_cond_wait (&ext2_journal->j_flush_wait,
> +                              &ext2_journal->j_state_lock);
> +           goto retry;
> +         }
> +       else if (!jb_com_real->jb_is_written)
> +         jb_com_real->jb_is_flushing = 1;    /* Claim it for the Pager */
> +     }
>      }
>  
> -  return intercepted;
> +  chk = ext2_journal->j_checkpoint_list;
> +  while (chk)
> +    {
> +      journal_buffer_t *jb_chk = journal_map_lookup (&chk->t_buffer_map, b);
> +      if (jb_chk)
> +     {
> +       if (jb_chk->jb_is_flushing)
> +         {
> +           pthread_cond_wait (&ext2_journal->j_flush_wait,
> +                              &ext2_journal->j_state_lock);
> +           goto retry;
> +         }
> +       else if (!jb_chk->jb_is_written)
> +         {
> +           /* Claim the buffer so the Checkpoint thread yields if it wakes 
> up. */
> +           jb_chk->jb_is_flushing = 1;
> +         }
> +     }
> +      chk = chk->t_checkpoint_next;
> +    }
> +
> +  return 0;
>  }
>  
>  /**
> - * Checks if a block is part of an active transaction.
> - * Used by the coalescing loop to stop before a hazard block.
> + * Checks if a block is safe to coalesce into a physical batch write.
> + * If it is safe, it preemptively claims the block in the checkpoint lists
> + * so that background checkpoint threads yield to the Pager.
>   * MUST be called with JOURNAL_LOCK held.
>   */
>  static int
> -journal_has_active_transaction_locked (block_t b)
> +journal_claim_safe_block_locked (block_t b)
>  {
> -  int active = 0;
> -
>    diskfs_transaction_t *commit = ext2_journal->j_committing_transaction;
>    diskfs_transaction_t *run = ext2_journal->j_running_transaction;
>  
> @@ -2279,10 +2485,77 @@ journal_has_active_transaction_locked (block_t b)
>    journal_buffer_t *jb_commit =
>      commit ? journal_map_lookup (&commit->t_buffer_map, b) : NULL;
>  
> +  /* If commit crossed WAL barrier, it's safe to coalesce and write */
> +  if (commit && commit->t_state == T_COMMITTED)
> +    jb_commit = NULL;
> +
> +  /* If it is an active hazard, we cannot claim it for physical coalescing */
>    if (jb_run || jb_commit)
> -    active = 1;
> +    return 0;
> +
> +  /* Stop coalescing if ANY thread is actively flushing this block to disk */
> +  journal_buffer_t *jb_commit_flush =
> +    commit ? journal_map_lookup (&commit->t_buffer_map, b) : NULL;
> +  if (jb_commit_flush && jb_commit_flush->jb_is_flushing)
> +    return 0;
>  
> -  return active;
> +  diskfs_transaction_t *chk = ext2_journal->j_checkpoint_list;
> +  while (chk)
> +    {
> +      journal_buffer_t *jb_chk = journal_map_lookup (&chk->t_buffer_map, b);
> +      if (jb_chk && jb_chk->jb_is_flushing)
> +     return 0;
> +      chk = chk->t_checkpoint_next;
> +    }
> +
> +  /* It is safe to write to disk. CLAIM IT in the lists! */
> +  if (commit && commit->t_state == T_COMMITTED)
> +    {
> +      if (jb_commit_flush && !jb_commit_flush->jb_is_written)
> +     jb_commit_flush->jb_is_flushing = 1;
> +    }
> +
> +  chk = ext2_journal->j_checkpoint_list;
> +  while (chk)
> +    {
> +      journal_buffer_t *jb_chk = journal_map_lookup (&chk->t_buffer_map, b);
> +      if (jb_chk && !jb_chk->jb_is_written)
> +     jb_chk->jb_is_flushing = 1;
> +      chk = chk->t_checkpoint_next;
> +    }
> +
> +  return 1;
> +}
> +
> +static void
> +journal_clear_flushing_locked (block_t start_block, size_t n_blocks)
> +{
> +  for (size_t i = 0; i < n_blocks; i++)
> +    {
> +      block_t b = start_block + i;
> +      diskfs_transaction_t *txn = ext2_journal->j_checkpoint_list;
> +      while (txn)
> +     {
> +       journal_buffer_t *jb = journal_map_lookup (&txn->t_buffer_map, b);
> +       if (jb && jb->jb_is_flushing)
> +         {
> +           jb->jb_is_flushing = 0;
> +           pthread_cond_broadcast (&ext2_journal->j_flush_wait);
> +         }
> +       txn = txn->t_checkpoint_next;
> +     }
> +
> +      txn = ext2_journal->j_committing_transaction;
> +      if (txn)
> +     {
> +       journal_buffer_t *jb = journal_map_lookup (&txn->t_buffer_map, b);
> +       if (jb && jb->jb_is_flushing)
> +         {
> +           jb->jb_is_flushing = 0;
> +           pthread_cond_broadcast (&ext2_journal->j_flush_wait);
> +         }
> +     }
> +    }
>  }
>  
>  /**
> @@ -2339,8 +2612,8 @@ journal_store_write (block_t start_block, size_t 
> length, void *buf,
>           {
>             block_t next_b = start_block + i + flush_count;
>  
> -           if (journal_has_active_transaction_locked (next_b))
> -             break;          /* Stop coalescing; this next block might need 
> hazard handling */
> +           if (!journal_claim_safe_block_locked (next_b))
> +             break;
>  
>             flush_count++;
>           }
> @@ -2363,6 +2636,10 @@ journal_store_write (block_t start_block, size_t 
> length, void *buf,
>           flush_needed =
>             journal_notify_blocks_written_locked (b, actual_blocks);
>  
> +       if (actual_blocks < flush_count)
> +         journal_clear_flushing_locked (b + actual_blocks,
> +                                        flush_count - actual_blocks);
> +
>         total_written += chunk_amount;
>         i += actual_blocks;
>  
> @@ -2468,16 +2745,6 @@ journal_notify_block_changed (block_t block)
>    if (!ext2_journal)
>      return;
>  
> -  if (thread_is_checkpointing)
> -    {
> -      /* We are in a recursive trap! Defer this block for later. */
> -      if (deferred_count < MAX_DEFERRED_BLOCKS)
> -     deferred_blocks[deferred_count++] = block;
> -      else
> -     JRNL_LOG_WARN ("Deferred block queue full! Dropping block %u", block);
> -      return;
> -    }
> -
>    JOURNAL_LOCK (ext2_journal);
>    diskfs_transaction_t *txn =
>      diskfs_journal_start_transaction_locked (ext2_journal);
> @@ -2487,3 +2754,28 @@ journal_notify_block_changed (block_t block)
>    diskfs_journal_stop_transaction_locked (ext2_journal, txn);
>    JOURNAL_UNLOCK (ext2_journal);
>  }
> +
> +void
> +diskfs_journal_shutdown (void)
> +{
> +  if (!ext2_journal)
> +    return;
> +
> +  ext2_journal->j_must_exit = 1;     /* Signal kjournald so it doesn't wake 
> up */
> +  /* Commit any pending transaction (e.g. the superblock clean-state
> +     flags written by diskfs_set_hypermetadata).  */
> +  journal_commit_running_transaction ();
> +
> +  /* Sync the disk pager to ensure all shadow data is on disk.  */
> +  sync_global (1);
> +
> +  /* Checkpoint all remaining transactions and mark the journal clean.  */
> +  journal_quiesce_checkpoints ();
> +
> +  ext2_journal = NULL;
> +
> +  /* Final hardware flush.  */
> +  error_t err = store_sync (store);
> +  if (err && err != EOPNOTSUPP && err != D_INVALID_OPERATION)
> +    ext2_warning ("device flush failed: %s", strerror (err));
> +}
> diff --git a/ext2fs/journal.h b/ext2fs/journal.h
> index 0f98c262b..a96739f9f 100644
> --- a/ext2fs/journal.h
> +++ b/ext2fs/journal.h
> @@ -80,13 +80,6 @@ error_t
>  journal_store_read (block_t start_block, size_t length, void **buf,
>                   size_t *read_amount);
>  
> -/**
> - * Safely marks the journal as clean on disk.
> - * MUST only be called after sync_global(1) ensures no pager I/O is in 
> flight,
> - * otherwise asynchronous pager notifications will cause a Use-After-Free!
> - */
> -void journal_quiesce_checkpoints (void);
> -
>  /**
>   * Records a range of deleted blocks so they can be unpinned from older
>   * checkpoint lists AFTER this transaction safely commits.
> diff --git a/ext2fs/pager.c b/ext2fs/pager.c
> index 70c555bb8..73c7d0b3c 100644
> --- a/ext2fs/pager.c
> +++ b/ext2fs/pager.c
> @@ -964,7 +964,7 @@ pager_report_extent (struct user_pager_info *pager,
>  void
>  pager_clear_user_data (struct user_pager_info *upi)
>  {
> -  if (upi->type == FILE_DATA)
> +  if (upi->type == FILE_DATA && upi->node)
>      {
>        struct pager *pager;
>  
> @@ -1561,50 +1561,92 @@ diskfs_get_filemap_pager_struct (struct node *node)
>  void
>  diskfs_shutdown_pager (void)
>  {
> -  error_t shutdown_one (void *v_p)
> +  /* TODO: Implement the Ext3/Ext4 Orphan Inode List (s_last_orphan).
> +            Currently, if a file is unlinked (nlink=0) but still held open 
> by a
> +            Mach pager, it will be abandoned on disk without dtime=0 if the
> +            system halts, causing fsck to complain. This manual teardown 
> forces
> +            the nodes to drop synchronously before the final journal commit.
> +            Once the Orphan List is implemented, this entire manual pager 
> cleanup
> +            can be safely removed. Unlinked files will be added to the 
> superblock's
> +            orphan list, and the OS can just pull the power. The next boot 
> will
> +            silently clean them up. */
> +  error_t shutdown_and_clear (void *v_p)
>      {
>        struct pager *p = v_p;
> +      struct user_pager_info *upi = pager_get_upi (p);
> +
> +      /* First, shutdown the pager: sync and flush all dirty pages,
> +         then destroy the port right.  This must happen before we
> +         release the node reference, because pager_sync/pager_flush
> +         may need to access the node's allocsize and alloc_lock.  */
>        pager_shutdown (p);
> +
> +      /* After pager_shutdown, the pager has been removed from the
> +         bucket's hash table (via ports_destroy_right).  But we can
> +         still access it because ports_bucket_iterate holds a hard
> +         reference on our behalf.
> +
> +         Now release the pager's weak node reference, mimicking what
> +         pager_dropweak + pager_clear_user_data would do.  This
> +         ensures diskfs_drop_node runs synchronously for any unlinked
> +         nodes before we commit the final journal transaction.
> +
> +         Without this, unlinked nodes would be left in a half-deleted
> +         state: nlink=0 on disk but dtime unset, bitmap not cleared,
> +         and free-counts not updated — all in an uncommitted journal
> +         transaction lost on exit(0).  */
> +      if (upi->type == FILE_DATA && upi->node)
> +        {
> +          int cleared = 0;
> +
> +          /* Clear the node->pager back-pointer (as pager_dropweak does)
> +             so the assert in pager_clear_user_data is satisfied.  */
> +          pthread_spin_lock (&node_to_page_lock);
> +          if (diskfs_node_disknode (upi->node)->pager
> +              && pager_get_upi (diskfs_node_disknode (upi->node)->pager) == 
> upi)
> +            {
> +              diskfs_node_disknode (upi->node)->pager = NULL;
> +              cleared = 1;
> +            }
> +          pthread_spin_unlock (&node_to_page_lock);
> +
> +          if (cleared)
> +            ports_port_deref_weak (p);
> +
> +          /* Release the weak node reference acquired in diskfs_get_filemap.
> +             If this is the last reference, diskfs_drop_node is called
> +             synchronously, which sets dtime, clears the inode bitmap,
> +             and updates free-counts.  */
> +          diskfs_nrele_light (upi->node);
> +
> +          /* Prevent pager_clear_user_data (which fires when the iterator
> +             drops its hard ref) from double-releasing the node.  */
> +          upi->node = NULL;
> +        }
> +
>        return 0;
>      }
>  
> +  ports_bucket_iterate (file_pager_bucket, shutdown_and_clear);
> +
> +  /* pager_shutdown + diskfs_nrele_light above may have triggered
> +     diskfs_drop_node for unlinked nodes, which writes dtime, clears
> +     the inode bitmap, updates free-counts, and starts a new journal
> +     transaction.  We MUST commit this transaction before quiescing. */
>    write_all_disknodes ();
>    journal_commit_running_transaction ();
>  
> -  ports_bucket_iterate (file_pager_bucket, shutdown_one);
> +  if (!ext2_journal)
> +    {
> +      error_t err = store_sync (store);
> +      if (err && err != EOPNOTSUPP && err != D_INVALID_OPERATION)
> +        ext2_warning ("device flush failed: %s", strerror (err));
> +    }
>  
> -  /* Sync everything on the the disk pager.  */
> -  sync_global (1);
> -  journal_quiesce_checkpoints ();
> -  store_sync (store);
>    /* Despite the name of this function, we never actually shutdown the disk
>       pager, just make sure it's synced. */
>  }
>  
> -static error_t
> -journal_sync_one (void *v_p)
> -{
> -  struct pager *p = v_p;
> -  pager_sync (p, 1);
> -  return 0;
> -}
> -
> -/**
> - * Sync all the pagers synchronously, but don't call
> - * journal_commit here. It would deadlock.
> - **/
> -void
> -journal_sync_everything (void)
> -{
> -  write_all_disknodes ();
> -  ports_bucket_iterate (file_pager_bucket, journal_sync_one);
> -  sync_global (1);
> -  error_t err = store_sync (store);
> -  /* Ignore EOPNOTSUPP (drivers), but warn on real I/O errors */
> -  if (err && err != EOPNOTSUPP && err != D_INVALID_OPERATION)
> -    ext2_warning ("device flush failed: %s", strerror (err));
> -}
> -
>  /* Sync all the pagers. */
>  void
>  diskfs_sync_everything (int wait)
> diff --git a/libdiskfs/diskfs.h b/libdiskfs/diskfs.h
> index 490e1ae07..d8dac1293 100644
> --- a/libdiskfs/diskfs.h
> +++ b/libdiskfs/diskfs.h
> @@ -576,6 +576,14 @@ void diskfs_journal_set_sync (diskfs_transaction_t *txn);
>     synchronous I/O guarantee. */
>  int diskfs_journal_needs_sync (diskfs_transaction_t *txn);
>  
> +/* The user may define this function.  It is called once at the very end
> +   of diskfs_shutdown, after all pagers have been shut down and
> +   hypermetadata has been written, to perform any final journal-specific
> +   cleanup (e.g. committing the last transaction, checkpointing all
> +   remaining transactions, and marking the journal clean on disk).
> +   The default definition does nothing.  */
> +void diskfs_journal_shutdown (void);
> +
>  /* The user must define this function.  Sync the info in NP->dn_stat
>     and any associated format-specific information to disk.  If WAIT is true,
>     then return only after the physicial media has been completely updated. */
> diff --git a/libdiskfs/init-startup.c b/libdiskfs/init-startup.c
> index 997429b31..d51cc2787 100644
> --- a/libdiskfs/init-startup.c
> +++ b/libdiskfs/init-startup.c
> @@ -180,6 +180,7 @@ diskfs_S_startup_dosync (mach_port_t handle)
>       {
>         diskfs_sync_everything (1);
>         diskfs_set_hypermetadata (1, 1);
> +       diskfs_journal_shutdown ();
>         _diskfs_diskdirty = 0;
>  
>         /* XXX: if some application writes something after that, we will
> diff --git a/libdiskfs/journal.c b/libdiskfs/journal.c
> index edc1a9370..eb93b8f69 100644
> --- a/libdiskfs/journal.c
> +++ b/libdiskfs/journal.c
> @@ -1,10 +1,11 @@
>  /* Default version of Journal in libdiskfs.
> -   It implements default implementations of 5 functions:
> +   It implements default implementations of 6 functions:
>       - diskfs_journal_start_transaction
>       - diskfs_journal_stop_transaction
>       - diskfs_journal_commit_transaction
>       - diskfs_journal_needs_sync
>       - diskfs_journal_set_sync
> +     - diskfs_journal_shutdown
>  
>     diskfs_journal_start_transaction returns NULL,
>     diskfs_journal_needs_sync returns 0.
> @@ -66,3 +67,9 @@ diskfs_journal_needs_sync (diskfs_transaction_t *txn)
>    /* Do nothing */
>    return 0;
>  }
> +
> +void __attribute__((weak))
> +diskfs_journal_shutdown (void)
> +{
> +  /* Do nothing */
> +}
> diff --git a/libdiskfs/shutdown.c b/libdiskfs/shutdown.c
> index 3f774c31b..fb677ffa8 100644
> --- a/libdiskfs/shutdown.c
> +++ b/libdiskfs/shutdown.c
> @@ -91,6 +91,7 @@ diskfs_shutdown (int flags)
>      {
>        diskfs_shutdown_pager ();
>        diskfs_set_hypermetadata (1, 1);
> +      diskfs_journal_shutdown ();
>      }
>  
>    return 0;
> -- 
> 2.55.0
> 


-- 
Samuel
 Créer une hiérarchie supplementaire pour remedier à un problème (?) de
 dispersion est d'une logique digne des Shadocks.
 * BT in: Guide du Cabaliste Usenet - La Cabale vote oui (les Shadocks aussi) *

Reply via email to