]> git.hungrycats.org Git - linux/commit
btrfs: stripe_alloc: settle a logged extent's stripes at log commit
authorZygo Blaxell <ce3g8jdj@umail.furryterror.org>
Thu, 30 Jul 2026 00:51:24 +0000 (20:51 -0400)
committerZygo Blaxell <ce3g8jdj@umail.furryterror.org>
Fri, 18 Sep 2026 21:36:21 +0000 (17:36 -0400)
commitb5eb67896a8e2778c70889587ea1ecd26380032e
tree830884bd0875d7ed28ea8ed362d6d8b03135f005
parent195289176548062cbf8c5db868751f9823cebf46
btrfs: stripe_alloc: settle a logged extent's stripes at log commit

Stripe-exclusive allocation bounds degraded-crash damage to the current
transaction, but within that window two log commits can still share a
stripe: fsync 1 writes the head of a stripe, the open run keeps filling
it, and a later write's read-modify-write rewrites the parity that
protects fsync 1's data.  A degraded crash during the second write tears
the first -- the write hole's shape, confined to the log window, and the
reason the fsync guarantee has so far been "loss is detectable" rather
than "completed fsyncs survive".

Close it with a case analysis.  A logged extent in a full stripe is
already safe: nothing ever writes a full stripe again, since COW never
overwrites and there are no free sectors left to allocate.  A logged
extent in a partial stripe is exposed only to future writes into that
stripe's remaining sectors -- so at log commit, close the stripe's open
run (nothing further allocates into it), kick any parked partial writes
for it, and wait for its in-flight data IO before the log super is
written.  This is the log-window analogue of invariant I2, which the
commit-time retirement provides for full commits.  Both log paths are
hooked: the fast path per extent map in log_one_extent(), and the
full-sync path where copy_items() walks new data extents (old-transaction
extents are skipped there, and are exactly the ones already settled by
their own commit).

The drain is bounded and join-free: run inflight is counted from
allocation, which happens during writeback with submission following in
the same pass, so the wait is bio flight time plus the parked-write
deadline that the flush short-circuits; write_done reporting needs no
transaction join, so waiting under the inode log mutex and a running
transaction handle is safe.

The cost is spatial: every fsync that logs an extent in a partial stripe
retires that stripe early, trapping its unwritten tail like any other
partially filled stripe until it frees or balance reclaims it.
Fsync-heavy workloads therefore burn a stripe tail per touched stripe
per fsync; the follow-up per-inode log runs and copy-forward relocation
exist to reclaim exactly that cost, and are optimizations on top of the
guarantee this patch completes.

Assisted-by: Claude:claude-fable-5
(cherry picked from commit 53c44fdbca58bd1fa267a5648e236ed4238d99ed)
fs/btrfs/block-group.c
fs/btrfs/block-group.h
fs/btrfs/tree-log.c