btrfs: stripe_alloc: allow-rmw policy, and write-in-place for isolated nocow extents
Re-enable write-in-place for nodatacow files and preallocated extents
under stripe-exclusive allocation, with the understood caveat that the
write hole cannot be prevented for data that opts out of COW: an
in-place write RMWs its stripe's parity, so a degraded crash can tear
the stripe. What CAN be guaranteed is the blast radius: in-place is
permitted only for extents whose full stripes are isolated to the
writing inode, so such a crash can tear only the writing file's own
data -- the nodatacow contract, no worse.
The gate is per extent, decided where nocow eligibility is already
checked: fast path, the stripe belongs to one of the inode's own
NOCOW-class runs (which only ever held its extents); slow path, the
committed extent tree proves sole ownership (single plain data ref,
count 1, matching root and objectid; anything shared, foreign or
metadata fails). The fully-free claim rule keeps uncommitted foreign
extents out of partially used stripes, so the committed tree is
authoritative. Extents that fail -- anything allocated before stripe
isolation existed, extents shared through reflink or snapshots,
relocated extents -- simply stay force-COWed, and because their rewrite
is steered into the inode's private NOCOW run, the next overwrite of
the same data passes: legacy nocow files migrate themselves to
isolation in one COW generation, with no tool and no flag day.
Log-commit settling skips NOCOW-class runs: nodatacow data gets no
fsync survival guarantee (its own later in-place writes can always
tear it), so closing the run would trap its tail for nothing. The
write-hole debug checker skips groups that hosted NOCOW-class runs,
like relocation-used groups, since isolated in-place writes land in
stripes whose runs have drained.
The interface is a word-list policy naming the cases in which
stripe_alloc may permit the legacy unsafe RMW, each independently:
nodatacow in-place writes for nodatacow files' extents
prealloc in-place writes into preallocated extents
fsync waive the close-at-log-commit guarantee: no settling, no
per-inode LOG steering, no carry-forward; logged stripes
may be extended and RMWed as before those patches
It is "allow_rmw", not "allow_overwrite": the fsync case overwrites
nothing, but all three permit read-modify-write of stripes that a
degraded crash can then tear. The nodatacow and prealloc cases require
the per-extent stripe isolation above (the blast radius stays confined
to the writing file); fsync restores the 3a-era exposure where a
degraded crash may cost just-fsynced data, detectably, in exchange for
none of the log-commit costs.
The policy is persistent as the btrfs.stripe_alloc_allow_rmw property
on the top-level root directory, following stripe_alloc's precedent
(applied when the root inode loads during mount, before any user IO),
with a mount option of the same name as a non-persistent override; the
effective policy is their union. Words are separated by comma, space
or colon -- the mount option form must use colon, since mount splits
options at commas. Everything defaults off: plain stripe_alloc keeps
forcing COW and keeps the full fsync guarantee.