]> git.hungrycats.org Git - bees/log
bees
21 months agobytevector: don't deadlock on operator<<
Zygo Blaxell [Wed, 4 Dec 2024 04:34:09 +0000 (23:34 -0500)]
bytevector: don't deadlock on operator<<

operator<< was a friend class that locked the ByteVector, then invoked
hexdump on the bytevector, which used ByteVector::operator[]...which
locked the ByteVector, resulting in a deadlock.

operator<< shouldn't be a friend class anyway.  Make hexdump use the
normal public access methods for ByteVector.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoroots: avoid copying a BtrfsIoctlSearchKey
Zygo Blaxell [Tue, 3 Dec 2024 21:51:24 +0000 (16:51 -0500)]
roots: avoid copying a BtrfsIoctlSearchKey

Although all the members of BtrfsExtentDataFetcher are theoretically
copiable, there's no need to actually make any such copy.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agocontext: spell "progress" correctly
Zygo Blaxell [Mon, 2 Dec 2024 14:49:58 +0000 (09:49 -0500)]
context: spell "progress" correctly

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agocontext: add a PROGRESS: header in $BEESSTATUS
Zygo Blaxell [Sun, 1 Dec 2024 16:32:36 +0000 (11:32 -0500)]
context: add a PROGRESS: header in $BEESSTATUS

Make it clearer where the progress information goes.

Also add placeholder text so the progress section isn't empty at startup,
when the progress hasn't been calculated yet.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: post-5.7 toxic extent handling
Zygo Blaxell [Wed, 27 Nov 2024 01:29:54 +0000 (20:29 -0500)]
docs: post-5.7 toxic extent handling

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agobees: post-kernel-5.7 toxic extent handling
Zygo Blaxell [Fri, 28 May 2021 06:13:33 +0000 (02:13 -0400)]
bees: post-kernel-5.7 toxic extent handling

Toxic extents are mostly gone in kernel 5.7 and later.  Increase the
timeout for toxic extent handling to reduce false positives, and remove
persistenly stored toxic hashes from the hash table.

Toxic hashes are still stored nonpersistently to help mitigate problems
due to any remaining kernel bugs.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoextent scan: don't serialize dedupe and LOGICAL_INO when using extent scan mode
Zygo Blaxell [Sat, 23 Nov 2024 04:26:37 +0000 (23:26 -0500)]
extent scan: don't serialize dedupe and LOGICAL_INO when using extent scan mode

The serialization doesn't seem to be necessary for the extent scan mode.
No infinite loops in the kernel have been observed in the past two years,
despite never having used MultiLock for the extent scanner.

Leave the serialization for now on the subvol scanners.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: default scan mode is 4, "extent"
Zygo Blaxell [Wed, 27 Nov 2024 02:59:20 +0000 (21:59 -0500)]
docs: default scan mode is 4, "extent"

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agomain: set default scan mode to mode 4 (EXTENT)
Zygo Blaxell [Mon, 16 Jan 2023 04:11:09 +0000 (23:11 -0500)]
main: set default scan mode to mode 4 (EXTENT)

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: old missing features are not missing any more
Zygo Blaxell [Sat, 25 Feb 2023 08:13:23 +0000 (03:13 -0500)]
docs: old missing features are not missing any more

The extent scan mode has been implemented (partially, but close enough
to win benchmarks).

New features include several nuisance dedupe countermeasures.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: add scan mode 4, "extent"
Zygo Blaxell [Thu, 5 Jan 2023 06:06:06 +0000 (01:06 -0500)]
docs: add scan mode 4, "extent"

Extent is a different kind of scan mode, so introduce the concept of
the two kinds of scan mode, and rearrange the description of scan modes
along the new boundaries.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoprogress: squeeze the progress table into 80 columns or less
Zygo Blaxell [Fri, 29 Nov 2024 06:51:33 +0000 (01:51 -0500)]
progress: squeeze the progress table into 80 columns or less

We don't need the subvol numbers since they're only interesting to
developers.

We don't need both max and min sizes, pick one and drop the other.

Replace "16E" with "max"--it is the same number of characters, but
doesn't require the user to know what 1<<64 is off the top of their head.

Shorten "remain" to "todo" because sometimes those extra two columns
matter.

Drop the seconds field in ETA timestamps.  Long scan arrival times are
years away, and short scan arrival times are only updated once every
5 minutes, so the extra precision isn't useful.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoprogress: put the progress table in the stats and status files
Zygo Blaxell [Wed, 27 Nov 2024 20:58:06 +0000 (15:58 -0500)]
progress: put the progress table in the stats and status files

Make the progress information more accessible, without having to
enable full debug log and fish it out of the stream with grep.

Also increase the progress log level to INFO.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoextent scan: fix crawl_map creation
Zygo Blaxell [Sat, 30 Nov 2024 17:45:58 +0000 (12:45 -0500)]
extent scan: fix crawl_map creation

There are two crawl_maps in extent scan's next_transid:  one gets
initialized, the other gets used.  This works OK as long as bees is
resuming an existing scan, because the two maps are identical; however,
but it fails if bees is starting without an existing set of crawl data,
and one of the two maps is empty or partially filled.

The failure is intermittent, as the crawl map is being populated at
the same time next_transid runs.  It will eventually be completed after
several transaction cycles, at which point bees runs normally.
It does add significant delays during startup for benchmarks.

There's only one crawl_map in extent scan, it always has the same
crawlers, and extent scan's `next_transid` creates it by itself.
Ignore the map from BeesRoots/BeesCrawl.

Also throw in some missing but helpful trace statements.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoprogress: estimate actual data sizes for progress report
Zygo Blaxell [Tue, 26 Nov 2024 22:47:33 +0000 (17:47 -0500)]
progress: estimate actual data sizes for progress report

Replace pointers in the "done" and "total" columns with estimated data
sizes for each size tier.  The estimation is based on statistics
collected from extents scanned during the current bees run.

Move the total size for the entire filesystem up to the heading.

Report the _completed_ position (i.e. the one that would be saved in
`beescrawl.dat`), not the _queued_ position (i.e. the one where the
next Task would be created in memory).

At the end of the data, the crawl pointer ends up at some random point
in the filesystem just after the newest extent, so the progress gets to
99.7% and then goes to some random value like 47% or 3%, not to 100%.
Report "deferred" in the "done" column when the crawler is waiting for
the next transid, and "finished" in the "%done" column when the crawler
has reached the end of the data.  Suppress the ETA when finished.  This
makes it clear that there's no further work to do for these crawlers.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: add event counters for extent scan
Zygo Blaxell [Thu, 6 Jul 2023 17:13:42 +0000 (13:13 -0400)]
docs: add event counters for extent scan

Add a section for all the new extent scan event counters.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoextent scan: refactor BeesScanMode so derived classes decide their own scan scheduling
Zygo Blaxell [Sat, 30 Nov 2024 04:57:15 +0000 (23:57 -0500)]
extent scan: refactor BeesScanMode so derived classes decide their own scan scheduling

BeesScanModeExtent uses six scan Tasks instead of one, which leads
to awkwardness like the do_scan method to tell crawl_roots how to do
what it shouldn't need to know how to do anyway.

Move the crawl_roots logic into the ::scan methods themselves.

This also deletes the very popular "crawl_more ran out of data" message.
Extent scan explicitly indicates when a scan is complete, so there's
no longer a need to fish this message out of the log.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoextent scan: put all the refs in a single Task, sort them, use idle task
Zygo Blaxell [Wed, 27 Nov 2024 20:46:29 +0000 (15:46 -0500)]
extent scan: put all the refs in a single Task, sort them, use idle task

The sorting avoids problematic read orders, like extent refs in the same
inode with descending offsets, that btrfs is not optimized for.

Putting everything in one Task keeps the queue sizes small, and
manages the lock contention much more calmly.

We only want to be mapping extent refs if there's not enough extents
already in the queue to keep worker threads busy, so use the `idle()`
method instead of `run()`.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoextent scan: introduce SCAN_MODE_EXTENT
Zygo Blaxell [Thu, 22 Dec 2022 03:42:39 +0000 (22:42 -0500)]
extent scan: introduce SCAN_MODE_EXTENT

The EXTENT scan mode reads the extent tree, splits it into tiers by
extent size, converts each tiers's extents into subvol/inode/offset refs,
then runs the legacy bees dedupe engine on the refs.

The extent scan mode can cheaply compute completion percentage and ETA,
so do that every time a new transid is observed.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agotask: add an idle queue
Zygo Blaxell [Thu, 28 Nov 2024 05:01:37 +0000 (00:01 -0500)]
task: add an idle queue

Add a second level queue which is only serviced when the local and global
queues are empty.

At some point there might be a need to implement a full priority queue,
but for now two classes are sufficient.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agofs: add some performance metrics for TREE_SEARCH_V2 calls
Zygo Blaxell [Sun, 24 Oct 2021 19:17:43 +0000 (15:17 -0400)]
fs: add some performance metrics for TREE_SEARCH_V2 calls

These give some visibility into how efficiently bees is using the
TREE_SEARCH_V2 ioctl.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agotable: add a simple text table renderer
Zygo Blaxell [Sun, 29 Jan 2023 02:26:51 +0000 (21:26 -0500)]
table: add a simple text table renderer

This should help clean up some of the uglier status outputs.

Supports:

 * multi-line table cells
 * character fills
 * sparse tables
 * insert, delete by row and column
 * vertical separators

and not much else.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: remove "matched_" prefix event counters
Zygo Blaxell [Thu, 28 Nov 2024 01:39:09 +0000 (20:39 -0500)]
docs: remove "matched_" prefix event counters

We can no longer reliably determine the number of hash table matches,
since we'll stop counting after the first one.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoscan_one_extent: remove the unreadahead after benchmark results
Zygo Blaxell [Sat, 30 Nov 2024 00:22:53 +0000 (19:22 -0500)]
scan_one_extent: remove the unreadahead after benchmark results

That unreadahead used to result in a 10% hit on benchmarks.  Now it's
closer to 75%.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoBeesRangePair: drop the _really_ expensive toxic extent workaround
Zygo Blaxell [Fri, 29 Nov 2024 06:42:28 +0000 (01:42 -0500)]
BeesRangePair: drop the _really_ expensive toxic extent workaround

We were doing a `LOGICAL_INO` ioctl on every _block_ of a matching extent,
just to see how long it takes.  It takes a while!

This could be modified to do an ioctl with the `IGNORE_OFFSET` flag,
once per new extent, but the kernel bug was fixed a long time ago, so
we can start removing all the toxic extent code.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoscan_one_extent: in skip/scan lines, log whether extent is compressed
Zygo Blaxell [Thu, 28 Nov 2024 19:02:57 +0000 (14:02 -0500)]
scan_one_extent: in skip/scan lines, log whether extent is compressed

Useful for debugging the compressed-zero-block cases.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoscan_one_extent: reduce the number of LOGICAL_INO calls before finding a duplicate...
Zygo Blaxell [Thu, 28 Nov 2024 01:29:09 +0000 (20:29 -0500)]
scan_one_extent: reduce the number of LOGICAL_INO calls before finding a duplicate block range

When we have multiple possible matches for a block, we proceed in three
phases:

1.  retrieve each match's extent refs and put them in a list,
2.  iterate over the list converting viable block matches into range matches,
3.  sort and flatten the list of range matches into a non-overlapping
list of ranges that cover all duplicate blocks exactly once.

The separation of phase 1 and 2 creates a performance issue when there
are many block matches in phase 1, and all the range matches in phase
2 are the same length.  Even though we might quickly find the longest
possible matching range early in phase 2, we first extract all of the
extent refs from every possible matching block in phase 1, even though
most of those refs will never be used.

Fix this by moving the extent ref retrieval in phase 1 into a single
loop in phase 2, and stop looping over matching blocks as soon as any
dedupe range is created.  This avoids iterating over a large list of
blocks with expensive `LOGICAL_INO` ioctls in an attempt to improve the
match when there is no hope of improvement, e.g. when all match ranges
are 4K and the content is extremely prevalent in the data.

If we find a matched block that is part of a short matching range,
we can replace it with a block that is part of a long matching range,
because there is a good chance we will find a matching hash block in
the long range by looking up hashes after the end of the short range.
In that case, overlapping dedupe ranges covering both blocks in the
target extent will be inserted into the dedupe list, and the longest
matches will be selected at phase 3.  This usually provides a similar
result to that of the loop in phase 1, but _much_ more efficiently.

Some operations are left in phase 1, but they are all using internal
functions, not ioctls.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: event counter updates after fixing counter names and scan_one_extent improvements
Zygo Blaxell [Wed, 27 Nov 2024 03:27:07 +0000 (22:27 -0500)]
docs: event counter updates after fixing counter names and scan_one_extent improvements

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoscan_one_extent: eliminate nuisance dedupes, drop caches after reading data
Zygo Blaxell [Sat, 23 Nov 2024 16:14:37 +0000 (11:14 -0500)]
scan_one_extent: eliminate nuisance dedupes, drop caches after reading data

A laundry list of problems fixed:

 * Track which physical blocks have been read recently without making
 any changes, and don't read them again.

 * Separate dedupe, split, and hole-punching operations into distinct
 planning and execution phases.

 * Keep the longest dedupe from overlapping dedupe matches, and flatten
 them into non-overlapping operations.

 * Don't scan extents that have blocks already in the hash table.
 We can't (yet) touch such an extent without making unreachable space.
 Let them go.

 * Give better information in the scan summary visualization:  show dedupe
 range start and end points (<ddd>), matching blocks (=), copy blocks
 (+), zero blocks (0), inserted blocks (.), unresolved match blocks
 (M), should-have-been-inserted-but-for-some-reason-wasn't blocks (i),
 and there's-a-bug-we-didn't-do-this-one blocks (#).

 * Drop cached data from extents that have been inserted into the hash
 table without modification.

 * Rewrite the hole punching for uncompressed extents, which apparently
 hasn't worked properly since the beginning.

Nuisance dedupe elimination:

 * Don't do more than 100 dedupe, copy, or hole-punch operations per
 extent ref.

 * Don't split an extent or punch a hole unless dedupe would save at
 least half of the extent ref's size.

 * Write a "skip:" summary showing the planned work when nuisance
 dedupe elimination decides to skip an extent.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agotypes: add shrink_begin and shrink_end methods for BeesFileRange and BeesRangePair
Zygo Blaxell [Tue, 19 Nov 2024 04:57:29 +0000 (23:57 -0500)]
types: add shrink_begin and shrink_end methods for BeesFileRange and BeesRangePair

These allow trimming of overlapping dedupes.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agocounters: fix counter names for scan_eof, scan_no_fd, scanf_deferred_inode
Zygo Blaxell [Tue, 19 Nov 2024 22:39:45 +0000 (17:39 -0500)]
counters: fix counter names for scan_eof, scan_no_fd, scanf_deferred_inode

This code gets moved around from time to time and ends up with the
wrong prefix.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agomultilock: allow turning it off
Zygo Blaxell [Thu, 21 Nov 2024 21:26:05 +0000 (16:26 -0500)]
multilock: allow turning it off

Add a master switch to turn off the entire MultiLock infrastructure for
testing, without having to remove and add all the individual entry points.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agofs: handle ENOENT within lib
Zygo Blaxell [Fri, 29 Nov 2024 06:07:55 +0000 (01:07 -0500)]
fs: handle ENOENT within lib

This prevents the storms of exceptions that occur when a subvol is
deleted.  We simply treat the entire tree as if it was empty.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agobtrfs-tree: accessors for TreeFetcher classes' type and tree values
Zygo Blaxell [Wed, 20 Nov 2024 16:30:28 +0000 (11:30 -0500)]
btrfs-tree: accessors for TreeFetcher classes' type and tree values

Sometimes we have a generic TreeFetcher and we need to know which tree
it came from.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agobtrfs-tree: add root refs and extent flags fields
Zygo Blaxell [Wed, 20 Nov 2024 15:15:53 +0000 (10:15 -0500)]
btrfs-tree: add root refs and extent flags fields

Lazily filling in accessor methods for btrfs objects as needed by bees.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agotask: fix try_lock argument description
Zygo Blaxell [Tue, 19 Nov 2024 20:29:41 +0000 (15:29 -0500)]
task: fix try_lock argument description

try_lock allows specification of a different Task to be run instead of
the current Task when the lock is busy.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agobees: increase file cache size limits
Zygo Blaxell [Thu, 15 Jul 2021 01:37:40 +0000 (21:37 -0400)]
bees: increase file cache size limits

With some extents having 9999 refs, we can use much larger caches for
file descriptors.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: resolve_overflow limit is only 655050 when BTRFS_MAX_EXTENT_REF_COUNT is
Zygo Blaxell [Thu, 28 Nov 2024 19:37:06 +0000 (14:37 -0500)]
docs: resolve_overflow limit is only 655050 when BTRFS_MAX_EXTENT_REF_COUNT is

Use the current header value in the doc.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agobees: reduce extent ref limit to 9999
Zygo Blaxell [Mon, 14 Jun 2021 15:50:18 +0000 (11:50 -0400)]
bees: reduce extent ref limit to 9999

Originally the limit was 2730 (64KiB worth of ref pointers).  This limit
was a little too low for some common workloads, so it was then raised by
a factor of 256 to 699050, but there are a lot of problems with extent
counts that large.  Most of those problems are memory usage and speed
problems, but some of them trigger subtle kernel MM issues.

699050 references is too many to be practical.  Set the limit to 9999,
only 3-4x larger than the original 2730, to give up on deduplication
when each deduped ref reduces the amount of space by no more than 0.01%.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: event counter updates after readahead sanity improvements
Zygo Blaxell [Wed, 27 Nov 2024 03:23:55 +0000 (22:23 -0500)]
docs: event counter updates after readahead sanity improvements

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoreadahead: inject more sanity at the foundation of an insane architecture
Zygo Blaxell [Mon, 18 Nov 2024 19:22:26 +0000 (14:22 -0500)]
readahead: inject more sanity at the foundation of an insane architecture

This solves a third bad problem with bees reads:

3.  The architecture above the read operations will issue read requests
for the same physical blocks over and over in a short period of time.

Fixing that properly requires rewriting the upper-level code, but a
simple small table of recent read requests can reduce the effect of the
problem by orders of magnitude.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoreadahead: inject some sanity at the foundation of an insane architecture
Zygo Blaxell [Fri, 15 Nov 2024 19:35:01 +0000 (14:35 -0500)]
readahead: inject some sanity at the foundation of an insane architecture

This solves some of the worst problems with bees reads:

1.  The kernel readahead doesn't work.  More precisely, it's much better
adapted for a very different use case:  a single thread alternating
between reading a file sequentially and processing the data that was read.
bees has multiple threads which compete for access to IO and then issue
reads in random order immediately after the call to readahead.  The kernel
uses idle ioprio scheduling for the readaheads, so the readaheads get
preempted by the random reads, or cancels the readaheads because the
data access pattern isn't sequential after the readahead was issued.

2.  Seeking drives perform terribly with multiple competing readers,
especially with btrfs striped profiles where the iops are broken into
tiny stripe-sized pieces.  At one point I intended to read the btrfs
device map and figure out which devices can be read in parallel, but to
make that useful, the user needs to have an array with multiple drives
in single profile, or 4+ drives in raid1 profile.  In all other cases,
the elaborate calculations always return the same result:  there can be
only one reader at a time.

This commit fixes both problems:

1.  Don't use the kernel readahead.  Use normal reads into a dummy
buffer instead.

2.  Allow only one thread to readahead at any time.  Once the read is
completed, the data is in the page cache, and all the random-order small
reads that bees does will hit the page cache, not a spinning disk.
In some cases we need to read two things close together, so add a
`bees_readahead_pair` which holds one lock across both reads.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agohash: use kernel readahead instead of bees_readahead to prefetch hash table
Zygo Blaxell [Mon, 25 Nov 2024 03:11:42 +0000 (22:11 -0500)]
hash: use kernel readahead instead of bees_readahead to prefetch hash table

The hash table is read sequentially and from a single thread, so
the kernel's implementation of readahead is appropriate here.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: add allocator regression in 6.0+ kernels
Zygo Blaxell [Sat, 7 Oct 2023 05:35:07 +0000 (01:35 -0400)]
docs: add allocator regression in 6.0+ kernels

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agocontext: when a task fails to acquire an extent lock, don't go ahead and scan the...
Zygo Blaxell [Tue, 19 Nov 2024 20:17:41 +0000 (15:17 -0500)]
context: when a task fails to acquire an extent lock, don't go ahead and scan the extent anyway

Commit c3b664fea54cfd8ac25411cbdb9536e4f24b008e ("context: don't forget
to retry locked extents") removed the critical return that prevents a
Task from processing an extent that is locked.

Put the return back.

Fixes: c3b664fea54cfd8ac25411cbdb9536e4f24b008e ("context: don't forget to retry locked extents")
Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agofs: get rid of 16 MiB limit on dedupe requests
Zygo Blaxell [Tue, 12 Nov 2024 01:47:36 +0000 (20:47 -0500)]
fs: get rid of 16 MiB limit on dedupe requests

The kernel has not required a 16 MiB limit on dedupe requests since
v4.18-rc1 b67287682688 ("Btrfs: dedupe_file_range ioctl: remove 16MiB
restriction").

Kernels before v4.18 would truncate the request and return the size
actually deduped in `bytes_deduped`.  Kernel v4.18 and later will loop
in the kernel until the entire request is satisfied (although still
in 16 MiB chunks, so larger extents will be split).

Modify the loop in userspace to measure the size the kernel actually
deduped, instead of assuming the kernel will only accept 16 MiB.
On current kernels this will always loop exactly once.

Since we now rely on `bytes_deduped`, make sure it has a sane value.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agoRevert "context: add experimental code for avoiding tiny extents"
Zygo Blaxell [Mon, 25 Nov 2024 01:16:46 +0000 (20:16 -0500)]
Revert "context: add experimental code for avoiding tiny extents"
because this problem is better solved elsewhere.

This reverts commit 11fabd66a84b6631fb39bb6b2c066c689351fc26.

21 months agousage: the default scan mode is 3 (recent)
Zygo Blaxell [Sat, 2 Sep 2023 16:19:43 +0000 (12:19 -0400)]
usage: the default scan mode is 3 (recent)

The code and docs were changed some time ago, but not the usage message.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agodocs: add the 6.10..6.12 delayed refs bug
Zygo Blaxell [Sat, 30 Nov 2024 00:30:42 +0000 (19:30 -0500)]
docs: add the 6.10..6.12 delayed refs bug

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agocrawl: rename next_transid() to avoid confusion with BeesScanMode::next_transid()
Zygo Blaxell [Sat, 30 Nov 2024 17:49:22 +0000 (12:49 -0500)]
crawl: rename next_transid() to avoid confusion with BeesScanMode::next_transid()

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
21 months agotrace: add file and line numbers all the way up the stack
Zygo Blaxell [Sat, 30 Nov 2024 15:18:58 +0000 (10:18 -0500)]
trace: add file and line numbers all the way up the stack

These were added to crucible all the way back in 2018 (1beb61fb78ba
"crucible: error: record location of exception in what() message")
but it's even more useful in the stack tracer in bees.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
2 years agocontext: reduce the size of LOGICAL_INO buffers
Zygo Blaxell [Mon, 17 Jul 2023 21:49:21 +0000 (17:49 -0400)]
context: reduce the size of LOGICAL_INO buffers

Since we'll never process more than BEES_MAX_EXTENT_REF_COUNT extent
references by definition, it follows that we should not allocate buffer
space for them when we perform the LOGICAL_INO ioctl.

There is some evidence (particularly
https://github.com/Zygo/bees/issues/260#issuecomment-1627598058) that
the kernel is subjecting the page cache to a lot of disruption when
trying allocate large buffers for LOGICAL_INO.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
2 years agousage: the default scan mode is 1 (independent)
Zygo Blaxell [Sat, 2 Sep 2023 16:19:18 +0000 (12:19 -0400)]
usage: the default scan mode is 1 (independent)

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
2 years agolib: fix btrfs_data_container pointer casts for 32-bit userspace on 64-bit kernels
Zygo Blaxell [Thu, 18 Apr 2024 03:07:41 +0000 (23:07 -0400)]
lib: fix btrfs_data_container pointer casts for 32-bit userspace on 64-bit kernels

Apparently reinterpret_cast<uint64_t> sign-extends 32-bit pointers.
This is OK when running on a 32-bit kernel that will truncate the pointer
to 32 bits, but when running on a 64-bit kernel, the extra bits are
interpreted as part of the (now very invalid) address.

Use <uintptr_t> instead, which is unsigned, integer, and the same word
size as the arch's pointer type.  Ordinary numeric conversion can take
it from there, filling the rest of the word with zeros.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: add vmalloc bug to kernel bugs list
Zygo Blaxell [Thu, 6 Jul 2023 17:45:09 +0000 (13:45 -0400)]
docs: add vmalloc bug to kernel bugs list

The bug is:

v6.3-rc6: f349b15e183d mm: vmalloc: avoid warn_alloc noise caused by fatal signal

The fixes are:

v6.4: 95a301eefa82 mm/vmalloc: do not output a spurious warning when huge vmalloc() fails
v6.3.10: c189994b5dd3 mm/vmalloc: do not output a spurious warning when huge vmalloc() fails

The bug has been backported to LTS, but the fix has not:

v6.2.11: 61334bc29781 mm: vmalloc: avoid warn_alloc noise caused by fatal signal
v6.1.24: ef6bd8f64ce0 mm: vmalloc: avoid warn_alloc noise caused by fatal signal
v5.15.107: a184df0de132 mm: vmalloc: avoid warn_alloc noise caused by fatal signal

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agocontext: log when LOGICAL_INO returns 0 refs
Zygo Blaxell [Mon, 8 May 2023 12:09:08 +0000 (08:09 -0400)]
context: log when LOGICAL_INO returns 0 refs

There was a bug in kernel 6.3 where LOGICAL_INO with IGNORE_OFFSET
sometimes fails to ignore the offset.  That bug is now fixed, but
LOGICAL_INO still returns 0 refs much more often than seems appropriate.

This is most likely because bees frequently deletes extents while there
is still work waiting for them in Task queues.  In this case, LOGICAL_INO
correctly returns an empty list, because every reference to some extent
is deleted, but the new extent tree with that extent removed is not yet
committed in btrfs.

Add a DEBUG-level log message and an event counter to track these events.
In the absence of a kernel bug, the debug message may indicate CPU time
was wasted performing a search whose outcome could have been predicted.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: add IGNORE_OFFSET regression in 6.2..6.3 to kernel bugs list
Zygo Blaxell [Wed, 10 May 2023 02:24:00 +0000 (22:24 -0400)]
docs: add IGNORE_OFFSET regression in 6.2..6.3 to kernel bugs list

This doesn't impact the current bees master, but it does break bees-next.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agocontext: downgrade toxic extent workaround message
Zygo Blaxell [Thu, 6 Jul 2023 16:36:08 +0000 (12:36 -0400)]
context: downgrade toxic extent workaround message

Toxic extents are much less of a problem now than they were in kernels
before 5.7.  Downgrade the log message level to reflect their lesser
importance.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agotest: GCC 13 fix for limits.cc
Zygo Blaxell [Mon, 8 May 2023 01:02:58 +0000 (21:02 -0400)]
test: GCC 13 fix for limits.cc

GCC complains that #include <cstdint> is missing, so add that.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agobtrfs-tree: fix build on clang++16
Zygo Blaxell [Mon, 8 May 2023 01:16:40 +0000 (21:16 -0400)]
btrfs-tree: fix build on clang++16

The "loops" variable isn't read (only set) if not built with extra
debug code.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: working around `btrfs send` issues isn't really a feature
Zygo Blaxell [Tue, 7 Mar 2023 15:20:16 +0000 (10:20 -0500)]
docs: working around `btrfs send` issues isn't really a feature

The critical kernel bugs in send have been fixed for years.
The limitations that remain aren't bugs, and bees has no sustainable
workaround for them.

Also update copyright year range.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: fill in missing LTS backports for "1119a72e223f btrfs: tree-checker: do not...
Zygo Blaxell [Tue, 7 Mar 2023 15:15:53 +0000 (10:15 -0500)]
docs: fill in missing LTS backports for "1119a72e223f btrfs: tree-checker: do not error out if extent ref hash doesn't match"

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agoroots: make sure transid_max's computed value isn't max
Zygo Blaxell [Tue, 7 Feb 2023 03:46:46 +0000 (22:46 -0500)]
roots: make sure transid_max's computed value isn't max

We check the result of transid_max_nocache(), but not the result of
transid_max().  The latter is a computed result that is even more likely
to be wrong[citation needed].

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: add "missing" features that have been in development for some time already
Zygo Blaxell [Sat, 25 Feb 2023 08:13:23 +0000 (03:13 -0500)]
docs: add "missing" features that have been in development for some time already

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: update GCC versions list and clarify markdown statement
Zygo Blaxell [Sat, 25 Feb 2023 08:13:00 +0000 (03:13 -0500)]
docs: update GCC versions list and clarify markdown statement

I don't know if anyone else is testing GCC versions before 8.0 any more,
but I'm not.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: update front page
Zygo Blaxell [Sat, 25 Feb 2023 08:12:27 +0000 (03:12 -0500)]
docs: update front page

At least one user was significantly confused by "designed for large
filesystems".

The btrfs send workarounds aren't new any more.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: minor changes to how-it-works based on past user questions
Zygo Blaxell [Sat, 25 Feb 2023 08:12:15 +0000 (03:12 -0500)]
docs: minor changes to how-it-works based on past user questions

Clarify that "too large" and "too small" are some distance away from each other.
The Goldilocks zone is _wide_.

The interval between cache drops is now shorter.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: various gotcha updates
Zygo Blaxell [Sat, 25 Feb 2023 08:11:30 +0000 (03:11 -0500)]
docs: various gotcha updates

Fixing the obviously wrong and out of date stuff.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: simplify the exit-with-SIGTERM description
Zygo Blaxell [Sat, 25 Feb 2023 08:10:29 +0000 (03:10 -0500)]
docs: simplify the exit-with-SIGTERM description

The description now matches the code again.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: update the feature interactions page
Zygo Blaxell [Sat, 25 Feb 2023 08:09:25 +0000 (03:09 -0500)]
docs: update the feature interactions page

Fixing the obviously out-of-date and no-longer-tested things.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: update kernel bugs and workarounds list for 6.2.0
Zygo Blaxell [Sat, 25 Feb 2023 08:08:59 +0000 (03:08 -0500)]
docs: update kernel bugs and workarounds list for 6.2.0

Remove some of the repetition to make the document easier to edit.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agocontext: create a Pool of BtrfsIoctlLogicalInoArgs objects
Zygo Blaxell [Tue, 21 Feb 2023 05:04:31 +0000 (00:04 -0500)]
context: create a Pool of BtrfsIoctlLogicalInoArgs objects

Each object contains a 16 MiB buffer, which is very heavy for some
malloc implementations.

Keep the objects in a Pool so that their buffers are only allocated and
deallocated once in the process lifetime.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agofs: allow BtrfsIoctlLogicalInoArgs to be reused, remove virtual methods
Zygo Blaxell [Tue, 21 Feb 2023 04:44:20 +0000 (23:44 -0500)]
fs: allow BtrfsIoctlLogicalInoArgs to be reused, remove virtual methods

Some malloc implementations will try to mmap() and munmap() large buffers
every time they are used, causing a severe loss of performance.

Nothing ever overrode the virtual methods, and there was no virtual
destructor, so they cause compiler warnings at build time when used with
a template that tries to delete pointers to them.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agoProgressTracker: reduce memory usage with long-running work items
Zygo Blaxell [Fri, 24 Feb 2023 03:33:35 +0000 (22:33 -0500)]
ProgressTracker: reduce memory usage with long-running work items

ProgressTracker was only freeing memory for work items when they reach
the head of the work tracking queue.  If the first work item takes
hours to complete, and thousands of items are processed every second,
this leads to millions of completed items tracked in memory at a time,
wasting gigabytes of system RAM.

Rewrite ProgressHolderState methods to keep only incomplete work items
in memory, regardless of the order in which they are added or removed.

Also fix the unit tests which were relying on the memory leak to work,
and add test cases for code coverage.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agoseeker: fix the test for ILP32 platforms
Zygo Blaxell [Sun, 12 Feb 2023 18:00:00 +0000 (13:00 -0500)]
seeker: fix the test for ILP32 platforms

Not sure what I was thinking, but the argument here should clearly
be uint64_t.

Fixes: https://github.com/Zygo/bees/issues/248
Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agoroots: don't share a RootFetcher between threads
Zygo Blaxell [Mon, 20 Feb 2023 16:13:18 +0000 (11:13 -0500)]
roots: don't share a RootFetcher between threads

If the send workaround is enabled, it is possible for two threads (a
thread running the crawl_new task, and a thread attempting to apply the
send workaround) to access the same RootFetcher object at the same time.
That never ends well.

Give each function its own BtrfsRootFetcher object.

Fixes: https://github.com/Zygo/bees/issues/250
Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agoMakefile: also drop fiemap and fiewalk from main Makefile
Kai Krakow [Sat, 28 Jan 2023 10:21:51 +0000 (11:21 +0100)]
Makefile: also drop fiemap and fiewalk from main Makefile

Fixes: https://github.com/Zygo/bees/commit/ccd8dcd43f0c903a94367817a0a9e94034a5cce4
Signed-off-by: Kai Krakow <kai@kaishome.de>
3 years agohash: flush the table more slowly
Zygo Blaxell [Mon, 23 Jan 2023 04:44:51 +0000 (23:44 -0500)]
hash: flush the table more slowly

With SIGTERM and fast exit, the trickle writeback is less important.
We don't want to flood people's IO subsystems with continuous writes.
This really should be configurable at runtime.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agotest: simplify Makefile
Zygo Blaxell [Mon, 23 Jan 2023 02:52:51 +0000 (21:52 -0500)]
test: simplify Makefile

Make can build dependencies in parallel, so let Make do that.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agosrc: bees-version.cc cleanups
Zygo Blaxell [Mon, 23 Jan 2023 03:19:16 +0000 (22:19 -0500)]
src: bees-version.cc cleanups

Do rebuild bees-version.cc if libcrucible changes.
Don't rebuild bees-version.cc if it doesn't change.
Also use the standard suffix for new files.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agosrc: simplify Makefile
Zygo Blaxell [Mon, 23 Jan 2023 02:50:07 +0000 (21:50 -0500)]
src: simplify Makefile

Make can build dependencies in parallel, so let Make do that.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agolib: drop version.cc entirely
Zygo Blaxell [Sat, 19 Nov 2022 07:19:25 +0000 (02:19 -0500)]
lib: drop version.cc entirely

crucible::VERSION doesn't make much sense now that libcrucible no
longer exists as a shared library.  Nothing ever referenced it, so
it can go away.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agolib: simplify dependency generation
Zygo Blaxell [Mon, 23 Jan 2023 02:28:28 +0000 (21:28 -0500)]
lib: simplify dependency generation

We don't need to run all the dependencies first, Make can do those in parallel.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agofd: FS_IOC_SETFLAGS takes an int* argument not a long*
Zygo Blaxell [Tue, 14 Dec 2021 06:18:39 +0000 (01:18 -0500)]
fd: FS_IOC_SETFLAGS takes an int* argument not a long*

According to ioctl_iflags(2):

The type of the argument given to the FS_IOC_GETFLAGS and
FS_IOC_SETFLAGS  operations is int *, notwithstanding the
implication in the kernel source file include/uapi/linux/fs.h
that the argument is long *.

So this code doesn't work on be64 machines.

Also, Valgrind complains about it.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agofd: pwrite returns ssize_t not int
Zygo Blaxell [Thu, 7 Apr 2022 11:25:37 +0000 (07:25 -0400)]
fd: pwrite returns ssize_t not int

A subtle distinction, and not one that is particularly relevant to bees,
but it does make toolchains complain.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agofs: get rid of base class btrfs_ioctl_logical_ino_args
Zygo Blaxell [Sun, 23 Oct 2022 18:01:38 +0000 (14:01 -0400)]
fs: get rid of base class btrfs_ioctl_logical_ino_args

Another instance of the pattern where we derived a crucible class
from a btrfs struct.  Make it an automatic variable instead.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agofs: remove duplicate BTRFS_COMPRESS_ definitions
Zygo Blaxell [Sun, 25 Oct 2020 05:10:58 +0000 (01:10 -0400)]
fs: remove duplicate BTRFS_COMPRESS_ definitions

This was fixed in

7f660f50b lib: fs: stop using libbtrfs-dev helper functions to re-enable buffer length checks

but apparently some copies live on.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agofiemap, fiewalk: drop dead example/test code
Zygo Blaxell [Thu, 4 Nov 2021 05:48:53 +0000 (01:48 -0400)]
fiemap, fiewalk: drop dead example/test code

These tools are obsolete.  fiemap was a thin wrapper around FIEMAP,
but FIEMAP is not useful on btrfs.  fiewalk was a thin wrapper around
BtrfsExtentWalker, but development on BtrfsExtentWalker has been
abandoned.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agocontext: remove the one call to operator vector<> method in BtrfsIoctlLogicalInoArgs
Zygo Blaxell [Sun, 23 Oct 2022 21:51:18 +0000 (17:51 -0400)]
context: remove the one call to operator vector<> method in BtrfsIoctlLogicalInoArgs

There's only one user of this method.  Open-code it so we can kill the
method in libcrucible.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agohash: don't spin when writes fail
Zygo Blaxell [Thu, 19 Jan 2023 14:24:48 +0000 (09:24 -0500)]
hash: don't spin when writes fail

When a hash table write fails, we skip over the write throttling because
we didn't report that we successfully wrote an extent.  This can be bad
if the filesystem is full and the allocations for writes are burning a
lot of CPU time searching for free space.

We also don't retry the write later on since we assume the extent is
clean after a write attempt whether it was successful or not, so the
extent might not be written out later when writes are possible again.

Check whether a hash extent is dirty, and always throttle after
attempting the write.

If a write fails, leave the extent dirty so we attempt to write it out
the next time flush cycles through the hash table.  During shutdown
this will reattempt each failing write once, after that the updated hash
table data will be dropped.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: fix broken link in options.md
Zygo Blaxell [Mon, 16 Jan 2023 05:06:22 +0000 (00:06 -0500)]
docs: fix broken link in options.md

Links in docs/ are relative to docs/, not the top level.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agomain: catch exceptions and exit gracefully
Zygo Blaxell [Wed, 4 Jan 2023 01:37:52 +0000 (20:37 -0500)]
main: catch exceptions and exit gracefully

Calling 'bees -m4' should not call 'std::terminate()', but it does.

Use catch_all instead.  It will still pass the exit value to return
from main.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agofs: get rid of base class btrfs_ioctl_same_extent_info
Zygo Blaxell [Sun, 23 Oct 2022 18:01:38 +0000 (14:01 -0400)]
fs: get rid of base class btrfs_ioctl_same_extent_info

We only use BtrfsExtentInfo when it's exactly equivalent to the
base, so drop the derived class.

While we're here, fix BtrfsExtentSame::add so it uses a btrfs-compatible
uint64_t instead of an off_t.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agobtrfs-tree: fix whitespace and const
Zygo Blaxell [Sat, 25 Dec 2021 18:34:25 +0000 (13:34 -0500)]
btrfs-tree: fix whitespace and const

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agobtrfs-tree: translate item types for error messages
Zygo Blaxell [Sun, 4 Dec 2022 08:43:21 +0000 (03:43 -0500)]
btrfs-tree: translate item types for error messages

Look up the name when filling in the what() field for the exception.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agobtrfs-tree: add chunk items: length and type
Zygo Blaxell [Sun, 1 Jan 2023 01:15:12 +0000 (20:15 -0500)]
btrfs-tree: add chunk items: length and type

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agoreadahead: report the original size in BEESTOOLONG
Zygo Blaxell [Fri, 30 Dec 2022 19:36:55 +0000 (14:36 -0500)]
readahead: report the original size in BEESTOOLONG

BEESTOOLONG was always reporting a size of zero, and the offset of the
end of the readahead region.  Report the original size instead (and also
in BEESTRACE and BEESNOTE).

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agodocs: add crawl_again, drop crawl_restart
Zygo Blaxell [Thu, 29 Dec 2022 10:35:57 +0000 (05:35 -0500)]
docs: add crawl_again, drop crawl_restart

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agoroots: fix extent lock failure handling
Zygo Blaxell [Thu, 22 Dec 2022 05:12:10 +0000 (00:12 -0500)]
roots: fix extent lock failure handling

Drop the crawl_restart counter, it doesn't happen here (or anywhere else).

Add the crawl_again counter for extents that are restarted due to an
extent-level lock.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>
3 years agotrace: use pthread_setname wrapper
Zygo Blaxell [Thu, 29 Dec 2022 09:03:53 +0000 (04:03 -0500)]
trace: use pthread_setname wrapper

libcrucible can deal with the Linux kernel and/or libc's thread name
limitations.  No need to duplicate that work in bees.

Signed-off-by: Zygo Blaxell <bees@furryterror.org>