Commits · db346e7120a6dec1534184ea2abf9d22edbb9b8a · Kirill Smelkov / linux

An error occurred fetching the project authors.

22 Oct, 2023 40 commits

bcachefs: bch2_bucket_alloc_trans_early -> for_each_btree_key_norestart · db346e71

Kent Overstreet authored 2 years ago

Nested btree transactions require special care, and an upcoming patch is
going to add assertions to that effect. We don't want to be using them
unnecessarily, so this patch switches bch2_bucket_trans_early() to not
handle transaction restarts.

This patch also adds a cursor so that on transaction restart we can
continue scanning from where the previous search for an empty bucket
left off.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

db346e71

bcachefs: EINTR -> BCH_ERR_transaction_restart · 549d173c

Kent Overstreet authored 2 years ago

Now that we have error codes, with subtypes, we can switch to our own
error code for transaction restarts - and even better, a distinct error
code for each transaction restart reason: clearer code and better
debugging.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

549d173c

bcachefs: Prevent a btree iter overflow in alloc path · 90cecb92

Kent Overstreet authored 2 years ago

In bch2_bucket_alloc_trans(), we're iterating over buckets - but not
directly with an iterator, since we're iterating over the freespace
btree.

This means that we need to clear iter->path->preserve, otherwise we'll
end up retaining a btree_path for every alloc key we touched - which is
not what we want here.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

90cecb92

bcachefs: Improved errcodes · 615f867c

Kent Overstreet authored 2 years ago

Instead of overloading standard error codes (EINTR/EAGAIN), and defining
short lists of error codes in multiple places that potentially end up
overlapping & conflicting, we're now going to have one master list of
error codes.

Error codes are defined with an x-macro: thus we also have
bch2_err_str() now.

Also, error codes have a class field. Now, instead of checking for
errors with ==, code should use bch2_err_matches(), which returns true
if the error is equal to or a sub-error of the error class.

This means we can define unique errors for every source location where
an error is generated, which will help improve our error messages.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

615f867c

bcachefs: Improve bucket_alloc_fail tracepoint · 8ef98313

Kent Overstreet authored 2 years ago

We should be printing the number of free buckets, not just the number of
available buckets.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

8ef98313

bcachefs: Split out dev_buckets_free() · 30f0349d

Kent Overstreet authored 2 years ago

Previously, dev_buckets_available() only counted buckets that are
eligible to be allocated right now - i.e. buckets that don't have cached
data, or need discard, or need gc gens, etc.

But most users of this function want to know how many buckets are
eligible to be allocated from without moving data around - copygc,
allocator striping, which means we should be including cached data
buckets etc.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

30f0349d

bcachefs: Printbuf rework · 401ec4db

Kent Overstreet authored 2 years ago

This converts bcachefs to the modern printbuf interface/implementation,
synced with the version to be submitted upstream.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

401ec4db

bcachefs: Improve bch2_open_buckets_to_text() · 3518e6fa

Kent Overstreet authored 2 years ago

This patch updates bch2_open_buckets_to_text() to include the device and
bucket the open_bucket owns.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

3518e6fa

bcachefs: Fold bucket_state in to BCH_DATA_TYPES() · 822835ff

Kent Overstreet authored 2 years ago

Previously, we were missing accounting for buckets in need_gc_gens and
need_discard states. This matters because buckets in those states need
other btree operations done before they can be used, so they can't be
conuted when checking current number of free buckets against the
allocation watermark.

Also, we weren't directly counting free buckets at all. Now, data type 0
== BCH_DATA_free, and free buckets are counted; this means we can get
rid of the separate (poorly defined) count of unavailable buckets.

This is a new on disk format version, with upgrade and fsck required for
the accounting changes.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

822835ff

bcachefs: Kill allocator threads & freelists · f25d8215

Kent Overstreet authored 3 years ago

Now that we have new persistent data structures for the allocator, this
patch converts the allocator to use them.

Now, foreground bucket allocation uses the freespace btree to find
buckets to allocate, instead of popping buckets off the freelist.

The background allocator threads are no longer needed and are deleted,
as well as the allocator freelists. Now we only need background tasks
for invalidating buckets containing cached data (when we are low on
empty buckets), and for issuing discards.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f25d8215

bcachefs: Run btree updates after write out of write_point · b17d3cec

Kent Overstreet authored 2 years ago

In the write path, after the write to the block device(s) complete we
have to punt to process context to do the btree update.

Instead of using the work item embedded in op->cl, this patch switches
to a per write-point work item. This helps with two different issues:

 - lock contention: btree updates to the same writepoint will (usually)
   be updating the same alloc keys
 - context switch overhead: when we're bottlenecked on btree updates,
   having a thread (running out of a work item) checking the write point
   for completed ops is cheaper than queueing up a new work item and
   waking up a kworker.

In an arbitrary benchmark, 4k random writes with fio running inside a
VM, this patch resulted in a 10% improvement in total iops.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b17d3cec

bcachefs: x-macroize alloc_reserve enum · 3e154711

Kent Overstreet authored 2 years ago

This makes an array of strings available, like our other enums.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3e154711

bcachefs: Kill verify_not_stale() · fcf01959

Kent Overstreet authored 3 years ago

This is ancient code that's more effectively checked in other places
now.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

fcf01959

bcachefs: New in-memory array for bucket gens · a7860877

Kent Overstreet authored 3 years ago

The main in-memory bucket array is going away, but we'll still need to
keep bucket generations in memory, at least for now - ptr_stale() needs
to be an efficient operation.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

a7860877

bcachefs: Put open_buckets in a hashtable · 9ddffaf8

Kent Overstreet authored 3 years ago

This is so that the copygc code doesn't have to refer to
bucket_mark.owned_by_allocator - assisting in getting rid of the in
memory bucket array.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

9ddffaf8

bcachefs: Refactor open_bucket code · abe19d45

Kent Overstreet authored 3 years ago

Prep work for adding a hash table of open buckets - instead of embedding
a bch_extent_ptr, we need to refer to the bucket directly so that we're
not calling sector_to_bucket() in the hash table lookup code, which has
an expensive divide.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

abe19d45

bcachefs: bch2_alloc_sectors_append_ptrs() now takes cached flag · 57af63b2
Kent Overstreet authored 3 years ago
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
57af63b2

bcachefs: Rewrite bch2_bucket_alloc_new_fs() · 09943313

Kent Overstreet authored 3 years ago

This changes bch2_bucket_alloc_new_fs() to a simple bump allocator that
doesn't need to use the in memory bucket array, part of a larger patch
series to entirely get rid of the in memory bucket array, except for
gc/fsck.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

09943313

bcachefs: Make sure bch2_bucket_alloc_new_fs() obeys buckets_nouse · 6be1b6d9
Kent Overstreet authored 3 years ago
```
This fixes the filesystem migrate tool.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
```
6be1b6d9

bcachefs: Convert bucket_alloc_ret to negative error codes · fc6c01e2

Kent Overstreet authored 3 years ago

Start a new header, errcode.h, for bcachefs-private error codes - more
error codes will be converted later.

This patch just converts bucket_alloc_ret so that they can be mixed with
standard error codes and passed as ERR_PTR errors - the ec.c code was
doing this already, but incorrectly.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

fc6c01e2

bcachefs: Allocator refactoring · 89baec78

Kent Overstreet authored 3 years ago

This uses the kthread_wait_freezable() macro to simplify a lot of the
allocator thread code, along with cleaning up bch2_invalidate_bucket2().
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

89baec78

bcachefs: gc shouldn't care about owned_by_allocator · dac1525d

Kent Overstreet authored 3 years ago

The owned_by_allocator field is a purely in memory thing, even if/when
we bring back GC at runtime there's no need for it to be recalculating
this field. This is prep work for pulling it out of struct bucket, and
eventually getting rid of the bucket array.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

dac1525d

bcachefs: Fix an RCU splat · 3e07a730

Kent Overstreet authored 3 years ago

Writepoints are never deallocated so the rcu_read_lock() isn't really
needed, but we are doing lockless list traversal.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3e07a730

bcachefs: Fix copygc threshold · cb66fc5f

Kent Overstreet authored 3 years ago

Awhile back the meaning of is_available_bucket() and thus also
bch_dev_usage->buckets_unavailable changed to include buckets that are
owned by the allocator - this was so that the stat could be persisted
like other allocation information, and wouldn't have to be regenerated
by walking each bucket at mount time.

This broke copygc, which needs to consider buckets that are reclaimable
and haven't yet been grabbed by the allocator thread and moved onta
freelist. This patch fixes that by adding dev_buckets_reclaimable() for
copygc and the allocator thread, and cleans up some of the callers a bit.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

cb66fc5f

bcachefs: Refactor dev usage · 72eab8da

Kent Overstreet authored 4 years ago

This is to make it more amenable for serialization.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

72eab8da

bcachefs: Rework allocating buckets for stripes · 6c7585b0

Kent Overstreet authored 4 years ago

Allocating buckets for existing stripes was busted, in part because the
data structures were too contorted. This reworks new stripes so that we
have an array of open buckets that matches blocks in the stripe, and
it's sparse if we're reusing an existing stripe.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

6c7585b0

bcachefs: Reserve some open buckets for btree allocations · 890e3f5b

Kent Overstreet authored 4 years ago

This reverts part of the change from "bcachefs: Don't use
BTREE_INSERT_USE_RESERVE so much" - it turns out we still should be
reserving open buckets for btree node allocations, because otherwise
data bucket allocations (especially with erasure coding enabled) can use
up all our open buckets and we won't be able to do the metadata update
that lets us release those open bucket references. Oops.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

890e3f5b

bcachefs: Use separate new stripes for copygc and non-copygc · 8deed5f4

Kent Overstreet authored 4 years ago

Allocations for copygc have to be kept separate from everything else,
so that copygc doesn't get starved.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

8deed5f4

bcachefs: Change allocations for ec stripes to blocking · 2c40a240

Kent Overstreet authored 4 years ago

We don't want writes to not get erasure coded just because the allocator
temporarily wasn't keeping up.

However, it's not guaranteed that these allocations will ever succeed,
we can currently get stuck - especially if devices are different sizes -
we still have work to do in this area.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

2c40a240

bcachefs: Don't use BTREE_INSERT_USE_RESERVE so much · 3187aa8d

Kent Overstreet authored 4 years ago

Previously, we were using BTREE_INSERT_RESERVE in a lot of places where
it no longer makes sense.

 - we now have more open_buckets than we used to, and the reserves work
   better, so we shouldn't need to use BTREE_INSERT_RESERVE just because
   we're holding open_buckets pinned anymore.

 - We have the btree key cache for updates to the alloc btree, meaning
   we no longer need the btree reserve to ensure the allocator can make
   forward progress.

This means that we should only need a reserve for btree updates to
ensure that copygc can make forward progress.

Since it's now just for copygc, we can also fold RESERVE_BTREE into
RESERVE_MOVINGGC (the allocator's freelist reserve).
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3187aa8d

bcachefs: Don't write bucket IO time lazily · f30dd860

Kent Overstreet authored 4 years ago

With the btree key cache code, we don't need to update the alloc btree
lazily - and this will mean we can remove the bch2_alloc_write() call in
the shutdown path.

Future work: we really need to expend the bucket IO clocks from 16 to 64
bits, so that we don't have to rescale them.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f30dd860

bcachefs: Ensure we only allocate one EC bucket per writepoint · d3a2b5d8

Kent Overstreet authored 4 years ago

Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d3a2b5d8

bcachefs: Don't let copygc buckets be stolen by other threads · 74ed7e56

Kent Overstreet authored 4 years ago

And assorted other copygc fixes.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

74ed7e56

bcachefs: Delete unused arguments · 3d080aa5

Kent Overstreet authored 4 years ago

Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3d080aa5

bcachefs: Don't restrict copygc writes to the same device · 8f3b41ab

Kent Overstreet authored 4 years ago

This no longer makes any sense, since copygc is now one thread per
filesystem, not per device, with a single write point.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

8f3b41ab

bcachefs: Make copygc thread global · e6d11615

Kent Overstreet authored 4 years ago

Per device copygc threads don't move data to different devices and they
make fragmentation works - they don't make much sense anymore.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

e6d11615

bcachefs: Use x-macros for data types · 89fd25be
Kent Overstreet authored 4 years ago
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
89fd25be

bcachefs: Refactor stripe creation · f6b94a3b

Kent Overstreet authored 4 years ago

Prep work for the patch to update existing stripes with new data blocks.
This moves allocating new stripes into ec.c, and also sets up the data
structures so that we can handly only allocating some of the blocks in a
stripe.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f6b94a3b

bcachefs: Move stripe creation to workqueue · 703e2a43

Kent Overstreet authored 4 years ago

This is mainly to solve a lock ordering issue, and also simplifies the
code a bit.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

703e2a43

bcachefs: Make open bucket reserves more conservative · 6b5f9b29

Kent Overstreet authored 4 years ago

Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

6b5f9b29