Commits · 69d1f052d1675c2af7da496f0265f68673328afb · Kirill Smelkov / linux

22 Oct, 2023 40 commits

bcachefs: Correctly initialize new buckets on device resize · 69d1f052

Kent Overstreet authored Sep 28, 2023

bch2_dev_resize() was never updated for the allocator rewrite with
persistent freelists, and it wasn't noticed because the tests weren't
running fsck - oops.

Fix this by running bch2_dev_freespace_init() for the new buckets.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

69d1f052

bcachefs: Fix another smatch complaint · 4fc1f402

Kent Overstreet authored Sep 28, 2023

This should be harmless, but initialize last_seq anyways.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

4fc1f402

bcachefs: Use strsep() in split_devs() · dc08c661

Kent Overstreet authored Sep 28, 2023

Minor refactoring to fix a smatch complaint.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

dc08c661

bcachefs: Add iops fields to bch_member · 40f7914e

Hunter Shaffer authored Sep 25, 2023

Signed-off-by: Hunter Shaffer <huntershaffer182456@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

40f7914e

bcachefs: Rename bch_sb_field_members -> bch_sb_field_members_v1 · 9af26120

Hunter Shaffer authored Sep 25, 2023

Signed-off-by: Hunter Shaffer <huntershaffer182456@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

9af26120

bcachefs: New superblock section members_v2 · 3f7b9713

Hunter Shaffer authored Sep 25, 2023

members_v2 has dynamically resizable entries so that we can extend
bch_member. The members can no longer be accessed with simple array
indexing Instead members_v2_get is used to find a member's exact
location within the array and returns a copy of that member.
Alternatively member_v2_get_mut retrieves a mutable point to a member.
Signed-off-by: Hunter Shaffer <huntershaffer182456@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3f7b9713

bcachefs: Add new helper to retrieve bch_member from sb · 1241df58

Hunter Shaffer authored Sep 24, 2023

Prep work for introducing bch_sb_field_members_v2 - introduce new
helpers that will check for members_v2 if it exists, otherwise using v1
Signed-off-by: Hunter Shaffer <huntershaffer182456@gmail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

1241df58

bcachefs: bucket_lock() is now a sleepable lock · 73bbeaa2

Kent Overstreet authored Sep 27, 2023

fsck_err() may sleep - it takes a mutex and may allocate memory, so
bucket_lock() needs to be a sleepable lock.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

73bbeaa2

bcachefs: fix crc32c checksum merge byte order problem · 3c40841c

Brian Foster authored Sep 27, 2023

An fsstress task on a big endian system (s390x) quickly produces a
bunch of CRC errors in the system logs. Most of these are related to
the narrow CRCs path, but the fundamental problem can be reduced to
a single write and re-read (after dropping caches) of a previously
merged extent.

The key merge path that handles extent merges eventually calls into
bch2_checksum_merge() to combine the CRCs of the associated extents.
This code attempts to avoid a byte order swap by feeding the le64
values into the crc32c code, but the latter casts the resulting u64
value down to a u32, which truncates the high bytes where the actual
crc value ends up. This results in a CRC value that does not change
(since it is merged with a CRC of 0), and checksum failures ensue.

Fix the checksum merge code to swap to cpu byte order on the
boundaries to the external crc code such that any value casting is
handled properly.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3c40841c

bcachefs: Fix bch2_inode_delete_keys() · 42206663

Kent Overstreet authored Sep 27, 2023

bch2_inode_delete_keys() was using BTREE_ITER_NOT_EXTENTS, on the
assumption that it would never need to split extents.

But that caused a race with extents being split by other threads -
specifically, the data move path. Extents iterators have the iterator
position pointing to the start of the extent, which avoids the race.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

42206663

bcachefs: Make btree root read errors recoverable · 7dcf62c0

Kent Overstreet authored Sep 26, 2023

The entire btree will be lost, but that is better than the entire
filesystem not being recoverable.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

7dcf62c0

bcachefs: Fall back to requesting passphrase directly · 1ee608c6

Kent Overstreet authored Sep 26, 2023

We can only do this in userspace, unfortunately - but kernel keyrings
have never seemed to worked reliably, this is a useful fallback.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

1ee608c6

bcachefs: Fix looping around bch2_propagate_key_to_snapshot_leaves() · d281701b
Kent Overstreet authored Sep 26, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
d281701b

bcachefs: bch_err_msg(), bch_err_fn() now filters out transaction restart errors · d2a990d1

Kent Overstreet authored Sep 26, 2023

These errors aren't actual errors, and should never be printed - do this
in the common helpers.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d2a990d1

bcachefs: Silence transaction restart error message · a190cbcf
Kent Overstreet authored Sep 26, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
a190cbcf

bcachefs: More assertions for nocow locking · 1e3b4098

Kent Overstreet authored Sep 24, 2023

 - assert in shutdown path that no nocow locks are held
 - check for overflow when taking nocow locks
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

1e3b4098

bcachefs: nocow locking: Fix lock leak · efedfc2e
Kent Overstreet authored Sep 24, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
efedfc2e
bcachefs: Fixes for building in userspace · 793a06d9
Kent Overstreet authored Sep 23, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
793a06d9

bcachefs: Ignore unknown mount options · 03ef80b4

Kent Overstreet authored Sep 23, 2023

This makes mount option handling consistent with other filesystems -
options may be handled at different layers, so an option we don't know
about might not be intended for us.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

03ef80b4

bcachefs: Always check for invalid bkeys in main commit path · b560e32e

Kent Overstreet authored Sep 23, 2023

Previously, we would check for invalid bkeys at transaction commit time,
but only if CONFIG_BCACHEFS_DEBUG=y.

This check is important enough to always be on - it appears there's been
corruption making it into the journal that would have been caught by it.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b560e32e

bcachefs: Make sure to initialize equiv when creating new snapshots · eebe8a84

Kent Overstreet authored Sep 23, 2023

Previously, equiv was set in the snapshot deletion path, which is where
it's needed - equiv, for snapshot ID equivalence classes, would ideally
be a private data structure to the snapshot deletion path.

But if a new snapshot is created while snapshot deletion is running,
move_key_to_correct_snapshot() moves a key to snapshot id 0 - oops.

Fixes: https://github.com/koverstreet/bcachefs/issues/593Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

eebe8a84

bcachefs: Fix a null ptr deref in bch2_get_alloc_in_memory_pos() · 82142a55
Kent Overstreet authored Sep 22, 2023
```
Reported-by: smatch
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
82142a55

bcachefs: Fix changing durability using sysfs · d8b6f8c3

Torge Matthies authored Sep 21, 2023

Signed-off-by: Torge Matthies <openglfreak@googlemail.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d8b6f8c3

bcachefs: initial freeze/unfreeze support · 7239f8e0

Brian Foster authored Sep 15, 2023

Initial support for the vfs superblock freeze and unfreeze
operations. Superblock freeze occurs in stages, where the vfs
attempts to quiesce high level write operations, page faults, fs
internal operations, and then finally calls into the filesystem for
any last stage steps (i.e. log flushing, etc.) before marking the
superblock frozen.

The majority of write paths are covered by freeze protection (i.e.
sb_start_write() and friends) in higher level common code, with the
exception of the fs-internal SB_FREEZE_FS stage (i.e.
sb_start_intwrite()). This typically maps to active filesystem
transactions in a manner that allows the vfs to implement a barrier
of internal fs operations during the freeze sequence. This is not a
viable model for bcachefs, however, because it utilizes transactions
both to populate the journal as well as to perform journal reclaim.
This means that mapping intwrite protection to transaction lifecycle
or transaction commit is likely to deadlock freeze, as quiescing the
journal requires transactional operations blocked by the final stage
of freeze.

The flipside of this is that bcachefs does already maintain its own
internal sets of write references for similar purposes, currently
utilized for transitions from read-write to read-only mode. Since
this largely mirrors the high level sequence involved with freeze,
we can simply invoke this mechanism in the freeze callback to fully
quiesce the filesystem in the final stage. This means that while the
SB_FREEZE_FS stage is essentially a no-op, the ->freeze_fs()
callback that immediately follows begins by performing effectively
the same step by quiescing all internal write references.

One caveat to this approach is that without integration of internal
freeze protection, write operations gated on internal write refs
will fail with an internal -EROFS error rather than block on
acquiring freeze protection. IOW, this is roughly equivalent to only
having support for sb_start_intwrite_trylock(), and not the blocking
variant. Many of these paths already use non-blocking internal write
refs and so would map into an sb_start_intwrite_trylock() anyways.
The only instance of this I've been able to uncover that doesn't
explicitly rely on a higher level non-blocking write ref is the
bch2_rbio_narrow_crcs() path, which updates crcs in certain read
cases, and Kent has pointed out isn't critical if it happens to fail
due to read-only status.

Given that, implement basic freeze support as described above and
leave tighter integration with internal freeze protection as a
possible future enhancement. There are multiple potential ideas
worth exploring here. For example, we could implement a multi-stage
freeze callback that might allow bcachefs to quiesce its internal
write references without deadlocks, we could integrate intwrite
protection with bcachefs' internal write references somehow or
another, or perhaps consider implementing blocking support for
internal write refs to be used specifically for freeze, etc. In the
meantime, this enables functional freeze support and the associated
test coverage that comes with it.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

7239f8e0

bcachefs: More minor smatch fixes · 40a53b92

Kent Overstreet authored Sep 20, 2023

 - fix a few uninitialized return values
 - return a proper error code in lookup_lostfound()
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

40a53b92

bcachefs: Minor bch2_btree_node_get() smatch fixes · 51c801bc

Kent Overstreet authored Sep 20, 2023

 - it's no longer possible for trans to be NULL
 - also, move "wait for read to complete" to the slowpath,
   __bch2_btree_node_get().
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

51c801bc

bcachefs: snapshots: Use kvfree_rcu_mightsleep() · d04fdf5c
Kent Overstreet authored Sep 20, 2023
```
kvfree_rcu() was renamed - not removed.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
d04fdf5c

bcachefs: Fix strndup_user() error checking · 97ecc236

Kent Overstreet authored Sep 20, 2023

strndup_user() returns an error pointer, not NULL.
Reported-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

97ecc236

bcachefs: drop journal lock before calling journal_write · cfda31c0

Kent Overstreet authored Sep 19, 2023

bch2_journal_write() expects process context, it takes journal_lock as
needed.
Reported-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

cfda31c0

bcachefs: bch2_ioctl_disk_resize_journal(): check for integer truncation · 4b33a191
Kent Overstreet authored Sep 19, 2023
```
Reported-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
4b33a191

bcachefs: Fix error checks in bch2_chacha_encrypt_key() · 75e0c478

Kent Overstreet authored Sep 19, 2023

crypto_alloc_sync_skcipher() returns an ERR_PTR, not NULL.
Reported-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

75e0c478

bcachefs: Fix an overflow check · a55fc65e

Kent Overstreet authored Sep 19, 2023

When bucket sector counts were changed from u16s to u32s, a few things
were missed. This fixes an overflow check, and a truncation that
prevented the overflow check from firing.
Reported-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

a55fc65e

bcachefs: Fix copy_to_user() usage in flush_buf() · f7f6943a

Kent Overstreet authored Sep 19, 2023

copy_to_user() returns the number of bytes successfully copied - not an
errcode.
Reported-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f7f6943a

bcachefs: fix race between journal entry close and pin set · 3e55189b

Brian Foster authored Sep 15, 2023

bcachefs freeze testing via fstests generic/390 occasionally
reproduces the following BUG from bch2_fs_read_only():

  BUG_ON(atomic_long_read(&c->btree_key_cache.nr_dirty));

This indicates that one or more dirty key cache keys still exist
after the attempt to flush and quiesce the fs. The sequence that
leads to this problem actually occurs on unfreeze (ro->rw), and
looks something like the following:

- Task A begins a transaction commit and acquires journal_res for
  the current seq. This transaction intends to perform key cache
  insertion.
- Task B begins a bch2_journal_flush() via bch2_sync_fs(). This ends
  up in journal_entry_want_write(), which closes the current journal
  entry and drops the reference to the pin list created on entry open.
  The pin put pops the front of the journal via fast reclaim since the
  reference count has dropped to 0.
- Task A attempts to set the journal pin for the associated cached
  key, but bch2_journal_pin_set() skips the pin insert because the
  seq of the transaction reservation is behind the front of the pin
  list fifo.

The end result is that the pin associated with the cached key is not
added, which prevents a subsequent reclaim from processing the key
and thus leaves it dangling at freeze time. The fundamental cause of
this problem is that the front of the journal is allowed to pop
before a transaction with outstanding reservation on the associated
journal seq is able to add a pin. The count for the pin list
associated with the seq drops to zero and is prematurely reclaimed
as a result.

The logical fix for this problem lies in how the journal buffer is
managed in similar scenarios where the entry might have been closed
before a transaction with outstanding reservations happens to be
committed.

When a journal entry is opened, the current sequence number is
bumped, the associated pin list is initialized with a reference
count of 1, and the journal buffer reference count is bumped (via
journal_state_inc()). When a journal reservation is acquired, the
reservation also acquires a reference on the associated buffer. If
the journal entry is closed in the meantime, it drops both the pin
and buffer references held by the open entry, but the buffer still
has references held by outstanding reservation. After the associated
transaction commits, the reservation release drops the associated
buffer references and the buffer is written out once the reference
count has dropped to zero.

The fundamental problem here is that the lifecycle of the pin list
reference held by an open journal entry is too short to cover the
processing of transactions with outstanding reservations. The
simplest way to address this is to expand the pin list reference to
the lifecycle of the buffer vs. the shorter lifecycle of the open
journal entry. This ensures the pin list for a seq with outstanding
reservation cannot be popped and reclaimed before all outstanding
reservations have been released, even if the associated journal
entry has been closed for further reservations.

Move the pin put from journal entry close to where final processing
of the journal buffer occurs. Create a duplicate helper to cover the
case where the caller doesn't already hold the journal lock. This
allows generic/390 to pass reliably.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3e55189b

bcachefs: prepare journal buf put to handle pin put · fc08031b

Brian Foster authored Sep 15, 2023

bcachefs freeze testing has uncovered some raciness between journal
entry open/close and pin list reference count management. The
details of the problem are described in a separate patch. In
preparation for the associated fix, refactor the journal buffer put
path a bit to allow it to eventually handle dropping the pin list
reference currently held by an open journal entry.

Retain the journal write dispatch helper since the closure code is
inlined and we don't want to increase the amount of inline code in
the transaction commit path, but rename the function to reflect
the purpose of final processing of the journal buffer.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

fc08031b

bcachefs: refactor pin put helpers · 92b63f5b

Brian Foster authored Sep 15, 2023

We have a couple journal pin put helpers to handle cases where the
journal lock is already held or not. Refactor the helpers to lock
and reclaim from the highest level and open code the reclaim from
the one caller of the internal variant. The latter call will be
moved into the journal buf release helper in a later patch.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

92b63f5b

bcachefs: snapshot: Add missing assignment in bch2_delete_dead_snapshots() · d67a72bf

Dan Carpenter authored Sep 15, 2023

This code accidentally left out the "ret = " assignment so the errors
from for_each_btree_key2() are not checked.

Fixes: 53534482a250 ("bcachefs: for_each_btree_key2()")
Signed-off-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d67a72bf

bcachefs: fs-ioctl: Fix copy_to_user() error code · 1f12900a

Dan Carpenter authored Sep 15, 2023

The copy_to_user() function returns the number of bytes that it wasn't
able to copy but we want to return -EFAULT to the user.

Fixes: e0750d947352 ("bcachefs: Initial commit")
Signed-off-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

1f12900a

bcachefs: acl: Add missing check in bch2_acl_chmod() · b6c22147

Dan Carpenter authored Sep 15, 2023

The "ret = bkey_err(k);" assignment was accidentally left out so the
call to bch2_btree_iter_peek_slot() is not checked for errors.

Fixes: 53306e096d91 ("bcachefs: Always check for transaction restarts")
Signed-off-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b6c22147

bcachefs: acl: Uninitialized variable in bch2_acl_chmod() · e9a0a26e

Dan Carpenter authored Sep 15, 2023

The clean up code at the end of the function uses "acl" so it needs
to be initialized to NULL.

Fixes: 53306e096d91 ("bcachefs: Always check for transaction restarts")
Signed-off-by: Dan Carpenter <dan.carpenter@linaro.org>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

e9a0a26e