Commits · 9d0d1c8b1c9d80b17cfa86ecd50c8933a742585c · nexedi / linux

29 Mar, 2017 1 commit

Btrfs: bring back repair during read · 9d0d1c8b

Liu Bo authored Mar 24, 2017

Commit 20a7db8a ("btrfs: add dummy callback for readpage_io_failed
and drop checks") made a cleanup around readpage_io_failed_hook, and
it was supposed to keep the original sematics, but it also
unexpectedly disabled repair during read for dup, raid1 and raid10.

This fixes the problem by letting data's inode call the generic
readpage_io_failed callback by returning -EAGAIN from its
readpage_io_failed_hook in order to notify end_bio_extent_readpage to
do the rest.  We don't call it directly because the generic one takes
an offset from end_bio_extent_readpage() to calculate the index in the
checksum array and inode's readpage_io_failed_hook doesn't offer that
offset.

Cc: David Sterba <dsterba@suse.cz>
Signed-off-by: Liu Bo <bo.li.liu@oracle.com>
Reviewed-by: David Sterba <dsterba@suse.com>
[ keep the const function attribute ]
Signed-off-by: David Sterba <dsterba@suse.com>

9d0d1c8b

17 Mar, 2017 2 commits

btrfs: add missing memset while reading compressed inline extents · e1699d2d

Zygo Blaxell authored Mar 10, 2017

This is a story about 4 distinct (and very old) btrfs bugs.

Commit c8b97818 ("Btrfs: Add zlib compression support") added
three data corruption bugs for inline extents (bugs #1-3).

Commit 93c82d57 ("Btrfs: zero page past end of inline file items")
fixed bug #1:  uncompressed inline extents followed by a hole and more
extents could get non-zero data in the hole as they were read.  The fix
was to add a memset in btrfs_get_extent to zero out the hole.

Commit 166ae5a4 ("btrfs: fix inline compressed read err corruption")
fixed bug #2:  compressed inline extents which contained non-zero bytes
might be replaced with zero bytes in some cases.  This patch removed an
unhelpful memset from uncompress_inline, but the case where memset is
required was missed.

There is also a memset in the decompression code, but this only covers
decompressed data that is shorter than the ram_bytes from the extent
ref record.  This memset doesn't cover the region between the end of the
decompressed data and the end of the page.  It has also moved around a
few times over the years, so there's no single patch to refer to.

This patch fixes bug #3:  compressed inline extents followed by a hole
and more extents could get non-zero data in the hole as they were read
(i.e. bug #3 is the same as bug #1, but s/uncompressed/compressed/).
The fix is the same:  zero out the hole in the compressed case too,
by putting a memset back in uncompress_inline, but this time with
correct parameters.

The last and oldest bug, bug #0, is the cause of the offending inline
extent/hole/extent pattern.  Bug #0 is a subtle and mostly-harmless quirk
of behavior somewhere in the btrfs write code.  In a few special cases,
an inline extent and hole are allowed to persist where they normally
would be combined with later extents in the file.

A fast reproducer for bug #0 is presented below.  A few offending extents
are also created in the wild during large rsync transfers with the -S
flag.  A Linux kernel build (git checkout; make allyesconfig; make -j8)
will produce a handful of offending files as well.  Once an offending
file is created, it can present different content to userspace each
time it is read.

Bug #0 is at least 4 and possibly 8 years old.  I verified every vX.Y
kernel back to v3.5 has this behavior.  There are fossil records of this
bug's effects in commits all the way back to v2.6.32.  I have no reason
to believe bug #0 wasn't present at the beginning of btrfs compression
support in v2.6.29, but I can't easily test kernels that old to be sure.

It is not clear whether bug #0 is worth fixing.  A fix would likely
require injecting extra reads into currently write-only paths, and most
of the exceptional cases caused by bug #0 are already handled now.

Whether we like them or not, bug #0's inline extents followed by holes
are part of the btrfs de-facto disk format now, and we need to be able
to read them without data corruption or an infoleak.  So enough about
bug #0, let's get back to bug #3 (this patch).

An example of on-disk structure leading to data corruption found in
the wild:

        item 61 key (606890 INODE_ITEM 0) itemoff 9662 itemsize 160
                inode generation 50 transid 50 size 47424 nbytes 49141
                block group 0 mode 100644 links 1 uid 0 gid 0
                rdev 0 flags 0x0(none)
        item 62 key (606890 INODE_REF 603050) itemoff 9642 itemsize 20
                inode ref index 3 namelen 10 name: DB_File.so
        item 63 key (606890 EXTENT_DATA 0) itemoff 8280 itemsize 1362
                inline extent data size 1341 ram 4085 compress(zlib)
        item 64 key (606890 EXTENT_DATA 4096) itemoff 8227 itemsize 53
                extent data disk byte 5367308288 nr 20480
                extent data offset 0 nr 45056 ram 45056
                extent compression(zlib)

Different data appears in userspace during each read of the 11 bytes
between 4085 and 4096.  The extent in item 63 is not long enough to
fill the first page of the file, so a memset is required to fill the
space between item 63 (ending at 4085) and item 64 (beginning at 4096)
with zero.

Here is a reproducer from Liu Bo, which demonstrates another method
of creating the same inline extent and hole pattern:

Using 'page_poison=on' kernel command line (or enable
CONFIG_PAGE_POISONING) run the following:

	# touch foo
	# chattr +c foo
	# xfs_io -f -c "pwrite -W 0 1000" foo
	# xfs_io -f -c "falloc 4 8188" foo
	# od -x foo
	# echo 3 >/proc/sys/vm/drop_caches
	# od -x foo

This produce the following on my box:

Correct output:  file contains 1000 data bytes followed
by zeros:

	0000000 cdcd cdcd cdcd cdcd cdcd cdcd cdcd cdcd
	*
	0001740 cdcd cdcd cdcd cdcd 0000 0000 0000 0000
	0001760 0000 0000 0000 0000 0000 0000 0000 0000
	*
	0020000

Actual output:  the data after the first 1000 bytes
will be different each run:

	0000000 cdcd cdcd cdcd cdcd cdcd cdcd cdcd cdcd
	*
	0001740 cdcd cdcd cdcd cdcd 6c63 7400 635f 006d
	0001760 5f74 6f43 7400 435f 0053 5f74 7363 7400
	0002000 435f 0056 5f74 6164 7400 645f 0062 5f74
	(...)
Signed-off-by: Zygo Blaxell <ce3g8jdj@umail.furryterror.org>
Reviewed-by: Liu Bo <bo.li.liu@oracle.com>
Reviewed-by: Chris Mason <clm@fb.com>
Signed-off-by: Chris Mason <clm@fb.com>

e1699d2d

Btrfs: fix regression in lock_delalloc_pages · 49d4a334

Liu Bo authored Mar 06, 2017

The bug is a regression after commit
(da2c7009 "btrfs: teach __process_pages_contig about PAGE_LOCK operation")
and commit
(76c0021d "Btrfs: use helper to simplify lock/unlock pages").

So if the dirty pages which are under writeback got truncated partially
before we lock the dirty pages, we couldn't find all pages mapping to the
delalloc range, and the bug didn't return an error so it kept going on and
found that the delalloc range got truncated and got to unlock the dirty
pages, and then the ASSERT could caught the error, and showed

-----------------------------------------------------------------------------
assertion failed: page_ops & PAGE_LOCK, file: fs/btrfs/extent_io.c, line: 1716
-----------------------------------------------------------------------------

This fixes the bug by returning the proper -EAGAIN.

Cc: David Sterba <dsterba@suse.com>
Reported-by: Dave Jones <davej@codemonkey.org.uk>
Signed-off-by: Liu Bo <bo.li.liu@oracle.com>
Signed-off-by: David Sterba <dsterba@suse.com>

49d4a334

07 Mar, 2017 1 commit

btrfs: remove btrfs_err_str function from uapi/linux/btrfs.h · 68598d2e

Dmitry V. Levin authored Mar 01, 2017

btrfs_err_str function is not called from anywhere and is replicated
in the userspace headers for btrfs-progs.

It's removal also fixes the following linux/btrfs.h userspace
compilation error:

/usr/include/linux/btrfs.h: In function 'btrfs_err_str':
/usr/include/linux/btrfs.h:740:11: error: 'NULL' undeclared (first use in this function)
    return NULL;
Suggested-by: Jeff Mahoney <jeffm@suse.com>
Signed-off-by: Dmitry V. Levin <ldv@altlinux.org>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>

68598d2e

28 Feb, 2017 36 commits

Merge branch 'for-chris-4.11-part2' of... · e9f467d0

Chris Mason authored Feb 28, 2017

Merge branch 'for-chris-4.11-part2' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux into for-linus-4.11

e9f467d0

btrfs: add dummy callback for readpage_io_failed and drop checks · 20a7db8a

David Sterba authored Feb 17, 2017

Make extent_io_ops::readpage_io_failed_hook callback mandatory and
define a dummy function for btrfs_extent_io_ops. As the failed IO
callback is not performance critical, the branch vs extra trade off does
not hurt.
Signed-off-by: David Sterba <dsterba@suse.com>

20a7db8a

btrfs: drop checks for mandatory extent_io_ops callbacks · 20c9801d

David Sterba authored Feb 17, 2017

We know that eadpage_end_io_hook, submit_bio_hook and merge_bio_hook are
always defined so we can drop the checks before we call them.
Signed-off-by: David Sterba <dsterba@suse.com>

20c9801d

btrfs: document existence of extent_io ops callbacks · 4d53dddb

David Sterba authored Feb 17, 2017

Some of the callbacks defined in btree_extent_io_ops and
btrfs_extent_io_ops do always exist so we don't need to check the
existence before each call. This patch just reorders the definition and
documents which are mandatory/optional.
Signed-off-by: David Sterba <dsterba@suse.com>

4d53dddb

btrfs: let writepage_end_io_hook return void · c3988d63

David Sterba authored Feb 17, 2017

There's no error path in any of the instances, always return 0.
Reviewed-by: Liu Bo <bo.li.liu@oracle.com>
Signed-off-by: David Sterba <dsterba@suse.com>

c3988d63

btrfs: do proper error handling in btrfs_insert_xattr_item · b9d04c60

David Sterba authored Feb 17, 2017

The space check in btrfs_insert_xattr_item is duplicated in it's caller
(do_setxattr) so we won't hit the BUG_ON. Continuing without any check
could be disasterous so turn it to a proper error handling.
Reviewed-by: Liu Bo <bo.li.liu@oracle.com>
Signed-off-by: David Sterba <dsterba@suse.com>

b9d04c60

btrfs: handle allocation error in update_dev_stat_item · fa252992
David Sterba authored Feb 15, 2017
```
Reviewed-by: Liu Bo <bo.li.liu@oracle.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
fa252992

btrfs: remove BUG_ON from __tree_mod_log_insert · 047e5e17

David Sterba authored Feb 14, 2017

All callers dereference the 'tm' parameter before it gets to this
function, the NULL check does not make much sense here.
Reviewed-by: Liu Bo <bo.li.liu@oracle.com>
Signed-off-by: David Sterba <dsterba@suse.com>

047e5e17

btrfs: derive maximum output size in the compression implementation · e5d74902

David Sterba authored Feb 14, 2017

The value of max_out can be calculated from the parameters passed to the
compressors, which is number of pages and the page size, and we don't
have to needlessly pass it around.
Signed-off-by: David Sterba <dsterba@suse.com>

e5d74902

btrfs: use predefined limits for calculating maximum number of pages for compression · 069eac78
David Sterba authored Feb 14, 2017
```
Signed-off-by: David Sterba <dsterba@suse.com>
```
069eac78

btrfs: export compression buffer limits in a header · ff763866

David Sterba authored Feb 14, 2017

Move the buffer limit definitions out of compress_file_range.
Reviewed-by: Qu Wenruo <quwenruo@cn.fujitsu.com>
Signed-off-by: David Sterba <dsterba@suse.com>

ff763866

btrfs: merge nr_pages input and output parameter in compress_pages · 4d3a800e

David Sterba authored Feb 14, 2017

The parameter saying how many pages can be allocated at maximum can be
merged with the output page counter, to save some stack space.  The
compression implementation will sink the parameter to a local variable
so everything works as before.

The nr_pages variables can also be simply merged in compress_file_range
into one.
Signed-off-by: David Sterba <dsterba@suse.com>

4d3a800e

btrfs: merge length input and output parameter in compress_pages · 38c31464

David Sterba authored Feb 14, 2017

The length parameter is basically duplicated for input and output in the
top level caller of the compress_pages chain. We can simply use one
variable for that and reduce stack consumption. The compression
implementation will sink the parameter to a local variable so everything
works as before.
Signed-off-by: David Sterba <dsterba@suse.com>

38c31464

btrfs: constify name of subvolume in creation helpers · 52f75f4f
David Sterba authored Feb 14, 2017
```
Signed-off-by: David Sterba <dsterba@suse.com>
```
52f75f4f
btrfs: constify buffers used by compression helpers · 14a3357b
David Sterba authored Feb 14, 2017
```
Signed-off-by: David Sterba <dsterba@suse.com>
```
14a3357b

btrfs: constify input buffer of btrfs_csum_data · 9ed57367

David Sterba authored Feb 14, 2017

The function does not modify the input buffer, also update a typecast in
one caller.
Signed-off-by: David Sterba <dsterba@suse.com>

9ed57367

btrfs: constify device path passed to relevant helpers · da353f6b
David Sterba authored Feb 14, 2017
```
Signed-off-by: David Sterba <dsterba@suse.com>
```
da353f6b
btrfs: make btrfs_inode_resume_unlocked_dio take btrfs_inode · 0b581701
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
0b581701
btrfs: make btrfs_inode_block_unlocked_dio take btrfs_inode · abcefb1e
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
abcefb1e

btrfs: Make btrfs_add_nondir take btrfs_inode · cef415af

Nikolay Borisov authored Feb 20, 2017

Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>

cef415af

btrfs: Make btrfs_add_link take btrfs_inode · db0a669f

Nikolay Borisov authored Feb 20, 2017

Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>

db0a669f

btrfs: Make btrfs_del_delalloc_inode take btrfs_inode · 9e3e97f4
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
9e3e97f4

btrfs: Make get_extent_t take btrfs_inode · fc4f21b1

Nikolay Borisov authored Feb 20, 2017

In addition to changing the signature, this patch also switches
all the functions which are used as an argument to also take btrfs_inode.
Namely those are: btrfs_get_extent and btrfs_get_extent_filemap.
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>

fc4f21b1

btrfs: Make check_extent_to_block take btrfs_inode · 1c8c9c52
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
1c8c9c52
btrfs: Make clone_update_extent_map take btrfs_inode · a2f392e4
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
a2f392e4
btrfs: Make btrfs_clear_bit_hook take btrfs_inode · 6fc0ef68
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
6fc0ef68
btrfs: Make btrfs_extent_item_to_extent_map take btrfs_inode · 9cdc5124
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
9cdc5124
btrfs: make btrfs_log_inode_parent take btrfs_inode · 19df27a9
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
19df27a9
btrfs: Make check_parent_dirs_for_sync take btrfs_inode · aefa6115
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
aefa6115

btrfs: Make btrfs_orphan_add take btrfs_inode · 73f2e545

Nikolay Borisov authored Feb 20, 2017

Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>

73f2e545

btrfs: make btrfs_orphan_del take btrfs_inode · 3d6ae7bb

Nikolay Borisov authored Feb 20, 2017

Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>

3d6ae7bb

btrfs: make btrfs_free_io_failure_record take btrfs_inode · 7ab7956e
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
7ab7956e

btrfs: make clean_io_failure take btrfs_inode · b30cb441

Nikolay Borisov authored Feb 20, 2017

Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>

b30cb441

btrfs: make repair_io_failure take btrfs_inode · 9d4f7f8a

Nikolay Borisov authored Feb 20, 2017

Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>

9d4f7f8a

btrfs: make check_compressed_csum take btrfs_inode · f898ac6a
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
f898ac6a
btrfs: make btrfs_print_data_csum_error take btrfs_inode · 0970a22e
Nikolay Borisov authored Feb 20, 2017
```
Signed-off-by: Nikolay Borisov <nborisov@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
```
0970a22e