Commits · fc187514d8af3a676c7bd7922439f9f5e5c6223f · Kirill Smelkov / linux

18 Oct, 2018 3 commits

nfs: remove redundant call to nfs_context_set_write_error() · fc187514

Benjamin Coddington authored Oct 18, 2018

We don't need to call this in the direct, read, or pnfs resend paths and
the only other caller is the write path in nfs_page_async_flush() which
already checks and sets the pg_error on the context.
Signed-off-by: Benjamin Coddington <bcodding@redhat.com>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

fc187514

nfs: Fix a missed page unlock after pg_doio() · fdbd1a2e

Benjamin Coddington authored Oct 18, 2018

We must check pg_error and call error_cleanup after any call to pg_doio.
Currently, we are skipping the unlock of a page if we encounter an error in
nfs_pageio_complete() before handing off the work to the RPC layer.
Signed-off-by: Benjamin Coddington <bcodding@redhat.com>
Cc: stable@vger.kernel.org
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

fdbd1a2e

SUNRPC: Fix a compile warning for cmpxchg64() · e732f448
Trond Myklebust authored Oct 18, 2018
```
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
```
e732f448

05 Oct, 2018 2 commits

NFSv4.x: fix lock recovery during delegation recall · 44f411c3

Olga Kornievskaia authored Oct 04, 2018

Running "./nfstest_delegation --runtest recall26" uncovers that
client doesn't recover the lock when we have an appending open,
where the initial open got a write delegation.

Instead of checking for the passed in open context against
the file lock's open context. Check that the state is the same.
Signed-off-by: Olga Kornievskaia <kolga@netapp.com>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

44f411c3

SUNRPC: use cmpxchg64() in gss_seq_send64_fetch_and_inc() · 21924765

Arnd Bergmann authored Oct 02, 2018

The newly introduced gss_seq_send64_fetch_and_inc() fails to build on
32-bit architectures:

net/sunrpc/auth_gss/gss_krb5_seal.c:144:14: note: in expansion of macro 'cmpxchg'
   seq_send = cmpxchg(&ctx->seq_send64, old, old + 1);
              ^~~~~~~
arch/x86/include/asm/cmpxchg.h:128:3: error: call to '__cmpxchg_wrong_size' declared with attribute error: Bad argument size for cmpxchg
   __cmpxchg_wrong_size();     \

As the message tells us, cmpxchg() cannot be used on 64-bit arguments,
that's what cmpxchg64() does.

Fixes: 571ed1fd ("SUNRPC: Replace krb5_seq_lock with a lockless scheme")
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

21924765

30 Sep, 2018 35 commits

NFSv4: Fix lookup revalidate of regular files · c7944ebb

Trond Myklebust authored Sep 28, 2018

If we're revalidating an existing dentry in order to open a file, we need
to ensure that we check the directory has not changed before we optimise
away the lookup.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

c7944ebb

NFS: Refactor nfs_lookup_revalidate() · 5ceb9d7f

Trond Myklebust authored Sep 28, 2018

Refactor the code in nfs_lookup_revalidate() as a stepping stone towards
optimising and fixing nfs4_lookup_revalidate().
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

5ceb9d7f

NFS: Fix dentry revalidation on NFSv4 lookup · be189f7e

Trond Myklebust authored Sep 27, 2018

We need to ensure that inode and dentry revalidation occurs correctly
on reopen of a file that is already open. Currently, we can end up
not revalidating either in the case of NFSv4.0, due to the 'cached open'
path.
Let's fix that by ensuring that we only do cached open for the special
cases of open recovery and delegation return.
Reported-by: Stan Hu <stanhu@gmail.com>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

be189f7e

SUNRPC: Replace krb5_seq_lock with a lockless scheme · 571ed1fd
Trond Myklebust authored Sep 29, 2018
```
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
```
571ed1fd

SUNRPC: Lockless lookup of RPCSEC_GSS mechanisms · 0c1c19f4

Trond Myklebust authored Sep 29, 2018

Use RCU protected lookups for discovering the supported mechanisms.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

0c1c19f4

SUNRPC: Remove rpc_authflavor_lock in favour of RCU locking · 4e4c3bef

Trond Myklebust authored Sep 27, 2018

Module removal is RCU safe by design, so we really have no need to
lock the auth_flavors[] array. Substitute a lockless scheme to
add/remove entries in the array, and then use rcu.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

4e4c3bef

NFS: Remove private spinlock in struct nfs_pgio_header · 1c6c4b74

Trond Myklebust authored Sep 25, 2018

Now that each struct nfs_pgio_header corresponds to one RPC call, we
only have one writer to the struct nfs_pgio_header.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

1c6c4b74

NFSv4: Save a few bytes in the nfs_pgio_args/res · 28d52235

Trond Myklebust authored Sep 24, 2018

Save a few bytes by allowing the read/write specific fields of the
structures to share storage.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

28d52235

NFSv3: Improve NFSv3 performance when server returns no post-op attributes · 8d8928d8

Trond Myklebust authored Mar 05, 2018

When the server fails to return post-op attributes, the client's
attempt to place read data directly in the page cache fails, and
so we have to do an extra copy in order to realign the data with
page borders.
This patch attempts to detect servers that don't return post-op
attributes on read (e.g. for pNFS) and adjusts the placement
calculation accordingly.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

8d8928d8

NFSv4: Split out NFS v4.2 copy completion functions · 80f42368

Anna Schumaker authored Sep 20, 2018

The convention in the rest of the code is to have a separate function
for anything that might be ifdef-ed out.
Signed-off-by: Anna Schumaker <Anna.Schumaker@Netapp.com>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

80f42368

NFS: Reduce indentation of nfs4_recovery_handle_error() · 000d3f95

Anna Schumaker authored Sep 11, 2018

This is to match kernel coding style for switch statements.
Signed-off-by: Anna Schumaker <Anna.Schumaker@Netapp.com>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

000d3f95

NFS: Reduce indentation of the switch statement in nfs4_reclaim_open_state() · 35a61606

Anna Schumaker authored Sep 11, 2018

Most places in the kernel tend to line up cases with the switch to
reduce indentation, so move this over to match that style.
Additionally, I handle the (status >= 0) case in the switch so that we
only "goto restart" from a single place after error handling.
Signed-off-by: Anna Schumaker <Anna.Schumaker@Netapp.com>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

35a61606

NFS: Split out the body of nfs4_reclaim_open_state() · cb7a8384

Anna Schumaker authored Sep 11, 2018

Moving all of this into a new function removes the need for cramped
indentation, making the code overall easier to look at.   I also take
this chance to switch copy recovery over to using
nfs4_stateid_match_other()
Signed-off-by: Anna Schumaker <Anna.Schumaker@Netapp.com>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

cb7a8384

nfs4: flex_file: ignore synthetic uid/gid for tightly coupled DSes · 10ec57e4

Tigran Mkrtchyan authored Aug 20, 2018

for tightly coupled DSes client must ignore provided synthetic uid and
gid as stated in draft-ietf-nfsv4-flex-files-19#section-5.1.
Signed-off-by: Tigran Mkrtchyan <tigran.mkrtchyan@desy.de>
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

10ec57e4

NFSv4.1: Fix the r/wsize checking · 943cff67

Trond Myklebust authored Sep 18, 2018

The intention of nfs4_session_set_rwsize() was to cap the r/wsize to the
buffer sizes negotiated by the CREATE_SESSION. The initial code had a
bug whereby we would not check the values negotiated by nfs_probe_fsinfo()
(the assumption being that CREATE_SESSION will always negotiate buffer values
that are sane w.r.t. the server's preferred r/wsizes) but would only check
values set by the user in the 'mount' command.

The code was changed in 4.11 to _always_ set the r/wsize, meaning that we
now never use the server preferred r/wsizes. This is the regression that
this patch fixes.
Also rename the function to nfs4_session_limit_rwsize() in order to avoid
future confusion.

Fixes: 03385332 (NFSv4.1 respect server's max size in CREATE_SESSION")
Cc: stable@vger.kernel.org # v4.11+
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

943cff67

NFSv4: Convert struct nfs4_state to use refcount_t · ace9fad4
Trond Myklebust authored Sep 02, 2018
```
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
```
ace9fad4

NFSv4: Convert open state lookup to use RCU · 9ae075fd

Trond Myklebust authored Sep 02, 2018

Further reduce contention on the inode->i_lock.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

9ae075fd

NFS: Convert lookups of the open context to RCU · 0de43976

Trond Myklebust authored Sep 02, 2018

Reduce contention on the inode->i_lock by ensuring that we use RCU
when looking up the NFS open context.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

0de43976

NFS: Simplify internal check for whether file is open for write · 6ba0c4e5
Trond Myklebust authored Sep 02, 2018
```
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
```
6ba0c4e5

NFS: Convert lookups of the lock context to RCU · 1db97eaa

Trond Myklebust authored Sep 02, 2018

Speed up lookups of an existing lock context by avoiding the inode->i_lock,
and using RCU instead.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

1db97eaa

pNFS: Don't allocate more pages than we need to fit a layoutget response · 28ced9a8

Trond Myklebust authored Sep 03, 2018

For the 'files' and 'flexfiles' layout types, we do not expect the reply
to be any larger than 4k. The block and scsi layout types are a little more
greedy, so we keep allocating the maximum response size for now.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

28ced9a8

pNFS: Don't zero out the array in nfs4_alloc_pages() · a2791d3a

Trond Myklebust authored Sep 03, 2018

We don't need a zeroed out array, since it is immediately being filled.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

a2791d3a

SUNRPC: Unexport xdr_partial_copy_from_skb() · ec846469

Trond Myklebust authored Sep 14, 2018

It is no longer used outside of net/sunrpc/socklib.c
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

ec846469

SUNRPC: Clean up xs_udp_data_receive() · 4f546149

Trond Myklebust authored Sep 14, 2018

Simplify the retry logic.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

4f546149

SUNRPC: Allow AF_LOCAL sockets to use the generic stream receive · 550aebfe
Trond Myklebust authored Sep 14, 2018
```
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
```
550aebfe
SUNRPC: Clean up - rename xs_tcp_data_receive() to xs_stream_data_receive() · c50b8ee0
Trond Myklebust authored Sep 14, 2018
```
In preparation for sharing with AF_LOCAL.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
```
c50b8ee0

SUNRPC: Simplify TCP receive code by switching to using iterators · 277e4ab7

Trond Myklebust authored Sep 14, 2018

Most of this code should also be reusable with other socket types.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

277e4ab7

SUNRPC: Add a bvec array to struct xdr_buf for use with iovec_iter() · 9d96acbc

Trond Myklebust authored Sep 13, 2018

Add a bvec array to struct xdr_buf, and have the client allocate it
when we need to receive data into pages.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

9d96acbc

SUNRPC: Add a label for RPC calls that require allocation on receive · 431f6eb3

Trond Myklebust authored Sep 16, 2018

If the RPC call relies on the receive call allocating pages as buffers,
then let's label it so that we
a) Don't leak memory by allocating pages for requests that do not expect
   this behaviour
b) Can optimise for the common case where calls do not require allocation.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

431f6eb3

SUNRPC: Convert the xprt->sending queue back to an ordinary wait queue · 79c99152

Trond Myklebust authored Sep 09, 2018

We no longer need priority semantics on the xprt->sending queue, because
the order in which tasks are sent is now dictated by their position in
the send queue.
Note that the backlog queue remains a priority queue, meaning that
slot resources are still managed in order of task priority.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

79c99152

SUNRPC: Fix priority queue fairness · f42f7c28

Trond Myklebust authored Sep 08, 2018

Fix up the priority queue to not batch by owner, but by queue, so that
we allow '1 << priority' elements to be dequeued before switching to
the next priority queue.
The owner field is still used to wake up requests in round robin order
by owner to avoid single processes hogging the RPC layer by loading the
queues.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

f42f7c28

SUNRPC: Convert xprt receive queue to use an rbtree · 95f7691d

Trond Myklebust authored Sep 07, 2018

If the server is slow, we can find ourselves with quite a lot of entries
on the receive queue. Converting the search from an O(n) to O(log(n))
can make a significant difference, particularly since we have to hold
a number of locks while searching.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

95f7691d

SUNRPC: Don't take transport->lock unnecessarily when taking XPRT_LOCK · bd79bc57
Trond Myklebust authored Sep 07, 2018
```
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
```
bd79bc57
SUNRPC: Cleanup: remove the unused 'task' argument from the request_send() · adfa7144
Trond Myklebust authored Sep 03, 2018
```
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>
```
adfa7144

SUNRPC: Clean up transport write space handling · c544577d

Trond Myklebust authored Sep 03, 2018

Treat socket write space handling in the same way we now treat transport
congestion: by denying the XPRT_LOCK until the transport signals that it
has free buffer space.
Signed-off-by: Trond Myklebust <trond.myklebust@hammerspace.com>

c544577d