linux.git/fs/proc, branch v5.10-rc2

Merge branch 'work.set_fs' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs

2020-10-22T16:59:21+00:00

Pull initial set_fs() removal from Al Viro:
 "Christoph's set_fs base series + fixups"

* 'work.set_fs' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs:
  fs: Allow a NULL pos pointer to __kernel_read
  fs: Allow a NULL pos pointer to __kernel_write
  powerpc: remove address space overrides using set_fs()
  powerpc: use non-set_fs based maccess routines
  x86: remove address space overrides using set_fs()
  x86: make TASK_SIZE_MAX usable from assembly code
  x86: move PAGE_OFFSET, TASK_SIZE & friends to page_{32,64}_types.h
  lkdtm: remove set_fs-based tests
  test_bitmap: remove user bitmap tests
  uaccess: add infrastructure for kernel builds with set_fs()
  fs: don't allow splice read/write without explicit ops
  fs: don't allow kernel reads and writes without iter ops
  sysctl: Convert to iter interfaces
  proc: add a read_iter method to proc proc_ops
  proc: cleanup the compat vs no compat file ops
  proc: remove a level of indentation in proc_get_inode

io-wq: inherit audit loginuid and sessionid

2020-10-17T15:25:47+00:00

Make sure the async io-wq workers inherit the loginuid and sessionid from
the original task, and restore them to unset once we're done with the
async work item.

While at it, disable the ability for kernel threads to write to their own
loginuid.

Signed-off-by: Jens Axboe

mm: remove the now-unnecessary mmget_still_valid() hack

2020-10-16T18:11:22+00:00

The preceding patches have ensured that core dumping properly takes the
mmap_lock.  Thanks to that, we can now remove mmget_still_valid() and all
its users.

Signed-off-by: Jann Horn 
Signed-off-by: Andrew Morton 
Acked-by: Linus Torvalds 
Cc: Christoph Hellwig 
Cc: Alexander Viro 
Cc: "Eric W . Biederman" 
Cc: Oleg Nesterov 
Cc: Hugh Dickins 
Link: http://lkml.kernel.org/r/20200827114932.3572699-8-jannh@google.com
Signed-off-by: Linus Torvalds

mm, oom_adj: don't loop through tasks in __set_oom_adj when not necessary

2020-10-14T01:38:35+00:00

Currently __set_oom_adj loops through all processes in the system to keep
oom_score_adj and oom_score_adj_min in sync between processes sharing
their mm.  This is done for any task with more that one mm_users, which
includes processes with multiple threads (sharing mm and signals).
However for such processes the loop is unnecessary because their signal
structure is shared as well.

Android updates oom_score_adj whenever a tasks changes its role
(background/foreground/...) or binds to/unbinds from a service, making it
more/less important.  Such operation can happen frequently.  We noticed
that updates to oom_score_adj became more expensive and after further
investigation found out that the patch mentioned in "Fixes" introduced a
regression.  Using Pixel 4 with a typical Android workload, write time to
oom_score_adj increased from ~3.57us to ~362us.  Moreover this regression
linearly depends on the number of multi-threaded processes running on the
system.

Mark the mm with a new MMF_MULTIPROCESS flag bit when task is created with
(CLONE_VM && !CLONE_THREAD && !CLONE_VFORK).  Change __set_oom_adj to use
MMF_MULTIPROCESS instead of mm_users to decide whether oom_score_adj
update should be synchronized between multiple processes.  To prevent
races between clone() and __set_oom_adj(), when oom_score_adj of the
process being cloned might be modified from userspace, we use
oom_adj_mutex.  Its scope is changed to global.

The combination of (CLONE_VM && !CLONE_THREAD) is rarely used except for
the case of vfork().  To prevent performance regressions of vfork(), we
skip taking oom_adj_mutex and setting MMF_MULTIPROCESS when CLONE_VFORK is
specified.  Clearing the MMF_MULTIPROCESS flag (when the last process
sharing the mm exits) is left out of this patch to keep it simple and
because it is believed that this threading model is rare.  Should there
ever be a need for optimizing that case as well, it can be done by hooking
into the exit path, likely following the mm_update_next_owner pattern.

With the combination of (CLONE_VM && !CLONE_THREAD && !CLONE_VFORK) being
quite rare, the regression is gone after the change is applied.

[surenb@google.com: v3]
  Link: https://lkml.kernel.org/r/20200902012558.2335613-1-surenb@google.com

Fixes: 44a70adec910 ("mm, oom_adj: make sure processes sharing mm have same view of oom_score_adj")
Reported-by: Tim Murray 
Suggested-by: Michal Hocko 
Signed-off-by: Suren Baghdasaryan 
Signed-off-by: Andrew Morton 
Acked-by: Christian Brauner 
Acked-by: Michal Hocko 
Acked-by: Oleg Nesterov 
Cc: Ingo Molnar 
Cc: Peter Zijlstra 
Cc: Thomas Gleixner 
Cc: Eugene Syromiatnikov 
Cc: Christian Kellner 
Cc: Adrian Reber 
Cc: Shakeel Butt 
Cc: Aleksa Sarai 
Cc: Alexey Dobriyan 
Cc: "Eric W. Biederman" 
Cc: Alexey Gladkov 
Cc: Michel Lespinasse 
Cc: Daniel Jordan 
Cc: Andrei Vagin 
Cc: Bernd Edlinger 
Cc: John Johansen 
Cc: Yafang Shao 
Link: https://lkml.kernel.org/r/20200824153036.3201505-1-surenb@google.com
Debugged-by: Minchan Kim 
Signed-off-by: Linus Torvalds

mm: proc: smaps_rollup: do not stall write attempts on mmap_lock

2020-10-14T01:38:31+00:00

smaps_rollup will try to grab mmap_lock and go through the whole vma list
until it finishes the iterating.  When encountering large processes, the
mmap_lock will be held for a longer time, which may block other write
requests like mmap and munmap from progressing smoothly.

There are upcoming mmap_lock optimizations like range-based locks, but the
lock applied to smaps_rollup would be the coarse type, which doesn't avoid
the occurrence of unpleasant contention.

To solve aforementioned issue, we add a check which detects whether anyone
wants to grab mmap_lock for write attempts.

Signed-off-by: Chinwen Chang 
Signed-off-by: Andrew Morton 
Cc: Steven Price 
Cc: Michel Lespinasse 
Cc: Matthias Brugger 
Cc: Vlastimil Babka 
Cc: Daniel Jordan 
Cc: Davidlohr Bueso 
Cc: Chinwen Chang 
Cc: Alexey Dobriyan 
Cc: "Matthew Wilcox (Oracle)" 
Cc: Jason Gunthorpe 
Cc: Song Liu 
Cc: Jimmy Assarsson 
Cc: Huang Ying 
Cc: Daniel Kiss 
Cc: Laurent Dufour 
Link: http://lkml.kernel.org/r/1597715898-3854-4-git-send-email-chinwen.chang@mediatek.com
Signed-off-by: Linus Torvalds

mm: smaps*: extend smap_gather_stats to support specified beginning

2020-10-14T01:38:31+00:00

Extend smap_gather_stats to support indicated beginning address at which
it should start gathering.  To achieve the goal, we add a new parameter
@start assigned by the caller and try to refactor it for simplicity.

If @start is 0, it will use the range of @vma for gathering.

Signed-off-by: Chinwen Chang 
Signed-off-by: Andrew Morton 
Reviewed-by: Steven Price 
Cc: Michel Lespinasse 
Cc: Alexey Dobriyan 
Cc: Daniel Jordan 
Cc: Daniel Kiss 
Cc: Davidlohr Bueso 
Cc: Huang Ying 
Cc: Jason Gunthorpe 
Cc: Jimmy Assarsson 
Cc: Laurent Dufour 
Cc: "Matthew Wilcox (Oracle)" 
Cc: Matthias Brugger 
Cc: Song Liu 
Cc: Vlastimil Babka 
Link: http://lkml.kernel.org/r/1597715898-3854-3-git-send-email-chinwen.chang@mediatek.com
Signed-off-by: Linus Torvalds

proc: optimise smaps for shmem entries

2020-10-14T01:38:29+00:00

Avoid bumping the refcount on pages when we're only interested in the
swap entries.

Signed-off-by: Matthew Wilcox (Oracle) 
Signed-off-by: Andrew Morton 
Acked-by: Johannes Weiner 
Cc: Alexey Dobriyan 
Cc: Chris Wilson 
Cc: Huang Ying 
Cc: Hugh Dickins 
Cc: Jani Nikula 
Cc: Matthew Auld 
Cc: William Kucharski 
Link: https://lkml.kernel.org/r/20200910183318.20139-5-willy@infradead.org
Signed-off-by: Linus Torvalds

sysctl: Convert to iter interfaces

2020-09-09T02:20:39+00:00

Using the read_iter/write_iter interfaces allows for in-kernel users
to set sysctls without using set_fs().  Also, the buffer is a string,
so give it the real type of 'char *', not void *.

[AV: Christoph's fixup folded in]

Signed-off-by: Matthew Wilcox (Oracle) 
Signed-off-by: Christoph Hellwig 
Signed-off-by: Al Viro

arm64: mte: Add PROT_MTE support to mmap() and mprotect()

2020-09-04T11:46:07+00:00

To enable tagging on a memory range, the user must explicitly opt in via
a new PROT_MTE flag passed to mmap() or mprotect(). Since this is a new
memory type in the AttrIndx field of a pte, simplify the or'ing of these
bits over the protection_map[] attributes by making MT_NORMAL index 0.

There are two conditions for arch_vm_get_page_prot() to return the
MT_NORMAL_TAGGED memory type: (1) the user requested it via PROT_MTE,
registered as VM_MTE in the vm_flags, and (2) the vma supports MTE,
decided during the mmap() call (only) and registered as VM_MTE_ALLOWED.

arch_calc_vm_prot_bits() is responsible for registering the user request
as VM_MTE. The newly introduced arch_calc_vm_flag_bits() sets
VM_MTE_ALLOWED if the mapping is MAP_ANONYMOUS. An MTE-capable
filesystem (RAM-based) may be able to set VM_MTE_ALLOWED during its
mmap() file ops call.

In addition, update VM_DATA_DEFAULT_FLAGS to allow mprotect(PROT_MTE) on
stack or brk area.

The Linux mmap() syscall currently ignores unknown PROT_* flags. In the
presence of MTE, an mmap(PROT_MTE) on a file which does not support MTE
will not report an error and the memory will not be mapped as Normal
Tagged. For consistency, mprotect(PROT_MTE) will not report an error
either if the memory range does not support MTE. Two subsequent patches
in the series will propose tightening of this behaviour.

Co-developed-by: Vincenzo Frascino 
Signed-off-by: Vincenzo Frascino 
Signed-off-by: Catalin Marinas 
Cc: Will Deacon

mm: Add PG_arch_2 page flag

2020-09-04T11:46:06+00:00

For arm64 MTE support it is necessary to be able to mark pages that
contain user space visible tags that will need to be saved/restored e.g.
when swapped out.

To support this add a new arch specific flag (PG_arch_2). This flag is
only available on 64-bit architectures due to the limited number of
spare page flags on the 32-bit ones.

Signed-off-by: Steven Price 
[catalin.marinas@arm.com: use CONFIG_64BIT for guarding this new flag]
Signed-off-by: Catalin Marinas 
Cc: Andrew Morton