linux-stable.git/include/linux/mm_types.h, branch v4.4.76

mm: Add a user_ns owner to mm_struct and fix ptrace permission checks

2017-01-06T10:16:11+00:00

commit bfedb589252c01fa505ac9f6f2a3d5d68d707ef4 upstream.

During exec dumpable is cleared if the file that is being executed is
not readable by the user executing the file.  A bug in
ptrace_may_access allows reading the file if the executable happens to
enter into a subordinate user namespace (aka clone(CLONE_NEWUSER),
unshare(CLONE_NEWUSER), or setns(fd, CLONE_NEWUSER).

This problem is fixed with only necessary userspace breakage by adding
a user namespace owner to mm_struct, captured at the time of exec, so
it is clear in which user namespace CAP_SYS_PTRACE must be present in
to be able to safely give read permission to the executable.

The function ptrace_may_access is modified to verify that the ptracer
has CAP_SYS_ADMIN in task->mm->user_ns instead of task->cred->user_ns.
This ensures that if the task changes it's cred into a subordinate
user namespace it does not become ptraceable.

The function ptrace_attach is modified to only set PT_PTRACE_CAP when
CAP_SYS_PTRACE is held over task->mm->user_ns.  The intent of
PT_PTRACE_CAP is to be a flag to note that whatever permission changes
the task might go through the tracer has sufficient permissions for
it not to be an issue.  task->cred->user_ns is always the same
as or descendent of mm->user_ns.  Which guarantees that having
CAP_SYS_PTRACE over mm->user_ns is the worst case for the tasks
credentials.

To prevent regressions mm->dumpable and mm->user_ns are not considered
when a task has no mm.  As simply failing ptrace_may_attach causes
regressions in privileged applications attempting to read things
such as /proc//stat

Acked-by: Kees Cook 
Tested-by: Cyrill Gorcunov 
Fixes: 8409cca70561 ("userns: allow ptrace from non-init user namespaces")
Signed-off-by: "Eric W. Biederman" 
Signed-off-by: Greg Kroah-Hartman

mm: use 'unsigned int' for compound_dtor/compound_order on 64BIT

2015-11-07T01:50:42+00:00

On 64 bit system we have enough space in struct page to encode
compound_dtor and compound_order with unsigned int.

On x86-64 it leads to slightly smaller code size due usesage of plain
MOV instead of MOVZX (zero-extended move) or similar effect.

allyesconfig:

   text	   data	    bss	    dec	    hex	filename
159520446	48146736	72196096	279863278	10ae5fee	vmlinux.pre
159520382	48146736	72196096	279863214	10ae5fae	vmlinux.post

On other architectures without native support of 16-bit data types the

Signed-off-by: Kirill A. Shutemov 
Acked-by: Michal Hocko 
Reviewed-by: Andrea Arcangeli 
Cc: "Paul E. McKenney" 
Cc: Andi Kleen 
Cc: Aneesh Kumar K.V 
Cc: Christoph Lameter 
Cc: David Rientjes 
Cc: Hugh Dickins 
Cc: Joonsoo Kim 
Cc: Sergey Senozhatsky 
Cc: Vlastimil Babka 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: make compound_head() robust

2015-11-07T01:50:42+00:00

Hugh has pointed that compound_head() call can be unsafe in some
context. There's one example:

	CPU0					CPU1

isolate_migratepages_block()
  page_count()
    compound_head()
      !!PageTail() == true
					put_page()
					  tail->first_page = NULL
      head = tail->first_page
					alloc_pages(__GFP_COMP)
					   prep_compound_page()
					     tail->first_page = head
					     __SetPageTail(p);
      !!PageTail() == true
    

The race is pure theoretical. I don't it's possible to trigger it in
practice. But who knows.

We can fix the race by changing how encode PageTail() and compound_head()
within struct page to be able to update them in one shot.

The patch introduces page->compound_head into third double word block in
front of compound_dtor and compound_order. Bit 0 encodes PageTail() and
the rest bits are pointer to head page if bit zero is set.

The patch moves page->pmd_huge_pte out of word, just in case if an
architecture defines pgtable_t into something what can have the bit 0
set.

hugetlb_cgroup uses page->lru.next in the second tail page to store
pointer struct hugetlb_cgroup. The patch switch it to use page->private
in the second tail page instead. The space is free since ->first_page is
removed from the union.

The patch also opens possibility to remove HUGETLB_CGROUP_MIN_ORDER
limitation, since there's now space in first tail page to store struct
hugetlb_cgroup pointer. But that's out of scope of the patch.

That means page->compound_head shares storage space with:

 - page->lru.next;
 - page->next;
 - page->rcu_head.next;

That's too long list to be absolutely sure, but looks like nobody uses
bit 0 of the word.

page->rcu_head.next guaranteed[1] to have bit 0 clean as long as we use
call_rcu(), call_rcu_bh(), call_rcu_sched(), or call_srcu(). But future
call_rcu_lazy() is not allowed as it makes use of the bit and we can
get false positive PageTail().

[1] http://lkml.kernel.org/g/20150827163634.GD4029@linux.vnet.ibm.com

Signed-off-by: Kirill A. Shutemov 
Acked-by: Michal Hocko 
Reviewed-by: Andrea Arcangeli 
Cc: Hugh Dickins 
Cc: David Rientjes 
Cc: Vlastimil Babka 
Acked-by: Paul E. McKenney 
Cc: Aneesh Kumar K.V 
Cc: Andi Kleen 
Cc: Christoph Lameter 
Cc: Joonsoo Kim 
Cc: Sergey Senozhatsky 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: pack compound_dtor and compound_order into one word in struct page

2015-11-07T01:50:42+00:00

The patch halves space occupied by compound_dtor and compound_order in
struct page.

For compound_order, it's trivial long -> short conversion.

For get_compound_page_dtor(), we now use hardcoded table for destructor
lookup and store its index in the struct page instead of direct pointer
to destructor. It shouldn't be a big trouble to maintain the table: we
have only two destructor and NULL currently.

This patch free up one word in tail pages for reuse. This is preparation
for the next patch.

Signed-off-by: Kirill A. Shutemov 
Reviewed-by: Michal Hocko 
Acked-by: Vlastimil Babka 
Reviewed-by: Andrea Arcangeli 
Cc: "Paul E. McKenney" 
Cc: Andi Kleen 
Cc: Aneesh Kumar K.V 
Cc: Christoph Lameter 
Cc: David Rientjes 
Cc: Hugh Dickins 
Cc: Joonsoo Kim 
Cc: Sergey Senozhatsky 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: drop page->slab_page

2015-11-07T01:50:42+00:00

Since 8456a648cf44 ("slab: use struct page for slab management") nobody
uses slab_page field in struct page.

Let's drop it.

Signed-off-by: Kirill A. Shutemov 
Acked-by: Christoph Lameter 
Acked-by: David Rientjes 
Acked-by: Vlastimil Babka 
Reviewed-by: Andrea Arcangeli 
Cc: Joonsoo Kim 
Cc: Andi Kleen 
Cc: "Paul E. McKenney" 
Cc: Aneesh Kumar K.V 
Cc: Hugh Dickins 
Cc: Michal Hocko 
Cc: Sergey Senozhatsky 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: hugetlb: proc: add HugetlbPages field to /proc/PID/status

2015-11-06T03:34:48+00:00

Currently there's no easy way to get per-process usage of hugetlb pages,
which is inconvenient because userspace applications which use hugetlb
typically want to control their processes on the basis of how much memory
(including hugetlb) they use.  So this patch simply provides easy access
to the info via /proc/PID/status.

Signed-off-by: Naoya Horiguchi 
Acked-by: Joern Engel 
Acked-by: David Rientjes 
Acked-by: Michal Hocko 
Cc: Mike Kravetz 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: drop __nocast from vm_flags_t definition

2015-09-08T22:35:28+00:00

__nocast does no good for vm_flags_t. It only produces useless sparse
warnings.

Let's drop it.

Signed-off-by: Kirill A. Shutemov 
Cc: Oleg Nesterov 
Acked-by: David Rientjes 
Cc: Johannes Weiner 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

x86, mm: trace when an IPI is about to be sent

2015-09-04T23:54:41+00:00

When unmapping pages it is necessary to flush the TLB.  If that page was
accessed by another CPU then an IPI is used to flush the remote CPU.  That
is a lot of IPIs if kswapd is scanning and unmapping >100K pages per
second.

There already is a window between when a page is unmapped and when it is
TLB flushed.  This series increases the window so multiple pages can be
flushed using a single IPI.  This should be safe or the kernel is hosed
already.

Patch 1 simply made the rest of the series easier to write as ftrace
        could identify all the senders of TLB flush IPIS.

Patch 2 tracks what CPUs potentially map a PFN and then sends an IPI
        to flush the entire TLB.

Patch 3 tracks when there potentially are writable TLB entries that
        need to be batched differently

Patch 4 increases SWAP_CLUSTER_MAX to further batch flushes

The performance impact is documented in the changelogs but in the optimistic
case on a 4-socket machine the full series reduces interrupts from 900K
interrupts/second to 60K interrupts/second.

This patch (of 4):

It is easy to trace when an IPI is received to flush a TLB but harder to
detect what event sent it.  This patch makes it easy to identify the
source of IPIs being transmitted for TLB flushes on x86.

Signed-off-by: Mel Gorman 
Reviewed-by: Rik van Riel 
Reviewed-by: Dave Hansen 
Acked-by: Ingo Molnar 
Cc: Linus Torvalds 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

userfaultfd: add vm_userfaultfd_ctx to the vm_area_struct

2015-09-04T23:54:41+00:00

This adds the vm_userfaultfd_ctx to the vm_area_struct.

Signed-off-by: Andrea Arcangeli 
Acked-by: Pavel Emelyanov 
Cc: Sanidhya Kashyap 
Cc: zhang.zhanghailiang@huawei.com
Cc: "Kirill A. Shutemov" 
Cc: Andres Lagar-Cavilla 
Cc: Dave Hansen 
Cc: Paolo Bonzini 
Cc: Rik van Riel 
Cc: Mel Gorman 
Cc: Andy Lutomirski 
Cc: Hugh Dickins 
Cc: Peter Feiner 
Cc: "Dr. David Alan Gilbert" 
Cc: Johannes Weiner 
Cc: "Huangpeng (Peter)" 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: make page pfmemalloc check more robust

2015-08-21T21:30:10+00:00

Commit c48a11c7ad26 ("netvm: propagate page->pfmemalloc to skb") added
checks for page->pfmemalloc to __skb_fill_page_desc():

        if (page->pfmemalloc && !page->mapping)
                skb->pfmemalloc = true;

It assumes page->mapping == NULL implies that page->pfmemalloc can be
trusted.  However, __delete_from_page_cache() can set set page->mapping
to NULL and leave page->index value alone.  Due to being in union, a
non-zero page->index will be interpreted as true page->pfmemalloc.

So the assumption is invalid if the networking code can see such a page.
And it seems it can.  We have encountered this with a NFS over loopback
setup when such a page is attached to a new skbuf.  There is no copying
going on in this case so the page confuses __skb_fill_page_desc which
interprets the index as pfmemalloc flag and the network stack drops
packets that have been allocated using the reserves unless they are to
be queued on sockets handling the swapping which is the case here and
that leads to hangs when the nfs client waits for a response from the
server which has been dropped and thus never arrive.

The struct page is already heavily packed so rather than finding another
hole to put it in, let's do a trick instead.  We can reuse the index
again but define it to an impossible value (-1UL).  This is the page
index so it should never see the value that large.  Replace all direct
users of page->pfmemalloc by page_is_pfmemalloc which will hide this
nastiness from unspoiled eyes.

The information will get lost if somebody wants to use page->index
obviously but that was the case before and the original code expected
that the information should be persisted somewhere else if that is
really needed (e.g.  what SLAB and SLUB do).

[akpm@linux-foundation.org: fix blooper in slub]
Fixes: c48a11c7ad26 ("netvm: propagate page->pfmemalloc to skb")
Signed-off-by: Michal Hocko 
Debugged-by: Vlastimil Babka 
Debugged-by: Jiri Bohac 
Cc: Eric Dumazet 
Cc: David Miller 
Acked-by: Mel Gorman 
Cc: 	[3.6+]
Signed-off-by: Andrew Morton 

Signed-off-by: Linus Torvalds