linux.git/mm/Kconfig, branch v4.17

mm: don't allow deferred pages with NEED_PER_CPU_KM

2018-05-19T00:17:12+00:00

It is unsafe to do virtual to physical translations before mm_init() is
called if struct page is needed in order to determine the memory section
number (see SECTION_IN_PAGE_FLAGS).  This is because only in mm_init()
we initialize struct pages for all the allocated memory when deferred
struct pages are used.

My recent fix in commit c9e97a1997 ("mm: initialize pages on demand
during boot") exposed this problem, because it greatly reduced number of
pages that are initialized before mm_init(), but the problem existed
even before my fix, as Fengguang Wu found.

Below is a more detailed explanation of the problem.

We initialize struct pages in four places:

1. Early in boot a small set of struct pages is initialized to fill the
   first section, and lower zones.

2. During mm_init() we initialize "struct pages" for all the memory that
   is allocated, i.e reserved in memblock.

3. Using on-demand logic when pages are allocated after mm_init call
   (when memblock is finished)

4. After smp_init() when the rest free deferred pages are initialized.

The problem occurs if we try to do va to phys translation of a memory
between steps 1 and 2.  Because we have not yet initialized struct pages
for all the reserved pages, it is inherently unsafe to do va to phys if
the translation itself requires access of "struct page" as in case of
this combination: CONFIG_SPARSE && !CONFIG_SPARSE_VMEMMAP

The following path exposes the problem:

  start_kernel()
   trap_init()
    setup_cpu_entry_areas()
     setup_cpu_entry_area(cpu)
      get_cpu_gdt_paddr(cpu)
       per_cpu_ptr_to_phys(addr)
        pcpu_addr_to_page(addr)
         virt_to_page(addr)
          pfn_to_page(__pa(addr) >> PAGE_SHIFT)

We disable this path by not allowing NEED_PER_CPU_KM with deferred
struct pages feature.

The problems are discussed in these threads:
  http://lkml.kernel.org/r/20180418135300.inazvpxjxowogyge@wfg-t540p.sh.intel.com
  http://lkml.kernel.org/r/20180419013128.iurzouiqxvcnpbvz@wfg-t540p.sh.intel.com
  http://lkml.kernel.org/r/20180426202619.2768-1-pasha.tatashin@oracle.com

Link: http://lkml.kernel.org/r/20180515175124.1770-1-pasha.tatashin@oracle.com
Fixes: 3a80a7fa7989 ("mm: meminit: initialise a subset of struct pages if CONFIG_DEFERRED_STRUCT_PAGE_INIT is set")
Signed-off-by: Pavel Tatashin 
Acked-by: Michal Hocko 
Reviewed-by: Andrew Morton 
Cc: Steven Sistare 
Cc: Daniel Jordan 
Cc: Mel Gorman 
Cc: Fengguang Wu 
Cc: Dennis Zhou 
Cc: 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

treewide: simplify Kconfig dependencies for removed archs

2018-03-26T13:55:57+00:00

A lot of Kconfig symbols have architecture specific dependencies.
In those cases that depend on architectures we have already removed,
they can be omitted.

Acked-by: Kalle Valo 
Acked-by: Alexandre Belloni 
Signed-off-by: Arnd Bergmann

Drop a bunch of metag references

2018-02-23T14:29:59+00:00

Now that arch/metag/ has been removed, drop a bunch of metag references
in various codes across the whole tree:
 - VM_GROWSUP and __VM_ARCH_SPECIFIC_1.
 - MT_METAG_* ELF note types.
 - METAG Kconfig dependencies (FRAME_POINTER) and ranges
   (MAX_STACK_SIZE_MB).
 - metag cases in tools (checkstack.pl, recordmcount.c, perf).

Signed-off-by: James Hogan 
Acked-by: Steven Rostedt (VMware) 
Acked-by: Peter Zijlstra (Intel) 
Reviewed-by: Guenter Roeck 
Cc: Ingo Molnar 
Cc: Arnaldo Carvalho de Melo 
Cc: Alexander Shishkin 
Cc: Jiri Olsa 
Cc: Namhyung Kim 
Cc: linux-mm@kvack.org
Cc: linux-metag@vger.kernel.org

mm: relax deferred struct page requirements

2018-02-01T01:18:36+00:00

There is no need to have ARCH_SUPPORTS_DEFERRED_STRUCT_PAGE_INIT, as all
the page initialization code is in common code.

Also, there is no need to depend on MEMORY_HOTPLUG, as initialization
code does not really use hotplug memory functionality.  So, we can
remove this requirement as well.

This patch allows to use deferred struct page initialization on all
platforms with memblock allocator.

Tested on x86, arm64, and sparc.  Also, verified that code compiles on
PPC with CONFIG_MEMORY_HOTPLUG disabled.

Link: http://lkml.kernel.org/r/20171117014601.31606-1-pasha.tatashin@oracle.com
Signed-off-by: Pavel Tatashin 
Acked-by: Heiko Carstens 	[s390]
Reviewed-by: Khalid Aziz 
Acked-by: Michael Ellerman 
Acked-by: Michal Hocko 
Cc: Steven Sistare 
Cc: Daniel Jordan 
Cc: Benjamin Herrenschmidt 
Cc: Paul Mackerras 
Cc: Kirill A. Shutemov 
Cc: Reza Arbab 
Cc: Martin Schwidefsky 
Cc: Thomas Gleixner 
Cc: Mel Gorman 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm: add infrastructure for get_user_pages_fast() benchmarking

2017-11-18T00:10:04+00:00

Performance of get_user_pages_fast() is critical for some workloads, but
it's tricky to test it directly.

This patch provides /sys/kernel/debug/gup_benchmark that helps with
testing performance of it.

See tools/testing/selftests/vm/gup_benchmark.c for userspace
counterpart.

Link: http://lkml.kernel.org/r/20170908215603.9189-2-kirill.shutemov@linux.intel.com
Signed-off-by: Kirill A. Shutemov 
Cc: Shuah Khan 
Cc: Ingo Molnar 
Cc: Thorsten Leemhuis 
Cc: Jonathan Corbet 
Cc: Huang Ying 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm/hmm: avoid bloating arch that do not make use of HMM

2017-09-09T01:26:46+00:00

This moves all new code including new page migration helper behind kernel
Kconfig option so that there is no codee bloat for arch or user that do
not want to use HMM or any of its associated features.

arm allyesconfig (without all the patchset, then with and this patch):
   text	   data	    bss	    dec	    hex	filename
83721896	46511131	27582964	157815991	96814b7	../without/vmlinux
83722364	46511131	27582964	157816459	968168b	vmlinux

[jglisse@redhat.com: struct hmm is only use by HMM mirror functionality]
  Link: http://lkml.kernel.org/r/20170825213133.27286-1-jglisse@redhat.com
[sfr@canb.auug.org.au: fix build (arm multi_v7_defconfig)]
  Link: http://lkml.kernel.org/r/20170828181849.323ab81b@canb.auug.org.au
Link: http://lkml.kernel.org/r/20170818032858.7447-1-jglisse@redhat.com
Signed-off-by: Jérôme Glisse 
Signed-off-by: Stephen Rothwell 
Cc: Dan Williams 
Cc: Arnd Bergmann 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm/device-public-memory: device memory cache coherent with CPU

2017-09-09T01:26:46+00:00

Platform with advance system bus (like CAPI or CCIX) allow device memory
to be accessible from CPU in a cache coherent fashion.  Add a new type of
ZONE_DEVICE to represent such memory.  The use case are the same as for
the un-addressable device memory but without all the corners cases.

Link: http://lkml.kernel.org/r/20170817000548.32038-19-jglisse@redhat.com
Signed-off-by: Jérôme Glisse 
Cc: Aneesh Kumar 
Cc: Paul E. McKenney 
Cc: Benjamin Herrenschmidt 
Cc: Dan Williams 
Cc: Ross Zwisler 
Cc: Balbir Singh 
Cc: David Nellans 
Cc: Evgeny Baskakov 
Cc: Johannes Weiner 
Cc: John Hubbard 
Cc: Kirill A. Shutemov 
Cc: Mark Hairgrove 
Cc: Michal Hocko 
Cc: Sherry Cheung 
Cc: Subhash Gutti 
Cc: Vladimir Davydov 
Cc: Bob Liu 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm/ZONE_DEVICE: new type of ZONE_DEVICE for unaddressable memory

2017-09-09T01:26:46+00:00

HMM (heterogeneous memory management) need struct page to support
migration from system main memory to device memory.  Reasons for HMM and
migration to device memory is explained with HMM core patch.

This patch deals with device memory that is un-addressable memory (ie CPU
can not access it).  Hence we do not want those struct page to be manage
like regular memory.  That is why we extend ZONE_DEVICE to support
different types of memory.

A persistent memory type is define for existing user of ZONE_DEVICE and a
new device un-addressable type is added for the un-addressable memory
type.  There is a clear separation between what is expected from each
memory type and existing user of ZONE_DEVICE are un-affected by new
requirement and new use of the un-addressable type.  All specific code
path are protect with test against the memory type.

Because memory is un-addressable we use a new special swap type for when a
page is migrated to device memory (this reduces the number of maximum swap
file).

The main two additions beside memory type to ZONE_DEVICE is two callbacks.
First one, page_free() is call whenever page refcount reach 1 (which
means the page is free as ZONE_DEVICE page never reach a refcount of 0).
This allow device driver to manage its memory and associated struct page.

The second callback page_fault() happens when there is a CPU access to an
address that is back by a device page (which are un-addressable by the
CPU).  This callback is responsible to migrate the page back to system
main memory.  Device driver can not block migration back to system memory,
HMM make sure that such page can not be pin into device memory.

If device is in some error condition and can not migrate memory back then
a CPU page fault to device memory should end with SIGBUS.

[arnd@arndb.de: fix warning]
  Link: http://lkml.kernel.org/r/20170823133213.712917-1-arnd@arndb.de
Link: http://lkml.kernel.org/r/20170817000548.32038-8-jglisse@redhat.com
Signed-off-by: Jérôme Glisse 
Signed-off-by: Arnd Bergmann 
Acked-by: Dan Williams 
Cc: Ross Zwisler 
Cc: Aneesh Kumar 
Cc: Balbir Singh 
Cc: Benjamin Herrenschmidt 
Cc: David Nellans 
Cc: Evgeny Baskakov 
Cc: Johannes Weiner 
Cc: John Hubbard 
Cc: Kirill A. Shutemov 
Cc: Mark Hairgrove 
Cc: Michal Hocko 
Cc: Paul E. McKenney 
Cc: Sherry Cheung 
Cc: Subhash Gutti 
Cc: Vladimir Davydov 
Cc: Bob Liu 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm/hmm/mirror: mirror process address space on device with HMM helpers

2017-09-09T01:26:45+00:00

This is a heterogeneous memory management (HMM) process address space
mirroring.  In a nutshell this provide an API to mirror process address
space on a device.  This boils down to keeping CPU and device page table
synchronize (we assume that both device and CPU are cache coherent like
PCIe device can be).

This patch provide a simple API for device driver to achieve address space
mirroring thus avoiding each device driver to grow its own CPU page table
walker and its own CPU page table synchronization mechanism.

This is useful for NVidia GPU >= Pascal, Mellanox IB >= mlx5 and more
hardware in the future.

[jglisse@redhat.com: fix hmm for "mmu_notifier kill invalidate_page callback"]
  Link: http://lkml.kernel.org/r/20170830231955.GD9445@redhat.com
Link: http://lkml.kernel.org/r/20170817000548.32038-4-jglisse@redhat.com
Signed-off-by: Jérôme Glisse 
Signed-off-by: Evgeny Baskakov 
Signed-off-by: John Hubbard 
Signed-off-by: Mark Hairgrove 
Signed-off-by: Sherry Cheung 
Signed-off-by: Subhash Gutti 
Cc: Aneesh Kumar 
Cc: Balbir Singh 
Cc: Benjamin Herrenschmidt 
Cc: Dan Williams 
Cc: David Nellans 
Cc: Johannes Weiner 
Cc: Kirill A. Shutemov 
Cc: Michal Hocko 
Cc: Paul E. McKenney 
Cc: Ross Zwisler 
Cc: Vladimir Davydov 
Cc: Bob Liu 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

mm/hmm: heterogeneous memory management (HMM for short)

2017-09-09T01:26:45+00:00

HMM provides 3 separate types of functionality:
    - Mirroring: synchronize CPU page table and device page table
    - Device memory: allocating struct page for device memory
    - Migration: migrating regular memory to device memory

This patch introduces some common helpers and definitions to all of
those 3 functionality.

Link: http://lkml.kernel.org/r/20170817000548.32038-3-jglisse@redhat.com
Signed-off-by: Jérôme Glisse 
Signed-off-by: Evgeny Baskakov 
Signed-off-by: John Hubbard 
Signed-off-by: Mark Hairgrove 
Signed-off-by: Sherry Cheung 
Signed-off-by: Subhash Gutti 
Cc: Aneesh Kumar 
Cc: Balbir Singh 
Cc: Benjamin Herrenschmidt 
Cc: Dan Williams 
Cc: David Nellans 
Cc: Johannes Weiner 
Cc: Kirill A. Shutemov 
Cc: Michal Hocko 
Cc: Paul E. McKenney 
Cc: Ross Zwisler 
Cc: Vladimir Davydov 
Cc: Bob Liu 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds