linux-stable.git/mm/sparse.c, branch linux-2.6.20.y

[PATCH] x86_64: allocate sparsemem memmap above 4G

2007-08-15T08:02:24+00:00

On systems with huge amount of physical memory, VFS cache and memory memmap
may eat all available system memory under 4G, then the system may fail to
allocate swiotlb bounce buffer.

There was a fix for this issue in arch/x86_64/mm/numa.c, but that fix dose
not cover sparsemem model.

This patch add fix to sparsemem model by first try to allocate memmap above
4G.

Signed-off-by: Zou Nan hai 
Acked-by: Suresh Siddha 
Cc: Andi Kleen 
Cc: 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds 
[chrisw: trivial backport]
Signed-off-by: Chris Wright 
Signed-off-by: Greg Kroah-Hartman

[PATCH] numa node ids are int, page_to_nid and zone_to_nid should return int

2006-12-07T16:39:23+00:00

NUMA node ids are passed as either int or unsigned int almost exclusivly
page_to_nid and zone_to_nid both return unsigned long.  This is a throw
back to when page_to_nid was a #define and was thus exposing the real type
of the page flags field.

In addition to fixing up the definitions of page_to_nid and zone_to_nid I
audited the users of these functions identifying the following incorrect
uses:

1) mm/page_alloc.c show_node() -- printk dumping the node id,
2) include/asm-ia64/pgalloc.h pgtable_quicklist_free() -- comparison
   against numa_node_id() which returns an int from cpu_to_node(), and
3) mm/mpolicy.c check_pte_range -- used as an index in node_isset which
   uses bit_set which in generic code takes an int.

Signed-off-by: Andy Whitcroft 
Cc: Christoph Lameter 
Cc: "Luck, Tony" 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

[PATCH] Get rid of zone_table[]

2006-12-07T16:39:20+00:00

The zone table is mostly not needed.  If we have a node in the page flags
then we can get to the zone via NODE_DATA() which is much more likely to be
already in the cpu cache.

In case of SMP and UP NODE_DATA() is a constant pointer which allows us to
access an exact replica of zonetable in the node_zones field.  In all of
the above cases there will be no need at all for the zone table.

The only remaining case is if in a NUMA system the node numbers do not fit
into the page flags.  In that case we make sparse generate a table that
maps sections to nodes and use that table to to figure out the node number.
 This table is sized to fit in a single cache line for the known 32 bit
NUMA platform which makes it very likely that the information can be
obtained without a cache miss.

For sparsemem the zone table seems to be have been fairly large based on
the maximum possible number of sections and the number of zones per node.
There is some memory saving by removing zone_table.  The main benefit is to
reduce the cache foootprint of the VM from the frequent lookups of zones.
Plus it simplifies the page allocator.

[akpm@osdl.org: build fix]
Signed-off-by: Christoph Lameter 
Cc: Dave Hansen 
Cc: Andy Whitcroft 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

[PATCH] memory hotplug: __GFP_NOWARN is better for __kmalloc_section_memmap()

2006-10-28T18:30:52+00:00

Add __GFP_NOWARN flag to calling of __alloc_pages() in
__kmalloc_section_memmap().  It can reduce noisy failure message.

In ia64, section size is 1 GB, this means that order 8 pages are necessary
for each section's memmap.  It is often very hard requirement under heavy
memory pressure as you know.  So, __alloc_pages() gives up allocation and
shows many noisy stack traces which means no page for each sections.
(Current my environment shows 32 times of stack trace....)

But, __kmalloc_section_memmap() calls vmalloc() after failure of it, and it
can succeed allocation of memmap.  So, its stack trace warning becomes just
noisy.  I suppose it shouldn't be shown.

Signed-off-by: Yasunori Goto 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

Remove obsolete #include

2006-06-30T17:25:36+00:00

Signed-off-by: Jörn Engel 
Signed-off-by: Adrian Bunk

[PATCH] spin/rwlock init cleanups

2006-06-28T00:32:39+00:00

locking init cleanups:

 - convert " = SPIN_LOCK_UNLOCKED" to spin_lock_init() or DEFINE_SPINLOCK()
 - convert rwlocks in a similar manner

this patch was generated automatically.

Motivation:

 - cleanliness
 - lockdep needs control of lock initialization, which the open-coded
   variants do not give
 - it's also useful for -rt and for lock debugging in general

Signed-off-by: Ingo Molnar 
Signed-off-by: Arjan van de Ven 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

[PATCH] sparsemem: record nid during memory present

2006-06-23T14:42:51+00:00

Record the node id as we mark sections for instantiation.  Use this nid
during instantiation to direct allocations.

Signed-off-by: Andy Whitcroft 
Cc: Mike Kravetz 
Cc: Dave Hansen 
Cc: Mel Gorman 
Cc: Bob Picco 
Cc: Jack Steiner 
Cc: Yasunori Goto 
Cc: Martin Bligh 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

[PATCH] SPARSEMEM incorrectly calculates section number

2006-05-21T19:59:17+00:00

A bad calculation/loop in __section_nr() could result in incorrect section
information being put into sysfs memory entries.  This primarily impacts
memory add operations as the sysfs information is used while onlining new
memory.

Fix suggested by Dave Hansen.

Note that the bug may not be obvious from the patch.  It actually occurs in
the function's return statement:

	return (root_nr * SECTIONS_PER_ROOT) + (ms - root);

In the existing code, root_nr has already been multiplied by
SECTIONS_PER_ROOT.

Signed-off-by: Mike Kravetz 
Cc: Dave Hansen 
Cc: Andy Whitcroft 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

[PATCH] add slab_is_available() routine for boot code

2006-05-15T18:20:56+00:00

slab_is_available() indicates slab based allocators are available for use.
SPARSEMEM code needs to know this as it can be called at various times
during the boot process.

Signed-off-by: Mike Kravetz 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds

[PATCH] sparsemem interaction with memory add bug fixes

2006-05-02T01:17:46+00:00

This patch fixes two bugs with the way sparsemem interacts with memory add.
They are:

- memory leak if memmap for section already exists

- calling alloc_bootmem_node() after boot

These bugs were discovered and a first cut at the fixes were provided by
Arnd Bergmann  and Joel Schopp .

Signed-off-by: Mike Kravetz 
Signed-off-by: Joel Schopp 
Signed-off-by: Andrew Morton 
Signed-off-by: Linus Torvalds