]> git.hungrycats.org Git - linux/log
linux
22 years ago[PATCH] rmaplock: mm lock ordering
Hugh Dickins [Tue, 24 Aug 2004 04:24:34 +0000 (21:24 -0700)]
[PATCH] rmaplock: mm lock ordering

With page_map_lock out of the way, there's no need for page_referenced and
try_to_unmap to use trylocks - provided we switch anon_vma->lock and
mm->page_table_lock around in anon_vma_prepare.  Though I suppose it's
possible that we'll find that vmscan makes better progress with trylocks than
spinning - we're free to choose trylocks again if so.

Try to update the mm lock ordering documentation in filemap.c.  But I still
find it confusing, and I've no idea of where to stop.  So add an mm lock
ordering list I can understand to rmap.c.

Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] rmaplock: SLAB_DESTROY_BY_RCU
Hugh Dickins [Tue, 24 Aug 2004 04:24:22 +0000 (21:24 -0700)]
[PATCH] rmaplock: SLAB_DESTROY_BY_RCU

With page_map_lock gone, how to stabilize page->mapping's anon_vma while
acquiring anon_vma->lock in page_referenced_anon and try_to_unmap_anon?

The page cannot actually be freed (vmscan holds reference), but however much
we check page_mapped (which guarantees that anon_vma is in use - or would
guarantee that if we added suitable barriers), there's no locking against page
becoming unmapped the instant after, then anon_vma freed.

It's okay to take anon_vma->lock after it's freed, so long as it remains a
struct anon_vma (its list would become empty, or perhaps reused for an
unrelated anon_vma: but no problem since we always check that the page located
is the right one); but corruption if that memory gets reused for some other
purpose.

This is not unique: it's liable to be problem whenever the kernel tries to
approach a structure obliquely.  It's generally solved with an atomic
reference count; but one advantage of anon_vma over anonmm is that it does not
have such a count, and it would be a backward step to add one.

Therefore...  implement SLAB_DESTROY_BY_RCU flag, to guarantee that such a
kmem_cache_alloc'ed structure cannot get freed to other use while the
rcu_read_lock is held i.e.  preempt disabled; and use that for anon_vma.

Fix concerns raised by Manfred: this flag is incompatible with poisoning and
destructor, and kmem_cache_destroy needs to synchronize_kernel.

I hope SLAB_DESTROY_BY_RCU may be useful elsewhere; but though it's safe for
little anon_vma, I'd be reluctant to use it on any caches whose immediate
shrinkage under pressure is important to the system.

Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] rmaplock: kill page_map_lock
Hugh Dickins [Tue, 24 Aug 2004 04:24:11 +0000 (21:24 -0700)]
[PATCH] rmaplock: kill page_map_lock

The pte_chains rmap used pte_chain_lock (bit_spin_lock on PG_chainlock) to
lock its pte_chains.  We kept this (as page_map_lock: bit_spin_lock on
PG_maplock) when we moved to objrmap.  But the file objrmap locks its vma tree
with mapping->i_mmap_lock, and the anon objrmap locks its vma list with
anon_vma->lock: so isn't the page_map_lock superfluous?

Pretty much, yes.  The mapcount was protected by it, and needs to become an
atomic: starting at -1 like page _count, so nr_mapped can be tracked precisely
up and down.  The last page_remove_rmap can't clear anon page mapping any
more, because of races with page_add_rmap; from which some BUG_ONs must go for
the same reason, but they've served their purpose.

vmscan decisions are naturally racy, little change there beyond removing
page_map_lock/unlock.  But to stabilize the file-backed page->mapping against
truncation while acquiring i_mmap_lock, page_referenced_file now needs page
lock to be held even for refill_inactive_zone.  There's a similar issue in
acquiring anon_vma->lock, where page lock doesn't help: which this patch
pretends to handle, but actually it needs the next.

Roughly 10% cut off lmbench fork numbers on my 2*HT*P4.  Must confess my
testing failed to show the races even while they were knowingly exposed: would
benefit from testing on racier equipment.

Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] rmaplock: PageAnon in mapping
Hugh Dickins [Tue, 24 Aug 2004 04:23:59 +0000 (21:23 -0700)]
[PATCH] rmaplock: PageAnon in mapping

First of a batch of five patches to eliminate rmap's page_map_lock, replace
its trylocking by spinlocking, and use anon_vma to speed up swapoff.

Patches updated from the originals against 2.6.7-mm7: nothing new so I won't
spam the list, but including Manfred's SLAB_DESTROY_BY_RCU fixes, and omitting
the unuse_process mmap_sem fix already in 2.6.8-rc3.

This patch:

Replace the PG_anon page->flags bit by setting the lower bit of the pointer in
page->mapping when it's anon_vma: PAGE_MAPPING_ANON bit.

We're about to eliminate the locking which kept the flags and mapping in
synch: it's much easier to work on a local copy of page->mapping, than worry
about whether flags and mapping are in synch (though I imagine it could be
done, at greater cost, with some barriers).

Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] Fix /proc/pid/statm documentation
Roger Luethi [Tue, 24 Aug 2004 04:23:48 +0000 (21:23 -0700)]
[PATCH] Fix /proc/pid/statm documentation

I really wanted /proc/pid/statm to die and I still believe the
reasoning is valid.  As it doesn't look like that is going to happen,
though, I offer this fix for the respective documentation.  Note: lrs/drs
fields are switched.

Signed-off-by: Roger Luethi <rl@hellgate.ch>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] Automatically enable bigsmp on big HP machines
Arjan van de Ven [Tue, 24 Aug 2004 04:23:35 +0000 (21:23 -0700)]
[PATCH] Automatically enable bigsmp on big HP machines

This enables apic=bigsmp automatically on some big HP machines that need
it.  This makes them boot without kernel parameters on a generic arch
kernel.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] ia64: dma_mapping fix
William Lee Irwin III [Tue, 24 Aug 2004 04:23:25 +0000 (21:23 -0700)]
[PATCH] ia64: dma_mapping fix

We need to be able to dereference struct device in
include/asm-ia64/dma-mapping.h.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] md: make MD no device warning KERN_WARNING
Andi Kleen [Tue, 24 Aug 2004 04:23:14 +0000 (21:23 -0700)]
[PATCH] md: make MD no device warning KERN_WARNING

Prevents some noise during boot up when no MD volumes are found.

I think I picked it up from someone else, but I cannot remember from whom
(sorry)

Cc: Neil Brown <neilb@cse.unsw.edu.au>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] Make MAX_INIT_ARGS 32
Pete Zaitcev [Tue, 24 Aug 2004 04:23:04 +0000 (21:23 -0700)]
[PATCH] Make MAX_INIT_ARGS 32

We at Red Hat shipped a larger number of arguments for quite some time, it
was required for installations on IBM mainframe (s390), which doesn't have
a good way to pass arguments.

There are a number of reasonable situations that go past the current limits
of 8.  One that comes to mind is when you want to perform a manual vnc
install on a headless machine using anaconda.  This requires passing in a
number of parameters to get anaconda past the initial (no-gui) loader
screens.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] AIO: workqueue context switch reduction
Suparna Bhattacharya [Tue, 24 Aug 2004 04:22:52 +0000 (21:22 -0700)]
[PATCH] AIO: workqueue context switch reduction

From: Chris Mason

I compared the 2.6 pipetest results with the 2.4 suse kernel, and 2.6 was
roughly 40% slower.  During the pipetest run, 2.6 generates ~600,000
context switches per second while 2.4 generates 30 or so.

aio-context-switch (attached) has a few changes that reduces our context
switch rate, and bring performance back up to 2.4 levels.  These have only
really been tested against pipetest, they might make other workloads worse.

The basic theory behind the patch is that it is better for the userland
process to call run_iocbs than it is to schedule away and let the worker
thread do it.

1) on io_submit, use run_iocbs instead of run_iocb
2) on io_getevents, call run_iocbs if no events were available.

3) don't let two procs call run_iocbs for the same context at the same
   time.  They just end up bouncing on spinlocks.

The first three optimizations got me down to 360,000 context switches per
second, and they help build a little structure to allow optimization #4,
which uses queue_delayed_work(HZ/10) instead of queue_work.

That brings down the number of context switches to 2.4 levels.

Adds aio_run_all_iocbs so that normal processes can run all the pending
retries on the run list.  This allows worker threads to keep using list
splicing, but regular procs get to run the list until it stays empty.  The
end result should be less work for the worker threads.

I was able to trigger short stalls (1sec) with aio-stress, and with the
current patch they are gone.  Could be wishful thinking on my part though,
please let me know how this works for you.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] AIO: Splice runlist for fairness across io contexts
Suparna Bhattacharya [Tue, 24 Aug 2004 04:22:40 +0000 (21:22 -0700)]
[PATCH] AIO: Splice runlist for fairness across io contexts

This patch tries be a little fairer across multiple io contexts in handling
retries, helping make sure progress happens uniformly across different io
contexts (especially if they are acting on independent queues).

It splices the ioctx runlist before processing it in __aio_run_iocbs.  If
new iocbs get added to the ctx in meantime, it queues a fresh workqueue
entry instead of handling them righaway, so that other ioctxs' retries get
a chance to be processed before the newer entries in the queue.

This might make a difference in a situation where retries are getting
queued very fast on one ioctx, while the workqueue entry for another ioctx
is stuck behind it.  I've only seen this occasionally earlier and can't
recreate it consistently, but may be worth including anyway.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] AIO: retry infrastructure fixes and enhancements
Suparna Bhattacharya [Tue, 24 Aug 2004 04:22:28 +0000 (21:22 -0700)]
[PATCH] AIO: retry infrastructure fixes and enhancements

From: Daniel McNeil <daniel@osdl.org>
From: Chris Mason <mason@suse.com>

 AIO: retry infrastructure fixes and enhancements

 Reorganises, comments and fixes the AIO retry logic. Fixes
 and enhancements include:

   - Split iocb setup and execution in io_submit
        (also fixes io_submit error reporting)
   - Use aio workqueue instead of keventd for retries
   - Default high level retry methods
   - Subtle use_mm/unuse_mm fix
   - Code commenting
   - Fix aio process hang on EINVAL (Daniel McNeil)
   - Hold the context lock across unuse_mm
   - Acquire task_lock in use_mm()
   - Allow fops to override the retry method with their own
   - Elevated ref count for AIO retries (Daniel McNeil)
   - set_fs needed when calling use_mm
   - Flush workqueue on __put_ioctx (Chris Mason)
   - Fix io_cancel to work with retries (Chris Mason)
   - Read-immediate option for socket/pipe retry support

 Note on default high-level retry methods support
 ================================================

 High-level retry methods allows an AIO request to be executed as a series of
 non-blocking iterations, where each iteration retries the remaining part of
 the request from where the last iteration left off, by reissuing the
 corresponding AIO fop routine with modified arguments representing the
 remaining I/O.  The retries are "kicked" via the AIO waitqueue callback
 aio_wake_function() which replaces the default wait queue entry used for
 blocking waits.

 The high level retry infrastructure is responsible for running the
 iterations in the mm context (address space) of the caller, and ensures that
 only one retry instance is active at a given time, thus relieving the fops
 themselves from having to deal with potential races of that sort.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] cpqfc: add missing pci_enable_device()
Bjorn Helgaas [Tue, 24 Aug 2004 04:22:16 +0000 (21:22 -0700)]
[PATCH] cpqfc: add missing pci_enable_device()

Add pci_enable_device()/pci_disable_device().  In the past, drivers
often worked without this, but it is now required in order to route
PCI interrupts correctly.

Signed-off-by: Bjorn Helgaas <bjorn.helgaas@hp.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] de4x5.c: add missing pci_enable_device()
Bjorn Helgaas [Tue, 24 Aug 2004 04:22:05 +0000 (21:22 -0700)]
[PATCH] de4x5.c: add missing pci_enable_device()

Add pci_enable_device()/pci_disable_device().  In the past, drivers
often worked without this, but it is now required in order to route
PCI interrupts correctly.

Signed-off-by: Bjorn Helgaas <bjorn.helgaas@hp.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] ioc3-eth.c: add missing pci_enable_device()
Bjorn Helgaas [Tue, 24 Aug 2004 04:21:53 +0000 (21:21 -0700)]
[PATCH] ioc3-eth.c: add missing pci_enable_device()

Add pci_enable_device()/pci_disable_device().  In the past, drivers often
worked without this, but it is now required in order to route PCI interrupts
correctly.

Signed-off-by: Bjorn Helgaas <bjorn.helgaas@hp.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] hp100.c: add missing pci_enable_device()
Bjorn Helgaas [Tue, 24 Aug 2004 04:21:42 +0000 (21:21 -0700)]
[PATCH] hp100.c: add missing pci_enable_device()

Add pci_enable_device()/pci_disable_device().  In the past, drivers often
worked without this, but it is now required in order to route PCI interrupts
correctly.

Signed-off-by: Bjorn Helgaas <bjorn.helgaas@hp.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] ibmasm: add missing pci_enable_device()
Bjorn Helgaas [Tue, 24 Aug 2004 04:21:30 +0000 (21:21 -0700)]
[PATCH] ibmasm: add missing pci_enable_device()

Add pci_enable_device()/pci_disable_device().  In the past, drivers often
worked without this, but it is now required in order to route PCI
interrupts correctly.

Signed-off-by: Bjorn Helgaas <bjorn.helgaas@hp.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] tpam_main.c: add missing pci_enable_device()
Bjorn Helgaas [Tue, 24 Aug 2004 04:21:20 +0000 (21:21 -0700)]
[PATCH] tpam_main.c: add missing pci_enable_device()

Add pci_enable_device()/pci_disable_device().  In the past, drivers
often worked without this, but it is now required in order to route
PCI interrupts correctly.

Signed-off-by: Bjorn Helgaas <bjorn.helgaas@hp.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] ip2main.c: add missing pci_enable_device()
Bjorn Helgaas [Tue, 24 Aug 2004 04:21:08 +0000 (21:21 -0700)]
[PATCH] ip2main.c: add missing pci_enable_device()

I don't have this hardware, so this has been compiled but not tested.

Add pci_enable_device()/pci_disable_device In the past, drivers often worked
without this, but it is now required in order to route PCI interrupts
correctly.  In addition, this driver incorrectly used the IRQ value from PCI
config space rather than the one in the struct pci_dev.

Signed-off-by: Bjorn Helgaas <bjorn.helgaas@hp.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] idt77252.c: add missing pci_enable_device()
Bjorn Helgaas [Tue, 24 Aug 2004 04:20:57 +0000 (21:20 -0700)]
[PATCH] idt77252.c: add missing pci_enable_device()

Add pci_enable_device()/pci_disable_device().  In the past, drivers often
worked without this, but it is now required in order to route PCI
interrupts correctly.

Signed-off-by: Bjorn Helgaas <bjorn.helgaas@hp.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] don't pass mem_map into init functions
Dave Hansen [Tue, 24 Aug 2004 04:20:45 +0000 (21:20 -0700)]
[PATCH] don't pass mem_map into init functions

  When using CONFIG_NONLINEAR, a zone's mem_map isn't contiguous, and isn't
  allocated in the same place.  This means that nonlinear doesn't really have
  a mem_map[] to pass into free_area_init_node() or memmap_init_zone() which
  makes any sense.

  So, this patch removes the 'struct page *mem_map' argument to both of
  those functions.  All non-NUMA architectures just pass a NULL in there,
  which is ignored.  The solution on the NUMA arches is to pass the mem_map in
  via the pgdat, which works just fine.

  To replace the removed arguments, a call to pfn_to_page(node_start_pfn) is
  made.  This is valid because all of the pfn_to_page() implementations rely
  only on the pgdats, which are already set up at this time.  Plus, the
  pfn_to_page() method should work for any future nonlinear-type code.

  Finally, the patch creates a function: node_alloc_mem_map(), which I plan
  to effectively #ifdef out for nonlinear at some future date.

  Compile tested and booted on SMP x86, NUMAQ, and ppc64.

From: Jesse Barnes <jbarnes@engr.sgi.com>

  Fix up ia64 specific memory map init function in light of Dave's
  memmap_init cleanups.

Signed-off-by: Jesse Barnes <jbarnes@sgi.com>
From: Dave Hansen <haveblue@us.ibm.com>

  Looks like I missed a couple of architectures.  This patch, on top of my
  previous one and Jesse's should clean up the rest.

From: William Lee Irwin III <wli@holomorphy.com>

  x86-64 wouldn't compile with NUMA support on, as node_alloc_mem_map()
  references mem_map outside #ifdefs on CONFIG_NUMA/CONFIG_DISCONTIGMEM.  This
  patch wraps that reference in such an #ifdef.

From: William Lee Irwin III <wli@holomorphy.com>

  Initializing NODE_DATA(nid)->node_mem_map prior to calling it should do.

From: Dave Hansen <haveblue@us.ibm.com>

  Rick, I bet you didn't think your nerf weapons would be so effective in
  getting that compile error fixed, did you?

  Applying the attached patch and commenting out this line:

  arch/i386/kernel/nmi.c: In function `proc_unknown_nmi_panic':
  arch/i386/kernel/nmi.c:558: too few arguments to function `proc_dointvec'

  will let it compile.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] watchdog: fix warning "defined but not used"
Guillaume Thouvenin [Tue, 24 Aug 2004 04:20:32 +0000 (21:20 -0700)]
[PATCH] watchdog: fix warning "defined but not used"

Function wdtpci_init_one() in file wdt_pci.c generates a warning when
compiling the watchdog driver.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] first/next_cpu returns values > NR_CPUS
William Lee Irwin III [Tue, 24 Aug 2004 04:20:21 +0000 (21:20 -0700)]
[PATCH] first/next_cpu returns values > NR_CPUS

Zwane Mwaikambo <zwane@fsmlabs.com> wrote:

  The following caused some fireworks whilst merging i386 cpu hotplug.
  any_online_cpu(0x2) returns 32 on i386 if we're forced to continue past the
  only set bit due to the additional find_first_bit in the find_next_bit i386
  implementation.  Not wanting to change current behaviour in the bitops
  primitives and since the NR_CPUS thing is a cpumask issue, i've opted to fix
  next_cpu() and first_cpu() instead.

This might save a couple of lines of code.

From: <akpm@osdl.org>

  Fix cross-arch ulong/int disaster with find_next_bit().

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] New x86-64 merge
Andi Kleen [Tue, 24 Aug 2004 04:20:09 +0000 (21:20 -0700)]
[PATCH] New x86-64 merge

This fixes various issues in the previous update, in particular
a kernel without CONFIG_GART_IOMMU should boot now again,

The kernel discoverys PCI BUS<->CPU affinity on AMD systems
now.  It is so far used by dma_alloc_coherent to allocate memory
Experimental patches to add this to sysfs exist, but they're not
included yet. On systems with no memory on a CPU this information may
be wrong.

It has a new experimental CONFIG_UNORDERED_IO option. When enabled
it uses write combining for stores to device iomemory mapping. This
may give better performance with some device drivers, but has a slight
risk of breaking drivers (in general if a driver works on ia64,ppc64,sparc64
it should also work). Based on some discussions with Grant Grundler.

It requires the driver to use memory barriers properly. I would be interested
in feedback on any performance changes you're seeing. For a production system I
would recommend to keep it turned off(although I run it on all my systems and
haven't run into any problems yet)

ACPI and Centrino speedstep is enabled now for Nocona systems.

The IOMMU code does lazy merging by default now, which should be safe
and may increase performance on block IO.  It also avoids SAC force by default
now.

The machine check code has been improved again, hopefully it is good
now. It will log now machine check events from before the last reset.
And various other fixes.

The x86-64 parts are now gcc 3.5 clean.

And various other fixes

- Update defconfig
- Reset lost ticks on lost time warning, print RIP.
- Make TASK_SIZE test for 32bit (Arjan van de Ven)
- Work around bug in generic code that broke pcibus_to_cpumask
- Actually fix dummy iommu code
- Compile i386 acpi and speedstep-centrino cpufreq modules
- Export cpu_khz
- Fix compilation without GART_IOMMU
- Optimize find_*_bit functions for small fields
- Discover nodes near PCI busses on K8 (Travis Betak, changed by me)
- Optimize gart tlb flush slightly
- Add experimental CONFIG_UNORDERED_IO for unordered IO stores
- Add 32bit emulation for PTRACE_GETEVENTMSG
- Fix kernel_fpu_{begin,end} for preemptive kernels (Alexander Nyberg)
- Readd proper check for biomerge (got lost)
- Set up 32bit vsyscall page for ptrace early
- Add 32bit emulation for lookup_dcookie() for oprofile
- Export copy_page / clear_page
- Use rex prefix in save_init_fpu fxsave (Jan Beulich)
- Make it compile again
- Fix handling of hwdev == NULL (= ISA/LPC devices) in swiotlb
- Convert PCI DMA code to dma devices
- Change IOMMU code to use dummy fallback device instead of hardcoded
  NULL tests everywhere.
- Test iommu_sac_force instead of nommu for DAC supported macro
  (will cause more drivers to use DAC)
- Harden non IOMMU dma_alloc_consistent code to fail less likely.
- Remove use of strsep in option parsers
- Remove duplicated exports (Arjan van der Ven)
- Fix EFAULT checking in ptrace (John Blackwood)
- Update defconfig
- Remove dead URL from boot/setup.S (R.J. Wysocki)
- Use compat_sigval_t instead of sigval_t32 (Al Viro)
- Nanooptimization in 32bit ptregs calls
- Fix gcc 3.5 compilation in mtrr.h
- Pass pt_regs as pointer to avoid illegal pass by reference (for gcc 3.5)
- Make set_bit take int not long (Harald Dunkel)
- Avoid panic on pci_map_sg and pci_alloc_consistent overflow in GART IOMMU
- Handle large lost time delays in HPET code (Suresh B. Siddha)
- Work around theoretical bugs in prefetch handling (suggested by Jamie Lokier)
- Remove mtrr_strings declaration for gcc 3.5
- Set KBUILD_IMAGE for make rpm (William Lee Irwin III)
- Add iommu=noaperture to not touch the aperture
- Clean up argument parsing for iommu= option
- Export symbols for xchgadd based rwsems (still disabled)
- Define iommu_bio_merge for !CONFIG_GART_IOMMU
- Don't use backwards rep ; movsb for memmove
- Out line bitmap search functions (saves 8k .text, from i386)
- Convert bitmap search functions to 64bit accesses and optimize them
  a bit.
- Handle corrupted page tables in page fault handler
- Set iommu_merge (without force) to on by default again.
- Don't do bio merging by default for iommu=merge. This should make it
  safe to use again
- Add iommu=biomerge option to enable BIO merging (like old iommu=merge)
- Fix iommu=memaper=... parsing
- More MCE fixes (based on a patch by Eric Morton, heavily changed by me)
- Fix check for banks causing exceptions
- Allow to reinit MCEs later even after mce=off, fix wrong
  use of __initdata
  to disable at boot, but reenable later.
- Log left over machine checks after boot and resume
- Fix missing prototype warning with CPU_FREQ on
- Fix parsing of noexec=on (Ian Hastie)
- Fix warning in ia32_binfmt.c
- Resync time variable cpu frequency handling with i386
- Resync msr.c with i386
- Add 0x60 level 1 intel cache descriptor (from i386)
- Remove duplicated 32bit ioctls (Arnd Bergmann)
- Enable -msoft-float (from i386)
- Use faster version of FPU hang fix - handle the exception
  * a bit experimental, if you see "kernel ... math error" events
    in the log please report.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] preset loops_per_jiffy for faster booting
Adam Kropelin [Tue, 24 Aug 2004 04:19:45 +0000 (21:19 -0700)]
[PATCH] preset loops_per_jiffy for faster booting

Adds a kernel boot parameter "lpj=NNN" which allows the operator to specify
the loops-per-jiffy value.  This shaves up to a quarter of a second off
boot times, which are critical for embedded appliances.

It's a bit thin, but the code is in __init.

Signed-off-by: Adam Kropelin <akropel1@rochester.rr.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] Fix drivers/isdn/hisax/avm_pci.c build warning when !CONFIG_ISAPNP
Mika Kukkonen [Tue, 24 Aug 2004 04:19:34 +0000 (21:19 -0700)]
[PATCH] Fix drivers/isdn/hisax/avm_pci.c build warning when !CONFIG_ISAPNP

  CC [M]  drivers/isdn/hisax/avm_pci.o
drivers/isdn/hisax/avm_pci.c: In function `setup_avm_pcipnp':
drivers/isdn/hisax/avm_pci.c:817: warning: label `ready' defined but not used

Patch is big because I replaced the '} else { ...  }' with 'goto ready; }'
and so had to remove one level of indentation from code.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] Make UML build and run
Jeff Dike [Tue, 24 Aug 2004 04:19:22 +0000 (21:19 -0700)]
[PATCH] Make UML build and run

This patch includes the following -
updated defconfig
move uml.lds.S and main.c from arch/um to arch/um/kernel per Sam's suggestions
steal bitops.c from arch/i386
convert all calls to open_private_file to dentry_open

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] UML fixes
Jeff Dike [Tue, 24 Aug 2004 04:18:53 +0000 (21:18 -0700)]
[PATCH] UML fixes

The patch below fixes a few UML-specific bugs not related to the rest of the
kernel
a bogus error return and some formatting in the fork code
correct calculation of task.thread.kernel_stack
remove a bogus panic
a couple of fixes to allow UML to boot in the presence of exec-shield

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] UML updates
Jeff Dike [Tue, 24 Aug 2004 04:18:42 +0000 (21:18 -0700)]
[PATCH] UML updates

The patch below brings UML up to date with interface changes and the like
irq.c includes profile.h to bring in a missing definition
use the cpu_{set,clear} interface
use the new get_signal_to_deliver interface
define instruction_pointer

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: remove a group of unused bh functions
Coywolf Qi Hunt [Tue, 24 Aug 2004 04:18:30 +0000 (21:18 -0700)]
[PATCH] uml: remove a group of unused bh functions

This patch removes a group of unused bh functions in um.  This 2.2 legacy
code should be cleaned up.

Signed-off-by: Coywolf Qi Hunt <coywolf@greatcn.org>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Fix os_process_pc and os_process_parent for corner cases.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:18:19 +0000 (21:18 -0700)]
[PATCH] uml: Fix os_process_pc and os_process_parent for corner cases.

Update os_process_pc and os_process_parent: now a PID can be > 32768 (so
increase number of digits) and make it work even with spaces in the command
name.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: little-kmalloc
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:18:07 +0000 (21:18 -0700)]
[PATCH] uml: little-kmalloc

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Make malloc() call vmalloc if needed. Needed for hostfs on 2.6 host.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:17:56 +0000 (21:17 -0700)]
[PATCH] uml: Make malloc() call vmalloc if needed. Needed for hostfs on 2.6 host.

From: Oleg Drokin <green@linuxhacker.ru>, Jeff Dike <jdike@addtoit.com>, and
me

If size > 128K, with this patch malloc will call vmalloc; free will detect
whether to call vfree or kfree or __real_free().  The 2.4 version could forget
free()ing something; this has been fixed.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Removes dead code in trap_kern.c
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:17:44 +0000 (21:17 -0700)]
[PATCH] uml: Removes dead code in trap_kern.c

That code comes from the out_of_memory section; in 2.4 it was correct to put
it for "default:", since it was called when handle_mm_fault() return value was
!= 0, 1, 2, i.e.  it was 3, OOM (but the i386 code put it out of line, for
better performance).  Here, instead, the OOM case is handled on its own, so if
handle_mm_fault() != from the listed cases we must BUG().

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Avoids a panic for a legal situation
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:17:33 +0000 (21:17 -0700)]
[PATCH] uml: Avoids a panic for a legal situation

From: Alex Züpke <azu@sysgo.de>, and me

SKAS mode is like 4G/4G (here we have actually 3G/3G) for guest processes, so
when checking for kernel stack overflow, we must first make sure we are
checking a kernel-space address.  Also, correctly test for stack overflows
(i.e.  check if there is less than 1k of stack left; see
arch/i386/kernel/irq.c:do_IRQ()).  And also, THREAD_SIZE != PAGE_SIZE * 2, in
general (though this setting is almost never changed, so we didn't notice
this1).  Thanks to the good eye of Alex Züpke <azu@sysgo.de> for first seeing
this bug, and providing a test program:

/*
 * trigger.c - triggers panic("Kernel stack overflow") in UML
 *
 * 20040630, azu@sysgo.de
 */

#include <stdio.h>
#include <setjmp.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/types.h>
#include <sys/stat.h>
#include <sys/mman.h>

#define LOW  0xa0000000
#define HIGH 0xb0000000

int main(int argc, char **argv)
{
unsigned long addr;
int fd;

fd = open("/dev/zero", O_RDWR);

printf("This may take some time ... one more cup of coffee ...\n");

for(addr = LOW; addr < HIGH; addr += 0x1000)
{
pid_t p;
if(mmap((void*)addr, 0x1000, PROT_READ, MAP_SHARED | MAP_FIXED, fd, 0) == MAP_FAILED)
printf("mmap failed\n");

p = fork();
if(p == -1)
printf("fork failed\n");

if(p == 0)
{
/* child context */
int *p = (int *)addr;
volatile int x;

x = *p;
return 0;
}
/* father context */
waitpid(p, 0, 0);

if(munmap((void*)addr, 0x1000) == -1)
printf("munmap failed\n");
}

close(fd);
printf("done\n");
}

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Adds some exports
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:17:21 +0000 (21:17 -0700)]
[PATCH] uml: Adds some exports

Adds some exports

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Handles correctly errno == EINTR in lots of places.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:17:10 +0000 (21:17 -0700)]
[PATCH] uml: Handles correctly errno == EINTR in lots of places.

On various places (mostly waitpid() calls) this patch makes sure that if errno
== EINTR on return, then the syscall is endlessly retried.  It also defines a
simple generic way to do this.

Signed-off-by: <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Fix for sysemu patches
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:17:00 +0000 (21:17 -0700)]
[PATCH] uml: Fix for sysemu patches

- Correct some silly errors (dereferencing a pointer before checking if it's
  != NULL when creating /proc/sysemu, some error messages)

- separate using_sysemu from sysemu_supported (so to refuse to activate
  sysemu if it is not supported, avoiding panics)

- not probe sysemu if in tt mode.

Signed-off-by: <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Adds /proc/sysemu to toggle SYSEMU usage.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:16:48 +0000 (21:16 -0700)]
[PATCH] uml: Adds /proc/sysemu to toggle SYSEMU usage.

Adds /proc/sysemu to toggle SYSEMU usage.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Adds the "nosysemu" command line parameter to disable SYSEMU
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:16:37 +0000 (21:16 -0700)]
[PATCH] uml: Adds the "nosysemu" command line parameter to disable SYSEMU

Adds the "nosysemu" command line parameter to disable SYSEMU

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Use PTRACE_SCEMU (the so-called SYSEMU) to reduce syscall cost.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:16:25 +0000 (21:16 -0700)]
[PATCH] uml: Use PTRACE_SCEMU (the so-called SYSEMU) to reduce syscall cost.

Turns off syscall emulation patch for ptrace (SYSEMU) on.  SYSEMU is a
performance-patch introduced by Laurent Vivier.  It changes behaviour of
ptrace() and helps reducing host context switch rate.  To make it working, you
need a kernel patch for your host, too.  See
http://perso.wanadoo.fr/laurent.vivier/UML/ for further information.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Folds hostaudio_user.c into hostaudio_kern.c.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:16:14 +0000 (21:16 -0700)]
[PATCH] uml: Folds hostaudio_user.c into hostaudio_kern.c.

Folds hostaudio_user.c into hostaudio_kern.c.  A lot of code less.  Also note
that I no more update ppos(as I used to do in the 2.4 patch): I checked that
OSS never changes ppos, so hostaudio did the right thing.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Fixes raw() and uses it in check_one_sigio; also fixes a silly panic...
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:16:02 +0000 (21:16 -0700)]
[PATCH] uml: Fixes raw() and uses it in check_one_sigio; also fixes a silly panic (EINTR returned by call).

Fixes raw() and uses it in check_one_sigio; also fixes a silly panic (EINTR
returned by call).

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Reduces code in *_user files, by moving it in _kern files if already...
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:15:51 +0000 (21:15 -0700)]
[PATCH] uml: Reduces code in *_user files, by moving it in _kern files if already possible.

Reduces code in *_user files, by moving it in _kern files if already possible.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Avoids compile failure when host misses tkill().
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:15:40 +0000 (21:15 -0700)]
[PATCH] uml: Avoids compile failure when host misses tkill().

Avoids compile failure when host misses tkill(), by simply using kill() in
that case.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Kill useless warnings
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:15:29 +0000 (21:15 -0700)]
[PATCH] uml: Kill useless warnings

Fixes some little warnings about "Defined but not used ..." by #ifdef'ing
things

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Fixes "fixdep.c" to support arch/um/include/uml-config.h.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:15:17 +0000 (21:15 -0700)]
[PATCH] uml: Fixes "fixdep.c" to support arch/um/include/uml-config.h.

You probably saw that if you change one config option, even if
linux/autoconf.h (which is included by everything) changes, the kernel is
smart enough not to recompile everything.  But with UML this no more holds.
Why?  Because, as you see in this patch, fixdep avoids making anything depend
onto linux/autoconf.h *explicitly*, but nobody taught him to do the same for
arch/um/include/uml-config.h.  So apply this patch.  Do not say "I don't want
to change the generic Kbuild for one arch": this cannot hurt.  It's a bugfix
for us, a no-op for others.

Note: with this patch, fixdep will still add a dependency from a file
containing UML_CONFIG_BYE onto CONFIG_BYE.  Since someone could think that
fixdep should grep for [^A-Z_]CONFIG_ rather than simply for CONFIG_, I've
added a comment that ask *not to fix* this "bug".

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Makes "make help ARCH=um" work.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:15:06 +0000 (21:15 -0700)]
[PATCH] uml: Makes "make help ARCH=um" work.

Makes "make help ARCH=um" work.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Adds LEGACY_PTY config option
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:14:54 +0000 (21:14 -0700)]
[PATCH] uml: Adds LEGACY_PTY config option

The second adds the LEGACY_PTY config option. Without it, with late 2.6 kernels
/dev/ptyxx won't work. In fact, with those kernels, root_fs_toms does not
work, because it's "unable to allocate TTY pair". And removes the dead option
"UNIX98_PTY_COUNT" (just commented out for now).

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Fixes an host fd leak caused by hostfs.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:14:43 +0000 (21:14 -0700)]
[PATCH] uml: Fixes an host fd leak caused by hostfs.

In detail, on 2.4 we used force_delete() to make sure inode were not cached,
and we then close the host file when the inode is cleared; when porting to 2.6
the "force_delete" thing was dropped, and this patch adds a fix for this (by
setting drop_inode = generic_delete_inode).  Search for drop_inode in the 2.6
Documentation/filesystems/vfs.txt for info about this.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Avoid that gcc breaks UML with "unit at a time" compilation mode.
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:14:34 +0000 (21:14 -0700)]
[PATCH] uml: Avoid that gcc breaks UML with "unit at a time" compilation mode.

Avoid that gcc breaks UML with "unit at a time" compilation mode.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Readds (just for now) ghash.h for UML
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:14:22 +0000 (21:14 -0700)]
[PATCH] uml: Readds (just for now) ghash.h for UML

Just for now and just for UML; it will go away.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: rename console_device
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:14:10 +0000 (21:14 -0700)]
[PATCH] uml: rename console_device

In the -mm tree (in this moment) and not in 2.6.7 there is another
console_device in include/linux/console.h; so I renamed the UML one (it's
static).

Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: CPU scheduler update
Andrew Morton [Tue, 24 Aug 2004 04:13:58 +0000 (21:13 -0700)]
[PATCH] uml: CPU scheduler update

Update UML for CPU scheduler changes

Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] UML updates
Jeff Dike [Tue, 24 Aug 2004 04:13:46 +0000 (21:13 -0700)]
[PATCH] UML updates

The patch below brings UML up to date with some changes in the rest of the
kernel:
an updated defconfig
checksum.h includes in6.h to get a definition of in6_addr
added a missing cpu_{set,clear} change
removed include/asm-um/module.h since it's really a link

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] UML: remove the COW block driver
Jeff Dike [Tue, 24 Aug 2004 04:13:34 +0000 (21:13 -0700)]
[PATCH] UML: remove the COW block driver

The code is still there but it's not built.  Below is a patch which removes
it totally.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] uml: Uml base patch
Paolo \'Blaisorblade\' Giarrusso [Tue, 24 Aug 2004 04:13:23 +0000 (21:13 -0700)]
[PATCH] uml: Uml base patch

The main part of UML; it is the last distributed patch for 2.6.7 Removes skas
support from the main UML patch; apply or get conflicts.

Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] flexible-mmap for ppc64
Anton Blanchard [Tue, 24 Aug 2004 04:12:38 +0000 (21:12 -0700)]
[PATCH] flexible-mmap for ppc64

From: <arjanv@redhat.com>

Implement the new address space layout for 32-bit apps running on ppc64.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] flex mmap for s390(x)
Arjan van de Ven [Tue, 24 Aug 2004 04:12:26 +0000 (21:12 -0700)]
[PATCH] flex mmap for s390(x)

Below is a patch from Pete Zaitcev (zaitcev@redhat.com) to also use the
flex mmap infrastructure for s390(x).  The IBM Domino guys *really* seem to
want this.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sysctl tunable for flexmmap
Arjan van de Ven [Tue, 24 Aug 2004 04:12:13 +0000 (21:12 -0700)]
[PATCH] sysctl tunable for flexmmap

  Create /proc/sys/vm/legacy_va_layout.  If this is non-zero, the kernel
  will use the old mmap layout for all tasks.  it presently defaults to zero
  (the new layout).

From: William Lee Irwin III <wli@holomorphy.com>

  hugetlb CONFIG_SYSCTL=n fix

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] flexmmap patchkit: fix for 32 bit emu for 64 bit arches
Arjan van de Ven [Tue, 24 Aug 2004 04:12:01 +0000 (21:12 -0700)]
[PATCH] flexmmap patchkit: fix for 32 bit emu for 64 bit arches

Utz Lehmann <u.lehmann@de.tecosim.com> found a problem with the flexmmap
patches on x86-64, what he is seeing is that the 32 bit personality isn't
set at the first point of setting the allocator strategy.  The solution is
simple, in binfmt_elf the personality is set so put the pick-layout
function there.  Please consider,

Signed-off-by: Arjan van de Ven <arjanv@redhat.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] i386 virtual memory layout rework
Ingo Molnar [Tue, 24 Aug 2004 04:11:50 +0000 (21:11 -0700)]
[PATCH] i386 virtual memory layout rework

  Rework the i386 mm layout to allow applications to allocate more virtual
  memory, and larger contiguous chunks.

  - the patch is compatible with existing architectures that either make
    use of HAVE_ARCH_UNMAPPED_AREA or use the default mmap() allocator - there
    is no change in behavior.

  - 64-bit architectures can use the same mechanism to clean up 32-bit
    compatibility layouts: by defining HAVE_ARCH_PICK_MMAP_LAYOUT and
    providing a arch_pick_mmap_layout() function - which can then decide
    between various mmap() layout functions.

  - I also introduced a new personality bit (ADDR_COMPAT_LAYOUT) to signal
    older binaries that dont have PT_GNU_STACK.  x86 uses this to revert back
    to the stock layout.  I also changed x86 to not clear the personality bits
    upon exec(), like x86-64 already does.

  - once every architecture that uses HAVE_ARCH_UNMAPPED_AREA has defined
    its arch_pick_mmap_layout() function, we can get rid of
    HAVE_ARCH_UNMAPPED_AREA altogether, as a final cleanup.

  the new layout generation function (__get_unmapped_area()) got significant
  testing in FC1/2, so i'm pretty confident it's robust.

  Compiles & boots fine on an 'old' and on a 'new' x86 distro as well.

  The two known breakages were:

     http://www.redhatconfig.com/msg/67248.html

     [ 'cyzload' third-party utility broke. ]

     http://www.zipworld.com/au/~akpm/dde.tar.gz

     [ your editor broke :-) ]

  both were caused by application bugs that did:

int ret = malloc();

if (ret <= 0)
failure;

  such bugs are easy to spot if they happen, and if it happens it's possible
  to work it around immediately without having to change the binary, via the
  setarch patch.

  No other application has been found to be affected, and this particular
  change got pretty wide coverage already over RHEL3 and exec-shield, it's in
  use for more than a year.

  The setarch utility can be used to trigger the compatibility layout on
  x86, the following version has been patched to take the `-L' option:

  http://people.redhat.com/mingo/flexible-mmap/setarch-1.4-2.tar.gz

  "setarch -L i386 <command>" will run the command with the old layout.

From: Hugh Dickins <hugh@veritas.com>

  The problem is in the flexible mmap patch: arch_get_unmapped_area_topdown
  is liable to give your mmap vm_start above TASK_SIZE with vm_end wrapped;
  which is confusing, and ends up as that BUG_ON(mm->map_count).

  The patch below stops that behaviour, but it's not the full solution:
  wilson_mmap_test -s 1000 then simply cannot allocate memory for the large
  mmap, whereas it works fine non-top-down.

  I think it's wrong to interpret a large or rlim_infinite stack rlimit as
  an inviolable request to reserve that much for the stack: it makes much less
  VM available than bottom up, not what was intended.  Perhaps top down should
  go bottom up (instead of belly up) when it fails - but I'd probably better
  leave that to Ingo.

  Or perhaps the default should place stack below text (as WLI suggested and
  ELF intended, with its text defaulting to 0x08048000, small progs sharing
  page table between stack and text and data); with a further personality for
  those needing bigger stack.

From: Ingo Molnar <mingo@elte.hu>

  - fall back to the bottom-up layout if the stack can grow unlimited (if
  the stack ulimit has been set to RLIM_INFINITY)

  - try the bottom-up allocator if the top-down allocator fails - this can
  utilize the hole between the true bottom of the stack and its ulimit, as a
  last-resort effort.

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: smt fixes
Ingo Molnar [Tue, 24 Aug 2004 04:11:37 +0000 (21:11 -0700)]
[PATCH] sched: smt fixes

while looking at HT scheduler bugreports and boot failures i discovered a
bad assumption in most of the HT scheduling code: that resched_task() can
be called without holding the task's runqueue.

This is most definitely not valid - doing it without locking can lead to
the task on that CPU exiting, and this CPU corrupting the (ex-) task_info
struct.  It can also lead to HT-wakeup races with task switching on that
other CPU.  (this_CPU marking the wrong task on that_CPU as need_resched -
resulting in e.g.  idle wakeups not working.)

The attached patch against fixes it all up. Changes:

- resched_task() needs to touch the task so the runqueue lock of that CPU
  must be held: resched_task() now enforces this rule.

- wake_priority_sleeper() was called without holding the runqueue lock.

- wake_sleeping_dependent() needs to hold the runqueue locks of all
  siblings (2 typically).  Effects of this ripples back to schedule() as
  well - in the non-SMT case it gets compiled out so it's fine.

- dependent_sleeper() needs the runqueue locks too - and it's slightly
  harder because it wants to know the 'next task' info which might change
  during the lock-drop/reacquire.  Ripple effect on schedule() => compiled
  out on non-SMT so fine.

- resched_task() was disabling preemption for no good reason - all paths
  that called this function had either a spinlock held or irqs disabled.

Compiled & booted on x86 SMP and UP, with and without SMT. Booted the
SMT kernel on a real SMP+HT box as well. (Unpatched kernel wouldn't even
boot with the resched_task() assert in place.)

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: self-reaping atomicity fix
Ingo Molnar [Tue, 24 Aug 2004 04:11:26 +0000 (21:11 -0700)]
[PATCH] sched: self-reaping atomicity fix

disable preemption in the self-reap codepath, as such tasks may not be on
the tasklist anymore and CPU-hotplug relies on the tasklist to migrate
tasks.

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] permit sleeping in release_task()
Ingo Molnar [Tue, 24 Aug 2004 04:11:14 +0000 (21:11 -0700)]
[PATCH] permit sleeping in release_task()

release_task() calls proc_pid_flush() call dput(), which can sleep.  But
that's a late-in-exit no-preempt path with CONFIG_PREEMPT.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: new task fix
Ingo Molnar [Tue, 24 Aug 2004 04:11:02 +0000 (21:11 -0700)]
[PATCH] sched: new task fix

Rusty noticed that we update the parent ->avg_sleep without holding the
runqueue lock. Also the code needed cleanups.

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: nonlinear timeslices
Ingo Molnar [Tue, 24 Aug 2004 04:10:51 +0000 (21:10 -0700)]
[PATCH] sched: nonlinear timeslices

* Nick Piggin <nickpiggin@yahoo.com.au> wrote:

> Increasing priority (negative nice) doesn't have much impact. -20 CPU
> hog only gets about double the CPU of a 0 priority CPU hog and only
> about 120% the CPU time of a nice -10 hog.

this is a property of the base scheduler as well.

We can do a nonlinear timeslice distribution trivially - the attached
patch implements the following timeslice distribution ontop of
2.6.8-rc3-mm1:

   [ -20 ... 0 ... 19 ] => [800ms ... 100ms ... 5ms]

the nice-20/nice+19 ratio is now 1:160 - sufficient for all aspects.

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: whitespace cleanups
Ingo Molnar [Tue, 24 Aug 2004 04:10:39 +0000 (21:10 -0700)]
[PATCH] sched: whitespace cleanups

- whitespace and style cleanups

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] schedstat: UP fix
Andrew Morton [Tue, 24 Aug 2004 04:10:27 +0000 (21:10 -0700)]
[PATCH] schedstat: UP fix

SMP fix --
    for_each_domain() is not defined if not CONFIG_SMP, so show_schedstat
    needed a couple of extra ifdefs.

Signed-off-by: Rick Lindsley <ricklind@us.ibm.com>
Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: sparc32 fixes
William Lee Irwin III [Tue, 24 Aug 2004 04:10:16 +0000 (21:10 -0700)]
[PATCH] sched: sparc32 fixes

Fix up sparc32 properly.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: consolidate init_idle() and fork_by_hand()
William Lee Irwin III [Tue, 24 Aug 2004 04:10:03 +0000 (21:10 -0700)]
[PATCH] sched: consolidate init_idle() and fork_by_hand()

It appears that init_idle() and fork_by_hand() could be combined into a
single method that calls init_idle() on behalf of the caller.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] move CONFIG_SCHEDSTATS to arch/ppc64/Kconfig.debug
Nathan Lynch [Tue, 24 Aug 2004 04:09:52 +0000 (21:09 -0700)]
[PATCH] move CONFIG_SCHEDSTATS to arch/ppc64/Kconfig.debug

Otherwise it shows up under "iSeries device drivers", which doesn't seem
right.

Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] scheduler statistics
Rick Lindsley [Tue, 24 Aug 2004 04:09:41 +0000 (21:09 -0700)]
[PATCH] scheduler statistics

It adds lots of CPU scheduler stats in /proc/pid/stat.  They are described in
the new Documentation//sched-stats.txt

We were carrying this patch offline for some time, but as there's still
considerable ongoing work in this area, and as the new stats are a
configuration option, I think it's best that this capability be in the base
kernel.

Nick removed a fair amount of statistics that he wasn't using.  The full patch
gathers more information.  In particular, his patch doesn't include the code
to measure the latency between the time a process is made runnable and the
time it hits a processor which will be key to measuring interactivity changes.

He passed his changes back to me and I got finished merging his changes with
the current statistics patches just before OLS.  I believe this is largely a
superset of the patch you grabbed and should port relatively easily too.

Versions also exist for

    2.6.8-rc2
    2.6.8-rc2-mm1
    2.6.8-rc2-mm2

at
    http://eaglet.rain.com/rick/linux/schedstat/patches/

and within 24 hours at

    http://oss.software.ibm.com/linux/patches/?patch_id=730&show=all

The version below is for 2.6.8-rc2-mm2 without the staircase code and has
been compiled cleanly but not yet run.

From: Ingo Molnar <mingo@elte.hu>

this code needs a couple of cleanups before it can go into mainline:

fs/proc/array.c, fs/proc/base.c, fs/proc/proc_misc.c:

 - moved the new /proc/<PID>/stat fields to /proc/<PID>/schedstat,
   because the new fields break older procps. It's cleaner this way
   anyway. This moving of fields necessiated a bump to version 10.

Documentation/sched-stats.txt:

 - updated sched-stats.txt for version 10

 - wake_up_forked_thread() => wake_up_new_task()

 - updated the per-process field description

Kconfig:

 - removed the default y and made the option dependent on DEBUG_KERNEL.
   This is really for scheduler analysis, normal users dont need the
   overhead.

include/linux/sched.h:

 - moved the definitions into kernel/sched.c - this fixes UP compilation
   and is cleaner.

 - also moved the sched-domain definitions to sched.c - now that the
   sched-domains internals are not exposed to architectures this is
   doable. It's also necessary due to the previous change.

kernel/fork.c:

 - moved the ->sched_info init to sched_fork() where it belongs.

kernel/sched.c:

 - wake_up_forked_thread() -> wake_up_new_task(), wuft_cnt -> wunt_cnt,
   wuft_moved -> wunt_moved.

 - wunt_cnt and wunt_moved were defined by never updated - added the
   missing code to wake_up_new_task().

 - whitespace/style police

 - removed whitespace changes done to code not related to schedstats -
   i'll send a separate patch for these (and more).

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: adjust p4 per-cpu gain
Con Kolivas [Tue, 24 Aug 2004 04:09:28 +0000 (21:09 -0700)]
[PATCH] sched: adjust p4 per-cpu gain

The smt-nice handling is a little too aggressive by not estimating the per cpu
gain as high enough for pentium4 hyperthread.  This patch changes the per
sibling cpu gain from 15% to 25%.  The true per cpu gain is entirely dependant
on the workload but overall the 2 species of Pentium4 that support
hyperthreading have about 20-30% gain.

P.S: Anton - For the power processors that are now using this SMT nice
infrastructure it would be worth setting this value separately at 40%.

Signed-off-by: Con Kolivas <kernel@kolivas.org>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] Create cpu_sibling_map for PPC64
Matthew Dobson [Tue, 24 Aug 2004 04:09:16 +0000 (21:09 -0700)]
[PATCH] Create cpu_sibling_map for PPC64

In light of some proposed changes in the sched_domains code, I coded up
this little ditty that simply creates and populates a cpu_sibling_map for
PPC64 machines.  The patch just checks the CPU flags to determine if the
CPU supports SMT (aka Hyper-Threading aka Multi-Threading aka ...) and
fills in a mask of the siblings for each CPU in the system.  This should
allow us to build sched_domains for PPC64 with generic code in
kernel/sched.c for the SMT systems.  SMT is becoming more popular and is
turning up in more and more architectures.  I don't think it will be too
long until this feature is supported by most arches...

Signed-off-by: Matthew Dobson <colpatch@us.ibm.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: isolated sched domains
Dimitri Sivanich [Tue, 24 Aug 2004 04:09:04 +0000 (21:09 -0700)]
[PATCH] sched: isolated sched domains

Here's a version of the isolated scheduler domain code that I mentioned in
an RFC on 7/22.  This patch applies on top of 2.6.8-rc2-mm1 (to include all
of the new arch_init_sched_domain code).  This patch also contains the 2
line fix to remove the check of first_cpu(sd->groups->cpumask)) that Jesse
sent in earlier.

Note that this has not been tested with CONFIG_SCHED_SMT.  I hope that my
handling of those instances is OK.

Signed-off-by: Dimitri Sivanich <sivanich@sgi.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: limit cpuspan of node scheduler domains
Jesse Barnes [Tue, 24 Aug 2004 04:08:53 +0000 (21:08 -0700)]
[PATCH] sched: limit cpuspan of node scheduler domains

  This patch limits the cpu span of each node's scheduler domain to prevent
  balancing across too many cpus.  The cpus included in a node's domain are
  determined by the SD_NODES_PER_DOMAIN define and the arch specific
  sched_domain_node_span routine if ARCH_HAS_SCHED_DOMAIN is defined.  If
  ARCH_HAS_SCHED_DOMAIN is not defined, behavior is unchanged--all possible
  cpus will be included in each node's scheduling domain.  Currently, only
  ia64 provides an arch specific sched_domain_node_span routine.

From: Jesse Barnes <jbarnes@engr.sgi.com>

  This patch adds some more NUMA specific logic to the creation of scheduler
  domains.  Domains spanning all CPUs in a large system are too large to
  schedule across efficiently, leading to livelocks and inordinate amounts of
  time being spent in scheduler routines.  With this patch applied, the node
  scheduling domains for NUMA platforms will only contain a specified number
  of nearby CPUs, based on the value of SD_NODES_PER_DOMAIN.  It also allows
  arches to override SD_NODE_INIT, which sets the domain scheduling parameters
  for each node's domain.  This is necessary especially for large systems.

  Possible future directions:

  o multilevel node hierarchy (e.g.  node domains could contain 4 nodes
    worth of CPUs, supernode domains could contain 32 nodes worth, etc.  each
    with their own SD_NODE_INIT values)

  o more tweaking of SD_NODE_INIT values for good load balancing vs.
    overhead tradeoffs

From: mita akinobu <amgta@yacht.ocn.ne.jp>

  Compile fix

Signed-off-by: Jesse Barnes <jbarnes@sgi.com>
Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: consolidate sched domains
Nick Piggin [Tue, 24 Aug 2004 04:08:41 +0000 (21:08 -0700)]
[PATCH] sched: consolidate sched domains

  Teach the generic domains builder about SMT, and consolidate all
  architecture specific domain code into that.  Also, the SD_*_INIT macros can
  now be redefined by arch code without duplicating the entire setup code.
  This can be done by defining ARCH_HASH_SCHED_TUNE.

  The generic builder has been simplified with the addition of a helper
  macro which will probably prove to be useful to arch specific code as well
  and should be exported if that is the case.

Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au>
From: Matthew Dobson <colpatch@us.ibm.com>

  The attached patch is against 2.6.8-rc2-mm2, and removes Nick's
  conditional definition & population of cpu_sibling_map[] in favor of my
  unconditional ones.  This does not affect how cpu_sibling_map is used, just
  gives it broader scope.

From: Nick Piggin <nickpiggin@yahoo.com.au>

  Small fix to sched-consolidate-domains.patch picked up by

From: Suresh <suresh.b.siddha@intel.com>

  another sched consolidate domains fix

From: Nick Piggin <nickpiggin@yahoo.com.au>

  Don't use cpu_sibling_map if !CONFIG_SCHED_SMT

  This one spotted by Dimitri Sivanich <sivanich@sgi.com>

Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: fork hotplug hanling cleanup
Ingo Molnar [Tue, 24 Aug 2004 04:08:29 +0000 (21:08 -0700)]
[PATCH] sched: fork hotplug hanling cleanup

- remove the hotplug lock from around much of fork(), and re-copy the
  cpus_allowed mask to solve the hotplug race cleanly.

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Srivatsa Vaddagiri <vatsa@in.ibm.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: remove balance on clone
Nick Piggin [Tue, 24 Aug 2004 04:08:17 +0000 (21:08 -0700)]
[PATCH] sched: remove balance on clone

This removes balance on clone capability altogether.  I told Andi we wouldn't
remove it yet, but provided it is in a single small patch, he mightn't get too
upset.

Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: disable balance on clone
Nick Piggin [Tue, 24 Aug 2004 04:08:06 +0000 (21:08 -0700)]
[PATCH] sched: disable balance on clone

Don't balance on clone by default.

Balance on clone has a number of trivial performance failure cases, but it was
needed to get decent OpenMP performance on NUMA (Opteron) systems.  Not doing
child-runs-first for new threads also solves this problem in a nicer way
(implemented in a previous patch).

Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: sched misc changes
Nick Piggin [Tue, 24 Aug 2004 04:07:54 +0000 (21:07 -0700)]
[PATCH] sched: sched misc changes

Add some likely/unliklies, a for_each_cpu => for_each_cpu_online, and close
the sched_exit race.

From: Ingo Molnar <mingo@elte.hu>

  fix a typo in a previous patch breaking RT scheduling & interactivity.

Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au>
Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: make rt_task unlikely
Nick Piggin [Tue, 24 Aug 2004 04:07:42 +0000 (21:07 -0700)]
[PATCH] sched: make rt_task unlikely

From: Ingo Molnar <mingo@elte.hu>

RT tasks are unlikely, move this into rt_task() instead of open-coding it.

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: misc cleanups #2
Ingo Molnar [Tue, 24 Aug 2004 04:07:30 +0000 (21:07 -0700)]
[PATCH] sched: misc cleanups #2

 - fix two stale comments
 - cleanup

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] kernel thread idle fix
Nick Piggin [Tue, 24 Aug 2004 04:07:19 +0000 (21:07 -0700)]
[PATCH] kernel thread idle fix

Now that init_idle does not remove tasks from the runqueue, those
architectures that use kernel_thread instead of copy_process for the idle
task will break.  To fix, ensure that CLONE_IDLETASK tasks are not put on
the runqueue in the first place.

Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: cleanup, improve sched <=> fork APIs
Nick Piggin [Tue, 24 Aug 2004 04:07:08 +0000 (21:07 -0700)]
[PATCH] sched: cleanup, improve sched <=> fork APIs

Move balancing and child-runs-first logic from fork.c into sched.c where
it belongs.

* Consolidate wake_up_forked_process and wake_up_forked_thread into
  wake_up_new_process, and pass in clone_flags as suggested by Linus.  This
  removes a lot of code duplication and allows all logic to be handled in that
  function.

* Don't do balance-on-clone balancing for vfork'ed threads.

* Don't do set_task_cpu or balance one clone in wake_up_new_process.
  Instead do it in sched_fork to fix set_cpus_allowed races.

* Don't do child-runs-first for CLONE_VM processes, as there is obviously no
  COW benifit to be had.  This is a big one, it enables Andi's workload to run
  well without clone balancing, because the OpenMP child threads can get
  balanced off to other nodes *before* they start running and allocating
  memory.

* Rename sched_balance_exec to sched_exec: hide the policy from the API.

From: Ingo Molnar <mingo@elte.hu>

  rename wake_up_new_process -> wake_up_new_task.

  in sched.c we are gradually moving away from the overloaded 'process' or
  'thread' notion to the traditional task (or context) naming.

Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au>
Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: cleanup init_idle()
Nick Piggin [Tue, 24 Aug 2004 04:06:56 +0000 (21:06 -0700)]
[PATCH] sched: cleanup init_idle()

Clean up init_idle to not use wake_up_forked_process, then undo all the stuff
that call does.  Instead, do everything in init_idle.

Make double_rq_lock depend on CONFIG_SMP because it is no longer used on UP.

Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] sched: fix timeslice calculations for HZ=1000.
Ingo Molnar [Tue, 24 Aug 2004 04:06:43 +0000 (21:06 -0700)]
[PATCH] sched: fix timeslice calculations for HZ=1000.

The main benefit is that with the default HZ=1000 nice +19 tasks now get 5
msecs of timeslices, so the ratio of CPU use is linear.  (nice 0 task gets
20 times more CPU time than a nice 19 task.  Prior this change the ratio
was 1:10)

another effect is that nice 0 tasks now get a round 100 msecs of timeslices
(as intended), instead of 102 msecs.

here's a table of old/new timeslice values, for HZ=1000 and 100:

                      HZ=1000         (   HZ=100   )
                    old    new        ( old    new )

        nice -20:   200    200        ( 200    200 )
        nice -19:   195    195        ( 190    190 )
        ...
        nice 0:     102    100        ( 100    100 )
        nice 1:      97     95        (  90     90 )
        nice 2:      92     90        (  90     90 )
        ...
        nice 17:     19     15        (  10     10 )
        nice 18:     14     10        (  10     10 )
        nice 19:     10      5        (  10     10 )

i've tested the patch on x86.

Signed-off-by: Ingo Molnar <mingo@elte.hu>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years agoLinux 2.6.9-rc1 v2.6.9-rc1
Linus Torvalds [Mon, 23 Aug 2004 16:59:58 +0000 (09:59 -0700)]
Linux 2.6.9-rc1

22 years ago[PATCH] ppc64: use struct list_head for hose_list
Paul Mackerras [Mon, 23 Aug 2004 16:39:00 +0000 (09:39 -0700)]
[PATCH] ppc64: use struct list_head for hose_list

This patch changes hose_list from a simple linked list to a
"list.h"-style list.  This is in preparation for the runtime
addition/removal of PCI Host Bridges.

Signed-off-by: John Rose <johnrose@austin.ibm.com>
Signed-off-by: Paul Mackerras <paulus@samba.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years ago[PATCH] ppc64: fix enable_surveillance() for power5
Nathan Fontenot [Mon, 23 Aug 2004 16:38:48 +0000 (09:38 -0700)]
[PATCH] ppc64: fix enable_surveillance() for power5

On some platforms (notably power5) you can't enable surveillance
(firmware/service processor watchdog) from the kernel - you have to do
it in the firmware.

This patch changes enable_surveillance() to make the message that is
printed in this situation more informative.  Additionaly, the rtas_call
was changed to rtas_set_indicator so as to avoid having to handle
RTAS_BUSY returns.

Signed-off-by: Nathan Fontenot <nfont@austin.ibm.com>
Signed-off-by: Paul Mackerras <paulus@samba.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
22 years agoMerge bk://ppc.bkbits.net/for-linus-ppc64
Linus Torvalds [Mon, 23 Aug 2004 14:24:32 +0000 (07:24 -0700)]
Merge bk://ppc.bkbits.net/for-linus-ppc64
into ppc970.osdl.org:/home/torvalds/v2.6/linux

22 years agoUse F_SETLK instead of F_SETLK64 in nfs locking code.
Linus Torvalds [Mon, 23 Aug 2004 14:13:03 +0000 (07:13 -0700)]
Use F_SETLK instead of F_SETLK64 in nfs locking code.

The code doesn't actually _care_ about 32/64-bit issues,
only about F_SETLK vs F_SETLKW, and the F_SETLK64 doesn't
exist except as a compatibility thing on 64-bit architectures
(since the regular one already _is_ 64-bit, of course).

22 years agoMerge http://nfsclient.bkbits.net/linux-2.6
Trond Myklebust [Mon, 23 Aug 2004 17:41:30 +0000 (13:41 -0400)]
Merge http://nfsclient.bkbits.net/linux-2.6
into fys.uio.no:/home/linux/bitkeeper/nfsclient-2.6

22 years agoRPC,NFSv4: NFSv4 operations that create or destroy state on the
Trond Myklebust [Mon, 23 Aug 2004 16:02:36 +0000 (12:02 -0400)]
RPC,NFSv4: NFSv4 operations that create or destroy state on the
   server are not allowed to be interrupted as that may result in the
   client and server disagreeing.

22 years agoNFSv4: Enable delegations by actually telling the server about our
Trond Myklebust [Mon, 23 Aug 2004 16:01:42 +0000 (12:01 -0400)]
NFSv4: Enable delegations by actually telling the server about our
   recall ability.

Signed-off-by: Trond Myklebust <trond.myklebust@fys.uio.no>
22 years agoNFSv4: return all delegations we hold if the server issues a
Trond Myklebust [Mon, 23 Aug 2004 16:00:57 +0000 (12:00 -0400)]
NFSv4: return all delegations we hold if the server issues a
   NFS4ERR_CB_PATH_DOWN error.

22 years agoNFSv4: More aggressive caching if we have a delegation.
Trond Myklebust [Mon, 23 Aug 2004 16:00:15 +0000 (12:00 -0400)]
NFSv4: More aggressive caching if we have a delegation.

Signed-off-by: Trond Myklebust <trond.myklebust@fys.uio.no>
22 years agoNFSv4: Delegated open.
Trond Myklebust [Mon, 23 Aug 2004 15:59:17 +0000 (11:59 -0400)]
NFSv4: Delegated open.

Signed-off-by: Trond Myklebust <trond.myklebust@fys.uio.no>
22 years agoNFSv4: Recover delegations on server reboot.
Trond Myklebust [Mon, 23 Aug 2004 15:58:37 +0000 (11:58 -0400)]
NFSv4: Recover delegations on server reboot.

Signed-off-by: Trond Myklebust <trond.myklebust@fys.uio.no>