People are mainly concerned with showing off their total bogomips, not
per-cpu bogomips, so turn it into a KERN_DEBUG message for the benefit of
systems with lots of CPUs.
James Morris [Tue, 24 Aug 2004 04:31:13 +0000 (21:31 -0700)]
[PATCH] libfs: move transaction file ops into libfs
Below is an updated version of the patch which moves duplicated
transaction-based file operation code into libfs. Since the last post, the
patch has been through a couple of iterations with Al, who suggested a
number of cleanups including locking and interface simplification.
For filesystem writers, the interface is now much simpler. The
simple_transaction_get() helper should be part of the file op write method.
This safely obtains the transaction request data during write(), allocates
a page for it and stores it there. The data is returned to the caller for
potential further processing, which then makes it available for the next
read() call via simple_transaction_set(). See the selinuxfs and nfsctl
code for examples of use.
Signed-off-by: James Morris <jmorris@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ben Leslie [Tue, 24 Aug 2004 04:30:39 +0000 (21:30 -0700)]
[PATCH] Use posix headers in sumversion.c
When compiling Linux on Mac OSX I had trouble with scripts/sumversion.c.
It includes <netinet/in.h> to obtain to definitions of htonl and ntohl.
On Mac OSX these are found in <arpa/inet.h>. After checking the POSIX
specification it appears that this is the correct place to get the
definitons for these functions.
Using this header also appears to work on Linux (at least with
Glibc-2.3.2).
It seems clearer to me to go with the POSIX standard than implementing
#if __APPLE__ style macros, but if such an approach is preferred I can
supply patches for that instead.
A patch against 2.6.7 which change <netinet/in.h> -> <arpa/inet.h> is
attached.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
only call high2lowuid in the case of trying to put a bigger (32 bit, say)
uid/gid in a smaller (16 bit, in this case) word. Gcc is smart enough to see
that the comparison in high2lowuid() macro is silly if called with a 16 bit
source uid, but not smart enough to understand from the __convert_uid() logic
that this is exactly the case that high2lowuid() won't be called.
So replace the logical "<" operator with the bit op "&~". This obfuscates
things enough to shut gcc up.
Only build the half-dozen files that use SET_UID/SET_GID, on arch i386 and
ia64. Only the file fs/smbfs/inode.c showed the warning, both arch's, and
this patch fixed both. Untested further, past staring at the code long enough
to convince myself the change has no actual affect on the code's results.
Signed-off-by: Paul Jackson <pj@sgi.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Dave Jones [Tue, 24 Aug 2004 04:29:52 +0000 (21:29 -0700)]
[PATCH] fix inlining failures
arch/i386/mach-generic/summit.c: In function `send_IPI_all':
include/asm/mach-summit/mach_ipi.h:4: sorry, unimplemented: inlining failed in call to 'send_IPI_mask_sequence': function body not available
arch/i386/mach-generic/summit.c:8: sorry, unimplemented: called from here
make[1]: *** [arch/i386/mach-generic/summit.o] Error 1
make: *** [arch/i386/mach-generic] Error 2
arch/i386/mach-generic/bigsmp.c: In function `send_IPI_all':
include/asm/mach-bigsmp/mach_ipi.h:4: sorry, unimplemented: inlining failed in call to 'send_IPI_mask_sequence': function body not available
arch/i386/mach-generic/bigsmp.c:8: sorry, unimplemented: called from here
make[1]: *** [arch/i386/mach-generic/bigsmp.o] Error 1
make: *** [arch/i386/mach-generic] Error 2
arch/i386/mach-generic/es7000.c: In function `send_IPI_all':
include/asm/mach-es7000/mach_ipi.h:4: sorry, unimplemented: inlining failed in call to 'send_IPI_mask_sequence': function body not available
arch/i386/mach-generic/es7000.c:8: sorry, unimplemented: called from here
make[1]: *** [arch/i386/mach-generic/es7000.o] Error 1
make: *** [arch/i386/mach-generic] Error 2
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Pawel Sikora [Tue, 24 Aug 2004 04:29:41 +0000 (21:29 -0700)]
[PATCH] apm_info.disabled fix
This minor fix is required to proper init "APM emulation" on HP-OmniBooks.
(An external patch). "APM emulation" is very useful if you want to use a tool
which looks into /proc/apm for getting informations about battery charging.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Josh Aas [Tue, 24 Aug 2004 04:29:29 +0000 (21:29 -0700)]
[PATCH] Reduce bkl usage in do_coredump
A patch that reduces bkl usage in do_coredump. I don't see anywhere that
it is necessary except for the call to format_corename, which is controlled
via sysctl (sys_sysctl holds the bkl).
Neil Brown [Tue, 24 Aug 2004 04:29:06 +0000 (21:29 -0700)]
[PATCH] md: RAID10 module
This patch adds a 'raid10' module which provides features similar to both
raid0 and raid1 in the one array. Various combinations of layout are
supported.
This code is still "experimental", but appears to work.
Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Tue, 24 Aug 2004 04:28:54 +0000 (21:28 -0700)]
[PATCH] md: remove most calls to __bdevname from md.c
__bdevname now only prints major/minor number which isn't much help. So
remove most calls to it from md.c, replacing those that are useful by calls
to bdevname (often printing the message when the error is first detected
rather than higher up the call tree).
Also discard hot_generate_error which doesn't do anything useful and never
has.
Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Tue, 24 Aug 2004 04:28:31 +0000 (21:28 -0700)]
[PATCH] md: assorted fixes/improvemnet to generic md resync code.
1/ Introduce "mddev->resync_max_sectors" so that an md personality
can ask for resync to cover a different address range than that of a
single drive. raid10 will use this.
2/ fix is_mddev_idle so that if there seem to be a negative number
of events, it doesn't immediately assume activity.
3/ make "sync_io" (the count of IO sectors used for array resync)
an atomic_t to avoid SMP races.
4/ Pass md_sync_acct a "block_device" rather than the containing "rdev",
as the whole rdev isn't needed. Also make this an inline function.
5/ Make sure recovery gets interrupted on any error.
Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
During the kernel summit, some discussion was had about the support
requirements for a userspace program loader that loads executables into
hugetlb on behalf of a major application (Oracle). In order to support
this in a robust fashion, the cleanup of the hugetlb must be robust in the
presence of disorderly termination of the programs (e.g. kill -9). Hence,
the cleanup semantics are those of System V shared memory, but Linux'
System V shared memory needs one critical extension for this use:
executability.
The following microscopic patch enables this major application to provide
robust hugetlb cleanup.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
PAE is artificially limited in terms of swapspace to the same bitsplit as
ordinary i386, a 5/24 split (32 swapfiles, 64GB max swapfile size), when a
5/27 split (32 swapfiles, 512GB max swapfile size) is feasible. This patch
transparently removes that limitation by using more of the space available
in PAE's wider ptes for swap ptes.
While this is obviously not likely to be used directly, it is important
from the standpoint of strict non-overcommit, where the swapspace must be
potentially usable in order to be reserved for non-overcommit. There are
workloads with Committed_AS of over 256GB on ia32 PAE wanting strict
non-overcommit to prevent being OOM killed.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Zwane Mwaikambo [Tue, 24 Aug 2004 04:27:55 +0000 (21:27 -0700)]
[PATCH] fix i386/x86_64 idle routine selection
This was broken when the mwait stuff went in since it executes after the
initial idle_setup() has already selected an idle routine and overrides it
with default_idle.
Signed-off-by: Venkatesh Pallipadi <venkatesh.pallipadi@intel.com> Signed-off-by: Zwane Mwaikambo <zwane@linuxpower.ca> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Manfred Spraul [Tue, 24 Aug 2004 04:27:43 +0000 (21:27 -0700)]
[PATCH] remove magic +1 from shm segment count
Michael Kerrisk found a bug in the shm accounting code: sysv shm allows to
create SHMMNI+1 shared memory segments, instead of SHMMNI segments. The +1
is probably from the first shared anonymous mapping implementation that
used the sysv code to implement shared anon mappings.
The implementation got replaced, it's now the other way around (sysv uses
the shared anon code), but the +1 remained.
Zwane Mwaikambo [Tue, 24 Aug 2004 04:27:32 +0000 (21:27 -0700)]
[PATCH] OProfile/XScale fixes for PXA270/XScale2
The incorrect mask was being used when writing back to PMNC write-only-zero
bits as well as only ticking the CCNT every 64 processor cycles. Tested on
IOP331 and PXA270, i'm still looking for XScale1 users...
Signed-off-by: Luca Rossato <l.rossato@tiscali.it> Signed-off-by: Zwane Mwaikambo <zwane@arm.linux.org.uk> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The sole remaining usage of CLONE_IDLETASK is to determine whether pid
allocation should be performed in copy_process(). This patch eliminates
that last branch on CLONE_IDLETASK in the normal process creation path,
removes the masking of CLONE_IDLETASK from clone_flags as it's now ignored
under all circumstances, and furthermore eliminates the symbol
CLONE_IDLETASK entirely.
From: William Lee Irwin III <wli@holomorphy.com>
Fix the fork-idle consolidation. During that consolidation, the generic
code was made to pass a pointer to on-stack pt_regs that had been memset()
to 0. ia64, however, requires a NULL pt_regs pointer argument and
dispatches on that in its copy_thread() function to do SMP
trampoline-specific RSE -related setup. Passing pointers to zeroed pt_regs
resulted in SMP wakeup -time deadlocks and exceptions.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Every arch now bears the burden of sanitizing CLONE_IDLETASK out of the
clone_flags passed to do_fork() by userspace. This patch hoists the
masking of CLONE_IDLETASK out of the system call entrypoints into
do_fork(), and thereby removes some small overheads from do_fork(), as
do_fork() may now assume that CLONE_IDLETASK has been cleared.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Josh Aas [Tue, 24 Aug 2004 04:26:54 +0000 (21:26 -0700)]
[PATCH] improve speed of freeing bootmem
Attached is a patch that greatly improves the speed of freeing boot memory.
On ia64 machines with 2GB or more memory (I didn't test with less, but I
can't imagine there being a problem), the speed improvement is about 75%
for the function free_all_bootmem_core. This translates to savings on the
order of 1 minute / TB of memory during boot time. That number comes from
testing on a machine with 512GB, and extrapolating based on profiling of an
unpatched 4TB machine. For 4 and 8 TB machines, the time spent in this
function is about 1 minutes/TB, which is painful especially given that
there is no indication of what is going on put to the console (this issue
to possibly be addressed later).
The basic idea is to free higher order pages instead of going through every
single one. Also, some unnecessary atomic operations are done away with
and replaced with non-atomic equivalents, and prefetching is done where it
helps the most. For a more in-depth discusion of this patch, please see
the linux-ia64 archives (topic is "free bootmem feedback patch").
The patch is originally Tony Luck's, and I added some further optimizations
(non-atomic ops improvements and prefetching).
Signed-off-by: Tony Luck <tony.luck@intel.com> Signed-off-by: Josh Aas <josha@sgi.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Badari Pulavarty [Tue, 24 Aug 2004 04:26:42 +0000 (21:26 -0700)]
[PATCH] Fix mpage_readpage() for big requests
The problem is, if we increase our readhead size arbitrarily (say 2M), we
call mpage_readpages() with 2M and when it tries to allocated a bio enough to
fit 2M it fails, then we kick it back to "confused" code - which does 4K at
a time.
The fix is to ask for the maxium the driver can handle.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland Dreier [Tue, 24 Aug 2004 04:26:31 +0000 (21:26 -0700)]
[PATCH] x86: remove hard-coded numbers from ptr_ok()
Looks like arch/i386/kernel/doublefault.c is one place in the code that
hardcodes the assumption that PAGE_OFFSET == 0xC0000000. Here's a patch
that fixes that.
Signed-off-by: Roland Dreier <roland@topspin.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andrea Arcangeli [Tue, 24 Aug 2004 04:26:07 +0000 (21:26 -0700)]
[PATCH] Correctly handle d_path error returns
There's some minor bug in the d_path handling (the nfsd one may not the the
correct fix, there's no failure path for it, so I just terminate the
string, and the last one in the audit subsystem is just a robustness
cleanup if somebody will extend d_path in the future, right now it's a
noop).
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andrew Morton [Tue, 24 Aug 2004 04:25:56 +0000 (21:25 -0700)]
[PATCH] alloc_pages priority tuning
Fix up the logic which decides when the caller can dip into page reserves.
- If the caller has realtime scheduling policy, or if the caller cannot run
direct reclaim, then allow the caller to use up to a quarter of the page
reserves.
- If the caller has __GFP_HIGH then allow the caller to use up to half of
the page reserves.
- If the caller has PF_MEMALLOC then the caller can use 100% of the page
reserves.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Tue, 24 Aug 2004 04:25:44 +0000 (21:25 -0700)]
[PATCH] vm: alloc_pages watermark fixes
Previously the ->protection[] logic was broken. It was difficult to follow
and basically didn't use the asynch reclaim watermarks (pages_min,
pages_low, pages_high) properly.
Now use ->protection *only* for lower-zone protection. So the allocator
now explicitly uses the ->pages_low, ->pages_min watermarks and adds
->protection on top of that, instead of trying to use ->protection for
everything.
Pages are allocated down to (->pages_low + ->protection), once this is
reached, kswapd the background reclaim is started; after this, we can
allocate down to (->pages_min + ->protection) without blocking; the memory
below pages_min is reserved for __GFP_HIGH and PF_MEMALLOC allocations.
kswapd attempts to reclaim memory until ->pages_high is reached.
Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Tue, 24 Aug 2004 04:25:33 +0000 (21:25 -0700)]
[PATCH] vm: writeout watermark tuning
Slightly change the writeout watermark calculations so we keep background
and synchronous writeout watermarks in the same ratios after adjusting them
for the amout of mapped memory. This ensures we should always attempt to
start background writeout before synchronous writeout and preserves the
admin's desired background-versus-forground ratios after we've
auto-adjusted one of them.
Signed-off-by: Nick Piggin <nickpiggin@cyberone.com.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Hugh Dickins [Tue, 24 Aug 2004 04:25:21 +0000 (21:25 -0700)]
[PATCH] simple fs stop -ve dentries
A tmpfs user reported increasingly slow directory reads when repeatedly
creating and unlinking in a mkstemp-like way. The negative dentries
accumulate alarmingly (until memory pressure finally frees them), and are
just a hindrance to any in-memory filesystem. simple_lookup set d_op to
arrange for negative dentries to be deleted immediately.
(But I failed to discover how it is that on-disk filesystems seem to keep
their negative dentries within manageable bounds: this effect was gross
with tmpfs or ramfs, but no problem at all with extN or reiser.)
Hugh Dickins [Tue, 24 Aug 2004 04:25:09 +0000 (21:25 -0700)]
[PATCH] clarify get_task_mm (mmgrab)
Clarify mmgrab by collapsing it into get_task_mm (in fork.c not inline),
and commenting on the special case it is guarding against: when use_mm in
an AIO daemon temporarily adopts the mm while it's on its way out.
Marcelo Tosatti [Tue, 24 Aug 2004 04:24:57 +0000 (21:24 -0700)]
[PATCH] x86 bitops.h commentary on instruction reordering
Back when we were discussing the need for a memory barrier in sync_page(),
it came to me (thanks Andrea!) that the bit operations can be perfectly
reordered on architectures other than x86.
I think the commentary on i386 bitops.h is misleading, its worth to note
that that these operations are not guaranteed not to be reordered on
different architectures.
clear_bit() already does that:
* clear_bit() is atomic and may not be reordered. However, it does
* not contain a memory barrier, so if it is used for locking purposes,
* you should call smp_mb__before_clear_bit() and/or smp_mb__after_clear_bit()
* in order to ensure changes are visible on other processors.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Hugh Dickins [Tue, 24 Aug 2004 04:24:46 +0000 (21:24 -0700)]
[PATCH] rmaplock: swapoff use anon_vma
Swapoff can make good use of a page's anon_vma and index, while it's still
left in swapcache, or once it's brought back in and the first pte mapped back:
unuse_vma go directly to just one page of only those vmas with the same
anon_vma. And unuse_process can skip any vmas without an anon_vma (extending
the hugetlb check: hugetlb vmas have no anon_vma).
This just hacks in on top of the existing procedure, still going through all
the vmas of all the mms in mmlist. A more elegant procedure might replace
mmlist by a list of anon_vmas: but that would be more work to implement, with
apparently more overhead in the common paths.
Hugh Dickins [Tue, 24 Aug 2004 04:24:34 +0000 (21:24 -0700)]
[PATCH] rmaplock: mm lock ordering
With page_map_lock out of the way, there's no need for page_referenced and
try_to_unmap to use trylocks - provided we switch anon_vma->lock and
mm->page_table_lock around in anon_vma_prepare. Though I suppose it's
possible that we'll find that vmscan makes better progress with trylocks than
spinning - we're free to choose trylocks again if so.
Try to update the mm lock ordering documentation in filemap.c. But I still
find it confusing, and I've no idea of where to stop. So add an mm lock
ordering list I can understand to rmap.c.
Hugh Dickins [Tue, 24 Aug 2004 04:24:22 +0000 (21:24 -0700)]
[PATCH] rmaplock: SLAB_DESTROY_BY_RCU
With page_map_lock gone, how to stabilize page->mapping's anon_vma while
acquiring anon_vma->lock in page_referenced_anon and try_to_unmap_anon?
The page cannot actually be freed (vmscan holds reference), but however much
we check page_mapped (which guarantees that anon_vma is in use - or would
guarantee that if we added suitable barriers), there's no locking against page
becoming unmapped the instant after, then anon_vma freed.
It's okay to take anon_vma->lock after it's freed, so long as it remains a
struct anon_vma (its list would become empty, or perhaps reused for an
unrelated anon_vma: but no problem since we always check that the page located
is the right one); but corruption if that memory gets reused for some other
purpose.
This is not unique: it's liable to be problem whenever the kernel tries to
approach a structure obliquely. It's generally solved with an atomic
reference count; but one advantage of anon_vma over anonmm is that it does not
have such a count, and it would be a backward step to add one.
Therefore... implement SLAB_DESTROY_BY_RCU flag, to guarantee that such a
kmem_cache_alloc'ed structure cannot get freed to other use while the
rcu_read_lock is held i.e. preempt disabled; and use that for anon_vma.
Fix concerns raised by Manfred: this flag is incompatible with poisoning and
destructor, and kmem_cache_destroy needs to synchronize_kernel.
I hope SLAB_DESTROY_BY_RCU may be useful elsewhere; but though it's safe for
little anon_vma, I'd be reluctant to use it on any caches whose immediate
shrinkage under pressure is important to the system.
Hugh Dickins [Tue, 24 Aug 2004 04:24:11 +0000 (21:24 -0700)]
[PATCH] rmaplock: kill page_map_lock
The pte_chains rmap used pte_chain_lock (bit_spin_lock on PG_chainlock) to
lock its pte_chains. We kept this (as page_map_lock: bit_spin_lock on
PG_maplock) when we moved to objrmap. But the file objrmap locks its vma tree
with mapping->i_mmap_lock, and the anon objrmap locks its vma list with
anon_vma->lock: so isn't the page_map_lock superfluous?
Pretty much, yes. The mapcount was protected by it, and needs to become an
atomic: starting at -1 like page _count, so nr_mapped can be tracked precisely
up and down. The last page_remove_rmap can't clear anon page mapping any
more, because of races with page_add_rmap; from which some BUG_ONs must go for
the same reason, but they've served their purpose.
vmscan decisions are naturally racy, little change there beyond removing
page_map_lock/unlock. But to stabilize the file-backed page->mapping against
truncation while acquiring i_mmap_lock, page_referenced_file now needs page
lock to be held even for refill_inactive_zone. There's a similar issue in
acquiring anon_vma->lock, where page lock doesn't help: which this patch
pretends to handle, but actually it needs the next.
Roughly 10% cut off lmbench fork numbers on my 2*HT*P4. Must confess my
testing failed to show the races even while they were knowingly exposed: would
benefit from testing on racier equipment.
Hugh Dickins [Tue, 24 Aug 2004 04:23:59 +0000 (21:23 -0700)]
[PATCH] rmaplock: PageAnon in mapping
First of a batch of five patches to eliminate rmap's page_map_lock, replace
its trylocking by spinlocking, and use anon_vma to speed up swapoff.
Patches updated from the originals against 2.6.7-mm7: nothing new so I won't
spam the list, but including Manfred's SLAB_DESTROY_BY_RCU fixes, and omitting
the unuse_process mmap_sem fix already in 2.6.8-rc3.
This patch:
Replace the PG_anon page->flags bit by setting the lower bit of the pointer in
page->mapping when it's anon_vma: PAGE_MAPPING_ANON bit.
We're about to eliminate the locking which kept the flags and mapping in
synch: it's much easier to work on a local copy of page->mapping, than worry
about whether flags and mapping are in synch (though I imagine it could be
done, at greater cost, with some barriers).
Roger Luethi [Tue, 24 Aug 2004 04:23:48 +0000 (21:23 -0700)]
[PATCH] Fix /proc/pid/statm documentation
I really wanted /proc/pid/statm to die and I still believe the
reasoning is valid. As it doesn't look like that is going to happen,
though, I offer this fix for the respective documentation. Note: lrs/drs
fields are switched.
Signed-off-by: Roger Luethi <rl@hellgate.ch> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Pete Zaitcev [Tue, 24 Aug 2004 04:23:04 +0000 (21:23 -0700)]
[PATCH] Make MAX_INIT_ARGS 32
We at Red Hat shipped a larger number of arguments for quite some time, it
was required for installations on IBM mainframe (s390), which doesn't have
a good way to pass arguments.
There are a number of reasonable situations that go past the current limits
of 8. One that comes to mind is when you want to perform a manual vnc
install on a headless machine using anaconda. This requires passing in a
number of parameters to get anaconda past the initial (no-gui) loader
screens.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
I compared the 2.6 pipetest results with the 2.4 suse kernel, and 2.6 was
roughly 40% slower. During the pipetest run, 2.6 generates ~600,000
context switches per second while 2.4 generates 30 or so.
aio-context-switch (attached) has a few changes that reduces our context
switch rate, and bring performance back up to 2.4 levels. These have only
really been tested against pipetest, they might make other workloads worse.
The basic theory behind the patch is that it is better for the userland
process to call run_iocbs than it is to schedule away and let the worker
thread do it.
1) on io_submit, use run_iocbs instead of run_iocb
2) on io_getevents, call run_iocbs if no events were available.
3) don't let two procs call run_iocbs for the same context at the same
time. They just end up bouncing on spinlocks.
The first three optimizations got me down to 360,000 context switches per
second, and they help build a little structure to allow optimization #4,
which uses queue_delayed_work(HZ/10) instead of queue_work.
That brings down the number of context switches to 2.4 levels.
Adds aio_run_all_iocbs so that normal processes can run all the pending
retries on the run list. This allows worker threads to keep using list
splicing, but regular procs get to run the list until it stays empty. The
end result should be less work for the worker threads.
I was able to trigger short stalls (1sec) with aio-stress, and with the
current patch they are gone. Could be wishful thinking on my part though,
please let me know how this works for you.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] AIO: Splice runlist for fairness across io contexts
This patch tries be a little fairer across multiple io contexts in handling
retries, helping make sure progress happens uniformly across different io
contexts (especially if they are acting on independent queues).
It splices the ioctx runlist before processing it in __aio_run_iocbs. If
new iocbs get added to the ctx in meantime, it queues a fresh workqueue
entry instead of handling them righaway, so that other ioctxs' retries get
a chance to be processed before the newer entries in the queue.
This might make a difference in a situation where retries are getting
queued very fast on one ioctx, while the workqueue entry for another ioctx
is stuck behind it. I've only seen this occasionally earlier and can't
recreate it consistently, but may be worth including anyway.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] AIO: retry infrastructure fixes and enhancements
From: Daniel McNeil <daniel@osdl.org>
From: Chris Mason <mason@suse.com>
AIO: retry infrastructure fixes and enhancements
Reorganises, comments and fixes the AIO retry logic. Fixes
and enhancements include:
- Split iocb setup and execution in io_submit
(also fixes io_submit error reporting)
- Use aio workqueue instead of keventd for retries
- Default high level retry methods
- Subtle use_mm/unuse_mm fix
- Code commenting
- Fix aio process hang on EINVAL (Daniel McNeil)
- Hold the context lock across unuse_mm
- Acquire task_lock in use_mm()
- Allow fops to override the retry method with their own
- Elevated ref count for AIO retries (Daniel McNeil)
- set_fs needed when calling use_mm
- Flush workqueue on __put_ioctx (Chris Mason)
- Fix io_cancel to work with retries (Chris Mason)
- Read-immediate option for socket/pipe retry support
Note on default high-level retry methods support
================================================
High-level retry methods allows an AIO request to be executed as a series of
non-blocking iterations, where each iteration retries the remaining part of
the request from where the last iteration left off, by reissuing the
corresponding AIO fop routine with modified arguments representing the
remaining I/O. The retries are "kicked" via the AIO waitqueue callback
aio_wake_function() which replaces the default wait queue entry used for
blocking waits.
The high level retry infrastructure is responsible for running the
iterations in the mm context (address space) of the caller, and ensures that
only one retry instance is active at a given time, thus relieving the fops
themselves from having to deal with potential races of that sort.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Bjorn Helgaas [Tue, 24 Aug 2004 04:22:16 +0000 (21:22 -0700)]
[PATCH] cpqfc: add missing pci_enable_device()
Add pci_enable_device()/pci_disable_device(). In the past, drivers
often worked without this, but it is now required in order to route
PCI interrupts correctly.
Bjorn Helgaas [Tue, 24 Aug 2004 04:22:05 +0000 (21:22 -0700)]
[PATCH] de4x5.c: add missing pci_enable_device()
Add pci_enable_device()/pci_disable_device(). In the past, drivers
often worked without this, but it is now required in order to route
PCI interrupts correctly.
Add pci_enable_device()/pci_disable_device(). In the past, drivers often
worked without this, but it is now required in order to route PCI interrupts
correctly.
Bjorn Helgaas [Tue, 24 Aug 2004 04:21:42 +0000 (21:21 -0700)]
[PATCH] hp100.c: add missing pci_enable_device()
Add pci_enable_device()/pci_disable_device(). In the past, drivers often
worked without this, but it is now required in order to route PCI interrupts
correctly.
Bjorn Helgaas [Tue, 24 Aug 2004 04:21:30 +0000 (21:21 -0700)]
[PATCH] ibmasm: add missing pci_enable_device()
Add pci_enable_device()/pci_disable_device(). In the past, drivers often
worked without this, but it is now required in order to route PCI
interrupts correctly.
Add pci_enable_device()/pci_disable_device(). In the past, drivers
often worked without this, but it is now required in order to route
PCI interrupts correctly.
I don't have this hardware, so this has been compiled but not tested.
Add pci_enable_device()/pci_disable_device In the past, drivers often worked
without this, but it is now required in order to route PCI interrupts
correctly. In addition, this driver incorrectly used the IRQ value from PCI
config space rather than the one in the struct pci_dev.
Add pci_enable_device()/pci_disable_device(). In the past, drivers often
worked without this, but it is now required in order to route PCI
interrupts correctly.
Dave Hansen [Tue, 24 Aug 2004 04:20:45 +0000 (21:20 -0700)]
[PATCH] don't pass mem_map into init functions
When using CONFIG_NONLINEAR, a zone's mem_map isn't contiguous, and isn't
allocated in the same place. This means that nonlinear doesn't really have
a mem_map[] to pass into free_area_init_node() or memmap_init_zone() which
makes any sense.
So, this patch removes the 'struct page *mem_map' argument to both of
those functions. All non-NUMA architectures just pass a NULL in there,
which is ignored. The solution on the NUMA arches is to pass the mem_map in
via the pgdat, which works just fine.
To replace the removed arguments, a call to pfn_to_page(node_start_pfn) is
made. This is valid because all of the pfn_to_page() implementations rely
only on the pgdats, which are already set up at this time. Plus, the
pfn_to_page() method should work for any future nonlinear-type code.
Finally, the patch creates a function: node_alloc_mem_map(), which I plan
to effectively #ifdef out for nonlinear at some future date.
Compile tested and booted on SMP x86, NUMAQ, and ppc64.
From: Jesse Barnes <jbarnes@engr.sgi.com>
Fix up ia64 specific memory map init function in light of Dave's
memmap_init cleanups.
Signed-off-by: Jesse Barnes <jbarnes@sgi.com>
From: Dave Hansen <haveblue@us.ibm.com>
Looks like I missed a couple of architectures. This patch, on top of my
previous one and Jesse's should clean up the rest.
From: William Lee Irwin III <wli@holomorphy.com>
x86-64 wouldn't compile with NUMA support on, as node_alloc_mem_map()
references mem_map outside #ifdefs on CONFIG_NUMA/CONFIG_DISCONTIGMEM. This
patch wraps that reference in such an #ifdef.
From: William Lee Irwin III <wli@holomorphy.com>
Initializing NODE_DATA(nid)->node_mem_map prior to calling it should do.
From: Dave Hansen <haveblue@us.ibm.com>
Rick, I bet you didn't think your nerf weapons would be so effective in
getting that compile error fixed, did you?
Applying the attached patch and commenting out this line:
arch/i386/kernel/nmi.c: In function `proc_unknown_nmi_panic':
arch/i386/kernel/nmi.c:558: too few arguments to function `proc_dointvec'
will let it compile.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The following caused some fireworks whilst merging i386 cpu hotplug.
any_online_cpu(0x2) returns 32 on i386 if we're forced to continue past the
only set bit due to the additional find_first_bit in the find_next_bit i386
implementation. Not wanting to change current behaviour in the bitops
primitives and since the NR_CPUS thing is a cpumask issue, i've opted to fix
next_cpu() and first_cpu() instead.
This might save a couple of lines of code.
From: <akpm@osdl.org>
Fix cross-arch ulong/int disaster with find_next_bit().
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andi Kleen [Tue, 24 Aug 2004 04:20:09 +0000 (21:20 -0700)]
[PATCH] New x86-64 merge
This fixes various issues in the previous update, in particular
a kernel without CONFIG_GART_IOMMU should boot now again,
The kernel discoverys PCI BUS<->CPU affinity on AMD systems
now. It is so far used by dma_alloc_coherent to allocate memory
Experimental patches to add this to sysfs exist, but they're not
included yet. On systems with no memory on a CPU this information may
be wrong.
It has a new experimental CONFIG_UNORDERED_IO option. When enabled
it uses write combining for stores to device iomemory mapping. This
may give better performance with some device drivers, but has a slight
risk of breaking drivers (in general if a driver works on ia64,ppc64,sparc64
it should also work). Based on some discussions with Grant Grundler.
It requires the driver to use memory barriers properly. I would be interested
in feedback on any performance changes you're seeing. For a production system I
would recommend to keep it turned off(although I run it on all my systems and
haven't run into any problems yet)
ACPI and Centrino speedstep is enabled now for Nocona systems.
The IOMMU code does lazy merging by default now, which should be safe
and may increase performance on block IO. It also avoids SAC force by default
now.
The machine check code has been improved again, hopefully it is good
now. It will log now machine check events from before the last reset.
And various other fixes.
The x86-64 parts are now gcc 3.5 clean.
And various other fixes
- Update defconfig
- Reset lost ticks on lost time warning, print RIP.
- Make TASK_SIZE test for 32bit (Arjan van de Ven)
- Work around bug in generic code that broke pcibus_to_cpumask
- Actually fix dummy iommu code
- Compile i386 acpi and speedstep-centrino cpufreq modules
- Export cpu_khz
- Fix compilation without GART_IOMMU
- Optimize find_*_bit functions for small fields
- Discover nodes near PCI busses on K8 (Travis Betak, changed by me)
- Optimize gart tlb flush slightly
- Add experimental CONFIG_UNORDERED_IO for unordered IO stores
- Add 32bit emulation for PTRACE_GETEVENTMSG
- Fix kernel_fpu_{begin,end} for preemptive kernels (Alexander Nyberg)
- Readd proper check for biomerge (got lost)
- Set up 32bit vsyscall page for ptrace early
- Add 32bit emulation for lookup_dcookie() for oprofile
- Export copy_page / clear_page
- Use rex prefix in save_init_fpu fxsave (Jan Beulich)
- Make it compile again
- Fix handling of hwdev == NULL (= ISA/LPC devices) in swiotlb
- Convert PCI DMA code to dma devices
- Change IOMMU code to use dummy fallback device instead of hardcoded
NULL tests everywhere.
- Test iommu_sac_force instead of nommu for DAC supported macro
(will cause more drivers to use DAC)
- Harden non IOMMU dma_alloc_consistent code to fail less likely.
- Remove use of strsep in option parsers
- Remove duplicated exports (Arjan van der Ven)
- Fix EFAULT checking in ptrace (John Blackwood)
- Update defconfig
- Remove dead URL from boot/setup.S (R.J. Wysocki)
- Use compat_sigval_t instead of sigval_t32 (Al Viro)
- Nanooptimization in 32bit ptregs calls
- Fix gcc 3.5 compilation in mtrr.h
- Pass pt_regs as pointer to avoid illegal pass by reference (for gcc 3.5)
- Make set_bit take int not long (Harald Dunkel)
- Avoid panic on pci_map_sg and pci_alloc_consistent overflow in GART IOMMU
- Handle large lost time delays in HPET code (Suresh B. Siddha)
- Work around theoretical bugs in prefetch handling (suggested by Jamie Lokier)
- Remove mtrr_strings declaration for gcc 3.5
- Set KBUILD_IMAGE for make rpm (William Lee Irwin III)
- Add iommu=noaperture to not touch the aperture
- Clean up argument parsing for iommu= option
- Export symbols for xchgadd based rwsems (still disabled)
- Define iommu_bio_merge for !CONFIG_GART_IOMMU
- Don't use backwards rep ; movsb for memmove
- Out line bitmap search functions (saves 8k .text, from i386)
- Convert bitmap search functions to 64bit accesses and optimize them
a bit.
- Handle corrupted page tables in page fault handler
- Set iommu_merge (without force) to on by default again.
- Don't do bio merging by default for iommu=merge. This should make it
safe to use again
- Add iommu=biomerge option to enable BIO merging (like old iommu=merge)
- Fix iommu=memaper=... parsing
- More MCE fixes (based on a patch by Eric Morton, heavily changed by me)
- Fix check for banks causing exceptions
- Allow to reinit MCEs later even after mce=off, fix wrong
use of __initdata
to disable at boot, but reenable later.
- Log left over machine checks after boot and resume
- Fix missing prototype warning with CPU_FREQ on
- Fix parsing of noexec=on (Ian Hastie)
- Fix warning in ia32_binfmt.c
- Resync time variable cpu frequency handling with i386
- Resync msr.c with i386
- Add 0x60 level 1 intel cache descriptor (from i386)
- Remove duplicated 32bit ioctls (Arnd Bergmann)
- Enable -msoft-float (from i386)
- Use faster version of FPU hang fix - handle the exception
* a bit experimental, if you see "kernel ... math error" events
in the log please report.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Adam Kropelin [Tue, 24 Aug 2004 04:19:45 +0000 (21:19 -0700)]
[PATCH] preset loops_per_jiffy for faster booting
Adds a kernel boot parameter "lpj=NNN" which allows the operator to specify
the loops-per-jiffy value. This shaves up to a quarter of a second off
boot times, which are critical for embedded appliances.
It's a bit thin, but the code is in __init.
Signed-off-by: Adam Kropelin <akropel1@rochester.rr.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Mika Kukkonen [Tue, 24 Aug 2004 04:19:34 +0000 (21:19 -0700)]
[PATCH] Fix drivers/isdn/hisax/avm_pci.c build warning when !CONFIG_ISAPNP
CC [M] drivers/isdn/hisax/avm_pci.o
drivers/isdn/hisax/avm_pci.c: In function `setup_avm_pcipnp':
drivers/isdn/hisax/avm_pci.c:817: warning: label `ready' defined but not used
Patch is big because I replaced the '} else { ... }' with 'goto ready; }'
and so had to remove one level of indentation from code.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 24 Aug 2004 04:19:22 +0000 (21:19 -0700)]
[PATCH] Make UML build and run
This patch includes the following -
updated defconfig
move uml.lds.S and main.c from arch/um to arch/um/kernel per Sam's suggestions
steal bitops.c from arch/i386
convert all calls to open_private_file to dentry_open
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 24 Aug 2004 04:18:53 +0000 (21:18 -0700)]
[PATCH] UML fixes
The patch below fixes a few UML-specific bugs not related to the rest of the
kernel
a bogus error return and some formatting in the fork code
correct calculation of task.thread.kernel_stack
remove a bogus panic
a couple of fixes to allow UML to boot in the presence of exec-shield
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 24 Aug 2004 04:18:42 +0000 (21:18 -0700)]
[PATCH] UML updates
The patch below brings UML up to date with interface changes and the like
irq.c includes profile.h to bring in a missing definition
use the cpu_{set,clear} interface
use the new get_signal_to_deliver interface
define instruction_pointer
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Fix os_process_pc and os_process_parent for corner cases.
Update os_process_pc and os_process_parent: now a PID can be > 32768 (so
increase number of digits) and make it work even with spaces in the command
name.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Make malloc() call vmalloc if needed. Needed for hostfs on 2.6 host.
From: Oleg Drokin <green@linuxhacker.ru>, Jeff Dike <jdike@addtoit.com>, and
me
If size > 128K, with this patch malloc will call vmalloc; free will detect
whether to call vfree or kfree or __real_free(). The 2.4 version could forget
free()ing something; this has been fixed.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
That code comes from the out_of_memory section; in 2.4 it was correct to put
it for "default:", since it was called when handle_mm_fault() return value was
!= 0, 1, 2, i.e. it was 3, OOM (but the i386 code put it out of line, for
better performance). Here, instead, the OOM case is handled on its own, so if
handle_mm_fault() != from the listed cases we must BUG().
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
SKAS mode is like 4G/4G (here we have actually 3G/3G) for guest processes, so
when checking for kernel stack overflow, we must first make sure we are
checking a kernel-space address. Also, correctly test for stack overflows
(i.e. check if there is less than 1k of stack left; see
arch/i386/kernel/irq.c:do_IRQ()). And also, THREAD_SIZE != PAGE_SIZE * 2, in
general (though this setting is almost never changed, so we didn't notice
this1). Thanks to the good eye of Alex Züpke <azu@sysgo.de> for first seeing
this bug, and providing a test program:
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Handles correctly errno == EINTR in lots of places.
On various places (mostly waitpid() calls) this patch makes sure that if errno
== EINTR on return, then the syscall is endlessly retried. It also defines a
simple generic way to do this.
Signed-off-by: <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
- Correct some silly errors (dereferencing a pointer before checking if it's
!= NULL when creating /proc/sysemu, some error messages)
- separate using_sysemu from sysemu_supported (so to refuse to activate
sysemu if it is not supported, avoiding panics)
- not probe sysemu if in tt mode.
Signed-off-by: <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Adds /proc/sysemu to toggle SYSEMU usage.
Adds /proc/sysemu to toggle SYSEMU usage.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Adds the "nosysemu" command line parameter to disable SYSEMU
Adds the "nosysemu" command line parameter to disable SYSEMU
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Use PTRACE_SCEMU (the so-called SYSEMU) to reduce syscall cost.
Turns off syscall emulation patch for ptrace (SYSEMU) on. SYSEMU is a
performance-patch introduced by Laurent Vivier. It changes behaviour of
ptrace() and helps reducing host context switch rate. To make it working, you
need a kernel patch for your host, too. See
http://perso.wanadoo.fr/laurent.vivier/UML/ for further information.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Folds hostaudio_user.c into hostaudio_kern.c.
Folds hostaudio_user.c into hostaudio_kern.c. A lot of code less. Also note
that I no more update ppos(as I used to do in the 2.4 patch): I checked that
OSS never changes ppos, so hostaudio did the right thing.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Fixes raw() and uses it in check_one_sigio; also fixes a silly panic (EINTR returned by call).
Fixes raw() and uses it in check_one_sigio; also fixes a silly panic (EINTR
returned by call).
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Reduces code in *_user files, by moving it in _kern files if already possible.
Reduces code in *_user files, by moving it in _kern files if already possible.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Avoids compile failure when host misses tkill().
Avoids compile failure when host misses tkill(), by simply using kill() in
that case.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Fixes some little warnings about "Defined but not used ..." by #ifdef'ing
things
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Fixes "fixdep.c" to support arch/um/include/uml-config.h.
You probably saw that if you change one config option, even if
linux/autoconf.h (which is included by everything) changes, the kernel is
smart enough not to recompile everything. But with UML this no more holds.
Why? Because, as you see in this patch, fixdep avoids making anything depend
onto linux/autoconf.h *explicitly*, but nobody taught him to do the same for
arch/um/include/uml-config.h. So apply this patch. Do not say "I don't want
to change the generic Kbuild for one arch": this cannot hurt. It's a bugfix
for us, a no-op for others.
Note: with this patch, fixdep will still add a dependency from a file
containing UML_CONFIG_BYE onto CONFIG_BYE. Since someone could think that
fixdep should grep for [^A-Z_]CONFIG_ rather than simply for CONFIG_, I've
added a comment that ask *not to fix* this "bug".
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The second adds the LEGACY_PTY config option. Without it, with late 2.6 kernels
/dev/ptyxx won't work. In fact, with those kernels, root_fs_toms does not
work, because it's "unable to allocate TTY pair". And removes the dead option
"UNIX98_PTY_COUNT" (just commented out for now).
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Fixes an host fd leak caused by hostfs.
In detail, on 2.4 we used force_delete() to make sure inode were not cached,
and we then close the host file when the inode is cleared; when porting to 2.6
the "force_delete" thing was dropped, and this patch adds a fix for this (by
setting drop_inode = generic_delete_inode). Search for drop_inode in the 2.6
Documentation/filesystems/vfs.txt for info about this.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Avoid that gcc breaks UML with "unit at a time" compilation mode.
Avoid that gcc breaks UML with "unit at a time" compilation mode.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Readds (just for now) ghash.h for UML
Just for now and just for UML; it will go away.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 24 Aug 2004 04:13:46 +0000 (21:13 -0700)]
[PATCH] UML updates
The patch below brings UML up to date with some changes in the rest of the
kernel:
an updated defconfig
checksum.h includes in6.h to get a definition of in6_addr
added a missing cpu_{set,clear} change
removed include/asm-um/module.h since it's really a link
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The main part of UML; it is the last distributed patch for 2.6.7 Removes skas
support from the main UML patch; apply or get conflicts.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Cc: Jeff Dike <jdike@addtoit.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Arjan van de Ven [Tue, 24 Aug 2004 04:12:26 +0000 (21:12 -0700)]
[PATCH] flex mmap for s390(x)
Below is a patch from Pete Zaitcev (zaitcev@redhat.com) to also use the
flex mmap infrastructure for s390(x). The IBM Domino guys *really* seem to
want this.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Arjan van de Ven [Tue, 24 Aug 2004 04:12:13 +0000 (21:12 -0700)]
[PATCH] sysctl tunable for flexmmap
Create /proc/sys/vm/legacy_va_layout. If this is non-zero, the kernel
will use the old mmap layout for all tasks. it presently defaults to zero
(the new layout).
Arjan van de Ven [Tue, 24 Aug 2004 04:12:01 +0000 (21:12 -0700)]
[PATCH] flexmmap patchkit: fix for 32 bit emu for 64 bit arches
Utz Lehmann <u.lehmann@de.tecosim.com> found a problem with the flexmmap
patches on x86-64, what he is seeing is that the 32 bit personality isn't
set at the first point of setting the allocator strategy. The solution is
simple, in binfmt_elf the personality is set so put the pick-layout
function there. Please consider,
Signed-off-by: Arjan van de Ven <arjanv@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ingo Molnar [Tue, 24 Aug 2004 04:11:50 +0000 (21:11 -0700)]
[PATCH] i386 virtual memory layout rework
Rework the i386 mm layout to allow applications to allocate more virtual
memory, and larger contiguous chunks.
- the patch is compatible with existing architectures that either make
use of HAVE_ARCH_UNMAPPED_AREA or use the default mmap() allocator - there
is no change in behavior.
- 64-bit architectures can use the same mechanism to clean up 32-bit
compatibility layouts: by defining HAVE_ARCH_PICK_MMAP_LAYOUT and
providing a arch_pick_mmap_layout() function - which can then decide
between various mmap() layout functions.
- I also introduced a new personality bit (ADDR_COMPAT_LAYOUT) to signal
older binaries that dont have PT_GNU_STACK. x86 uses this to revert back
to the stock layout. I also changed x86 to not clear the personality bits
upon exec(), like x86-64 already does.
- once every architecture that uses HAVE_ARCH_UNMAPPED_AREA has defined
its arch_pick_mmap_layout() function, we can get rid of
HAVE_ARCH_UNMAPPED_AREA altogether, as a final cleanup.
the new layout generation function (__get_unmapped_area()) got significant
testing in FC1/2, so i'm pretty confident it's robust.
Compiles & boots fine on an 'old' and on a 'new' x86 distro as well.
The two known breakages were:
http://www.redhatconfig.com/msg/67248.html
[ 'cyzload' third-party utility broke. ]
http://www.zipworld.com/au/~akpm/dde.tar.gz
[ your editor broke :-) ]
both were caused by application bugs that did:
int ret = malloc();
if (ret <= 0)
failure;
such bugs are easy to spot if they happen, and if it happens it's possible
to work it around immediately without having to change the binary, via the
setarch patch.
No other application has been found to be affected, and this particular
change got pretty wide coverage already over RHEL3 and exec-shield, it's in
use for more than a year.
The setarch utility can be used to trigger the compatibility layout on
x86, the following version has been patched to take the `-L' option:
"setarch -L i386 <command>" will run the command with the old layout.
From: Hugh Dickins <hugh@veritas.com>
The problem is in the flexible mmap patch: arch_get_unmapped_area_topdown
is liable to give your mmap vm_start above TASK_SIZE with vm_end wrapped;
which is confusing, and ends up as that BUG_ON(mm->map_count).
The patch below stops that behaviour, but it's not the full solution:
wilson_mmap_test -s 1000 then simply cannot allocate memory for the large
mmap, whereas it works fine non-top-down.
I think it's wrong to interpret a large or rlim_infinite stack rlimit as
an inviolable request to reserve that much for the stack: it makes much less
VM available than bottom up, not what was intended. Perhaps top down should
go bottom up (instead of belly up) when it fails - but I'd probably better
leave that to Ingo.
Or perhaps the default should place stack below text (as WLI suggested and
ELF intended, with its text defaulting to 0x08048000, small progs sharing
page table between stack and text and data); with a further personality for
those needing bigger stack.
From: Ingo Molnar <mingo@elte.hu>
- fall back to the bottom-up layout if the stack can grow unlimited (if
the stack ulimit has been set to RLIM_INFINITY)
- try the bottom-up allocator if the top-down allocator fails - this can
utilize the hole between the true bottom of the stack and its ulimit, as a
last-resort effort.
Ingo Molnar [Tue, 24 Aug 2004 04:11:37 +0000 (21:11 -0700)]
[PATCH] sched: smt fixes
while looking at HT scheduler bugreports and boot failures i discovered a
bad assumption in most of the HT scheduling code: that resched_task() can
be called without holding the task's runqueue.
This is most definitely not valid - doing it without locking can lead to
the task on that CPU exiting, and this CPU corrupting the (ex-) task_info
struct. It can also lead to HT-wakeup races with task switching on that
other CPU. (this_CPU marking the wrong task on that_CPU as need_resched -
resulting in e.g. idle wakeups not working.)
The attached patch against fixes it all up. Changes:
- resched_task() needs to touch the task so the runqueue lock of that CPU
must be held: resched_task() now enforces this rule.
- wake_priority_sleeper() was called without holding the runqueue lock.
- wake_sleeping_dependent() needs to hold the runqueue locks of all
siblings (2 typically). Effects of this ripples back to schedule() as
well - in the non-SMT case it gets compiled out so it's fine.
- dependent_sleeper() needs the runqueue locks too - and it's slightly
harder because it wants to know the 'next task' info which might change
during the lock-drop/reacquire. Ripple effect on schedule() => compiled
out on non-SMT so fine.
- resched_task() was disabling preemption for no good reason - all paths
that called this function had either a spinlock held or irqs disabled.
Compiled & booted on x86 SMP and UP, with and without SMT. Booted the
SMT kernel on a real SMP+HT box as well. (Unpatched kernel wouldn't even
boot with the resched_task() assert in place.)
Ingo Molnar [Tue, 24 Aug 2004 04:11:26 +0000 (21:11 -0700)]
[PATCH] sched: self-reaping atomicity fix
disable preemption in the self-reap codepath, as such tasks may not be on
the tasklist anymore and CPU-hotplug relies on the tasklist to migrate
tasks.