Ingo Molnar [Mon, 18 Oct 2004 15:55:37 +0000 (08:55 -0700)]
[PATCH] generic irq subsystem: core
The main goal of this patch is to consolidate all the different but still
fundamentally similar arch/*/kernel/irq.c code into the kernel/irq/ subsystem.
There are 4 new files in the kernel/irq/ directory:
- handle.c: core bits: __do_IRQ() and handle_IRQ_event(),
callable from arch-specific irq.c code.
- manage.c: the main driver apis
- spurious.c: the handling of buggy interrupt sources.
- autoprobe.c: probing of interrupts - older code but still in use.
- proc.c: /proc/irq/ code.
- internals.h for irq-core-internal interfaces not visible to drivers
nor arch PIC code.
An architecture enables the generic hardirq code by defining
CONFIG_GENERIC_HARDIRQS in its arch Kconfig. People doing this conversion
should check out the x86/x64/ppc/ppc64 patches for details - the conversion is
quite straightforward but every converted function (i.e. every function
removed from the arch irq.c) _must_ be matched to the generic version and if
there is any detail that the generic code should do it has to be added to the
generic code. All of the currently converted 4 architectures were converted
like that, and the generic code was extended/fixed along the way.
Other changes related to this patchset:
- clean up the irq include files (linux/irq.h, linux/interrupt.h,
linux/hardirq.h) and consolidate asm-*/[hard]irq.h. Note, to keep all
non-touched architectures in an untouched state this consolidation is
done carefully and strictly under CONFIG_GENERIC_HARDIRQS.
Once the consolidation is done we can do a couple of final cleanups
to reach the following logical splitup of 3 include files:
linux/interrupt.h: driver-visible APIs and details
linux/irq.h: core irq and arch-PIC code, internals
asm-*/irq.h: arch PIC and irq delivery details
the following include files will likely vanish:
linux/hardirq.h merges into linux/irq.h
asm-*/hardirq.h: merges into asm-*/irq.h
asm-*/hw_irq.h: merges into asm-*/irq.h
Christoph would like to do these once the current wave of
cleanups gets in.
Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Christoph Hellwig <hch@lst.de> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Gregory Kurz [Mon, 18 Oct 2004 15:55:24 +0000 (08:55 -0700)]
[PATCH] fork() bug invalidates file descriptors
Take a process P1 that spawns a thread T (aka. a clone with CLONE_FILES).
If P1 forks another process P2 (aka. not a clone) while T is blocked in a
open() that should return file descriptor FD, then FD will be unusable in
P2. This leads to strange behaviors in the context of P2: close(FD)
returns EBADF, while dup2(a_valid_fd, FD) returns EBUSY and of course FD is
never returned again by any syscall...
/*
* This program is meant to show that calling fork() while a clone spawned
* with CLONE_FILES is blocked in open() makes a fd number unusable in the
* child.
*
*
* Parent Clone Child
* |
* clone(CLONE_FILES)-
Hugh Dickins [Mon, 18 Oct 2004 15:54:48 +0000 (08:54 -0700)]
[PATCH] __set_page_dirty_nobuffers mappings
Marcelo noticed that the BUG_ON in __set_page_dirty_nobuffers doesn't make
much sense: it lost its way in 2.6.7, amidst so many page_mappings!
It's supposed to be checking that, although page->mapping may suddenly go NULL
from truncation, and although tmpfs swizzles page_mapping(page) between tmpfs
inode address_space and swapper_space, there's sufficient stabilization while
here in __set_page_dirty_nobuffers that the mapping after we locked
mapping->tree_lock is the same as the mapping before we locked
mapping->tree_lock i.e. the lock we hold is the right one.
Roland McGrath [Mon, 18 Oct 2004 15:54:38 +0000 (08:54 -0700)]
[PATCH] exec: fix posix-timers leak and pending signal loss
I've found some problems with exec and fixed them with this patch to
de_thread.
The second problem is that a multithreaded exec loses all pending signals.
This is violation of POSIX rules. But a moment's thought will show it's
also just not desireable: if you send a process a SIGTERM while it's in the
middle of calling exec, you expect either the original program in that
process or the new program being exec'd to handle that signal or be killed
by it. As it stands now, you can try to kill a process and have that
signal just evaporate if it's multithreaded and calls exec just then. I
really don't know what the rationale was behind the de_thread code that
allocates a new signal_struct. It doesn't make any sense now. The other
code there ensures that the old signal_struct is no longer shared. Except
for posix-timers, all the state there is stuff you want to keep. So my
changes just keep the old structs when they are no longer shared, and all
the right state is retained (after clearing out posix-timers).
The final bug is that the cumulative statistics of dead threads and dead
child processes are lost in the abandoned signal_struct. This is also
fixed by holding on to it instead of replacing it.
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Lev Makhlis [Mon, 18 Oct 2004 15:54:26 +0000 (08:54 -0700)]
[PATCH] show aggregate per-process counters in /proc/PID/stat 2
Add up resource usage counters for live and dead threads to show aggregate
per-process usage in /proc/<pid>/stat. This mirrors the new getrusage()
semantics. /proc/<pid>/task/<tid>/stat still has the per-thread usage.
After moving the counter aggregation loop inside a task->sighand lock to
avoid nasty race conditions, it has survived stress-testing with '(while
true; do sleep 1 & done) & top -d 0.1'
Signed-off-by: Lev Makhlis <mlev@despammed.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Arnd Bergmann [Mon, 18 Oct 2004 15:54:02 +0000 (08:54 -0700)]
[PATCH] add missing linux/syscalls.h includes
I found that the prototypes for sys_waitid and sys_fcntl in
<linux/syscalls.h> don't match the implementation. In order to keep all
prototypes in sync in the future, now include the header from each file
implementing any syscall.
Ingo Molnar [Mon, 18 Oct 2004 15:53:48 +0000 (08:53 -0700)]
[PATCH] softirqs: fix latency of softirq processing
The attached patch fixes a local_bh_enable() buglet: we first enabled
softirqs then did we do local_softirq_pending() - often this is preemptible
code. So this task could be preempted and there's no guarantee that
softirq processing will occur (except the periodic timer tick).
The race window is small but existent. This could result in packet
processing latencies or timer expiration latencies - hard to detect and
annoying bugs.
The fix is to invoke softirqs with softirqs enabled but preemption still
disabled. Patch is against 2.6.9-rc2-mm1.
Roland McGrath [Mon, 18 Oct 2004 15:53:35 +0000 (08:53 -0700)]
[PATCH] fix PTRACE_ATTACH race with real parent's wait calls
There is a race between PTRACE_ATTACH and the real parent calling wait.
For a moment, the task is put in PT_PTRACED but with its parent still
pointing to its real_parent. In this circumstance, if the real parent
calls wait without the WUNTRACED flag, he can see a stopped child status,
which wait should never return without WUNTRACED when the caller is not
using ptrace. Here it is not the caller that is using ptrace, but some
third party.
This patch avoids this race condition by adding the PT_ATTACHED flag to
distinguish a real parent from a ptrace_attach parent when PT_PTRACED is
set, and then having wait use this flag to confirm that things are in order
and not consider the child ptraced when its ->ptrace flags are set but its
parent links have not yet been switched. (ptrace_check_attach also uses it
similarly to rule out a possible race with a bogus ptrace call by the real
parent during ptrace_attach.)
While looking into this, I noticed that every arch's sys_execve has:
current->ptrace &= ~PT_DTRACE;
with no locking at all. So, if an exec happens in a race with
PTRACE_ATTACH, you could wind up with ->ptrace not having PT_PTRACED set
because this store clobbered it. That will cause later BUG hits because
the parent links indicate ptracedness but the flag is not set. The patch
corrects all the places I found to use task_lock around diddling ->ptrace
when it's possible to be racing with ptrace_attach. (The ptrace operation
code itself doesn't have this issue because it already excludes anyone else
being in ptrace_attach.)
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Mon, 18 Oct 2004 15:53:22 +0000 (08:53 -0700)]
[PATCH] add WCONTINUED support to wait4 syscall
POSIX specifies the new WCONTINUED flag for waitpid, not just for waitid.
I overlooked this addition when I implemented waitid. The real work was
already done to support waitid, but waitpid needs to report the results
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Mon, 18 Oct 2004 15:53:09 +0000 (08:53 -0700)]
[PATCH] make rlimit settings per-process instead of per-thread
POSIX specifies that the limit settings provided by getrlimit/setrlimit are
shared by the whole process, not specific to individual threads. This
patch changes the behavior of those calls to comply with POSIX.
I've moved the struct rlimit array from task_struct to signal_struct, as it
has the correct sharing properties. (This reduces kernel memory usage per
thread in multithreaded processes by around 100/200 bytes for 32/64
machines respectively.) I took a fairly minimal approach to the locking
issues with the newly shared struct rlimit array. It turns out that all
the code that is checking limits really just needs to look at one word at a
time (one rlim_cur field, usually). It's only the few places like
getrlimit itself (and fork), that require atomicity in accessing a whole
struct rlimit, so I just used a spin lock for them and no locking for most
of the checks. If it turns out that readers of struct rlimit need more
atomicity where they are now cheap, or less overhead where they are now
atomic (e.g. fork), then seqcount is certainly the right thing to use for
them instead of readers using the spin lock. Though it's in signal_struct,
I didn't use siglock since the access to rlimits never needs to disable
irqs and doesn't overlap with other siglock uses. Instead of adding
something new, I overloaded task_lock(task->group_leader) for this; it is
used for other things that are not likely to happen simultaneously with
limit tweaking. To me that seems preferable to adding a word, but it would
be trivial (and arguably cleaner) to add a separate lock for these users
(or e.g. just use seqlock, which adds two words but is optimal for readers).
Most of the changes here are just the trivial s/->rlim/->signal->rlim/.
I stumbled across what must be a long-standing bug, in reparent_to_init.
It does:
memcpy(current->rlim, init_task.rlim, sizeof(*(current->rlim)));
when surely it was intended to be:
memcpy(current->rlim, init_task.rlim, sizeof(current->rlim));
As rlim is an array, the * in the sizeof expression gets the size of the
first element, so this just changes the first limit (RLIMIT_CPU). This is
for kernel threads, where it's clear that resetting all the rlimits is what
you want. With that fixed, the setting of RLIMIT_FSIZE in nfsd is
superfluous since it will now already have been reset to RLIM_INFINITY.
The other subtlety is removing:
tsk->rlim[RLIMIT_CPU].rlim_cur = RLIM_INFINITY;
in exit_notify, which was to avoid a race signalling during self-reaping
exit. As the limit is now shared, a dying thread should not change it for
others. Instead, I avoid that race by checking current->state before the
RLIMIT_CPU check. (Adding one new conditional in that path is now required
one way or another, since if not for this check there would also be a new
race with self-reaping exit later on clearing current->signal that would
have to be checked for.)
The one loose end left by this patch is with process accounting.
do_acct_process temporarily resets the RLIMIT_FSIZE limit while writing the
accounting record. I left this as it was, but it is now changing a limit
that might be shared by other threads still running. I left this in a
dubious state because it seems to me that processing accounting may already
be more generally a dubious state when it comes to NPTL threads. I would
think you would want one record per process, with aggregate data about all
threads that ever lived in it, not a separate record for each thread.
I don't use process accounting myself, but if anyone is interested in
testing it out I could provide a patch to change it this way.
One final note, this is not 100% to POSIX compliance in regards to rlimits.
POSIX specifies that RLIMIT_CPU refers to a whole process in aggregate, not
to each individual thread. I will provide patches later on to achieve that
change, assuming this patch goes in first.
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Pavel Machek [Mon, 18 Oct 2004 15:52:31 +0000 (08:52 -0700)]
[PATCH] swsusp: progress in percent
swsusp currently has very poor progress indication. Thanks to Erik Rigtorp
<erik@rigtorp.com>, we have percentages there, so people know how long wait
to expect. Please apply,
From: Erik Rigtorp <erik@rigtorp.com> Signed-off-by: Pavel Machek <pavel@suse.cz> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andrea Arcangeli [Mon, 18 Oct 2004 15:52:19 +0000 (08:52 -0700)]
[PATCH] parport_pc superio chip fixes
This patch fixes some troubles that somebody reported me with the superio
chips.
In short rmmod parport_pc && cat /proc/iomem was good enough for crashing
the box hard on some machine (and hwscan --printer was doing just that).
The way the oops triggers is that iomem tries to vsprintf the p->name, but
the p->name was a static string in the module address (now unloaded).
The reason is that the superio chip scanning leaves up to two persistent
ranges claimed. But the second (legacy) pass has no way to notice the
resources are already reclaimed. Plus if the superio->io was different
than the "io" variable (the range to scan for superio chips) the "io" range
would generate a leak of the original "io" range too.
I simply make sure to always release the requested space during the superio
scan, and I make sure not to istantiate new ranges in the p->base that
would cause the later parport scan to fail too (plus leaving up to leaked
resources).
The previous code that was returning values and was leaving garbage in
there made no sense to me. My best guess (assuming I didn't misread it ;)
is that probably somebody added the request_region without realizing
they're pointing to the very same address that would be requested later
(and nobody does accesses on those ranges until later, so it was very safe
to claim it later).
Disclaimer: I don't have the specs of the winbond and smsc at hand, I just
guessed what they do from the code (nothing checks superio->io except
get_superio_dma get_superio_irq, which made the thing enough self
explainatory to fix it without specs)
Signed-off-by: Andrea Arcangeli <andrea@novell.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Seth Rohit [Mon, 18 Oct 2004 15:52:07 +0000 (08:52 -0700)]
[PATCH] add sys_setaltroot()
Add a new system call setaltroot(2).
Currently, using the altroot feature is accessible only via the
set_personality() system call. It is accessible to user space only if there
is more than one exec domain in the system. This patch allows using the
altroot feature on systems where there is only one exec domain.
It is possible to work around the issue by adding a dummy exec domain, but it
was rejected for not being very elegant.
If this feature is implemented in userspace, it adds a 16% overhead on a test
case which greps for a single word in the kernel source tree.
Signed-off-by: Zou Nanhai <nanhai.zou@intel.com> Signed-off-by: Gordon Jin <gordon.jin@intel.com> Signed-off-by: Arun Sharma <arun.sharma@intel.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
None of the compatibility defines make sense for assembly
files, and gcc has trouble with vararg macros when using
"-traditional" (which is used for asm), to the point of
ICE'ing.
[PATCH] ppc32/64: FPU/vector register restore after signal
This fixes some issues with restoring the altivec and/or FPU registers
upon return from a signal or when setting a context. It also add a
proper stack backlink to the signal frames created for 64 bits
applications.
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Mike Miller [Mon, 18 Oct 2004 09:52:20 +0000 (04:52 -0500)]
[PATCH] cciss: fixes for clustering
This patch changes our open specifically for clustering software. We must
allow root to access any volume or device with a LUN ID. We also modified
our revalidate function for this reason.
If a logical is reserved, we must register it with the OS with size=0. Then
the backup system can call BLKRRPART after breaking the reservation to
set the device to the correct size.
We also must register a controller with no logical volumes for the online
utilities to function. This is the way we've done it since the 2.2 kernel.
Which doesn't neccesarily make it right, but we have legacy apps to consider.
Signed off by: Mike Miller <mike.miller@hp.com> Signed-off-by: James Bottomley <James.Bottomley@SteelEye.com>
Ben Dooks [Mon, 18 Oct 2004 23:56:50 +0000 (00:56 +0100)]
[ARM PATCH] 2144/1: S3C2410 - s3c2440 fixes and clock updates
Patch from Ben Dooks
Fixes the following problems and ommisions:
- added variable for base crystal rate
- moved clock variables into clock.c
- fixed bug in identifying s3c2440 cpus
- added initial support for new uart registration
- removed base blocks from include/asm/arch/hardware.h
Ben Dooks [Mon, 18 Oct 2004 23:48:44 +0000 (00:48 +0100)]
[ARM PATCH] 2131/1: Add _iomem to the IO string functions
Patch from Ben Dooks
This patch stops mtd from generating problems of
casting pointers to ints, due to the memcpy_fromio
and related functions all taking `unsigned long`
for their IO addresses.
this also found a real bug, qla2xxx isn't iounmapping at host removal at
all currently - and if the right cpp macro would have been set it'd be
too late.
Signed-off-by: James Bottomley <James.Bottomley@SteelEye.com>
Oliver Neukum [Mon, 18 Oct 2004 01:22:09 +0000 (18:22 -0700)]
[PATCH] security issue in firmware system
The firmware loader has a security issue. Firmware on some devices can
write to all memory through DMA. Therefore the ability to feed firmware
to the kernel is equivalent to writing to /dev/kmem. CAP_SYS_RAWIO is
needed to protect itself.
[ Editors note: the firmware file is 0644, and owned by root, so this
"security issue" is really only an issue for people who use
capabilities explicitly, rather than the regular Unix permissions.
This patch makes it do the same checks we do for /dev/mem etc. ]
Signed-Off-By: Oliver Neukum <oliver@neukum.name> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Adrian Bunk <bunk@stusta.de> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nathan Lynch [Sun, 17 Oct 2004 02:21:08 +0000 (19:21 -0700)]
[PATCH] ppc64: fix smp_startup_cpu for cpu hotplug
This change is needed in order to allow cpus to be onlined after
boot. This used to work but the declaration of
pseries_secondary_smp_init in this file was changed in Ben's big
cleanup patch a while back, so the cpu would start at a bad address.
Nick Piggin [Sun, 17 Oct 2004 02:20:56 +0000 (19:20 -0700)]
[PATCH] kswapd lockup fix
Fix some bugs in the kswapd logic which can cause kswapd lockups.
The balance_pgdat() logic is supposed to cause kswapd to loop across all zones
in the node until each zone either
a) has enough pages free or
b) is deemed to be in an "all pages unreclaimable" state.
In the latter case, we just give the zone a light scan on each balance_pgdat()
scan and wait for the zone to come back to life again.
But the zone->all_unreclaimable logic is broken - if the zone has no pages on
the LRU at all, we perform no scanning of that zone (of course). So the
zone->pages_scanned is not incremented and the expression
if (zone->pages_scanned > zone->present_pages * 2)
zone->all_unreclaimable = 1;
so if the zone has no LRU pages it will still enter the all_unreclaimable
state.
Another problem is that if the zone has no LRU pages we will tell
shrink_slab() that we scanned zero LRU pages. This causes shrink_slab() to
scan zero slab objects, which is obviously wrong. So change shrink_slab() to
perform a decent chunk of slab scanning in this situation.
And put a cond_resched() into the balance_pgdat() outer loop. Probably
unnecessary, but that's what Jeff had in place when he confirmed that this
patch fixed the lockup :(
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Pavel Machek [Sun, 17 Oct 2004 02:20:42 +0000 (19:20 -0700)]
[PATCH] swsusp: fix x86-64 - do not use memory in copy loop
In assembly code, there are some problems with "nosave" section (linker was
doing something stupid, like duplicating the section). We attempted to fix
it, but fix was worse then first problem. This fixes is for good: We no
longer use any memory in the copy loop. (Plus it fixes indentation and
uses meaningful labels.)
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
As Milton noticed, Anton actually broke the logic if the memory isn't
aligned in the first place. Sorry about this mess for such a little
piece of code. This _really_ fixes is it all
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Sat, 16 Oct 2004 08:03:14 +0000 (01:03 -0700)]
[PATCH] ppc64: fix some issues with mem_reserve
I found a couple of issues with reserve_mem:
- If we try and mem_reserve something of zero length, everything
reserved after it would get ignored. This is because early_reserve_mem
sees a zero length as a terminator.
- The code rounded the top down instead of up.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nowadays, it's possible to build CONFIG_PPC_PMAC without CONFIG_PPC_PSERIES,
in which case, eeh will not be included in the build (and the eeh checks are
turned into no-ops). However, we then "lose" the iomap functions. This patch
moves them to a separate file.
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andrew Morton [Sat, 16 Oct 2004 08:02:35 +0000 (01:02 -0700)]
[PATCH] ext3 direct io assert fix
Fix bug identified by Badari Pulavarty <pbadari@us.ibm.com>
Local variable `handle' will become stale if ext3_direct_io_get_blocks()
closes off the current transaction and starts a new one. This causes a BUG in
journal_stop().
So reacquire the handle from *current after performing the I/O.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Alexander Viro [Fri, 15 Oct 2004 11:01:38 +0000 (07:01 -0400)]
[PATCH] typhoon.c missing include
DMA_32BIT_MASK is declared in linux/dma-mapping.h; not all platforms get
it from already included headers, so we need explicit include here (fixes
breakage at least on alpha and sparc64).
Signed-off-by: Al Viro <viro@parcelfarce.linux.theplanet.co.uk>
Alan Stern [Fri, 15 Oct 2004 09:20:00 +0000 (04:20 -0500)]
[PATCH] Let LLD specify INQUIRY length
That sounds like a good suggestion. Even better, instead of adding a new
field we can simply use the existing inquiry_length.
This patch changes scsi_probe_lun() to use the value in
sdev->inquiry_length for the first INQUIRY attempt, if that value is
nonzero. Subsequent attempts are based, as before, on the blacklist flags
and the Additional Length field in the INQUIRY data.
The patch also contains a fairly extensive reorganization of the
subroutine. All the code that was duplicated for sending the INQUIRY
command twice has been consolidated. The routine now makes up to three
passes:
In the first pass, the transfer length is the value initially
found in sdev->inquiry_length if that has been set, otherwise
it is the current conservative 36 bytes.
If the first pass succeeds, the routine retrieves the blist flags
for the device and checks the Additional Length field. The blist
flags take precedence over sdev->inquiry_length, which in turn
takes precedence over the Additional Length. If it turns out
there is more data available than we transferred the first time,
a second pass tries to get it.
If the second pass succeeds the INQUIRY data may have changed,
so the blist flags are looked up again and the Additional Length
is checked again. If not, a third pass tries to get the data
back, using the same transfer length as the first pass.
Finally, the value stored in sdev->inquiry_length is set to the amount
actually transferred or the size computed from the Additional Length,
whichever is smaller.
Although the net change in the source file size is small, the new routine
has more comments and less code. Overall I think it's an improvement.
Signed-off-by: Alan Stern <stern@rowland.harvard.edu> Signed-off-by: James Bottomley <James.Bottomley@SteelEye.com>
John Rose [Fri, 15 Oct 2004 05:11:26 +0000 (22:11 -0700)]
[PATCH] PCI Hotplug: rpaphp safe list traversal
Hoping you will accept this fix. The bug can cause a crash upon hotplug
remove. The bug involves unsafe traversal of a list while deleting list
members. The fix uses list_for_each_safe() rather than
list_for_each(). Also threw in an initialization to get rid of a
compiler warning.
Signed-off-by: John Rose <johnrose@austin.ibm.com> Signed-off-by: Greg Kroah-Hartman <greg@kroah.com>
James Bottomley [Fri, 15 Oct 2004 04:46:27 +0000 (23:46 -0500)]
SCSI: Fix problems with non-power-of-two sector size discs
We can't support them, but the system should disable them cleanly
and continue when they're detected (at the moment it
dumps a stack trace).
The fix (hack) is to set them to zero size and 512 byte
sectors. This means they're still amenable to ioctls (like
to reformat them with a useful block size) but cannot
be read from or written to.
Signed-off-by: James Bottomley <James.Bottomley@SteelEye.com>
Linus Torvalds [Thu, 14 Oct 2004 04:03:12 +0000 (21:03 -0700)]
Take the whole PCI bus range into account when scanning PCI bridges.
A bridge that has been set up by firmware to cover multiple PCI
buses but doesn't actually have anything connected behind some of
them caused us to use the incorrect maxmimum bus number span when
scanning the bridge chip.
Problem reported by Tim Saunders, with Russell King suggesting
the fix.
Linus Torvalds [Thu, 14 Oct 2004 04:00:06 +0000 (21:00 -0700)]
Fix threaded user page write memory ordering
Make sure we order the writes to a newly created page
with the page table update that potentially exposes the
page to another CPU.
This is a no-op on any architecture where getting the
page table spinlock will already do the ordering (notably
x86), but other architectures can care.
Add a memory barrier to the assembly checksum code - the code was copied
straight from the i386 one, and the patch resyncs the code with the
original. I'll check if the original code can be included directly (i.e.
"#include") after 2.6.9.
Without this patch, every 2.6 UML release corrupts the checksum of every
UDP fragmented packet with size >= MTU (verified by various people, we all
agree on this issue; nobody reported "Works fine here"). The corrupted
packets are not accepted, thus blocking any kind of communication with
large-sized UDP packets.
In fact, I've even dissected the UML -> host traffic before and after this
patch with Ethereal - and it always reported an incorrect checksum for
fragmented UDP packets before and always correct after applying the patch.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: use always a separate io thread for UBD
Currently, ubd=sync is different from replacing ubd#= with ubd#s=. This is
against Principle of Least Surprise, so remove this difference.
Also the current ubd=sync behaviour is completely useless: it is to make sure
that when the kernel has synched its I/O to the virtual disk, the host does
not invalidate this with his caching; this causes ReiserFS corruption.
But since actually we call end_request() only after the io_thread has done its
work, we never lie to the block layer. Using O_SYNC as we do when replacing
ubd#= with ubd#s= is enough.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
From: BlaisorBlade <blaisorblade_spam@yahoo.it>, Chris Wright <chrisw@osdl.org>
Avoid deadlocking onto the request lock in the UBD driver, i.e. don't lock
the queue spinlock when called from the request function.
In detail:
Rename ubd_finish() to __ubd_finish() and remove ubd_io_lock from it. Add
wrapper, ubd_finish(), which grabs lock before calling __ubd_finish(). Update
do_ubd_request to use the lock free __ubd_finish() to avoid deadlock. Also,
apparently prepare_request is called with ubd_io_lock held, so remove locks
there.
Signed-off-by: Chris Wright <chrisw@osdl.org> Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Makes the UML build system work well even under parallel make (tested, so far,
even with -j50). Please notice that it must be updated for every makefile
change. Or better, every makefile change must use correct dependencies (and
they are easy to miss).
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Uml-specific patch (which requires a mainline hook, mailed separately).
This patch avoid the linking kludge which leaves kbuild link vmlinux and then
link it with libc inside linux. This kludge has the big problem of making
kallsyms break, since the kallsyms pass is done on a completely
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: no extraversion in arch/um/Makefile for mainline
Extraversion in arch/um/Makefile is not needed in mainline, but just for
separate patches; also, they should set it in the main Makefile, not elsewhere
(Jeff Garzik has just complained). Also remove the dependency from version.h
on arch/um/Makefile: it was added because arch/um/Makefile could change the
kernel version number.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
This forces make to use bash rather than whatever /bin/sh is linked to.
Without this, since there are some bash extensions used in the build and when
/bin/sh isn't bash, then the build fails without a clear error message.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: Set cflags before including arch Makefile
If arch/$(ARCH)/Makefile is included before adding -O2 (and the rest) to
CFLAGS, I must duplicate the addition of it to USER_CFLAGS for UML. So let's
fix this. Also, the below code is useless, since if CONFIG_DEBUG_INFO is y,
then CONFIG_FRAME_POINTER is always y.
Add some updates for API changes in 2.6.8 which were not included in the
original UML patch; these fixes were detected by some warnings, so I probably
missed some more ones.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
James Morris [Wed, 13 Oct 2004 14:28:10 +0000 (07:28 -0700)]
[PATCH] SELinux: fix bugs in mprotect hook
The patch below by Roland McGrath fixes two bugs in the implementation of
the selinux_file_mprotect hook:
It calls selinux_file_mmap, which has two problems. First, the stacked
security module will get both mmap and mprotect callbacks for an
mprotect call, which is wrong. Secondly, the vm_flags value contains
VM_* bits, and these do not match the MAP_* bits of the same name or
function, so it passes bogus flags and causes every mprotect to be
treated as if MAP_SHARED were in use.
The patch shares the common code while not having one function call the
other, and fixes these two bugs.
Signed-off-by: James Morris <jmorris@redhat.com> Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
This fixes a bug in SELinux to retain the ptracer SID (if any) across fork.
Otherwise, SELinux will always deny attempts by traced children to exec
domain-changing programs even if the policy would have allowed the tracer
to trace the new domains as well.
Signed-off-by: Stephen Smalley <sds@epoch.ncsc.mil> Signed-off-by: James Morris <jmorris@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Tim Schmielau [Wed, 13 Oct 2004 14:27:49 +0000 (07:27 -0700)]
[PATCH] Fix reporting of process start times
Derive process start times from the posix_clock_monotonic notion of uptime
instead of "jiffies", consistent with the earlier change to /proc/uptime
itself.
(http://linus.bkbits.net:8080/linux-2.5/cset@3ef4851dGg0fxX58R9Zv8SIq9fzNmQ?na%0Av=index.html|src/.|src/fs|src/fs/proc|related/fs/proc/proc_misc.c)
Process start times are reported to userspace in units of 1/USER_HZ since
boot, thus applications as procps need the value of "uptime" to convert
them into absolute time.
Currently "uptime" is derived from an ntp-corrected time base, but process
start time is derived from the free-running "jiffies" counter. This
results in inaccurate, drifting process start times as seen by the user,
even if the exported number stays constant, because the users notion of
"jiffies" changes in time.
It's John Stultz's patch anyways, which I only messed up a bit, but since
people started trading signed-off lines on lkml:
Signed-off-by: Tim Schmielau <tim@physik3.uni-rostock.de> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Olaf Kirch [Wed, 13 Oct 2004 14:27:25 +0000 (07:27 -0700)]
[PATCH] auth_domain_lookup fix
This patch makes sure that auth_domain_lookup returns NULL when it doesn't
find a matching entry, rather than the last entry in the hash chain.
Signed-off-by: Olaf Kirch <okir@suse.de> Acked-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>