Andrew Morton [Thu, 27 May 2004 00:36:14 +0000 (17:36 -0700)]
[PATCH] CPU Hotplug: restore Idle task's priority during CPU_DEAD notification
From: Srivatsa Vaddagiri <vatsa@in.ibm.com>
Fix a CPU Hotplug problem wherein idle task's "->prio" value is not
restored to MAX_PRIO during CPU_DEAD handling. Without this patch, once a
CPU is offlined and then later onlined, it becomes "more or less" useless
(does not run any task other than its idle task!)
Ingo said:
The __setscheduler() call is (technically) incorrect because in the
SCHED_NORMAL case the prio should be zero. So it's a bit cleaner to set up
the static priority to MAX_PRIO and then revert the policy to SCHED_NORMAL
via __setscheduler().
Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Andrew Morton <akpm@osdl.org>
Andrew Morton [Thu, 27 May 2004 00:35:42 +0000 (17:35 -0700)]
[PATCH] Fix the setting of file->f_ra on block-special files
We need to set file->f_ra _after_ calling blkdev_open(), when inode->i_mapping
points at the right thing. And we need to get it from
inode->i_mapping->host->i_mapping too, which represents the underlying device.
Also, don't test for null file->f_mapping in the O_DIRECT checks.
Andrew Morton [Thu, 27 May 2004 00:35:31 +0000 (17:35 -0700)]
[PATCH] Set d_bucket correctly for anonymous dentries
From: Neil Brown <neilb@cse.unsw.edu.au>
In researching the oopses reported in bug #2761, Neil came up with:
I have found one problem, but it isn't particularly new and I cannot
see how it would be related.
When d_alloc_anon creates an anonymous dentry, it is put on a special hash
chain for anonymous dentries (sb->s_anon), but d_bucket is set to
d_hash(parent, name_hash)
If, when it is eventually moved to a proper name, that hash value is the same
as the final hash value, it will not be moved to the right bucket, and so it
not be accessible by name. This patch should fix it.
anonymous dentries have their own private hash "bucket" (sb->s_anon) and so
d_bucket should be set to a unique (impossible) address, else d_move will
get confused.
Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org>
Andrew Morton [Thu, 27 May 2004 00:35:21 +0000 (17:35 -0700)]
[PATCH] posix locks oops fix
From: Andreas Gruenbacher <agruen@suse.de>
There is a race between unshare_files() and the following steal_locks().
As a consequence, steal_locks() may steal some additional FL_POSIX locks
that don't belong to the current thread. This triggers a BUG in
locks_remove_flock().
In detail, the current thread shares its files struct with other threads.
This causes unshare_files() to associate the current thread with a copy of
its files_struct. The copy shares all file objects with the original files
struct. In the time between unshare_files() and steal_locks(), another
thread creates a new file and a FL_POSIX lock on it. The current thread
gets into steal_locks() and takes over all FL_POSIX locks that refer to the
previous files_struct, including the new lock. We do
put_files_struct(original files_struct). This causes the file handle to
the new file to be closed. We get into locks_remove_posix() and miss the
lock, because its fl_owner field now refers to the new files_struct.
Finally we get into locks_remove_flock(), and stumble upon the lock.
While looking into this bug report I gathered the following data with a
SUSE kernel (oops and LKCD dump from Chris):
Here's a proposed fix. As a side effect, steal_locks no longer walks the
global list of locks, but only the locks of all open inodes.
What are the reasons (other than historic ones) for not getting rid of
fl_owner and using fl_pid instead, by the way? I think that would clean up
the whole mess with file locks a bit.
Andrew Morton [Thu, 27 May 2004 00:34:58 +0000 (17:34 -0700)]
[PATCH] ppc64 kernel hackers can't spell
From: Anton Blanchard <anton@samba.org>
From: Dave Hansen
This patch is obviously of the utmost importance. It probably doesn't matter
as much for kernel error messages, but one of these mistakes is in a
user-readable /proc file.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Dave Hansen <haveblue@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org>
Paul Mackerras [Wed, 26 May 2004 02:03:18 +0000 (19:03 -0700)]
[PATCH] ppc64: fix nonexistent irq affinity
This fixes a bug where, if we try to set the affinity on an unused
virtual IRQ number on a logically-partitioned pSeries system, we call
the firmware with physical IRQ number = -1, which it doesn't like.
With this patch we just ignore the attempt.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Linus Torvalds [Wed, 26 May 2004 00:56:23 +0000 (17:56 -0700)]
Split ptep_establish into "establish" and "update_access_flags"
ptep_establish() is used to establish a new mapping at COW time,
and it always replaces a non-writable page mapping with a totally
new page mapping that is dirty (and likely writable, although ptrace
may cause a non-writable new mapping). Because it was nonwritable,
we don't have to worry about losing concurrent dirty page bit updates.
ptep_update_access_flags() leaves the same page mapping, but updates
the accessed/dirty/writable bits (it only ever sets them, and never
removes any permissions). Often easier, but it may race with a dirty
bit update on another CPU.
Booted on x86 and ppc64.
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Steven King [Tue, 25 May 2004 06:09:27 +0000 (23:09 -0700)]
[IPSEC]: Fix buglet in AF_KEY spddelete
When trying to spddelete individual entries using setkey, spddelete always
fails. The culprit is in net/af_key.c; spdadd sets the family field of the
selector when creating an entry, but spddelete doesn't when building a
selector to match for xfrm_policy_bysel. Trivial fix is to have spddelete
set the family field in the selector in same way spdadd does.
Linus Torvalds [Tue, 25 May 2004 06:04:59 +0000 (23:04 -0700)]
Introduce architecture-specific "ptep_update_dirty_accessed()"
helper function to write-back the dirty and accessed bits from
ptep_establish().
Right now this defaults to the same old "set_pte()" that we've
always done, except for x86 where we now fix the (unlikely)
race in updating accessed bits and dropping a concurrent dirty
bit.
Herbert Xu [Tue, 25 May 2004 04:01:22 +0000 (21:01 -0700)]
[IPSEC]: Do not leak entries in xfrm_state_find.
In xfrm_state_find, the larval state never actually matures with
Openswan so it only ever gets deleted by the timer which means
that the time crash can't happen :) It becomes a (possible) memory
leak instead.
Paul Mackerras [Tue, 25 May 2004 03:27:46 +0000 (20:27 -0700)]
[PATCH] IRQ stacks for PPC64
Even with a 16kB stack, we have been seeing stack overflows on PPC64
under stress. This patch implements separate per-cpu stacks for
processing interrupts and softirqs, along the lines of the
CONFIG_4KSTACKS stuff on x86. At the moment the stacks are still 16kB
but I hope we can reduce that to 8kB in future. (Gcc is capable of
adding instructions to the function prolog to check the stack pointer
whenever it moves it downwards, and I want to use that when I try
using 8kB stacks so I can be confident that we aren't overflowing the
stack.)
Ingo Molnar [Tue, 25 May 2004 03:06:18 +0000 (20:06 -0700)]
[PATCH] x86-bigsmp: use fixed interrupt delivery
This patch, from Venkatesh Pallipadi, changes x86 IO-APICs to use fixed
interrupt delivery instead of lowest priority to support larger number
of CPUs. Only bigsmp is affected by this cleanup.
Andrew Morton [Tue, 25 May 2004 01:45:35 +0000 (18:45 -0700)]
[PATCH] sched_yield() microoptimisation
Signed-off-by: Ingo Molnar <mingo@elte.hu>
We can avoid the local_irq_enable() in sched_yield() because schedule()
unconditionally enables interrupts anyway.
Andrew Morton [Tue, 25 May 2004 01:45:03 +0000 (18:45 -0700)]
[PATCH] No interpretation of HD spindown timeout in laptop mode ACPI binding script.
From: Bart Samwel <bart@samwel.tk>
Currently the ACPI binding script in the Laptop Mode doc always says "20
seconds" and "2 hours" for the timeouts it uses. This is incorrect if the
user changed the config values, so we print something more general.
Andrew Morton [Tue, 25 May 2004 01:44:52 +0000 (18:44 -0700)]
[PATCH] rmap build fix
From: William Lee Irwin III <wli@holomorphy.com>
PMD_SIZE is not a compile-time constant on sparc. Use min() in there so
that the cluster size will be evaluated at runtime if the architecture
insists on doing that.
Andrew Morton [Tue, 25 May 2004 01:44:29 +0000 (18:44 -0700)]
[PATCH] Revert bogus x86-64 change
From: Andi Kleen <ak@muc.de>
The 32bit generic nops added with a previous patch to x86-64 alternative()
are not completely 64bit clean. This caused crashes in some cases. This
patch reverts this broken change.
Andrew Morton [Tue, 25 May 2004 01:43:49 +0000 (18:43 -0700)]
[PATCH] minor sched.c cleanup
Signed-off-by: Christian Meder <chris@onestepahead.de> Signed-off-by: Ingo Molnar <mingo@elte.hu>
The following obviously correct patch from Christian Meder simplifies the
DELTA() define.
Andrew Morton [Tue, 25 May 2004 01:43:28 +0000 (18:43 -0700)]
[PATCH] remap_file_pages: fix syscall declaration
Signed-off-by: Hugh Dickins <hugh@veritas.com>
sys_remap_file_pages is declared as asmlinkage in mm/fremap.c, but is the one
syscall declared without asmlinkage in include/linux/syscalls.h.
Andrew Morton [Tue, 25 May 2004 01:43:17 +0000 (18:43 -0700)]
[PATCH] remap_file_pages: implement MAP_POPULATE for all protections
Signed-off-by: Hugh Dickins <hugh@veritas.com>
It seems eccentric to implement MAP_POPULATE only on PROT_NONE mappings:
do_mmap_pgoff is passing down prot, then sys_remap_file_pages verifies it's
not set. I guess that's an oversight from when we realized that the prot arg
to sys_remap_file_pages was misdesigned.
There's another oddity whose heritage is harder for me to understand, so
please let me leave it to you: sys_remap_file_pages is declared as asmlinkage
in mm/fremap.c, but is the one syscall declared without asmlinkage in
include/linux/syscalls.h.
Andrew Morton [Tue, 25 May 2004 01:43:06 +0000 (18:43 -0700)]
[PATCH] Fix for lockup in reiserfs acl/xattrs
From: Jeff Mahoney <jeffm@suse.com>
The following is a patch to fix a locking problem in ACL/xattr code. It
manifests when a user attempts to set an xattr on a file which they do
not own, and on which an ACL is applied.
What happens is this:
reiserfs_setxattr [write lock inode xattr sem]
->xattr_set
-> lookup
-> __reiserfs_permission [if conditions above are met, and need_lock=
is
unset, read lock inode xattr sem] *lockup*
Since we already keep track of when to lock during permission calls, the
fix is simple: just make the locking conditional as it was before.
Andrew Morton [Tue, 25 May 2004 01:42:56 +0000 (18:42 -0700)]
[PATCH] UDF: directory reading fix
From: Ben Fennema <bfennema@falcon.csc.calpoly.edu>
The problem occured when files were stored on the disc in 16-bit per
character mode when all the upper bits were 0. The fs module
converted the file name given by the user to a 8-bit per character
string to compare, so the comparison always failed.
The patch maps the file from disc into the current locale and then
compares it directly to the file name given by the user.
Andrew Morton [Tue, 25 May 2004 01:41:49 +0000 (18:41 -0700)]
[PATCH] v4l: use saa7111 i2c module in V4L MXB driver
From: Michael Hunold <hunold@convergence.de>
The attached patch changes my "Multimedia eXtension Board" (MXB)
Video4Linux-driver to use the standard saa7111 video decoder infrastructure
(to which I recently submitted changes through Ronald Bultje) instead of
some home-brewn direct-access stuff.
Nothing serious, but it removes code duplication and makes the code use the
video decoder api.
Andrew Morton [Tue, 25 May 2004 01:41:40 +0000 (18:41 -0700)]
[PATCH] initramfs uncpio fix
From: <viro@parcelfarce.linux.theplanet.co.uk>
init/initramfs.c::do_skip() has an off-by-one that leads to unpacking
failures for some gzipped cpio images. We have
static int __init do_skip(void)
{
if (this_header + count <= next_header) {
eat(count);
return 1;
} else {
eat(next_header - this_header);
state = next_state;
return 0;
}
}
and that <= should actually be <. It almost never matters, since if we hit
the boundary case (header ending exactly on the gunzip window end) the
current variant will simply end up doing extra call of do_skip() when we
get to the next window and that will finish the work (assign state). The
only exception is when we hit that in the last window. That is, if there's
nothing after the final header (trailer). Then we miss the final state
transition (Skip -> Reset) and get "junk in archive" panic. Normally
cpio(1) pads the image to multiple of 512, so we actually have a bunch of
zeroes after the trailer. And that almost always saves our butts - trailer
is followed by zeroes, so we get to Reset state just fine.
So we never see that on small in-kernel image (it's less than 512 bytes, so
it gets a lot of padding) and we almost never see that on external ones
(1:127 odds of hitting the bug).
Andrew Morton [Tue, 25 May 2004 01:41:18 +0000 (18:41 -0700)]
[PATCH] swsusp: fix swsusp with intel-agp
From: Pavel Machek <pavel@suse.cz>
swsusp contained rather nasty bug where it killed machine when intel-agp or
anything else split kernel 4MB mapping. Herbert Xu diagnosed this. Fixed by
switching to "known good" mapping for during suspend/resume.
Andrew Morton [Tue, 25 May 2004 01:40:54 +0000 (18:40 -0700)]
[PATCH] matroxfb: Add support for mapping CRTC<->outputs at boot time
Signed-off-by: Petr Vandrovec <vandrove@vc.cvut.cz>
Some people expressed interest in having possibility to set CRTC <->
outputs mapping at boot time, without having to use 'matroxset' later after
kernel boots.
This patch adds option 'video=matroxfb:outputs:XYZ', where X sets which
CRTC will connect to primary output, Y sets secondary output and Z sets DVI
output.
In addition to that I also added missing memset() into maven, which was
broken since i2c was kobjectified.
Ivan Kokshaysky [Tue, 25 May 2004 01:38:44 +0000 (18:38 -0700)]
[PATCH] fix system clock on ruffian
Unlike most other alphas, ruffian uses i8253 timer instead of RTC
as the system clock source. However, the PIT clock divisor (LATCH)
is bogus since CLOCK_TICK_RATE has been changed to 32 KHz.
Fixed using recently introduced PIT_TICK_RATE macro.
Andrew Morton [Tue, 25 May 2004 01:37:14 +0000 (18:37 -0700)]
[PATCH] Fix race condition with current->group_info
From: Olaf Kirch <okir@suse.de>
I have been chasing a corruption of current->group_info on PPC during NFS
stress tests. The problem seems to be that nfsd is messing with its
group_info quite a bit, while some monitoring processes look at
/proc/<pid>/status and do a get_group_info/put_group_info without any locking.
This problem can be reproduced on ppc platforms within a few seconds if you
generate some NFS load and do a "cat /proc/XXX/status" of an nfsd thread in a
tight loop.
I therefore think changes to current->group_info, and querying it from a
different process, needs to be protected using the task_lock.
(akpm: task->group_info here is safe against exit() because the task holds a
ref on group_info which is released in __put_task_struct, and the /proc file
has a ref on the task_struct).
Andrew Morton [Tue, 25 May 2004 01:36:57 +0000 (18:36 -0700)]
[PATCH] ep_send_events() stack reduction
ep_send_events() uses ~350 bytes of stack for a local buffer of events to send
to userspace. The patch fixes that by removing the double-buffering
altogether. A pipe-based microbenchmark from Davide Libenzi
<davidel@xmailserver.org> was sped up by 1-2%.
Andrew Morton [Tue, 25 May 2004 01:36:46 +0000 (18:36 -0700)]
[PATCH] Fix the mangled-oops-output-on-SMP problem
From: Ingo Molnar <mingo@elte.hu>
printk currently does
if (oops_in_progres)
bust_printk_locks();
which means that once we oops, the printk locking is 100% ineffective and
multiple CPUs make an unreadable mess on a serial console. It's a significant
development hassle.
Fix that up by only popping locks once per ten seconds.
akpm@osdl.org did:
- Bump the timeout to 30 seconds - 9600 baud is slow.
- Handle jiffy wraps: change the logic so that we only skip the lockbust
if the current time is within 30 seconds of the previous lockbusting
attempt.
Andrew Morton [Tue, 25 May 2004 01:36:31 +0000 (18:36 -0700)]
[PATCH] Prevent scary warnings from knfsd
From: "J. Bruce Fields" <bfields@fieldses.org>
The kernel currently prints:
nfsd: nobody listening for auth.unix.ip upcall; has some daemon not been started?
on every bootup, during initscripts.
Neil Brown <neilb@cse.unsw.edu.au> says:
It was part of the recent set of idmapper patches. Bruce wanted the admin
to get a warning when the idmapper daemon wasn't running. I thought the
same warning should apply to any daemon that responded to upcalls.
In the case of auth.unix.ip it isn't strictly necessary for a daemon to be
running (for comparability with 2.4).
You can get rid of the warning by doing:
mount -t nfsd nfsd /proc/fs/nfs
before mountd is started (init scripts should start doing this I hope, but
distributions don't tend to use the init script from nfs-utils, so it is
hard to push it). This will trigger mountd to listen on auth.unix.ip and
others.
That's a hassle, so Bruce's patch limits the warning purely to the new
idmapper cache. It provides a callback in the cache_detail that individual
caches can use to log messages when upcalls fail because a userspace daemon
not running. Implement this method for the idmapping caches.
Andrew Morton [Tue, 25 May 2004 01:35:48 +0000 (18:35 -0700)]
[PATCH] ppc64: avoid bogus real IRQ numbers
Signed-off-by: Paul Mackerras <paulus@samba.org>
Early in the boot process on pSeries machines, we look in the Open Firmware
device tree for information about the interrupt assignments, and assign
virtual IRQ numbers for each physical IRQ. There is currently a couple of
bugs in this code which result in us assigning virtual IRQs for nonexistent
physical IRQs. This causes problems when we call the firmware to enable or
disable those nonexistent physical IRQs. Some versions at least of the
firmware will hit an assertion failure and crash the machine when this
happens.
This patch fixes the bugs and ensures that we don't try and use nonexistent
physical IRQ numbers. One bug was that we were mapping ISA interrupts,
which is unnecessary since virtual IRQ numbers 0 - 15 are reserved for
them. The other was that when we had a PCI interrupt (which is always in
the range 1 to 4, corresponding to INTA to INTD) which didn't have a
mapping in the PCI host bridge above it, we were just using the original
number (usually 1) rather than ignoring it.
Andrew Morton [Tue, 25 May 2004 01:35:31 +0000 (18:35 -0700)]
[PATCH] ppc64: bump IOMMU_MAX_ORDER
Signed-off-by: Anton Blanchard <anton@samba.org>
We have cards that want over 2MB of PCI consistent memory. The
IOMAP_MAX_ORDER limit is just to catch bad drivers early, so we can bump
this a bit.
We want some room to grow but our maximum get_free_pages allocation on
ppc64 is currently 16MB, so it doesnt make sense to go above that.
Andrew Morton [Tue, 25 May 2004 01:34:57 +0000 (18:34 -0700)]
[PATCH] dynamic addition of virtual disks on PPC64 iSeries
From: Stephen Rothwell <sfr@canb.auug.org.au>
This patch allows us to dynamically add virtual disks to an iSeries partition.
It works like this: after you have created the virtual disk file on OS/400
and attached it to the Linux partition, you need to write to
/sys/bus/vio/drivers/viodasd/probe (it doesn't matter what you write). This
will do the probe. It calls add_disk() for each new disk, so we get hotplug
events as a side effect.
This was the nicest way I could think of doing this as the interface to the
hypervisor is polled ...
Andrew Morton [Tue, 25 May 2004 01:34:44 +0000 (18:34 -0700)]
[PATCH] ppc64: fix to viopath.c
From: Anton Blanchard <anton@samba.org>
From: Olaf Hering and Nathan Lynch:
Fix a couple of nasty lurking bugs in viopath.c and add information
required to know if the iseries_veth module should be loaded on legacy
iSeries systems.