When the kmem_bufctl_t typedef got added to include/asm-ppc64/types.h,
it got added outside the #ifndef __ASSEMBLY__ section, causing
assembler errors. This patch, from David Gibson, moves it inside the
#ifndef __ASSEMBLY__ region.
Andrew Morton [Wed, 19 May 2004 13:10:35 +0000 (06:10 -0700)]
[PATCH] raid locking fix.
From: Neil Brown <neilb@cse.unsw.edu.au>
Fix bug #2661
Raid currently calls ->unplug_fn under spin_lock_irqsave(), but unplug_fns
can sleep.
After a morning of scratching my head and trying to come up with some that
does less locking, the following is the best I can come up with. I'm not
proud of it but it should work.
If I move "nr_pending" out or rdev into the per-personality structures
(e.g. mirror_info), and if I had "atomic_inc_if_nonzero" I could do with
without locking so much, but random atomic* functions don't seem trivial
Andrew Morton [Wed, 19 May 2004 09:42:48 +0000 (02:42 -0700)]
[PATCH] sir_dev locking fix
From: Martin Diehl <lists@mdiehl.de>
There was a spin_unlock missing in the raw mode tx-completion path. Probably
it slipped through because the raw mode stuff is never reached with my Actisys
hardware.
Andrew Morton [Wed, 19 May 2004 09:42:37 +0000 (02:42 -0700)]
[PATCH] s390: network driver
From: Martin Schwidefsky <schwidefsky@de.ibm.com>
Network driver changes:
- iucv: Make grab_param function SMP safe.
- lcs: Fix null-pointer dereference after unsuccessful set_online.
- qeth: Fix kmalloc flags in qeth_alloc_reply.
- qeth: Show broadcase capability also in route4/6 sysfs attributes.
- qeth: Remove debug code.
- qeth: Add option to qetharp user space interface to strip unused
fields from query arp records.
- qeth: Add shortcut in outbound path for HiperSockets.
- qeth: Add more info to qeth_perf_stats.
- qeth: Add support for direct SNMP interface to OSA express cards.
Andrew Morton [Wed, 19 May 2004 09:41:54 +0000 (02:41 -0700)]
[PATCH] use-before-uninitialized value in ext3(2)_find_ goal
From: Mingming Cao <cmm@us.ibm.com>
There is a uninitialized goal value being referenced in both ext3 and ext2
find goal block functions (ext3_find_goal() and ext2_find_goal()).
In the non-sequential write case, these functions check the goal value(non
zero) before calling ext3(2)_find_near() to find the goal block to
allocate.
Since the goal value is uninitialized(non zero), the ext3(2)_find_near() is
never being called in the non-sequential write, thus ext3(2)_find_goal()
failed to guide a goal block in the random write case.
ext3(2)_new_block() takes the junk goal value and will turn it to goal 0
since it's normally beyond the filesystem block number limit. The fix is
trivial.
Andrew Morton [Wed, 19 May 2004 09:41:43 +0000 (02:41 -0700)]
[PATCH] Fix overzealous use of online cpu iterators
From: Rusty Russell <rusty@rustcorp.com.au>
The IA64 hotplug CPU merge seems to have included some core changes: in
particular the recalc_bh_state() needs to sum for all (including offline)
cpus, since we don't empty the counters on CPU down. The totals printed by
/proc/stat (the first loop) should include offline cpus, too (apparently
printing out the per-cpu lines for offline cpus confuses top).
Andrew Morton [Wed, 19 May 2004 09:41:22 +0000 (02:41 -0700)]
[PATCH] VFS cache sizing fix for small machines
From: Matt Mackall <mpm@selenic.com>
Doing the algebra:
c = (a - b) * 3/2
a' = a - c = a - 3/2(a - b) = (2a - 3a + 3b)/2 = (3b - a)/2
a' >= 0
3b - a >= 0
3b >= a
b >= a/3
nr_free_pages() >= mempages/3
We can indeed get into trouble if we try to load a large kernel on a very
small box (ie kernel reserves more than 2/3 of usable memory). Surprisingly I
haven't hit this, but here's a fix.
Andrew Morton [Wed, 19 May 2004 09:40:49 +0000 (02:40 -0700)]
[PATCH] kNFSd: Remove check on number of threads waiting on user-space.
From: NeilBrown <neilb@cse.unsw.edu.au>
From: "J. Bruce Fields" <bfields@fieldses.org>
Currently we are counting the number of threads already asleep and returning
an immediate NFS4ERR_DELAY (==JUKEBOX) error if more than half are already
asleep.
This patch removes that logic, so instead we only return NFS4ERR_DELAY if an
upcall times out (if it takes more than a second to return).
With the thread counting there is the risk that even when all the relevant
subsystems are responsive, the client may still see occasional NFS4ERR_DELAY
returns just because, by coincidence, several upcalls were initiated at the
same time. I expect clients will delay several seconds before retrying after
NFS4ERR_DELAY, so this will be quite noticeable to users. Sporadic long
delays like this are likely to lead users to suspect a problem somewhere, when
in fact there is none.
The current scheme ensures that we can still process requests not depending on
upcalls, even when all threads would otherwise be tied up waiting on upcalls.
However, this is not something that should happen under normal circumstances;
if a server spends a significant portion of its time with all threads waiting
for upcalls, this a sign that something is seriously wrong.
In such a circumstance (e.g., an ldap server dies), we can, at least, bound
the waiting time to a second without the need for counting threads.
In short, removing the thread-counting will allow us to behave predictably
when things are working, while still allowing some progress when they don't.
It would be a worthwhile project to measure the amount of time threads spend
waiting for upcalls (or for reads, for that matter); if a significant portion
of the time they spend handling requests is spent sleeping, then there's an
opportunity to improve nfsd performance: if we can break the one-to-one
mapping between requests and threads, then we can lower the number of threads
required to keep the nfs server busy.
However, both the currently available options for doing this are problematic:
returning JUKEBOX/DELAY errors at random times will lead to unpredictable
performance, and saving a copy of the request to be processed from scratch
again later is wasteful and makes it difficult to provide correct semantics,
especially in the NFSv4 case.
So for now I believe waits with short timeouts are the best option.
Andrew Morton [Wed, 19 May 2004 09:40:39 +0000 (02:40 -0700)]
[PATCH] kNFSd: Reduce timeout when waiting for idmapper userspace daemon.
From: NeilBrown <neilb@cse.unsw.edu.au>
From: "J. Bruce Fields" <bfields@fieldses.org>
1 second should be plenty of time; if we're going to take longer than that
it's probably better just to return NFS4ERR_DELAY and let the client retry
anyway.
Andrew Morton [Wed, 19 May 2004 09:40:28 +0000 (02:40 -0700)]
[PATCH] kNFSd: Improve idmapper behaviour on failure.
From: NeilBrown <neilb@cse.unsw.edu.au>
From: "J. Bruce Fields" <bfields@fieldses.org>
Slightly better behavior on failed mapping (which may happen either because
idmapd is not running, or because there it has told us it doesn't know the
mapping.):
on name->id (setattr), return BADNAME. (I used ESRCH to
communicate BADNAME, just because it was the first error in
include/asm-generic/errno-base.h that had something to
do with nonexistance of something, and that we weren't
already using.)
id->name (getattr), return a string representation of the numerical
id. This is probably useless to the client, especially
since we're unlikely to accept such a string on a setattr,
but perhaps some client will find it mildly helpful.
Andrew Morton [Wed, 19 May 2004 09:40:08 +0000 (02:40 -0700)]
[PATCH] kNFSd: Protect reference to exp across calls to nfsd_cross_mnt
From: NeilBrown <neilb@cse.unsw.edu.au>
nfsd_cross_mnt can release the reference to the passed svc_export structure
when it returns a different svc_export structure. So we need to make sure we
have a counted reference before, and drop the reference afterwards.
Andrew Morton [Wed, 19 May 2004 09:39:57 +0000 (02:39 -0700)]
[PATCH] kNFSd: Change fh_compose to NOT consume a reference to the dentry.
From: NeilBrown <neilb@cse.unsw.edu.au>
fh_compose currently consumes a reference to the dentry but not the export
point. This is both inconsistent and confusing.
It is better if a routine like this doesn't consume reference points, so with
this patch, it doesn't. This fixes a couple of very subtle and unusual
reference counting errors.
Andrew Morton [Wed, 19 May 2004 09:38:21 +0000 (02:38 -0700)]
[PATCH] SubmittingDrivers completeness
From: Jonathan Corbet <corbet@lwn.net>
I noticed a patch went in to Documentation/SubmittingDrivers which tweaked
the URL for KernelTraffic. Here's a self-serving patch which makes that
section more complete; to be fair, I added two other sites too. Just in
case it's useful.
Andrew Morton [Wed, 19 May 2004 09:38:10 +0000 (02:38 -0700)]
[PATCH] Quota fix 3 - quota file corruption
From: Jan Kara <jack@ucw.cz>
This patch fixes possible quota files corruption which could happen when root
did not have any inodes&space allocated.
Originally this could not happen as structure would not be written to disk in
that case but with journalled quota we need to write even all-zero structure.
The fix is not very nice but change of the format on disk is probably worse (I
made a mistake with not including the usage-bitmaps into format :().
Andrew Morton [Wed, 19 May 2004 09:37:59 +0000 (02:37 -0700)]
[PATCH] SELinux: fix error handling in selinuxfs
From: Stephen Smalley <sds@epoch.ncsc.mil>
This patch against 2.6.6 fixes error handling for two out-of-memory conditions
in selinuxfs, avoiding potential deadlock due to returning without releasing a
semaphore. The patch was submitted by Karl MacMillan of Tresys.
Andrew Morton [Wed, 19 May 2004 09:37:28 +0000 (02:37 -0700)]
[PATCH] mark the `planb' video driver broken
From: Christoph Hellwig <hch@lst.de>
This one is missing updates from the v4l1 interfaces in 2.4 to the 2.6ish
v4l2 and thus doesn't compile. While we're at it also remove the
MOD_{INC,DEC}_USE_COUNT calls in it that were bogus even in 2.4 to avoid
false positives in grep.
If you use O=/someotherdir or KBUILD_OUTPUT=/someotherdir on the following
architectures: alpha, mips, sh and cris, the build process is probably
going to fail at one point or another, depending on the target you used,
because make can't find scripts/Makefile.build or scripts/Makefile.clean.
The following patch fixes this, I greped the whole tree and these four were
the only "offenders" I found.
Andrew Morton [Wed, 19 May 2004 09:36:00 +0000 (02:36 -0700)]
[PATCH] Work around gcc 3.3.3-hammer sched miscompilation on x86-64
From: Andi Kleen <ak@muc.de>
The new domain scheduler got miscompiled on x86-64 with gcc 3.3.3-hammer,
which is shipping with some distributions. The kernel deadlocks eventually
under light stress on SMP systems with the right options.
After some experiments it seems this simple change avoids the
miscompilation. It also doesn't pessimize the code unduly for other
architectures.
Andrew Morton [Wed, 19 May 2004 09:35:49 +0000 (02:35 -0700)]
[PATCH] slab: add kmem_cache_alloc_node
From: Manfred Spraul <manfred@colorfullife.com>
The attached patch adds a simple kmem_cache_alloc_node function: allocate
memory on a given node. The function is intended for cpu bound structures.
It's used for alloc_percpu and for the slab-internal per-cpu structures.
Jack Steiner reported a ~3% performance increase for AIM7 on a 64-way
Itanium 2.
Port maintainers: The patch could cause problems if CPU_UP_PREPARE is
called for a cpu on a node before the corresponding memory is attached
and/or if alloc_pages_node doesn't fall back to memory from another node if
there is no memory in the requested node. I think noone does that, but I'm
not sure.
Andrew Morton [Wed, 19 May 2004 09:35:39 +0000 (02:35 -0700)]
[PATCH] slab: allow arch override for kmem_bufctl_t
From: Manfred Spraul <manfred@colorfullife.com>
The slab allocator keeps track of the free objects in a slab with a linked
list of integers (typedef'ed to kmem_bufctl_t). Right now unsigned int is
used for kmem_bufctl_t, i.e. 4 bytes per-object overhead.
The attached patch implements a per-arch definition of for this type:
Theoretically, unsigned short is sufficient for kmem_bufctl_t and this would
reduce the per-object overhead to 2 bytes. But some archs cannot operate on
16-bit values efficiently, thus it's not possible to switch everyone to
ushort.
The chosen types are a result of dicussions with the various arch maintainers.
Andrew Morton [Wed, 19 May 2004 09:35:27 +0000 (02:35 -0700)]
[PATCH] slab: enable runtime cache line size on i386
From: Manfred Spraul <manfred@colorfullife.com>
the attached patch switches the SLAB_HWCACHE_ALIGN alignment from the
compile time L1 cache line size to the runtime detected value for i386.
x86-64 already uses the runtime detection.
Andrew Morton [Wed, 19 May 2004 09:35:17 +0000 (02:35 -0700)]
[PATCH] Fix arithmetic in shrink_zone()
From: Nick Piggin <nickpiggin@yahoo.com.au>
If the zone has a very small number of inactive pages, local variable
`ratio' can be huge and we do way too much scanning. So much so that Ingo
hit an NMI watchdog expiry, although that was because the zone would have a
had a single refcount-zero page in it, and that logic recently got fixed up
via get_page_testone().
Nick's patch simply puts a sane-looking upper bound on the number of pages
which we'll scan in this round.
It fixes another failure case: if the inactive list becomes very small
compared to the size of the active list, active list scanning (and therefore
inactive list refilling) also becomes small.
This patch causes inactive list scanning to be keyed off the size of the
active+inactive lists. It has the plus of hiding active and inactive
balancing implementation from the higher level scanning code. It will
slightly change other aspects of scanning behaviour, but probably not
significantly.
Andrew Morton [Wed, 19 May 2004 09:34:23 +0000 (02:34 -0700)]
[PATCH] security: add disable param to capabilities module
From: Chris Wright <chrisw@osdl.org>
Add disable param to capabilities module. Similar to the SELinux param for
disabling at boot time. This allows vendors to ship single binary image with
capabilities compiled statically, and disable it if they provide another
security model compiled as module.
Andrew Morton [Wed, 19 May 2004 09:34:12 +0000 (02:34 -0700)]
[PATCH] speed up readahead for seeky loads
From: Ram Pai <linuxram@us.ibm.com>
Currently the readahead code tends to read one more page than it should with
seeky database-style loads. This was to prevent bogus readahead triggering
when we step into the last page of the current window.
The patch removes that workaround and fixes up the suboptimal logic instead.
wrt the "rounding errors" mentioned in this patch, Ram provided the following
description:
Say the i/o size is 20 pages.
Our algorithm starts by a initial average i/o size of 'ra_pages/2' which
is mostly say 16.
Now every time we take a average, the 'average' progresses as follows
(16+20)/2=18
(18+20)/2=19
(19+20)/2=19
(19+20)/2=19.....
and the rounding error makes it never touch 20
Benchmarking sitrep:
IOZONE
run on a nfs mounted filesystem:
client machine 2proc, 733MHz, 2GB memory
server machine 8proc, 700Mhz, 8GB memory
Andrew Morton [Wed, 19 May 2004 09:34:02 +0000 (02:34 -0700)]
[PATCH] dpt_i2o warning fixes
drivers/scsi/dpt_i2o.c: In function `adpt_queue':
drivers/scsi/dpt_i2o.c:442: warning: use of cast expressions as lvalues is deprecated
drivers/scsi/dpt_i2o.c: In function `adpt_scsi_register':
drivers/scsi/dpt_i2o.c:2213: warning: use of cast expressions as lvalues is deprecated
This patch stops the iseries_veth driver trying to send every packet to too
many logical partitions. Consequently, the number of transmit errors falls to
(about) zero from a very large number. This should also improve performance a
bit as the driver is no longer doing 31 extra skb_clone()s and skb_free()s for
each packet.
Andrew Morton [Wed, 19 May 2004 09:33:08 +0000 (02:33 -0700)]
[PATCH] PPC32: Get full register set on bad kernel accesses
From: Paul Mackerras <paulus@samba.org>
At present on ppc32, if the kernel accesses a bad address and causes an
oops, or drops into the xmon debugger, we only have the contents of the
volatile registers available to print. The reason is that we only save the
volatile registers on entry for a page fault.
This patch restructures the code a bit so that if do_page_fault()
determines that the page fault is caused by a bad kernel access, it returns
to the caller, which then saves the full register set into the exception
frame before calling bad_page_fault(). This way we get the full set of
registers printed in the oops message.
Andrew Morton [Wed, 19 May 2004 09:32:36 +0000 (02:32 -0700)]
[PATCH] system_state splitup
Split the system_state state `SYSTEM_SHUTDOWN' into SYSTEM_HALT,
SYSTEM_POWER_OFF and SYSTEM_RESTART and export system_state to modules.
This allows driver shutdown routines to know why they are being shutdown. The
IDE subsystem wants this so that it knows to not spin the disks down across a
reboot.
This allows building the math-emu code as a module only when
CONFIG_SMP is not set. The fp trap handler cannot be preempted
on a single-CPU (as CONFIG_PREEMPT is not going to be supported
on alpha), so the module can be safely unloaded at any time.
Chris Mason [Wed, 19 May 2004 00:18:07 +0000 (17:18 -0700)]
[PATCH] Fix reiserfs inode size update race
reiserfs_file_write unlocks the pages it operated on before updating
i_size. This can lead to races with writepage, who checks i_size when
deciding how much of the file to zero out.
This patch also replaces SetPageReferenced with mark_page_accessed() in
reiserfs_file_write
This was verified to fix the BitKeeper data corruption problems that
Steven Cole has been debugging, where concurrent writes to a file and
writebacks to disk would cause zeroes in the file when CONFIG_PREEMPT
was enabled.
- fix arch/arm/Kconfig to allow IDE only on platforms supporting it
- introduce IDE_ARCH_OBSOLETE_INIT and ide_default_io_ctl() so
we can use generic ide_init_hwif_ports() and kill no longer needed
<asm-arm/arch-*/ide.h> (leave broken lh7a40x and sa1100 versions)
Add drivers/ide/arm/ide_arm.c for simple default IDE interfaces
and clean obsolete ide_init_default_hwifs() implementations
in asm-arm/arch-{cl7500,rpc,shark}/ide.h and asm-arm26/ide.h.
This allows us to kill ide_init_default_hwifs() completely
in the next patch (because lh7a40x and sa1100 are broken).
Scott Feldman [Tue, 18 May 2004 16:39:57 +0000 (12:39 -0400)]
[PATCH] e100: fix for incoherent arches
* Changed mapping on Rx skb to bi-directional. skb->data holds both the
RFD structure and the packet data, and the RFD is read/written by
HW. Issue found on Xscale HW that doesn't handle cache syncs auto-
matically. Other changes in patch are whitespace/spelling.