Russell King [Tue, 19 Oct 2004 17:11:22 +0000 (18:11 +0100)]
[SERIAL] Keep trying to register our console device.
Some serial drivers receive their serial port device information via
the device model. This unfortunately means that the selected port
may not be available when the console subsystem initialises, so we
must keep trying to register the console after each port is added.
Russell King [Tue, 19 Oct 2004 16:50:24 +0000 (17:50 +0100)]
[SERIAL] Add new port registration/unregistration functions.
serial8250_register_port()/serial8250_unregister_port() has the
capability of registering ports with their struct device nodes,
which allows sysfs to indicate which tty devices belong to which
hardware devices.
We also add a serial8250 platform device driver in an initial
attempt at PM for ISA ports. However, I'm leaving out the
platform device for the time being since adding that would cause
potential oops issues.
Russell King [Tue, 19 Oct 2004 16:22:21 +0000 (17:22 +0100)]
[SERIAL] Make port autoprobing set up->capabilities.
Convert port autoprobing to set up->capabilities as it discovers
various capabilities of the port. Warn when the detected
capabilities do not match those in the uart_config table.
Russell King [Tue, 19 Oct 2004 16:00:01 +0000 (17:00 +0100)]
[SERIAL] Clean up handling of LSR in receive function.
It's pointless accessing the LSR value via a pointer all the time -
it prevents the compiler optimising it. Also, ensure that we
recognise a break sent during a kernel printk correctly.
* split off ->dma_exec_cmd() from ->ide_dma_[read,write] functions
* choose command to execute by ->dma_exec_cmd() in higher layers
and remove ->ide_dma_[read,write]
* in Etrax ide.c driver REQ_DRIVE_TASKFILE requests weren't
handled properly for drive->addressing == 0
* in trm290.c read and write commands were interchanged
* in sgiioc4.c commands weren't sent to disk devices
* tag REQ_DRIVE_TASKFILE write requests with REQ_RW
* split off ->dma_setup() from ->ide_dma_[read,write] functions
* use ->dma_setup() directly in ATAPI drivers and remove media
checks from ->ide_dma_[read,write]
* ->ide_dma_[read,write,begin] cannot fail now
* in Etrax ide.c setup DMA for ATAPI devices before sending
command to drive (so setup order is the same as for disks)
Russell King [Tue, 19 Oct 2004 20:36:16 +0000 (21:36 +0100)]
[ARM] Add generic RTC implementation.
This provides a number of helper functions and data structures
for RTC implementations to make use of, including a standard
implemention for /proc/driver/rtc and the rtc miscdevice. It
supports runtime registration of RTC timekeeping sources.
Russell King [Tue, 19 Oct 2004 19:02:21 +0000 (20:02 +0100)]
[ARM] Sanitise Footbridge machine class.
Footbridge was suffering from a little lack of care and attention;
it still had the nasty arch.c file with all the associated #ifdef
gross-ness that entailed.
Re-jig footbridge support so that each machine type contains all
the necessary support code, with a separate common implementation
which they all share.
Russell King [Tue, 19 Oct 2004 17:53:05 +0000 (18:53 +0100)]
[ARM] Rehash iwmmxt signal handling.
In the near future, VFP will want to save state onto the user stack.
Therefore, separate out the iwmmxt specific parts, and implement
a generic "safe copy to user space using random CPU instructions".
This is necessary because iwmmxt and VFP both use special CPU
instructions to load and/or save their state.
I've been informed that /proc/profile livelocks some systems in the timer
interrupt, usually at boot. The following patch attempts to amortize the
atomic operations done on the profile buffer to address this stability
concern. This patch has nothing to do with performance; kernels using
periodic timer interrupts are under realtime constraints to complete
whatever work they perform within timer interrupts before the next timer
interrupt arrives lest they livelock, performing no work whatsoever apart
from servicing timer interrupts. The latency of the cacheline bounce for
prof_buffer contributes to the time spent in the timer interrupt, hence it
must be amortized when remote access latencies or deviations from fair
exclusive cacheline acquisition may cause cacheline bounces to take longer
than the interval between timer ticks.
What this patch does is to create a pair of per-cpu open-addressed
hashtables indexed by profile buffer slot holding values representing the
number of pending profile buffer hits for the profile buffer slot. When
this hashtable overflows, one iterates over the hashtable accounting each
of the pairs of profile buffer slots and hit counts to the global profile
buffer. Zero is a legitimate profile buffer slot, so zero hit counts
represent unused hashtable entries. The hashtable is furthermore protected
from flush IPI's by interrupt disablement.
In order to flush the pending profile hits for read_profile(), this patch
flips betweeen the pairs of per-cpu profile buffer by signalling all cpus
to flip via IPI at the time of read_profile(), followed by doing all the
work to flush the profile hits from the older per-cpu buffers in the
context of the caller of read_profile(), with exclusion provided by a
semaphore ensuring that only one caller of profile_flip_buffers() may
execute at a time, and using interrupt disablement to prevent buffer flip
IPI's from altering the hashtables or flip state while an update is in
progress. The flip state is per-cpu so that remote cpus need only disable
interrupts locally for synchronization, which is both simple and
busywait-free for remote cpus. The flip states all change in tandem when
some cpu requests the hashtables be flipped, and the requester waits for
the completion of smp_call_function() for notification that all cpus have
finished flipping between their hashtables. The IPI handler merely toggles
the flip state (which is an array index) between 0 and 1.
This is expected to be a much stronger amortization than merely reducing
the frequency of profile buffer access by a factor of the size of the
hashtable because numerous hits may be held for each of its entries. This
reduces what was before the patch a number of atomic increments equal to
what after the patch becomes the sum of the hits held for each entry in the
hashtable, to a number of atomic_add()'s equal to the number of entries in
the per_cpu hashtable. This is nondeterministic, but as the profile hits
tend to be concentrated in a very small number of profile buffer slots
during any given timing interval, is likely to represent a very large
number of atomic increments. This amortization of atomic increments does
not depend on the hash function, only the sharp peakedness of the
distribution of profile buffer hits.
This algorithm has two advantages over full-size per-cpu profile buffers.
The first is that the space footprint is much smaller. Per-cpu profile
buffers would increase the space requirements by a factor of
num_online_cpus(), where this algorithm only requires one page per cpu.
The second is that reading the profile state is much faster, because the
state that must be traversed is exactly the above space consumers, and the
relative reduction in size concomitantly reduces the time required for a
read operation.
I also took the liberty of adding some commentary to the comments at the
beginning of the file reflecting the major work done on profile.c in recent
months and describing what the file implements.
The reporters of this issue have verified that this resolves their timer
interrupt livelock on 512x Altixen. In my own testing on 4x logical
x86-64, this patch saw a rate of about 18 flushes per minute under load, or
about one flush every 3 seconds, for about 38.4 atomic accesses to the
profile buffer per second per cpu in one of the algorithm's worst cases,
about 3.84% of the number of atomic profile buffer accesses per second per
cpu as a normal kernel would commit. This represents a twenty-six-fold
increase in the scalability on SMP systems with 4KB PAGE_SIZE, i.e. with a
4KB PAGE_SIZE, the number of atomic profile buffer accesses per second per
cpu is reduced by a factor of 26, thereby increasing the number of cpus a
system must have before it would experience a timer interrupt livelock by a
factor of 26, with the proviso that cacheline bounces must take the same
amount of time to service. This increase in the scalability of the kernel
is expected to be much larger for ia64, which has a large PAGE_SIZE,
because the distribution of profile buffer hits is so sharply peaked that
doubling the hashtable size will much more than double the amortization
factor. In fact, only 19 flushes were observed on a 64x Altix over an
approximately 10 minute AIM7 run, and 1 flush on a 512x Altix over the
course of an entire AIM7 run, for truly vast effective amortization
factors.
A prior version of this patch, which did not include the node-local
hashtable allocation and bounded collision chains has been successfully
tested on 64x and 512x ia64 vs 2.6.9-rc2, 8x ia64 vs. 2.6.9-rc2-mm1, 4x
x86-64 vs. 2.6.9-rc2-mm1, and 6x sparc64 vs. 2.6.9-rc2-mm1. This patch
minus the hashtable initialization fix has been successfully tested on 2x
ppc64, 2x alpha, 8x ia64, 6x sparc64, and 4x x86-64, all vs.
2.6.9-rc2-mm1. This precise version of the patch has been successfully
tested on 8x ia32 against 2.6.9-rc2-mm1 and 6x sparc64 vs. both
2.6.9-rc2-mm1 and 2.6.9-rc2-mm2.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
James Morris [Tue, 19 Oct 2004 01:17:18 +0000 (18:17 -0700)]
[PATCH] SELinux: allow all filesystems to specify fscreate mount option
The patch below allows all types of filesystems to specify the fscreate
mount option (which is used to specify the security context of the
filesystem itself). This was previously only available for filesystems
with full xattr security labeling, but is also potentially required for
filesystems with e.g. psuedo xattr labeling such as devpts and tmpfs.
An example of use is to specify at mount time the fs security context of a
tmpfs filesystem, overriding the default specified in policy for that
filesystem.
This patch has been in the Fedora kernel for some weeks with no problems.
Signed-off-by: James Morris <jmorris@redhat.com> Signed-off-by: Stephen Smalley <sds@epoch.ncsc.mil> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] xattr: re-introduce validity check before xattr cache insert
* ext[23]_xattr_list():
- Before inserting an xattr block into the cache, make sure that the
block is not corrupted. The check got moved after inserting into the
cache in the xattr consolidation patches, so corrupted blocks could become
visible to cache users.
- Take a variable out of the loop that calls the ->list handlers.
* A few cosmetic changes.
Signed-off-by: Andreas Gruenbacher <agruen@suse.de> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
James Morris [Tue, 19 Oct 2004 01:16:53 +0000 (18:16 -0700)]
[PATCH] xattr consolidation v3 - tmpfs
This patch adds xattr support to tmpfs, and a security xattr handler. The
purpose of this is to allow udev to be mounted on tmpfs, as used currently by
Fedora.
Original patch from: Luke Kenneth Casson Leighton <lkcl@lkcl.net>.
Signed-off-by: James Morris <jmorris@redhat.com> Signed-off-by: Stephen Smalley <sds@epoch.ncsc.mil> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
James Morris [Tue, 19 Oct 2004 01:16:03 +0000 (18:16 -0700)]
[PATCH] xattr consolidation v3 - LSM
This patch replaces the dentry parameter with an inode in the LSM
inode_{set|get|list}security hooks, in keeping with the ext2/ext3 code.
dentries are not needed here.
Signed-off-by: James Morris <jmorris@redhat.com> Signed-off-by: Stephen Smalley <sds@epoch.ncsc.mil> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
James Morris [Tue, 19 Oct 2004 01:15:51 +0000 (18:15 -0700)]
[PATCH] xattr consolidation v3 - generic xattr API
This patch consolidates common xattr handling logic into the core fs code,
with modifications suggested by Christoph Hellwig (hang off superblock, remove
locking, use generic code as methods), for use by ext2, ext3 and devpts, as
well as upcoming tmpfs xattr code.
Signed-off-by: James Morris <jmorris@redhat.com> Signed-off-by: Stephen Smalley <sds@epoch.ncsc.mil> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ulrich Drepper [Tue, 19 Oct 2004 01:15:27 +0000 (18:15 -0700)]
[PATCH] Simplify last lib/idr.c change
The last change to alloc_layer in lib/idr.c unnecessarily complicates
the code and depending on the definition of spin_unlock will cause worse
code to be generated than necessary. The following patch should improve
the situation.
Haroldo Gamal [Tue, 19 Oct 2004 01:15:14 +0000 (18:15 -0700)]
[PATCH] smbfs does not honor uid, gid, file_mode and dir_mode supplied by user mount
This patch fixes "Samba Bugzilla Bug 999". The last version (2.6.8.1) of
smbfs kernel module do not honor uid, gid, file_mode and dir_mode supplied
by user during mount. This bug is also logged as "Kernel Bug Tracker Bug
3330".
To fully work, some modifications are needed to samba smbmount.c and
smbmnt.c files. Those patches are available at Samba and Kernel Bug
Tracker pages.
After those patches, if the user do not supply any of the parameters above,
the uid, gid, file_mode and dir_mode on the server will be used by the
client.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andi Kleen [Tue, 19 Oct 2004 01:14:38 +0000 (18:14 -0700)]
[PATCH] x86-64/i386: add mce tainting
This patch adds machine check tainting. When a handled machine check
occurs the oops gets a new 'M' flag. This is useful to ignore machines
with hardware problems in oops reports.
On i386 a thermal failure also sets this flag.
Done for x86-64 and i386 so far.
Signed-off-by: Andi Kleen <ak@suse.de> Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>