Manfred Spraul [Mon, 23 Aug 2004 05:40:37 +0000 (22:40 -0700)]
[PATCH] ipc: Add refcount to ipc_rcu_alloc
The lifetime of the ipc objects (sem array, msg queue, shm mapping) is
controlled by kern_ipc_perms->lock - a spinlock. There is no simple way to
reacquire this spinlock after it was dropped to
schedule()/kmalloc/copy_{to,from}_user/whatever.
The attached patch adds a reference count as a preparation to get rid of
sem_revalidate().
Davide Libenzi [Mon, 23 Aug 2004 05:40:02 +0000 (22:40 -0700)]
[PATCH] Don't use SYSGOOD for ptrace singlestep
The ptrace single step mode should not use the SYSGOOD bit and should not
report SIGTRAP|0x80 to the ptrace parent. The following patch add an
explicit check and to not add 0x80 in this is a singlestep trap.
Fixes blk_queue_resize_tags to properly handle allocation failures.
Currently, if a memory allocation failure occurs during
blk_queue_resize_tags, the tag map ends up getting freed, which should
not happen. The old tag map should be preserved and only the resize
should fail.
Signed-off-by: Jens Axboe <axboe@suse.de> Signed-off-by: Brian King <brking@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Brian King [Mon, 23 Aug 2004 05:39:14 +0000 (22:39 -0700)]
[PATCH] blk_resize_tags() fix
init_tag_map should not initialize the busy_list, refcnt, or busy fields in
the tag map since blk_queue_resize_tags can call it while requests are
active. Patch moves this initialization into blk_queue_init_tags.
Signed-off-by: Jens Axboe <axboe@suse.de> Signed-off-by: Brian King <brking@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Brian King [Mon, 23 Aug 2004 05:39:02 +0000 (22:39 -0700)]
[PATCH] blk_queue_free_tags() fix
This is a resend of three ll_rw_blk patches related to tagged queuing.
Currently blk_queue_free_tags cannot be called with ops outstanding. The
scsi_tcq API defined to LLD scsi drivers allows for scsi_deactivate_tcq to
be called (which calls blk_queue_free_tags) with ops outstanding. Change
blk_queue_free_tags to no longer free the tags, but rather just disable
tagged queuing and also modify blk_queue_init_tags to handle re-enabling
tagged queuing after it has been disabled.
Signed-off-by: Jens Axboe <axboe@suse.de> Signed-off-by: Brian King <brking@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Chris Mason [Mon, 23 Aug 2004 05:38:24 +0000 (22:38 -0700)]
[PATCH] add BH_Eopnotsupp for testing async barrier failures
In order for filesystems to detect asynchronous ordered write failures for
buffers sent via submit_bh, they need a bit they can test for in the buffer
head. This adds BH_Eopnotsupp and the related buffer operations
end_buffer_write_sync is changed to avoid a printk for BH_Eoptnotsupp
related failures, since the FS is responsible for a retry.
sync_dirty_buffer is changed to test for BH_Eopnotsupp and return
-EOPNOTSUPP to the caller
Some of this came from Jens Axboe
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Davide Libenzi [Mon, 23 Aug 2004 05:36:54 +0000 (22:36 -0700)]
[PATCH] ptrace single-stepping fix
This patch permits a ptrace process on x86 to "see" the instruction following
the INT #80h op. This has been tested on 2.6.6 using the appended test
source. Running over this:
(Andi says: "I think this patch is a bad idea. The ptrace handling is
traditionally fragile (I remember when merging a rather simple patch from IBM
for DR allocation long ago into the suse it broke several debuggers). If you
really want to do that wait for 2.7.")
Keith Owens [Mon, 23 Aug 2004 05:36:42 +0000 (22:36 -0700)]
[PATCH] i386 oops output: dump preceding code
This teaches the i386 oops dumper to dump opcodes preceding and after the
offending EIP. Supporting code against ksymoops has been tested and produces
output like the below.
Support for this was added to ksymoops-2.4.9.
Note that ksymoops will guarantee that the disassembly after the <eip> value
is always in sync - if the disassembly from the start of the Code: line does
not sync up with the EIP address ksymoops will perform the resync.
Warning (merge_maps): no symbols in merged map
Mar 18 23:47:36 vmm kernel: kernel BUG at fs/open.c:802!
Mar 18 23:47:36 vmm kernel: invalid operand: 0000 [#1]
Mar 18 23:47:36 vmm kernel: CPU: 0
Mar 18 23:47:36 vmm kernel: EIP: 0060:[<c014fedf>] VLI Not tainted
Using defaults from ksymoops -t elf32-i386 -a i386
Mar 18 23:47:36 vmm kernel: EFLAGS: 00010246
Mar 18 23:47:36 vmm kernel: eax: ccdfb900 ebx: 4001020d ecx: 00000000 edx: 0000007b
Mar 18 23:47:36 vmm kernel: esi: 00000000 edi: bfffdd70 ebp: ccdfdfbc esp: ccdfdfb0
Mar 18 23:47:36 vmm kernel: ds: 007b es: 007b ss: 0068
Mar 18 23:47:36 vmm kernel: Stack: 4001020d00000000bfffdd70ccdfc000c01092134001020d0000000000000003
Mar 18 23:47:36 vmm kernel: 00000000bfffdd70bfffdc88000000050000007b0000007b000000054000ef94
Mar 18 23:47:36 vmm kernel: 0000007300000206bfffdbd80000007b
Mar 18 23:47:36 vmm kernel: Call Trace:
Mar 18 23:47:36 vmm kernel: [<c0109213>] syscall_call+0x7/0xb
Mar 18 23:47:36 vmm kernel: Code: 14 98 f0 81 41 04 00 00 00 01 5b 89 ec 5d c3 90 b8 00 e0 ff ff 21 e0 55 89 e5 57 56 53 8b 00 81 b8 e4 01 00 00 0f 27 00 00 75 08 <0f> 0b 22 03 85 18 2f c0 8b 45 08 50 e8 30 d4 00 00 89 c7 83 c4
This architecture has variable length instructions, decoding before eip
is unreliable, take these instructions with a pinch of salt.
Code; c014feb4 No symbols available 00000000 <_EIP>:
Code; c014feb4 No symbols available
0: 14 98 adc $0x98,%al
Code; c014feb6 No symbols available
2: f0 81 41 04 00 00 00 lock addl $0x1000000,0x4(%ecx)
Code; c014febd No symbols available
9: 01
Code; c014febe No symbols available
a: 5b pop %ebx
Code; c014febf No symbols available
b: 89 ec mov %ebp,%esp
Code; c014fec1 No symbols available
d: 5d pop %ebp
Code; c014fec2 No symbols available
e: c3 ret
Code; c014fec3 No symbols available
f: 90 nop
Code; c014fec4 No symbols available
10: b8 00 e0 ff ff mov $0xffffe000,%eax
Code; c014fec9 No symbols available
15: 21 e0 and %esp,%eax
Code; c014fecb No symbols available
17: 55 push %ebp
Code; c014fecc No symbols available
18: 89 e5 mov %esp,%ebp
Code; c014fece No symbols available
1a: 57 push %edi
Code; c014fecf No symbols available
1b: 56 push %esi
Code; c014fed0 No symbols available
1c: 53 push %ebx
Code; c014fed1 No symbols available
1d: 8b 00 mov (%eax),%eax
Code; c014fed3 No symbols available
1f: 81 b8 e4 01 00 00 0f cmpl $0x270f,0x1e4(%eax)
Code; c014feda No symbols available
26: 27 00 00
Code; c014fedd No symbols available
29: 75 08 jne 33 <_EIP+0x33> c014fee7 No symbols available
This decode from eip onwards should be reliable
Code; c014fedf No symbols available 00000000 <_EIP>:
Code; c014fedf No symbols available <=====
0: 0f 0b ud2a <=====
Code; c014fee1 No symbols available
2: 22 03 and (%ebx),%al
Code; c014fee3 No symbols available
4: 85 18 test %ebx,(%eax)
Code; c014fee5 No symbols available
6: 2f das
Code; c014fee6 No symbols available
7: c0 8b 45 08 50 e8 30 rorb $0x30,0xe8500845(%ebx)
Code; c014feed No symbols available
e: d4 00 aam $0x0
Code; c014feef No symbols available
10: 00 .byte 0x0
Code; c014fef0 No symbols available
11: 89 c7 mov %eax,%edi
Code; c014fef2 No symbols available
13: 83 .byte 0x83
Code; c014fef3 No symbols available
14: c4 .byte 0xc4
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andrey Panin [Mon, 23 Aug 2004 05:36:30 +0000 (22:36 -0700)]
[PATCH] fix visws kernel build
CC arch/i386/kernel/cpu/intel.o
In file included from arch/i386/kernel/cpu/intel.c:19:
include/asm-i386/mach-visws/mach_apic.h: In function `cpu_present_to_apicid':
include/asm-i386/mach-visws/mach_apic.h:67: error: `BAD_APICID' undeclared (first use in this function)
include/asm-i386/mach-visws/mach_apic.h:67: error: (Each undeclared identifier is reported only once
include/asm-i386/mach-visws/mach_apic.h:67: error: for each function it appears in.)
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Santiago Leon [Mon, 23 Aug 2004 05:36:18 +0000 (22:36 -0700)]
[PATCH] ibmveth: add memory barrier for hypervisor synchronisation
This patch adds a memory barrier to ensure synchronization with the
hypervisor (and avoid a panic when the hypervisor is halfway through
writing to the descriptor). It also removes an unnecessary check that is
flawed anyway because the value can change between the atomic_inc() and the
assert.
Signed-off-by: Santiago Leon <santil@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Santiago Leon [Mon, 23 Aug 2004 05:36:06 +0000 (22:36 -0700)]
[PATCH] ibmveth: hypervisor return value fix
This patch checks for the LongBusy return code from the hypervisor, and
retries the operation (which is what the hypervisor expects the driver
to do). Please apply.
Signed-off-by: Santiago Leon <santil@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Santiago Leon [Mon, 23 Aug 2004 05:35:43 +0000 (22:35 -0700)]
[PATCH] ibmveth: module tag fixes
This and the following three patches contain bug fixes found in the
stabilization of SLES9.
This patch adds a call to MODULE_VERSION and changes the MODULE_AUTHOR call
to me (obviously with Dave Larson's permission). It also increments the
version number to keep track of the bug fixes. Please apply.
Signed-off-by: Santiago Leon <santil@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ryan S. Arnold [Mon, 23 Aug 2004 05:35:31 +0000 (22:35 -0700)]
[PATCH] HVCS fixes
Here are a set of HVCS (drivers/char/hvcs.c) fixes that were suggested by Jeff
Garzik on July 29th in his review of this driver as well as some other fixes
for problems I found while reviewing the driver. These are all relatively
minor, but necessary.
- Cleaned up curly braces on single line conditional blocks.
- Replaced debug memset(...,0x3F,...) with memset(...,0x00,...).
- Removed explicit '= 0' after static int declarations since these default
to zero.
- Removed list_for_each_safe() instances and replaced with
list_for_each_entry() which cut down on amt of code. The 'safe' version is
un-needed now that the driver is using spinlocks.
- Changed spin_lock_irqsave() to spin_lock() when locking hvcs_structs_lock
and hvcs_pi_lock since these are not touched in an int handler.
- changed spin_lock_irqsave() to spin_lock() in interrupt handler.
- Initialized hvcs_structs_lock and hvcs_pi_lock to SPIN_LOCK_UNLOCKED at
declaration tiem rather than in hvcs_module_init().
- Added spin_lock around list_del() in destroy_hvcs_struct() to protect the
list traversal from deletion. The original omission was an oversight.
- Removed '= NULL' from pointer declarations since they are initialized NULL
by default.
- Removed wmb() instance from hvcs_try_write(). They probably aren't needed
with locking in place.
- Added check and cleanup for hvcs_pi_buff = kmalloc() in
hvcs_module_init().
- Exposed hvcs_struct.index via a sysfs attribute so that the coupling
between /dev/hvcs* and a vty-server can be systematically determined.
- Moved kobject_put() in hvcs_open() outside of the
spin_unlock_irqrestore().
- In hvcs_probe() changed kmalloc(sizeof(*hvcsd),...) to
kmalloc(sizeof(struct hvcs_struct)) because hvcsd references a NULL pointer
at the time of kmalloc.
- Incremented the HVCS_DRIVER_VERSION to 1.3.1
arch/ppc64/kernel/hvcserver.c:
- Changed function documentation of EXPORTed functions to comply with proper
kernel-doc documentation style.
- Changed 'unsigned int' types to 'uint32_t' to comply with how unit
addresses and partition IDs are handled in other arch/ppc64 vterm code.
- Cleaned up curly braces on single line conditional blocks.
include/asm-ppc64/hvcserver.h:
- Added kernel-doc style documentation for hvcs_partner_info struct.
- changed 'unsigned int' types to 'uint32_t' to comply with how unit
addresses and partition IDs are handled in other arch/ppc64 vterm code.
Signed-off-by: Ryan S. Arnold <rsa@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Dave Boutcher [Mon, 23 Aug 2004 05:35:19 +0000 (22:35 -0700)]
[PATCH] ppc64: mf_proc file position fix
arch/ppc64/kernel/mf_proc.c uses a bad interface for moving along file
position in a proc_write routine. This quit working altogether in 2.6.8.
Patch to fix. And I did a quick scan of the kernel to see if anyone else
was similarly broken...apparantly not :-)
Fixes a broken update of f_pos in a proc file write routine.
Signed-off-by: Dave Boutcher <sleddog@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:35:08 +0000 (22:35 -0700)]
[PATCH] ppc64: Use correct buffer size in RTAS call
Firmware expects the size of the buffer that you hand it when you ask it
for information about a hardware error to be of a very specific size, but
different versions of firmware appearently expect different sizes; using
the wrong size results in a painful, hard-to-debug crash in firmware. Benh
provided a patch for this some months ago, but appreantly missed this code
path. This patch sets up the log buffer size dynamically; it also fixes a
bug with the return code not being handled correctly.
Signed-off-by: Linas Vepstas <linas@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The patch below implements the ability to query outstanding imalloc regions
for a given virtual address range. (Imalloc is the allocator of virtual
space for ioremap.) The patch extends im_get_area() to allow a region
criterion of IM_REGION_SUPERSET. For a particular "superset" virtual
address and size passed into im_get_area(), the function returns the first
outstanding region that is contained within this superset region.
The patch also changes iounmap_explicit() to allow for the unmapping of all
regions that fit under a "superset".
This ability is necessary for dynamic (runtime) removal of pci host bridges
(PHBs). For a PHB removal, the platform specification (the RPA) requires
that all of its children slots already be dynamically removed. Each of
these slot-level removals has fractured the imalloc region assigned to the
PHB at boot. At PHB removal time, it is necessary to iounmap() the
remaining artifacts of the initial PHB region.
Signed-off-by: John Rose <johnrose@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Stephen Rothwell [Mon, 23 Aug 2004 05:34:40 +0000 (22:34 -0700)]
[PATCH] ppc64 iSeries virtual DVD-RAM
This patch adds the ability to use DVD-RAM drives to the iSeries virtual
cdrom driver. This version adresses (hopefully) Jens comments on the
previous one.
Signed-off-by: Stephen Rothwell <sfr@canb.auug.org.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:34:28 +0000 (22:34 -0700)]
[PATCH] ppc64: better little-endian bitops
Below patch reuses the big-endian bitops for the little endian ones, and
moves the ext2_{set,clear}_bit_atomic functions to be truly atomic instead
of lock based.
This requires that the bitmaps passed to the ext2_* bitop functions are
8-byte aligned. I have been assured that they will be 512-byte or
1024-byte aligned, and sparc and ppc32 also impose an alignment requirement
on the bitmap.
Signed-off-by: Olof Johansson <olof@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:34:16 +0000 (22:34 -0700)]
[PATCH] ppc64: rtas_call was calling kmalloc too early
At present rtas_call() can be called before the kmalloc subsystem is
initialized, and if RTAS reports a hardware error, the code tries to do a
kmalloc to make a copy of the error report. This patch changes it so that
we don't do the kmalloc in that situation.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nathan Lynch [Mon, 23 Aug 2004 05:33:51 +0000 (22:33 -0700)]
[PATCH] ppc64: tweak schedule_timeout in __cpu_die
The current code does schedule_timeout(HZ) when waiting for a cpu to die,
which is a bit coarse and tends to limit the "throughput" of my stress
tests :)
Change the HZ timeout to HZ/5, increase the number of tries to 25 so the
overall wait time is similar. In practice, I've never seen the loop need
more than two iterations.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Mon, 23 Aug 2004 05:33:27 +0000 (22:33 -0700)]
[PATCH] ppc64: bolted SLB entry for iSeries
Tested, at least basically, on Power4 iSeries with shared processors, on
Power4 pSeries and RS64 (non-SLB) iSeries machines.
On pSeries SLB machines we "bolt" an SLB entry for the first segment of the
vmalloc() area into the SLB, to reduce the SLB miss rate. This caused
problems, so was disabled, on iSeries because the bolted entry was not
restored properly on shared processor switch. This patch adds information
about the bolted vmalloc segment to the lpar map, which should be restored
on shared processor switch.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
During some signal test, we found that v_regs pointer was not setup
correctly. v_regs was made to point to itself, as a result of which the
pointer was corrupted when vec registers were copied over. When the signal
handler returned, restore_sigcontext tried derefering the invalid pointer
and in the process killed the app with SIGSEGV.
Paul Mackerras [Mon, 23 Aug 2004 05:32:53 +0000 (22:32 -0700)]
[PATCH] ppc64: Reduce verbosity of RTAS error logs
Currently on pSeries systems the kernel will print out a hex dump of any
error events reported by the platform at boot time. These can be rather
large and are practically incomprehensible to humans. With this patch, the
kernel will by default print a 1-line summary for each error reported with
the severity, type, etc. printed as text strings. The old behaviour is
still available by using the rtasmsgs=on kernel command line option. The
patch also renames some RTAS-specific symbols to start with "RTAS_".
Signed-off-by: Nathan Fontenot <nfont@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:32:41 +0000 (22:32 -0700)]
[PATCH] ppc64 Fix unbalanced pci_dev_put in EEH code
The EEH code currently can end up doing an extra pci_dev_put() in the case
where we hot-unplug a card for which we are ignoring EEH errors (e.g. a
graphics card). This patch fixes that problem by only maintaining a
reference to the PCI device if we have entered any of its resource
addresses into our address -> PCI device cache. This patch is based on an
earlier patch by Linas Vepstas.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:32:30 +0000 (22:32 -0700)]
[PATCH] ppc64: log firmware errors during boot
Firmware can report errors at any time, and not atypically during boot.
However, these reports were being discarded until th rtasd comes up, which
occurs fairly late in the boot cycle. As a result, firmware errors during
boot were being silently ignored.
Signed-off-by: Linas Vepstas <linas@linas.org> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Mon, 23 Aug 2004 05:32:18 +0000 (22:32 -0700)]
[PATCH] ppc64: C99 initializers in INIT_THREAD
Fairly trivial PPC64 cleanup. This patch makes the ppc64 INIT_THREAD
#define use C99 initializers, which will make it less likely to get broken
if we need to change thread_struct.
Signed-off-by: David Gibson <dwg@au1.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:32:07 +0000 (22:32 -0700)]
[PATCH] ppc64: fix idle loop for offline cpu
In the default_idle and dedicated_idle loops, there are some inner loops
out of which we should break if the cpu is marked offline. Otherwise, it
is possible for the cpu to get stuck and never actually go offline.
shared_idle is unaffected.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:31:55 +0000 (22:31 -0700)]
[PATCH] ppc64: Don't call scheduler on offline cpu
When taking a cpu offline, once the cpu has been removed from
cpu_online_map, it is not supposed to service any more interrupts. This
presents a problem on ppc64 because we cannot truly disable the
decrementer. There used to be cpu_is_offline() checks in several scheduler
functions (e.g. rebalance_tick()) which papered over this issue, but these
checks were removed recently. So with recent 2.6 kernels, an attempt to
offline a cpu can result in a crash in find_busiest_group(). This patch
prevents such crashes.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:31:43 +0000 (22:31 -0700)]
[PATCH] ppc64: set tbl->it_type in iommu code
Here is a patch that sets struct iommu_table->it_type to TCE_PCI in
pSeries_iommu.c. This is just for code completeness (and it is updated in
iSeries_iommu.c, but was somehow missed in pSeries_iommu.c).
Signed-off-by: Ananth N Mavinakayanahalli <ananth@in.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Mon, 23 Aug 2004 05:31:08 +0000 (22:31 -0700)]
[PATCH] ppc64: allow oprofile module to be safely unloaded
Allow the oprofile module to be unloaded, before we never removed the
oprofile specific interrupt handler. Handle the pending exception case in
the dummy interrupt handler instead.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:30:20 +0000 (22:30 -0700)]
[PATCH] ppc64: rework secondary SMT thread setup at boot
Our (ab)use of cpu_possible_map in setup_system to start secondary SMT
threads bothers me. Mark such threads in cpu_possible_map during early
boot; let RTAS tell us which present cpus are still offline later so we can
start them.
I'm not totally sure about this one, it might be better to set up
cpu_sibling_map in prom_hold_cpus and use that in setup_system.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:30:08 +0000 (22:30 -0700)]
[PATCH] ppc64: use cpu_present_map in ppc64
Adopt the "standard" cpu_present_map for describing cpus which are present
in the system, but not necessarily online. cpu_present_map is meant to be
a superset of cpu_online_map and a subset of cpu_possible_map.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:29:56 +0000 (22:29 -0700)]
[PATCH] ppc64: use platform numbering of cpus for hypervisor calls.
We were using Linux's cpu numbering for cpu-related hypervisor calls (e.g.
vpa registration, H_CONFER). It happened to work most of the time because
Linux and the hypervisor usually, but not always, have the same numbering
for cpus.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:29:32 +0000 (22:29 -0700)]
[PATCH] ppc64: set time-related systemcfg fields
Somewhere along the line we lost the code that updates some fields of the
systemcfg structure that are used for translating timebase values to time
of day. I want to get rid of the systemcfg structure eventually, but
applications are using it (and in particular these fields) and I don't want
to break the ABI in a stable kernel series.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:28:40 +0000 (22:28 -0700)]
[PATCH] ppc32: Fix bug in altivec emulation
This patch fixes a bug in the kernel emulation of altivec instructions with
denormalized operands. The emulation of the vmaddfp and vmnsubfp
instructions was giving the wrong answer because I had the wrong order of
operands to the fmadds and fnmsubs instructions. This patch fixes it for
both ppc32 and ppc64.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:28:04 +0000 (22:28 -0700)]
[PATCH] ppc32: emulate obsolete instructions
This patch adds emulation in the illegal instruction handler for a couple
of old instructions that are no longer implemented in the PPC970 and later
chips. This patch adds the code for both ppc32 and ppc64, and cleans up
the ppc64 traps.c a bit, along the lines of the ppc32 code. It also makes
sure that the ppc64 code generates a SIGTRAP after emulating an instruction
if single-stepping is enabled.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
This patch adds code to the ppc32 alignment exception handler to make it
handle the load/store string and load/store multiple word instructions.
This is an issue for older CPUs such as the PPC601, which traps on
load/store string instructions which cross a page boundary (newer CPUs
handle this in hardware). I have a little test program which exercises
this code, so I am reasonably confident it's correct.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Matt Porter [Mon, 23 Aug 2004 05:27:40 +0000 (22:27 -0700)]
[PATCH] ppc32: make PPC40x large tlb mapping optional
This makes the PPC40x lowmem large tlb mapping selectable via a cmdline
option. This allows use of the normal page-sized mapping so that kernel
text can be read only if desired.
Signed-off-by: Josh Boyer <jwboyer@charter.net> Signed-off-by: Matt Porter <mporter@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Matt Porter [Mon, 23 Aug 2004 05:27:29 +0000 (22:27 -0700)]
[PATCH] ppc32: optimize/fix timer_interrupt loop
The following patch fixes the situation where the loop condition could
generate a next_dec of zero while exiting the loop. This is suboptimal on
Classic PPC because it forces another interrupt to occur and reenter the
handler. It is fatal on Book E cores, because their decrementer is stopped
when writing a zero (Classic interrupts on a 0->-1 transition, Book E
interrupts on a 1->0 transition). Instead, stay in the loop on a
next_dec==0.
Signed-off-by: Matt Porter <mporter@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] ppc32: remove hardcoded offsets from ppc asm
This patch by Vincent Hanquez removes some hard coded offsets for accessing
thread info fields from assembly, uses the normal offset generation
mecanism that we already have for other things instead.
Signed-off-by: Vincent Hanquez <tab@snarc.org> Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Keith Owens [Mon, 23 Aug 2004 05:27:05 +0000 (22:27 -0700)]
[PATCH] Make i386 die() more resilient against recursive errors
Make i386 die() more resilient against recursive errors, almost a cut
and paste of the ia64 die() routine. Much of the patch is indentation
changes.
Mainly to make it easier to add crash, lcrash, kmsgdump or other RAS patches.
They are invoked from die() and if they crash themselves, we have to avoid
recursive loops in die().
Signed-off-by: Keith Owens <kaos@sgi.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Akiyama Nobuyuki [Mon, 23 Aug 2004 05:26:54 +0000 (22:26 -0700)]
[PATCH] NMI trigger switch support for debugging(updated)
I made a patch for debugging with the help of NMI trigger switch.
When kernel hangs severely, keyboard operation(e.g.Ctrl-Alt-Del)
doesn't work properly. This patch enables debugging information
to be displayed on console in this case.
I think this feature is necessary as standard functionality.
Please feel free to use this patch and let me know if you have
any comments.
Background:
When a trouble occurs in kernel, we usually begin to investigate
with following information:
- panic >> panic message.
- oops >> CPU registers and stack trace.
- hang >> **NONE** no standard method established.
How it works:
Most IA32 servers have a NMI switch that fires NMI interrupt up.
The NMI interrupt can interrupt even if kernel is serious state,
for example deadlock under the interrupt disabled.
When the NMI switch is pressed after this feature is activated,
CPU registers and stack trace are displayed on console and then
panic occurs.
This feature is activated or deactivated with sysctl.
On IA32 architecture, only the following are defined as reason
of NMI interrupt:
- memory parity error
- I/O check error
The reason code of NMI switch is not defined, so this patch assumes
that all undefined NMI interrupts are fired by MNI switch.
However, oprofile and NMI watchdog also use undefined NMI interrupt.
Therefore this feature cannot be used at the same time with oprofile
and NMI watchdog. This feature hands NMI interrupt over to oprofile
and NMI watchdog. So, when they have been activated, this feature
doesn't work even if it is activated.
Arnd Bergmann [Mon, 23 Aug 2004 05:26:41 +0000 (22:26 -0700)]
[PATCH] fix reading string module parameters in sysfs
Reading the contents of a module_param_string through sysfs currently
oopses because the param_get_charp() function cannot operate on a
kparam_string struct. This introduces the required param_get_string.
Mike Kravetz [Mon, 23 Aug 2004 05:26:30 +0000 (22:26 -0700)]
[PATCH] proc fs task name locking fix
Races have been observed between excec-time overwriting of task->comm and
/proc accesses to the same data. This causes environment string
information to appear in /proc.
Fix that up by taking task_lock() around updates to and accesses to
task->comm.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
it took more than 80 usecs for XFree86 to do a context-switch!
it turns out that the reason for this (massive) context-switching
overhead is the following change in 2.6.8:
[PATCH] larger IO bitmaps
To demonstrate the effect of this change i've written ioperm-latency.c
(attached), which gives the following on vanilla 2.6.8.1:
# ./ioperm-latency
default no ioperm: scheduling latency: 2528 cycles
turning on port 80 ioperm: scheduling latency: 10563 cycles
turning on port 65535 ioperm: scheduling latency: 10517 cycles
the ChangeSet says:
Now, with the lazy bitmap allocation and per-CPU TSS, this
will really not drain any resources I think.
this is plain wrong. An increase in the IO bitmap size introduces
per-context-switch overhead as well: we now have to copy an 8K bitmap
every time XFree86 context-switches - even though XFree86 never uses
ports higher than 1024! I've straced XFree86 on a number of x86 systems
and in every instance ioperm() was used - so i'd say the majority of x86
Linux systems running 2.6.8.1 are affected by this problem.
This not only causes lots of overhead, it also trashes ~16K out of the
L1 and L2 caches, on every context-switch. It's as if XFree86 did a L1
cache flush on every context-switch ...
the simple solution would be to revert IO_BITMAP_BITS back to 1024 and
release 2.6.8.2?
I've implemented another solution as well, which tracks the
highest-enabled port # for every task and does the copying of the bitmap
intelligently. (patch attached) The patched kernel gives:
# ./ioperm-latency
default no ioperm: scheduling latency: 2423 cycles
turning on port 80 ioperm: scheduling latency: 2503 cycles
turning on port 65535 ioperm: scheduling latency: 10607 cycles
this is much more acceptable - the full overhead only occurs in the very
unlikely event of a task using the high ioport range. X doesnt suffer
any significant overhead.
(tracking the maximum allowed port # also allows a simplification of
io_bitmap handling: e.g. we dont do the invalid-offset trick anymore -
the IO bitmap in the TSS is always valid and secure.)
I tested the patch on x86 SMP and UP, it works fine for me. I tested
boundary conditions as well, it all seems secure.
/*
* Use a pair of RT processes bound to the same CPU to measure
* context-switch overhead:
*/
static void measure(void)
{
unsigned long i, min = ~0UL, pid, mask = 1, t1, t2;
sched_set_affinity(0, sizeof(mask), &mask);
pid = fork();
if (!pid)
for (;;) {
asm volatile ("sti; nop; cli");
sched_yield();
}
sched_yield();
for (i = 0; i < 100; i++) {
asm volatile ("sti; nop; cli");
CYCLES(t1);
sched_yield();
CYCLES(t2);
if (i > 10) {
if (t2 - t1 < min)
min = t2 - t1;
}
}
asm volatile ("sti");
It seems that on some OldWolrd macs, we don't get the OF stdout device,
thus the new set_preferred_console() dies at boot trying to dereference
a NULL pointer.
In 2.5.18 some minix-specific stuff was moved to the minix subdirectory
where it belonged. However, a typo crept in, causing inode disk usage
to be incorrectly reported. A few people have complained, but so far
not sufficiently loudly.
Alan Cox [Sun, 22 Aug 2004 07:03:35 +0000 (00:03 -0700)]
[PATCH] missing CPU descriptors
There are a couple of cache descriptors in the current Intel manuals
missing from our tables at least one of which appears in an actual
processor in the real world.
Jesse Barnes [Sun, 22 Aug 2004 05:30:28 +0000 (22:30 -0700)]
[PATCH] ACPI for 2.6
Define acpi_noirq on ia64 since it's used now in pci_link.c. All ia64
machines use ACPI, so we can just define it to 0 like we do for acpi_disabled
and acpi_pci_disabled.
Arnd Bergmann [Fri, 20 Aug 2004 10:29:39 +0000 (12:29 +0200)]
[WATCHDOG] v2.6.8.1 compat_ioctl-patch
The watchdog ioctl interface is defined correctly for 32 bit emulation,
although WIOC_GETSUPPORT was not marked as such, for an unclear reason.
WDIOC_SETTIMEOUT and WDIOC_GETTIMEOUT were added in may 2002 to the
code but never to the ioctl list. This adds all three definitions.
Signed-off-by: Arnd Bergmann <arnd@arndb.de> Signed-off-by: Wim Van Sebroeck <wim@iguana.be>