The patch below implements the ability to query outstanding imalloc regions
for a given virtual address range. (Imalloc is the allocator of virtual
space for ioremap.) The patch extends im_get_area() to allow a region
criterion of IM_REGION_SUPERSET. For a particular "superset" virtual
address and size passed into im_get_area(), the function returns the first
outstanding region that is contained within this superset region.
The patch also changes iounmap_explicit() to allow for the unmapping of all
regions that fit under a "superset".
This ability is necessary for dynamic (runtime) removal of pci host bridges
(PHBs). For a PHB removal, the platform specification (the RPA) requires
that all of its children slots already be dynamically removed. Each of
these slot-level removals has fractured the imalloc region assigned to the
PHB at boot. At PHB removal time, it is necessary to iounmap() the
remaining artifacts of the initial PHB region.
Signed-off-by: John Rose <johnrose@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Stephen Rothwell [Mon, 23 Aug 2004 05:34:40 +0000 (22:34 -0700)]
[PATCH] ppc64 iSeries virtual DVD-RAM
This patch adds the ability to use DVD-RAM drives to the iSeries virtual
cdrom driver. This version adresses (hopefully) Jens comments on the
previous one.
Signed-off-by: Stephen Rothwell <sfr@canb.auug.org.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:34:28 +0000 (22:34 -0700)]
[PATCH] ppc64: better little-endian bitops
Below patch reuses the big-endian bitops for the little endian ones, and
moves the ext2_{set,clear}_bit_atomic functions to be truly atomic instead
of lock based.
This requires that the bitmaps passed to the ext2_* bitop functions are
8-byte aligned. I have been assured that they will be 512-byte or
1024-byte aligned, and sparc and ppc32 also impose an alignment requirement
on the bitmap.
Signed-off-by: Olof Johansson <olof@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:34:16 +0000 (22:34 -0700)]
[PATCH] ppc64: rtas_call was calling kmalloc too early
At present rtas_call() can be called before the kmalloc subsystem is
initialized, and if RTAS reports a hardware error, the code tries to do a
kmalloc to make a copy of the error report. This patch changes it so that
we don't do the kmalloc in that situation.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nathan Lynch [Mon, 23 Aug 2004 05:33:51 +0000 (22:33 -0700)]
[PATCH] ppc64: tweak schedule_timeout in __cpu_die
The current code does schedule_timeout(HZ) when waiting for a cpu to die,
which is a bit coarse and tends to limit the "throughput" of my stress
tests :)
Change the HZ timeout to HZ/5, increase the number of tries to 25 so the
overall wait time is similar. In practice, I've never seen the loop need
more than two iterations.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Mon, 23 Aug 2004 05:33:27 +0000 (22:33 -0700)]
[PATCH] ppc64: bolted SLB entry for iSeries
Tested, at least basically, on Power4 iSeries with shared processors, on
Power4 pSeries and RS64 (non-SLB) iSeries machines.
On pSeries SLB machines we "bolt" an SLB entry for the first segment of the
vmalloc() area into the SLB, to reduce the SLB miss rate. This caused
problems, so was disabled, on iSeries because the bolted entry was not
restored properly on shared processor switch. This patch adds information
about the bolted vmalloc segment to the lpar map, which should be restored
on shared processor switch.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
During some signal test, we found that v_regs pointer was not setup
correctly. v_regs was made to point to itself, as a result of which the
pointer was corrupted when vec registers were copied over. When the signal
handler returned, restore_sigcontext tried derefering the invalid pointer
and in the process killed the app with SIGSEGV.
Paul Mackerras [Mon, 23 Aug 2004 05:32:53 +0000 (22:32 -0700)]
[PATCH] ppc64: Reduce verbosity of RTAS error logs
Currently on pSeries systems the kernel will print out a hex dump of any
error events reported by the platform at boot time. These can be rather
large and are practically incomprehensible to humans. With this patch, the
kernel will by default print a 1-line summary for each error reported with
the severity, type, etc. printed as text strings. The old behaviour is
still available by using the rtasmsgs=on kernel command line option. The
patch also renames some RTAS-specific symbols to start with "RTAS_".
Signed-off-by: Nathan Fontenot <nfont@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:32:41 +0000 (22:32 -0700)]
[PATCH] ppc64 Fix unbalanced pci_dev_put in EEH code
The EEH code currently can end up doing an extra pci_dev_put() in the case
where we hot-unplug a card for which we are ignoring EEH errors (e.g. a
graphics card). This patch fixes that problem by only maintaining a
reference to the PCI device if we have entered any of its resource
addresses into our address -> PCI device cache. This patch is based on an
earlier patch by Linas Vepstas.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:32:30 +0000 (22:32 -0700)]
[PATCH] ppc64: log firmware errors during boot
Firmware can report errors at any time, and not atypically during boot.
However, these reports were being discarded until th rtasd comes up, which
occurs fairly late in the boot cycle. As a result, firmware errors during
boot were being silently ignored.
Signed-off-by: Linas Vepstas <linas@linas.org> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Mon, 23 Aug 2004 05:32:18 +0000 (22:32 -0700)]
[PATCH] ppc64: C99 initializers in INIT_THREAD
Fairly trivial PPC64 cleanup. This patch makes the ppc64 INIT_THREAD
#define use C99 initializers, which will make it less likely to get broken
if we need to change thread_struct.
Signed-off-by: David Gibson <dwg@au1.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:32:07 +0000 (22:32 -0700)]
[PATCH] ppc64: fix idle loop for offline cpu
In the default_idle and dedicated_idle loops, there are some inner loops
out of which we should break if the cpu is marked offline. Otherwise, it
is possible for the cpu to get stuck and never actually go offline.
shared_idle is unaffected.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:31:55 +0000 (22:31 -0700)]
[PATCH] ppc64: Don't call scheduler on offline cpu
When taking a cpu offline, once the cpu has been removed from
cpu_online_map, it is not supposed to service any more interrupts. This
presents a problem on ppc64 because we cannot truly disable the
decrementer. There used to be cpu_is_offline() checks in several scheduler
functions (e.g. rebalance_tick()) which papered over this issue, but these
checks were removed recently. So with recent 2.6 kernels, an attempt to
offline a cpu can result in a crash in find_busiest_group(). This patch
prevents such crashes.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:31:43 +0000 (22:31 -0700)]
[PATCH] ppc64: set tbl->it_type in iommu code
Here is a patch that sets struct iommu_table->it_type to TCE_PCI in
pSeries_iommu.c. This is just for code completeness (and it is updated in
iSeries_iommu.c, but was somehow missed in pSeries_iommu.c).
Signed-off-by: Ananth N Mavinakayanahalli <ananth@in.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Mon, 23 Aug 2004 05:31:08 +0000 (22:31 -0700)]
[PATCH] ppc64: allow oprofile module to be safely unloaded
Allow the oprofile module to be unloaded, before we never removed the
oprofile specific interrupt handler. Handle the pending exception case in
the dummy interrupt handler instead.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:30:20 +0000 (22:30 -0700)]
[PATCH] ppc64: rework secondary SMT thread setup at boot
Our (ab)use of cpu_possible_map in setup_system to start secondary SMT
threads bothers me. Mark such threads in cpu_possible_map during early
boot; let RTAS tell us which present cpus are still offline later so we can
start them.
I'm not totally sure about this one, it might be better to set up
cpu_sibling_map in prom_hold_cpus and use that in setup_system.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:30:08 +0000 (22:30 -0700)]
[PATCH] ppc64: use cpu_present_map in ppc64
Adopt the "standard" cpu_present_map for describing cpus which are present
in the system, but not necessarily online. cpu_present_map is meant to be
a superset of cpu_online_map and a subset of cpu_possible_map.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:29:56 +0000 (22:29 -0700)]
[PATCH] ppc64: use platform numbering of cpus for hypervisor calls.
We were using Linux's cpu numbering for cpu-related hypervisor calls (e.g.
vpa registration, H_CONFER). It happened to work most of the time because
Linux and the hypervisor usually, but not always, have the same numbering
for cpus.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:29:32 +0000 (22:29 -0700)]
[PATCH] ppc64: set time-related systemcfg fields
Somewhere along the line we lost the code that updates some fields of the
systemcfg structure that are used for translating timebase values to time
of day. I want to get rid of the systemcfg structure eventually, but
applications are using it (and in particular these fields) and I don't want
to break the ABI in a stable kernel series.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:28:40 +0000 (22:28 -0700)]
[PATCH] ppc32: Fix bug in altivec emulation
This patch fixes a bug in the kernel emulation of altivec instructions with
denormalized operands. The emulation of the vmaddfp and vmnsubfp
instructions was giving the wrong answer because I had the wrong order of
operands to the fmadds and fnmsubs instructions. This patch fixes it for
both ppc32 and ppc64.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Mon, 23 Aug 2004 05:28:04 +0000 (22:28 -0700)]
[PATCH] ppc32: emulate obsolete instructions
This patch adds emulation in the illegal instruction handler for a couple
of old instructions that are no longer implemented in the PPC970 and later
chips. This patch adds the code for both ppc32 and ppc64, and cleans up
the ppc64 traps.c a bit, along the lines of the ppc32 code. It also makes
sure that the ppc64 code generates a SIGTRAP after emulating an instruction
if single-stepping is enabled.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
This patch adds code to the ppc32 alignment exception handler to make it
handle the load/store string and load/store multiple word instructions.
This is an issue for older CPUs such as the PPC601, which traps on
load/store string instructions which cross a page boundary (newer CPUs
handle this in hardware). I have a little test program which exercises
this code, so I am reasonably confident it's correct.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Matt Porter [Mon, 23 Aug 2004 05:27:40 +0000 (22:27 -0700)]
[PATCH] ppc32: make PPC40x large tlb mapping optional
This makes the PPC40x lowmem large tlb mapping selectable via a cmdline
option. This allows use of the normal page-sized mapping so that kernel
text can be read only if desired.
Signed-off-by: Josh Boyer <jwboyer@charter.net> Signed-off-by: Matt Porter <mporter@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Matt Porter [Mon, 23 Aug 2004 05:27:29 +0000 (22:27 -0700)]
[PATCH] ppc32: optimize/fix timer_interrupt loop
The following patch fixes the situation where the loop condition could
generate a next_dec of zero while exiting the loop. This is suboptimal on
Classic PPC because it forces another interrupt to occur and reenter the
handler. It is fatal on Book E cores, because their decrementer is stopped
when writing a zero (Classic interrupts on a 0->-1 transition, Book E
interrupts on a 1->0 transition). Instead, stay in the loop on a
next_dec==0.
Signed-off-by: Matt Porter <mporter@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] ppc32: remove hardcoded offsets from ppc asm
This patch by Vincent Hanquez removes some hard coded offsets for accessing
thread info fields from assembly, uses the normal offset generation
mecanism that we already have for other things instead.
Signed-off-by: Vincent Hanquez <tab@snarc.org> Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Keith Owens [Mon, 23 Aug 2004 05:27:05 +0000 (22:27 -0700)]
[PATCH] Make i386 die() more resilient against recursive errors
Make i386 die() more resilient against recursive errors, almost a cut
and paste of the ia64 die() routine. Much of the patch is indentation
changes.
Mainly to make it easier to add crash, lcrash, kmsgdump or other RAS patches.
They are invoked from die() and if they crash themselves, we have to avoid
recursive loops in die().
Signed-off-by: Keith Owens <kaos@sgi.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Akiyama Nobuyuki [Mon, 23 Aug 2004 05:26:54 +0000 (22:26 -0700)]
[PATCH] NMI trigger switch support for debugging(updated)
I made a patch for debugging with the help of NMI trigger switch.
When kernel hangs severely, keyboard operation(e.g.Ctrl-Alt-Del)
doesn't work properly. This patch enables debugging information
to be displayed on console in this case.
I think this feature is necessary as standard functionality.
Please feel free to use this patch and let me know if you have
any comments.
Background:
When a trouble occurs in kernel, we usually begin to investigate
with following information:
- panic >> panic message.
- oops >> CPU registers and stack trace.
- hang >> **NONE** no standard method established.
How it works:
Most IA32 servers have a NMI switch that fires NMI interrupt up.
The NMI interrupt can interrupt even if kernel is serious state,
for example deadlock under the interrupt disabled.
When the NMI switch is pressed after this feature is activated,
CPU registers and stack trace are displayed on console and then
panic occurs.
This feature is activated or deactivated with sysctl.
On IA32 architecture, only the following are defined as reason
of NMI interrupt:
- memory parity error
- I/O check error
The reason code of NMI switch is not defined, so this patch assumes
that all undefined NMI interrupts are fired by MNI switch.
However, oprofile and NMI watchdog also use undefined NMI interrupt.
Therefore this feature cannot be used at the same time with oprofile
and NMI watchdog. This feature hands NMI interrupt over to oprofile
and NMI watchdog. So, when they have been activated, this feature
doesn't work even if it is activated.
Arnd Bergmann [Mon, 23 Aug 2004 05:26:41 +0000 (22:26 -0700)]
[PATCH] fix reading string module parameters in sysfs
Reading the contents of a module_param_string through sysfs currently
oopses because the param_get_charp() function cannot operate on a
kparam_string struct. This introduces the required param_get_string.
Mike Kravetz [Mon, 23 Aug 2004 05:26:30 +0000 (22:26 -0700)]
[PATCH] proc fs task name locking fix
Races have been observed between excec-time overwriting of task->comm and
/proc accesses to the same data. This causes environment string
information to appear in /proc.
Fix that up by taking task_lock() around updates to and accesses to
task->comm.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
it took more than 80 usecs for XFree86 to do a context-switch!
it turns out that the reason for this (massive) context-switching
overhead is the following change in 2.6.8:
[PATCH] larger IO bitmaps
To demonstrate the effect of this change i've written ioperm-latency.c
(attached), which gives the following on vanilla 2.6.8.1:
# ./ioperm-latency
default no ioperm: scheduling latency: 2528 cycles
turning on port 80 ioperm: scheduling latency: 10563 cycles
turning on port 65535 ioperm: scheduling latency: 10517 cycles
the ChangeSet says:
Now, with the lazy bitmap allocation and per-CPU TSS, this
will really not drain any resources I think.
this is plain wrong. An increase in the IO bitmap size introduces
per-context-switch overhead as well: we now have to copy an 8K bitmap
every time XFree86 context-switches - even though XFree86 never uses
ports higher than 1024! I've straced XFree86 on a number of x86 systems
and in every instance ioperm() was used - so i'd say the majority of x86
Linux systems running 2.6.8.1 are affected by this problem.
This not only causes lots of overhead, it also trashes ~16K out of the
L1 and L2 caches, on every context-switch. It's as if XFree86 did a L1
cache flush on every context-switch ...
the simple solution would be to revert IO_BITMAP_BITS back to 1024 and
release 2.6.8.2?
I've implemented another solution as well, which tracks the
highest-enabled port # for every task and does the copying of the bitmap
intelligently. (patch attached) The patched kernel gives:
# ./ioperm-latency
default no ioperm: scheduling latency: 2423 cycles
turning on port 80 ioperm: scheduling latency: 2503 cycles
turning on port 65535 ioperm: scheduling latency: 10607 cycles
this is much more acceptable - the full overhead only occurs in the very
unlikely event of a task using the high ioport range. X doesnt suffer
any significant overhead.
(tracking the maximum allowed port # also allows a simplification of
io_bitmap handling: e.g. we dont do the invalid-offset trick anymore -
the IO bitmap in the TSS is always valid and secure.)
I tested the patch on x86 SMP and UP, it works fine for me. I tested
boundary conditions as well, it all seems secure.
/*
* Use a pair of RT processes bound to the same CPU to measure
* context-switch overhead:
*/
static void measure(void)
{
unsigned long i, min = ~0UL, pid, mask = 1, t1, t2;
sched_set_affinity(0, sizeof(mask), &mask);
pid = fork();
if (!pid)
for (;;) {
asm volatile ("sti; nop; cli");
sched_yield();
}
sched_yield();
for (i = 0; i < 100; i++) {
asm volatile ("sti; nop; cli");
CYCLES(t1);
sched_yield();
CYCLES(t2);
if (i > 10) {
if (t2 - t1 < min)
min = t2 - t1;
}
}
asm volatile ("sti");
It seems that on some OldWolrd macs, we don't get the OF stdout device,
thus the new set_preferred_console() dies at boot trying to dereference
a NULL pointer.
In 2.5.18 some minix-specific stuff was moved to the minix subdirectory
where it belonged. However, a typo crept in, causing inode disk usage
to be incorrectly reported. A few people have complained, but so far
not sufficiently loudly.
Alan Cox [Sun, 22 Aug 2004 07:03:35 +0000 (00:03 -0700)]
[PATCH] missing CPU descriptors
There are a couple of cache descriptors in the current Intel manuals
missing from our tables at least one of which appears in an actual
processor in the real world.
Jesse Barnes [Sun, 22 Aug 2004 05:30:28 +0000 (22:30 -0700)]
[PATCH] ACPI for 2.6
Define acpi_noirq on ia64 since it's used now in pci_link.c. All ia64
machines use ACPI, so we can just define it to 0 like we do for acpi_disabled
and acpi_pci_disabled.
Arnd Bergmann [Fri, 20 Aug 2004 10:29:39 +0000 (12:29 +0200)]
[WATCHDOG] v2.6.8.1 compat_ioctl-patch
The watchdog ioctl interface is defined correctly for 32 bit emulation,
although WIOC_GETSUPPORT was not marked as such, for an unclear reason.
WDIOC_SETTIMEOUT and WDIOC_GETTIMEOUT were added in may 2002 to the
code but never to the ioctl list. This adds all three definitions.
Signed-off-by: Arnd Bergmann <arnd@arndb.de> Signed-off-by: Wim Van Sebroeck <wim@iguana.be>