[PATCH] x86-64: sibling map fix for clustered mode
From: James Cleverdon
The value that cpuinfo returns for command 1 in ebx is the physical APIC ID
latched when the system comes out of reset.
Ordinarily, this is identical to the value in the local APIC's ID register,
because nearly all BIOSes accept the HW assigned value.
Our systems, made up of individual building blocks, can't do that. Each
node boots as a separate system and is joined together by the BIOS. Thus,
the BIOS rewrites the local APIC ID register with a new value.
Potomac and Nocona chips have a mechanism by which the BIOS writer can
change bits 7:5 to match the assigned cluster ID. Bits 2:0 come from the
thread ID. However, bits 4:3 are still those latched at reset. Oops!
Summary: Large clustered systems can't use cpuid to derive the sibling
information.
Fix: Use the local APIC ID. That's the value we use to online the CPUs, so
it had better be OK. For non-clustered systems, cpuid == local APIC, so
nothing but large boxes should be affected.
David Gibson [Fri, 17 Sep 2004 04:59:04 +0000 (21:59 -0700)]
[PATCH] ppc64: remove LARGE_PAGE_SHIFT constant
For historical reasons, ppc64 has ended up with two #defines for the size
of a large (16M) page: LARGE_PAGE_SHIFT and HPAGE_SHIFT. This patch
removes LARGE_PAGE_SHIFT in favour of the more widely used HPAGE_SHIFT.
Signed-off-by: David Gibson <dwg@au1.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Fri, 17 Sep 2004 04:58:49 +0000 (21:58 -0700)]
[PATCH] ppc64: fix CONFIG_CMDLINE
When I cleaned up our cmdline parsing, I missed a RELOC of CONFIG_CMDLINE
itself. Without it we copy something random into cmd_line, but only when
CONFIG_CMDLINE is enabled.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Fri, 17 Sep 2004 04:58:23 +0000 (21:58 -0700)]
[PATCH] ppc64: fix hotplug CPU when building a pseries+pmac kernel
When a pseries+pmac kernel is built, the rtas stop-self token wasnt being
initialised. Since doing this will safely fail on pmac, remove the
!CONFIG_PPC_PMAC restriction
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Fri, 17 Sep 2004 04:58:10 +0000 (21:58 -0700)]
[PATCH] ppc64: don't use state == SYSTEM_BOOTING
From: Nathan Lynch <nathanl@austin.ibm.com>
Fedora has a patch which introduces a new system state during boot. Change
system_state == SYSTEM_BOOTING to system_state < SYSTEM_RUNNING to match
it.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Fri, 17 Sep 2004 04:56:50 +0000 (21:56 -0700)]
[PATCH] ppc64: replace mmu_context_queue with idr allocator
Replace the mmu_context_queue structure with the idr allocator. The
mmu_context_queue allocation was quite large (~200kB) so on most machines
we will have a reduction in usage.
We might put a single entry cache on the front of this so we are more
likely to reuse ppc64 MMU hashtable entries that are in the caches.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Fri, 17 Sep 2004 04:56:38 +0000 (21:56 -0700)]
[PATCH] ppc64: powersave_nap sysctl
Implement powersave_nap sysctl, like ppc32. This allows us to disable the
nap function which is useful when profiling with oprofile (to get an
accurate count of idle time).
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Olaf Hering [Fri, 17 Sep 2004 04:56:13 +0000 (21:56 -0700)]
[PATCH] ppc32: open_pic2.c build fix
arch/ppc/syslib/open_pic2.c: In function `init_openpic2_sysfs':
arch/ppc/syslib/open_pic2.c:694: error: `ENODEV' undeclared (first use in this function)
arch/ppc/syslib/open_pic2.c:694: error: (Each undeclared identifier is reported only once
arch/ppc/syslib/open_pic2.c:694: error: for each function it appears in.)
possible fix below.
Signed-off-by: Olaf Hering <olh@suse.de> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Tom Rini [Fri, 17 Sep 2004 04:55:59 +0000 (21:55 -0700)]
[PATCH] ppc32: Fix arch/ppc/boot/common/ns16550.c
When <linux/timex.h> started including <asm/io.h> this exposed one of the
fragilities of the code in arch/ppc/boot/, namely that it is tied to the
kernel headers for some information, yet not really the kernel. The
following starts us in the direction of being less tied to the kernel by
providing our own serial_state definition (all we care about is the ability
to grab information from SERIAL_PORT_DFNS).
Signed-off-by: Tom Rini <trini@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Fri, 17 Sep 2004 04:55:43 +0000 (21:55 -0700)]
[PATCH] fix posix-timers leak
Exec fails to clean up posix-timers. This manifests itself in two ways, one
worse than the other. In the single-threaded case, it just fails to clear out
the timers on exec. POSIX says that exec clears out the timers from
timer_create (though not the setitimer ones), so it's wrong that a lingering
timer could fire after exec and kill the process with a signal it's not
expecting. In the multi-threaded case, it not only leaves lingering timers,
but it leaks them entirely when it replaces signal_struct, so they will never
be freed by the process exiting after that exec. The new per-user
RLIMIT_SIGPENDING actually limits the damage here, because a UID will fill up
its quota with leaked timers and then never be able to use timer_create again
(that's what my test program does). But if you have many many untrusted UIDs,
this leak could be considered a DoS risk.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
There's one additional step we can do ontop of the ports-max code to get rid
of copying in X.org's case: cache the last task that set up the IO bitmap.
This means we can set the offset to invalid and keep the IO bitmap of that
task, and switch back to a valid offset (without any copying) when switching
back to that task. (or do a copy if there is another ioperm task we switch
to.)
I've attached ioport-cache-2.6.8.1.patch that implements this. When
there's a single active ioperm() using task in the system then the
context-switch overhead is very low and constant:
# ./ioperm-latency
default no ioperm: scheduling latency: 2478 cycles
turning on port 80 ioperm: scheduling latency: 2499 cycles
turning on port 65535 ioperm: scheduling latency: 2481 cycles
This single-ioperm-user situation matches 99% of the actual ioperm()
usage scenarios and gets rid of any copying whatsoever - without relying
on any fault mechanism. I can see no advantage of the GPF approach over
this patch.
Ryan S. Arnold [Fri, 17 Sep 2004 04:55:05 +0000 (21:55 -0700)]
[PATCH] hvc_console fix to protect hvc_write against ldisc write after hvc_close
Due to the tty ldisc code not stopping write operations against a driver
even after a tty has been closed I added a mechanism to hvc_console in my
previous patch to prevent this by nulling out the tty->driver_data in
hvc_close() but I forgot to check tty->driver_data in hvc_write(). Anton
Blanchard got several oops'es from hvc_write() accessing NULL as if it were
a pointer to an hvc_struct usually stored in tty->driver_data.
So this patch checks tty->driver_data in hvc_write() before it is used.
Hopefully once Alan Cox's patch is checked in ldisc writes won't continue
to happen after tty closes.
Anton Blanchard has tested this patch and is unable to reproduce the oops
with it applied.
Changelog:
drivers/char/hvc_console.c
- Added comment to hvc_close() to explain the reason for NULLing
tty->driver_data.
- Added check to hvc_write() to verify that tty->driver_data is valid
(NOT NULL) which would be the case if the write operation was invoked
after a tty close was initiated on the tty.
Signed-off-by: Ryan S. Arnold <rsa@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Markus Lidel [Fri, 17 Sep 2004 04:54:51 +0000 (21:54 -0700)]
[PATCH] reduce ioremap memory size for Adaptec I2O controllers
The I2O subsystem currently map all memory from the I2O controller for the
controller's in queue, even if it is not necessary. This is a problem,
because on some systems the size returned from pci_resource_len() could be
128MB and only 1-4MB is needed.
Changes:
- only ioremap as much memory as the controller is actually using.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Mark Goodwin [Thu, 16 Sep 2004 22:25:43 +0000 (22:25 +0000)]
[IA64] SGI Altix hardware performance monitoring API
The SGI Altix PROM supports a SAL call for performance monitoring and for
exporting NUMA topology. We need this in community kernels for diagnostic
and performance tools to use, especially on very large machines.
This patch registers a dynamic misc device "sn_hwperf" that supports an
ioctl interface for reading/writing memory mapped registers on Altix
nodes and routers via the new SAL call. It also creates a read-only
procfs file "/proc/sgi_sn/sn_topology" to export NUMA topology and Altix
hardware inventory.
> What tools are using this?
Performance Co-Pilot http://oss.sgi.com/projects/pcp in particular,
pmshub, shubstats and linkstat. Numerous other users include anything
that needs knowledge of numa topology/interconnect in order to perform
well, e.g. mpt. BTW I have not exported any API functions .. at this
point I don't think we need any modules to call the API.
Signed-off-by: Mark Goodwin <markgw@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
Keith Owens [Thu, 16 Sep 2004 19:11:00 +0000 (19:11 +0000)]
[IA64] ar.k[56] have virtual addresses already, don't convert
r.k[56] used to contain physical addresses but now contain virtual
addresses. There are code remnants which still believe that they are
physical and "convert" ar.k[56] to virtual. This breaks when current
is not in region 7 (e.g. the idle task on cpu 0).
Signed-off-by: Keith Owens <kaos@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
Signed-off-by: Bart De Schuymer <bdschuym@pandora.be> Signed-off-by: Patrick McHardy <kaber@trash.net> Signed-off-by: David S. Miller <davem@davemloft.net>
Thomas Graf [Thu, 16 Sep 2004 06:13:12 +0000 (23:13 -0700)]
[PKT_SCHED]: Fix slab corruption in cbq_destroy
Fixes slab corruption in cbq_destroy. cbq_destroy_filters and
qdisc_put_rtab(q->link.R_tab) are already called in cbq_destroy_class.
The latter lead to a slab corruption due to repeated freeing of
q->link.R_tab because q->link is part of q->classes. Problem introduced
in 1.21.
Signed-off-by: Thomas Graf <tgraf@suug.ch> Signed-off-by: Patrick McHardy <kaber@trash.net> Signed-off-by: David S. Miller <davem@davemloft.net>
Herbert Xu [Thu, 16 Sep 2004 06:12:04 +0000 (23:12 -0700)]
[IPSEC]: Implement DSCP decapsulation
This patch adds DSCP decapsulation for IPsec. This is enabled by
a per-state flag which is off by default. Leaving it off by default
maintains compatibility and is also good for performance reasons.
I decided to not implement a toggle on the output path since not
encapsulating the DSCP can and should be done by netfilter.
Signed-off-by: Herbert Xu <herbert@gondor.apana.org.au> Signed-off-by: David S. Miller <davem@davemloft.net>
[IPV6]: NDISC: ensure responding to NS without link-layer information
When sending NA in response to NS, we may not know the
link-layer address for the destination of the NA
since unicast NS is not required to include its link-layer information.
In this case, we first need to resolve the link-layer address.
(RFC 2461 7.2.4)
We now create neighbour entry for the destination
and link-layer information will be automatically solved
in the output path.
Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> Signed-off-by: David S. Miller <davem@davemloft.net>
Petr Vandrovec [Thu, 16 Sep 2004 05:11:21 +0000 (22:11 -0700)]
[PATCH] matroxfb update + sparse annotations
This change switches matroxfb on x86 and x86_64 from dereferencing
pointers to {read,write}[bwl], as __pa() are gone from them, and so gcc
does not need an additional register for preserving address between
consecutive {read,write}[bwl].
Then it switches only supported architecture left (ppc/ppc64/arm) from
dereferencing pointers to __raw_{read,write}[bwl].
Third part is fixing sparse complaints: add __iomem here and there, and
switch one 1bit bitfield from int to unsigned int.
After this there should be no sparse complaints in matroxfb.
Signed-off-by: Petr Vandrovec <vandrove@vc.cvut.cz> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Thu, 16 Sep 2004 00:12:54 +0000 (17:12 -0700)]
[PATCH] back out siginfo_t.si_rusage from waitid changes
As I explained in the waitid patches, I added the si_rusage field to
siginfo_t with the idea of having the siginfo_t waitid fills in contain all
the information that wait4 or any such call could ever tell you. Nowhere
in POSIX nor anywhere else specifies this field in siginfo_t.
When Ulrich and I hashed out the system call interface we wanted, we looked
at siginfo_t and decided there was plenty of space to throw in si_rusage.
Well, it turns out we didn't check the 64-bit platforms. There struct
rusage is ridiculously large (lots of longs for things that are never in a
million years going to hit 2^32), and my changes bumped up the size of
siginfo_t. Changing that size is more trouble than it's worth.
This patch reverts the changes to the siginfo_t structure types,
and no longer provides the rusage details in SIGCHLD signal data.
Instead, I added a fifth argument to the waitid system call to fill in rusage.
waitid is the name of the POSIX function with four arguments. It might
make sense to rename the system call `waitsys' to follow SGI's system call
with the same arguments, or `wait5' in the mindless tradition. But, feh.
I just added the argument to sys_waitid, rather than worrying about
changing the name in all the tables (and choosing a new stupid name).
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[IA64-SGI]: fix `qw' might be used uninitialized warning
The compiler has no way of knowing whether nentries will be greater than 0, so
it was generating a warning that qw might be used uninitialized. Fix it by
explicitly setting it to 0. Cc'ing Brian in case he has an internal version
he'd like to keep in sync.
Signed-off-by: Jesse Barnes <jbarnes@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
Jeff Garzik [Wed, 15 Sep 2004 22:45:22 +0000 (18:45 -0400)]
[libata] add hook, and export functions needed for sata2 drivers
* add dev_select hook, and default/noop implementations
* export ata_dev_classify
* fix a couple bugs that cropped up when building with
ATA_VERBOSE_DEBUG
* export __sata_phy_reset, a variant that does not call
ata_bus_reset
[IPV6]: Missing xfrm_lookup() in icmpv6_{send,echo_reply}()
net/ipv6/icmp.c was not converted in xfrm_lookup() extraction patch.
This patch converts it; adding the missing call to xfrm_lookup in
icmpv6_{send,echo_reply}().
Signed-off-by: Kazunori Miyazawa <kazunori@miyazawa.org> Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> SIgned-off-by: David S. Miller <davem@davemloft.net>
David S. Miller [Tue, 14 Sep 2004 15:03:10 +0000 (08:03 -0700)]
[IPV4]: Make fib_semantics algorithms scale better.
A singly linked list was previously used to
do fib_info object lookup for various actions
in the routing code. This does not scale very
well with many devices and even a moderate number
of routes. This was noted by Benjamin Lahaise.
To fix all of this we use 3 hash tables, two of
which grow dynamically as the number fib_info
objects increases while the final one is fixes in
size.
The statically sized table hashes on device index.
This is used for fib_sync_down, fib_sync_up, and
ip_fib_check_default.
The first dynamically sized table is keyed on
protocol, prefsrc, and priority. This is used
by fib_create_info() to look for existing fib_info
objects matching the new one being constructed.
The last dynamically sized table is keyed on
the preferred source of the route if it has one
specified. This is used by fib_sync_down when
a local address disappears.
There are still some scalability problems for
Bens test case in fib_hash.c and I will try to
attack those next.
Signed-off-by: David S. Miller <davem@davemloft.net>
Noticed by BenH, happily harmless (nothing that uses that
code has been committed yet, and PIO seems to be pretty much
unused on at least the Apple G5 machines: all the normal
hardware is set up purely for MMIO, to the point that I
couldn't even test this thing).
Nicolas Pitre [Tue, 14 Sep 2004 10:54:04 +0000 (11:54 +0100)]
[ARM PATCH] 2094/1: don't lose the system timer after resuming from sleep on SA11x0 and
PXA2xx
Patch from Nicolas Pitre
Let's make sure OSCR doesn't end up to be restored with a value
past OSMR0 otherwise the system timer won't start ticking until
OSCR wraps around (aprox 17 min.
Also set OSCR _after_ OIER is restored to avoid matching when
corresponding match interrupt is masked out.