Mark Goodwin [Thu, 16 Sep 2004 22:25:43 +0000 (22:25 +0000)]
[IA64] SGI Altix hardware performance monitoring API
The SGI Altix PROM supports a SAL call for performance monitoring and for
exporting NUMA topology. We need this in community kernels for diagnostic
and performance tools to use, especially on very large machines.
This patch registers a dynamic misc device "sn_hwperf" that supports an
ioctl interface for reading/writing memory mapped registers on Altix
nodes and routers via the new SAL call. It also creates a read-only
procfs file "/proc/sgi_sn/sn_topology" to export NUMA topology and Altix
hardware inventory.
> What tools are using this?
Performance Co-Pilot http://oss.sgi.com/projects/pcp in particular,
pmshub, shubstats and linkstat. Numerous other users include anything
that needs knowledge of numa topology/interconnect in order to perform
well, e.g. mpt. BTW I have not exported any API functions .. at this
point I don't think we need any modules to call the API.
Signed-off-by: Mark Goodwin <markgw@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
Keith Owens [Thu, 16 Sep 2004 19:11:00 +0000 (19:11 +0000)]
[IA64] ar.k[56] have virtual addresses already, don't convert
r.k[56] used to contain physical addresses but now contain virtual
addresses. There are code remnants which still believe that they are
physical and "convert" ar.k[56] to virtual. This breaks when current
is not in region 7 (e.g. the idle task on cpu 0).
Signed-off-by: Keith Owens <kaos@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
Signed-off-by: Bart De Schuymer <bdschuym@pandora.be> Signed-off-by: Patrick McHardy <kaber@trash.net> Signed-off-by: David S. Miller <davem@davemloft.net>
Thomas Graf [Thu, 16 Sep 2004 06:13:12 +0000 (23:13 -0700)]
[PKT_SCHED]: Fix slab corruption in cbq_destroy
Fixes slab corruption in cbq_destroy. cbq_destroy_filters and
qdisc_put_rtab(q->link.R_tab) are already called in cbq_destroy_class.
The latter lead to a slab corruption due to repeated freeing of
q->link.R_tab because q->link is part of q->classes. Problem introduced
in 1.21.
Signed-off-by: Thomas Graf <tgraf@suug.ch> Signed-off-by: Patrick McHardy <kaber@trash.net> Signed-off-by: David S. Miller <davem@davemloft.net>
Herbert Xu [Thu, 16 Sep 2004 06:12:04 +0000 (23:12 -0700)]
[IPSEC]: Implement DSCP decapsulation
This patch adds DSCP decapsulation for IPsec. This is enabled by
a per-state flag which is off by default. Leaving it off by default
maintains compatibility and is also good for performance reasons.
I decided to not implement a toggle on the output path since not
encapsulating the DSCP can and should be done by netfilter.
Signed-off-by: Herbert Xu <herbert@gondor.apana.org.au> Signed-off-by: David S. Miller <davem@davemloft.net>
[IPV6]: NDISC: ensure responding to NS without link-layer information
When sending NA in response to NS, we may not know the
link-layer address for the destination of the NA
since unicast NS is not required to include its link-layer information.
In this case, we first need to resolve the link-layer address.
(RFC 2461 7.2.4)
We now create neighbour entry for the destination
and link-layer information will be automatically solved
in the output path.
Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> Signed-off-by: David S. Miller <davem@davemloft.net>
Petr Vandrovec [Thu, 16 Sep 2004 05:11:21 +0000 (22:11 -0700)]
[PATCH] matroxfb update + sparse annotations
This change switches matroxfb on x86 and x86_64 from dereferencing
pointers to {read,write}[bwl], as __pa() are gone from them, and so gcc
does not need an additional register for preserving address between
consecutive {read,write}[bwl].
Then it switches only supported architecture left (ppc/ppc64/arm) from
dereferencing pointers to __raw_{read,write}[bwl].
Third part is fixing sparse complaints: add __iomem here and there, and
switch one 1bit bitfield from int to unsigned int.
After this there should be no sparse complaints in matroxfb.
Signed-off-by: Petr Vandrovec <vandrove@vc.cvut.cz> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Thu, 16 Sep 2004 00:12:54 +0000 (17:12 -0700)]
[PATCH] back out siginfo_t.si_rusage from waitid changes
As I explained in the waitid patches, I added the si_rusage field to
siginfo_t with the idea of having the siginfo_t waitid fills in contain all
the information that wait4 or any such call could ever tell you. Nowhere
in POSIX nor anywhere else specifies this field in siginfo_t.
When Ulrich and I hashed out the system call interface we wanted, we looked
at siginfo_t and decided there was plenty of space to throw in si_rusage.
Well, it turns out we didn't check the 64-bit platforms. There struct
rusage is ridiculously large (lots of longs for things that are never in a
million years going to hit 2^32), and my changes bumped up the size of
siginfo_t. Changing that size is more trouble than it's worth.
This patch reverts the changes to the siginfo_t structure types,
and no longer provides the rusage details in SIGCHLD signal data.
Instead, I added a fifth argument to the waitid system call to fill in rusage.
waitid is the name of the POSIX function with four arguments. It might
make sense to rename the system call `waitsys' to follow SGI's system call
with the same arguments, or `wait5' in the mindless tradition. But, feh.
I just added the argument to sys_waitid, rather than worrying about
changing the name in all the tables (and choosing a new stupid name).
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[IA64-SGI]: fix `qw' might be used uninitialized warning
The compiler has no way of knowing whether nentries will be greater than 0, so
it was generating a warning that qw might be used uninitialized. Fix it by
explicitly setting it to 0. Cc'ing Brian in case he has an internal version
he'd like to keep in sync.
Signed-off-by: Jesse Barnes <jbarnes@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
Jeff Garzik [Wed, 15 Sep 2004 22:45:22 +0000 (18:45 -0400)]
[libata] add hook, and export functions needed for sata2 drivers
* add dev_select hook, and default/noop implementations
* export ata_dev_classify
* fix a couple bugs that cropped up when building with
ATA_VERBOSE_DEBUG
* export __sata_phy_reset, a variant that does not call
ata_bus_reset
[IPV6]: Missing xfrm_lookup() in icmpv6_{send,echo_reply}()
net/ipv6/icmp.c was not converted in xfrm_lookup() extraction patch.
This patch converts it; adding the missing call to xfrm_lookup in
icmpv6_{send,echo_reply}().
Signed-off-by: Kazunori Miyazawa <kazunori@miyazawa.org> Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> SIgned-off-by: David S. Miller <davem@davemloft.net>
David S. Miller [Tue, 14 Sep 2004 15:03:10 +0000 (08:03 -0700)]
[IPV4]: Make fib_semantics algorithms scale better.
A singly linked list was previously used to
do fib_info object lookup for various actions
in the routing code. This does not scale very
well with many devices and even a moderate number
of routes. This was noted by Benjamin Lahaise.
To fix all of this we use 3 hash tables, two of
which grow dynamically as the number fib_info
objects increases while the final one is fixes in
size.
The statically sized table hashes on device index.
This is used for fib_sync_down, fib_sync_up, and
ip_fib_check_default.
The first dynamically sized table is keyed on
protocol, prefsrc, and priority. This is used
by fib_create_info() to look for existing fib_info
objects matching the new one being constructed.
The last dynamically sized table is keyed on
the preferred source of the route if it has one
specified. This is used by fib_sync_down when
a local address disappears.
There are still some scalability problems for
Bens test case in fib_hash.c and I will try to
attack those next.
Signed-off-by: David S. Miller <davem@davemloft.net>
Noticed by BenH, happily harmless (nothing that uses that
code has been committed yet, and PIO seems to be pretty much
unused on at least the Apple G5 machines: all the normal
hardware is set up purely for MMIO, to the point that I
couldn't even test this thing).
Nicolas Pitre [Tue, 14 Sep 2004 10:54:04 +0000 (11:54 +0100)]
[ARM PATCH] 2094/1: don't lose the system timer after resuming from sleep on SA11x0 and
PXA2xx
Patch from Nicolas Pitre
Let's make sure OSCR doesn't end up to be restored with a value
past OSMR0 otherwise the system timer won't start ticking until
OSCR wraps around (aprox 17 min.
Also set OSCR _after_ OIER is restored to avoid matching when
corresponding match interrupt is masked out.
The final ia64 related cleanup to elf_read_implies_exec() seems to have
broken it. The effect is that the READ_IMPLIES_EXEC flag is never set
for !pt_gnu_stack binaries!
That's a bit more secure than we need to be, and might break some legacy
app that doesn't expect it.
Fix ABI in set_mempolicy() that got broken by an earlier change.
Add a check for very big input values and prevent excessive looping in the
kernel.
Cc: Paul "nyer, nyer, your mother wears combat boots!" Jackson <pj@sgi.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
As it's using the obsolete MOD_{INC,DEC}_USE_COUNT it's implicitly locked
already, but let's remove them and make it explicit so these macros can go
away completely without breaking m68k compile.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Tue, 14 Sep 2004 00:51:24 +0000 (17:51 -0700)]
[PATCH] BSD disklabel: handle more than 8 partitions
NetBSD allows 16 partitions, not just 8. This patch both ups the number,
and makes the recognition code tell you if the count in the disklabel
exceeds the number supported by the kernel.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
- I misspelled CONFIG_PREEMPT CONFIG_PREEPT as various people noticed.
But in fact that ifdef should just go, else we'll get drivers that
compile with CONFIG_PREEMPT but not without sooner or later.
- remove unused hardirq_trylock and hardirq_endlock
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] Add support for word-length UART registers
UARTS on several Intel IXP2000 systems are connected in such a way that
they can only be addressed using full word accesses instead of bytes.
Following patch adds a UPIO_MEM32 io-type to identify these UARTs.
Drop a config option which has disappeared from all archs. Btw, this
shouldn't be in the UML-specific part, but since we cannot include generic
Kconfigs to avoid problem with hardware-related configs, it's duplicated
for now.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: refer to CONFIG_USERMODE, not to CONFIG_UM
Correct one Kconfig dependency, which should refer to CONFIG_USERMODE
rather than to CONFIG_UM.
We should also figure out how to make the config process work better for
UML. We would like to make UML able to "source drivers/Kconfig" and have
the right drivers selectable (i.e. LVM, ramdisk, and so on) and the ones
for actual hardware excluded. I've been reading such a request even from
Jeff Dike at the last Kernel Summit, (in the lwn.net coverage) but without
any followup.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 14 Sep 2004 00:49:28 +0000 (17:49 -0700)]
[PATCH] uml: disable pending signals across a reboot
On reboot, all signals and signal sources are disabled so that
late-arriving signals don't show up after the reboot exec, confusing the
new image, which is not expecting signals yet.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 14 Sep 2004 00:49:04 +0000 (17:49 -0700)]
[PATCH] uml: fix scheduler race
This fixes a use-after-free bug in the context switching. A process going
out of context after exiting wakes up the next process and then kills
itself. The problem is that when it gets around to killing itself is up to
the host and can happen a long time later, including after the incoming
process has freed its stack, and that memory is possibly being used for
something else.
The fix is to have the incoming process kill the exiting process just to
make sure it can't be running at the point that its stack is freed.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ryan S. Arnold [Tue, 14 Sep 2004 00:48:26 +0000 (17:48 -0700)]
[PATCH] HVCS fix to replace yield with tty_wait_until_sent in hvcs_close
Following the same advice you gave in a recent hvc_console patch I have
modified HVCS to remove a while() { yield(); } from hvcs_close() which may
cause problems where real time scheduling is concerned and replaced it with
tty_wait_until_sent() which uses a real wait queue and is the proper method
for blocking a tty operation while waiting for data to be sent. This patch
has been tested to verify that all the paths of code that were changed were
hit during the code run and performed as expected including hotplug remove
of hvcs adapters and hangup of ttys.
- Replaced yield() in hvcs_close() with tty_wait_until_sent() to prevent
possible lockup with realtime scheduling.
- Removed hvcs_final_close() and reordered cleanup operations to prevent
discarding of pending data during an hvcs_close() call.
- Removed spinlock protection of hvcs_struct data members in
hvcs_write_room() and hvcs_chars_in_buffer() because they aren't needed.
Signed-off-by: Ryan S. Arnold <rsa@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
max_hw_sectors_kb is the maximum that the driver can handle and is
readonly. max_sectors_kb is the current max_sectors value and can be tuned
by root. PAGE_SIZE granularity is enforced.
It's all locking-safe and all affected layered drivers have been updated as
well. The patch has been in testing for a couple of weeks already as part
of the voluntary-preempt patches and it works just fine - people use it to
reduce IDE IRQ handling latencies.
This patch adds a prctl to modify current->comm as shown in /proc. This
feature was requested by KDE developers. In KDE most programs are started by
forking from a kdeinit program that already has the libraries loaded and some
other state.
Problem is to give these forked programs the proper name. It already writes
the command line in the environment (as seen in ps), but top uses a different
field in /proc/pid/status that reports current->comm. And that was always
"kdeinit" instead of the real command name. So you ended up with lots of
kdeinits in your top listing, which was not very useful.
This patch adds a new prctl PR_SET_NAME to allow a program to change its comm
field.
I considered the potential security issues of a program obscuring itself with
this interface, but I don't think it matters much because a program can
already obscure itself when the admin uses ps instead of top. In case of a
KDE desktop calling everything kdeinit is much more obfuscation than the
alternative.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] __copy_to_user() check in cdrom_read_cdda_old()
akpm: really, reads are supposed to return the number-of-bytes-read on faults,
or -EFAULT of no bytes were read. This patch returns either zero or -EFAULT,
ignoring any successfully transferred data. But the user interface (whcih is
an ioctl()) was never set up to do that.
I _think_ shmem_file_setup is protected against negative loff_t size by the
TASK_SIZE in each arch, but prefer the security of an explicit test. Wipe
those parentheses off its return(file), and update our Copyright.