[IPV6]: Missing xfrm_lookup() in icmpv6_{send,echo_reply}()
net/ipv6/icmp.c was not converted in xfrm_lookup() extraction patch.
This patch converts it; adding the missing call to xfrm_lookup in
icmpv6_{send,echo_reply}().
Signed-off-by: Kazunori Miyazawa <kazunori@miyazawa.org> Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> SIgned-off-by: David S. Miller <davem@davemloft.net>
David S. Miller [Tue, 14 Sep 2004 15:03:10 +0000 (08:03 -0700)]
[IPV4]: Make fib_semantics algorithms scale better.
A singly linked list was previously used to
do fib_info object lookup for various actions
in the routing code. This does not scale very
well with many devices and even a moderate number
of routes. This was noted by Benjamin Lahaise.
To fix all of this we use 3 hash tables, two of
which grow dynamically as the number fib_info
objects increases while the final one is fixes in
size.
The statically sized table hashes on device index.
This is used for fib_sync_down, fib_sync_up, and
ip_fib_check_default.
The first dynamically sized table is keyed on
protocol, prefsrc, and priority. This is used
by fib_create_info() to look for existing fib_info
objects matching the new one being constructed.
The last dynamically sized table is keyed on
the preferred source of the route if it has one
specified. This is used by fib_sync_down when
a local address disappears.
There are still some scalability problems for
Bens test case in fib_hash.c and I will try to
attack those next.
Signed-off-by: David S. Miller <davem@davemloft.net>
The final ia64 related cleanup to elf_read_implies_exec() seems to have
broken it. The effect is that the READ_IMPLIES_EXEC flag is never set
for !pt_gnu_stack binaries!
That's a bit more secure than we need to be, and might break some legacy
app that doesn't expect it.
Fix ABI in set_mempolicy() that got broken by an earlier change.
Add a check for very big input values and prevent excessive looping in the
kernel.
Cc: Paul "nyer, nyer, your mother wears combat boots!" Jackson <pj@sgi.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
As it's using the obsolete MOD_{INC,DEC}_USE_COUNT it's implicitly locked
already, but let's remove them and make it explicit so these macros can go
away completely without breaking m68k compile.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Tue, 14 Sep 2004 00:51:24 +0000 (17:51 -0700)]
[PATCH] BSD disklabel: handle more than 8 partitions
NetBSD allows 16 partitions, not just 8. This patch both ups the number,
and makes the recognition code tell you if the count in the disklabel
exceeds the number supported by the kernel.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
- I misspelled CONFIG_PREEMPT CONFIG_PREEPT as various people noticed.
But in fact that ifdef should just go, else we'll get drivers that
compile with CONFIG_PREEMPT but not without sooner or later.
- remove unused hardirq_trylock and hardirq_endlock
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] Add support for word-length UART registers
UARTS on several Intel IXP2000 systems are connected in such a way that
they can only be addressed using full word accesses instead of bytes.
Following patch adds a UPIO_MEM32 io-type to identify these UARTs.
Drop a config option which has disappeared from all archs. Btw, this
shouldn't be in the UML-specific part, but since we cannot include generic
Kconfigs to avoid problem with hardware-related configs, it's duplicated
for now.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: refer to CONFIG_USERMODE, not to CONFIG_UM
Correct one Kconfig dependency, which should refer to CONFIG_USERMODE
rather than to CONFIG_UM.
We should also figure out how to make the config process work better for
UML. We would like to make UML able to "source drivers/Kconfig" and have
the right drivers selectable (i.e. LVM, ramdisk, and so on) and the ones
for actual hardware excluded. I've been reading such a request even from
Jeff Dike at the last Kernel Summit, (in the lwn.net coverage) but without
any followup.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 14 Sep 2004 00:49:28 +0000 (17:49 -0700)]
[PATCH] uml: disable pending signals across a reboot
On reboot, all signals and signal sources are disabled so that
late-arriving signals don't show up after the reboot exec, confusing the
new image, which is not expecting signals yet.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 14 Sep 2004 00:49:04 +0000 (17:49 -0700)]
[PATCH] uml: fix scheduler race
This fixes a use-after-free bug in the context switching. A process going
out of context after exiting wakes up the next process and then kills
itself. The problem is that when it gets around to killing itself is up to
the host and can happen a long time later, including after the incoming
process has freed its stack, and that memory is possibly being used for
something else.
The fix is to have the incoming process kill the exiting process just to
make sure it can't be running at the point that its stack is freed.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ryan S. Arnold [Tue, 14 Sep 2004 00:48:26 +0000 (17:48 -0700)]
[PATCH] HVCS fix to replace yield with tty_wait_until_sent in hvcs_close
Following the same advice you gave in a recent hvc_console patch I have
modified HVCS to remove a while() { yield(); } from hvcs_close() which may
cause problems where real time scheduling is concerned and replaced it with
tty_wait_until_sent() which uses a real wait queue and is the proper method
for blocking a tty operation while waiting for data to be sent. This patch
has been tested to verify that all the paths of code that were changed were
hit during the code run and performed as expected including hotplug remove
of hvcs adapters and hangup of ttys.
- Replaced yield() in hvcs_close() with tty_wait_until_sent() to prevent
possible lockup with realtime scheduling.
- Removed hvcs_final_close() and reordered cleanup operations to prevent
discarding of pending data during an hvcs_close() call.
- Removed spinlock protection of hvcs_struct data members in
hvcs_write_room() and hvcs_chars_in_buffer() because they aren't needed.
Signed-off-by: Ryan S. Arnold <rsa@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
max_hw_sectors_kb is the maximum that the driver can handle and is
readonly. max_sectors_kb is the current max_sectors value and can be tuned
by root. PAGE_SIZE granularity is enforced.
It's all locking-safe and all affected layered drivers have been updated as
well. The patch has been in testing for a couple of weeks already as part
of the voluntary-preempt patches and it works just fine - people use it to
reduce IDE IRQ handling latencies.
This patch adds a prctl to modify current->comm as shown in /proc. This
feature was requested by KDE developers. In KDE most programs are started by
forking from a kdeinit program that already has the libraries loaded and some
other state.
Problem is to give these forked programs the proper name. It already writes
the command line in the environment (as seen in ps), but top uses a different
field in /proc/pid/status that reports current->comm. And that was always
"kdeinit" instead of the real command name. So you ended up with lots of
kdeinits in your top listing, which was not very useful.
This patch adds a new prctl PR_SET_NAME to allow a program to change its comm
field.
I considered the potential security issues of a program obscuring itself with
this interface, but I don't think it matters much because a program can
already obscure itself when the admin uses ps instead of top. In case of a
KDE desktop calling everything kdeinit is much more obfuscation than the
alternative.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] __copy_to_user() check in cdrom_read_cdda_old()
akpm: really, reads are supposed to return the number-of-bytes-read on faults,
or -EFAULT of no bytes were read. This patch returns either zero or -EFAULT,
ignoring any successfully transferred data. But the user interface (whcih is
an ioctl()) was never set up to do that.
I _think_ shmem_file_setup is protected against negative loff_t size by the
TASK_SIZE in each arch, but prefer the security of an explicit test. Wipe
those parentheses off its return(file), and update our Copyright.
Very minor adjustments to shmem_getpage return path: I now prefer it to return
NULL and let do_shmem_file_read use ZERO_PAGE(0) in that case; and we don't
need a local majmin variable, do_no_page initializes *type to VM_FAULT_MINOR
already.
If we're thinking about shmem scalability... isn't it silly that each shmem
object is added to the shmem_inodes list on creation, and removed on deletion,
yet the only use for that list is in swapoff (shmem_unuse)?
Call it shmem_swaplist; shmem_writepage add inode to swaplist when first swap
allocated (usually never); shmem_delete_inode remove inode from the list after
truncating (if called before, inode could be re-added to it).
Inode can remain on the swaplist after all its pages are swapped back in, just
be lazy about it; but if shmem_unuse finds swapped count now 0, save itself
time by then removing that inode from the swaplist.
Some might want a tmpfs mount with the improved scalability afforded by
omitting shmem superblock accounting; or some might just want to test it in an
externally-visible tmpfs mount instance.
Adopt the convention that mount option -o nr_blocks=0,nr_inodes=0 means
without resource limits, and hence no shmem_sb_info. Not recommended for
general use, but no worse than ramfs.
Disallow remounting from unlimited to limited (no accounting has been done so
far, so no idea whether it's permissible), and from limited to unlimited
(because we'd need then to free the sbinfo, and visit each inode to reset its
i_blocks to 0: why bother?).
SGI investigations have shown a dramatic contrast in scalability between
anonymous memory and shmem objects. Processes building distinct shmem objects
in parallel hit heavy contention on shmem superblock stat_lock. Across 256
cpus an intensive test runs 300 times slower than anonymous.
Jack Steiner has observed that all the shmem superblock free_blocks and
free_inodes accounting is redundant in the case of the internal mount used for
SysV shared memory and for shared writable /dev/zero objects (the cases which
most concern them): it specifically declines to limit.
Based upon Brent Casavant's SHMEM_NOSBINFO patch, this instead just removes
the shmem_sb_info structure from the internal kernel mount, testing where
necessary for null sbinfo pointer. shmem_set_size moved within CONFIG_TMPFS,
its arg named "sbinfo" as elsewhere.
This brings shmem object scalability up to that of anonymous memory, in the
case where distinct processes are building (faulting to allocate) distinct
objects. It significantly improves parallel building of a shared shmem object
(that test runs 14 times faster across 256 cpus), but other issues remain in
that case: to be addressed in later patches.
Keith Mannthey's Bugzilla #3268 drew attention to how tmpfs inodes and
dentries and long names and radix-tree nodes pin lowmem. Assuming about 1k of
lowmem per inode, we need to lower the default nr_inodes limit on machines
with significant highmem.
Be conservative, but more generous than in the original patch to Keith: limit
to number of lowmem pages, which works out around 200,000 on i386. Easily
overridden by giving the nr_inodes= mount option: those who want to sail
closer to the rocks should be allowed to do so.
Notice how tmpfs dentries cannot be reclaimed in the way that disk-based
dentries can: so even hard links need to be costed. They are cheaper than
inodes, but easier all round to charge the same. This way, the limit for hard
links is equally visible through "df -i": but expect occasional bugreports
that tmpfs links are being treated like this.
Would have been simpler just to move the free_inodes accounting from
shmem_delete_inode to shmem_unlink; but that would lose the charge on unlinked
but open files.
this patch fix a pnpbios problem with independant
resource(http://bugzilla.kernel.org/show_bug.cgi?id=3295) :
the old code assume that they are given at the beggining (before any
SMALL_TAG_STARTDEP entry), but in some case there are found after
SMALL_TAG_ENDDEP entry.
tag : 6 SMALL_TAG_STARTDEP
tag : 8 SMALL_TAG_PORT
tag : 6 SMALL_TAG_STARTDEP
tag : 8 SMALL_TAG_PORT
tag : 7 SMALL_TAG_ENDDEP
tag : 4 SMALL_TAG_IRQ <-- independant resource
tag : f SMALL_TAG_END
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Tue, 14 Sep 2004 00:44:21 +0000 (17:44 -0700)]
[PATCH] ppc64: improved VSID allocation algorithm
This patch has been tested both on SLB and segment table machines. This
new approach is far from the final word in VSID/context allocation, but
it's a noticeable improvement on the old method.
Replace the VSID allocation algorithm. The new algorithm first generates a
36-bit "proto-VSID" (with 0xfffffffff reserved). For kernel addresses this
is equal to the ESID (address >> 28), for user addresses it is:
(context << 15) | (esid & 0x7fff)
These are distinguishable from kernel proto-VSIDs because the top bit is
clear. Proto-VSIDs with the top two bits equal to 0b10 are reserved for
now.
The proto-VSIDs are then scrambled into real VSIDs with the multiplicative
hash:
This scramble is 1:1, because VSID_MULTIPLIER and VSID_MODULUS are co-prime
since VSID_MULTIPLIER is prime (the largest 28-bit prime, in fact).
This scheme has a number of advantages over the old one:
- We now have VSIDs for every kernel address (i.e. everything above
0xC000000000000000), except the very top segment. That simplifies a
number of things.
- We allow for 15 significant bits of ESID for user addresses with 20
bits of context. i.e. 8T (43 bits) of address space for up to 1M
contexts, significantly more than the old method (although we will need
changes in the hash path and context allocation to take advantage of
this).
- Because we use a real multiplicative hash function, we have better and
more robust hash scattering with this VSID algorithm (at least based on
some initial results).
Because the MODULUS is 2^n-1 we can use a trick to compute it efficiently
without a divide or extra multiply. This makes the new algorithm barely
slower than the old one.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:43:18 +0000 (17:43 -0700)]
[PATCH] ppc64: clean up idle loop code
Clean up our idle loop code:
- Remove a bunch of useless includes and make most functions static
- There were places where we werent disabling interrupts before checking
need_resched then calling the hypervisor to sleep our thread. We might
race with an IPI and end up missing a reschedule. Disable interrupts
around these regions to make them safe.
- We forgot to turn off the polling flag when exiting the dedicated_idle
idle loop. This could have resulted in all manner problems as other
cpus would avoid sending IPIs to force reschedules.
- Add a missing check for cpu_is_offline in the shared cpu idle loop.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:43:06 +0000 (17:43 -0700)]
[PATCH] ppc64: enable POWER5 low power mode in idle loop
Now that we understand (and have fixed) the problem with using low power mode
in the idle loop, lets enable it. It should save a fair amount of power.
(The problem was that our exceptions were inheriting the low power mode and so
were executing at a fraction of the normal cpu issue rate. We fixed it by
always bumping our priority to medium at the start of every exception).
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:54 +0000 (17:42 -0700)]
[PATCH] ppc64: restore smt-enabled=off kernel command line option
Restore the smt-enabled=off kernel command line functionality:
- Remove the SMT_DYNAMIC state now that smt_snooze_delay allows for the
same thing.
- Remove the early prom.c parsing for the option, put it into an
early_param instead.
- In setup_cpu_maps honour the smt-enabled setting
Note to Nathan: In order to allow cpu hotplug add of secondary threads after
booting with smt-enabled=off, I had to initialise cpu_present_map to
cpu_online_map in smp_cpus_done. Im not sure how you want to handle this but
it seems our present map currently does not allow cpus to be added into the
partition that werent there at boot (but were in the possible map).
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:28 +0000 (17:42 -0700)]
[PATCH] ppc64: remove EEH command line device matching code
We have had reports of people attempting to disable EEH on POWER5 boxes. This
is not supported, and the device will most likely not respond to config space
reads/writes. Remove the IBM location matching code that was being used to
disable devices as well as the global option.
We already have the ability to ignore EEH erros via the panic_on_oops sysctl
option, advanced users should make use of that instead.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:04 +0000 (17:42 -0700)]
[PATCH] ppc64: clean up kernel command line code
Clean up some of our command line code:
- We were copying the command line out of the device tree twice, but the
first time we forgot to add CONFIG_CMDLINE. Fix this and remove the
second copy.
- The command line birec code ran after we had done some command line
parsing in prom.c. This had the opportunity to really confuse the
user, with some options being parsed out of the device tree and the
other out of birecs. Luckily we could find no user of the command
line birecs, so remove them.
- remove duplicate printing of kernel command line;
- clean up iseries inits and create an iSeries_parse_cmdline.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:41:51 +0000 (17:41 -0700)]
[PATCH] ppc64: use nm --synthetic where available
On new toolchains we need to use nm --synthetic or we miss code symbols. Sam,
I'm not thrilled about this patch but Im not sure of an easier way. Any ideas?
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Will Schmidt [Tue, 14 Sep 2004 00:40:39 +0000 (17:40 -0700)]
[PATCH] ppc64: lparcfg fixes for processor counts
This patch corrects how the lparcfg interface was presenting the number of
active and potential processors. (As reported in LTC bugzilla number 10889).
- Correct output for partition_potential_processors and
system_active_processors.
- suppress pool related values in scenarios where they do not make
sense. (non-shared processor configurations)
- Display pool_capacity as a percentage, to match the behavior from
iSeries code.
Signed-off-by: Will Schmidt <willschm@us.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jason Davis [Tue, 14 Sep 2004 00:40:15 +0000 (17:40 -0700)]
[PATCH] ES7000 subarch update
The patch below implements an algorithm to determine an unique GSI override
for mapping GSIs to IO-APIC pins correctly. GSI overrides are required in
order for ES7000 machines to function properly since IRQ to pin mappings
are NOT all one-to-one. This patch applies only to the Unisys specific
ES7000 machines and has been tested thoroughly on several models of the
ES7000 line.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
In sched_exec, schedstat_inc will dereference a null pointer if no domain
is found with the SD_BALANCE_EXEC flag set. This was exposed during
testing of the previous patches where cpus are temporarily attached to a
dummy domain without SD_BALANCE_EXEC set.
With this patch we get two new slabcaches, for sctp socks, that previously
were being allocated on the default, that was tcp[6]_sock, i.e. wasting 288
bytes per sock in the IPv4 case and 256 bytes for the IPv6 version, this is in
preparation for DCCP (or any other new protocol :) ).
With this in place another nice side effect that is easier to achieve is to
get rid of struct sock sk->sk_slab, and instead use sk->sk_prot->slab, saving
sizeof(void *) on every struct sock instance, but this unfortunatly has to
wait for the conversion of all protocols that use per socket slabcaches to
use sk->sk_prot, AF_UNIX is the only one AFAIK, so I'll try to convert it to
use sk->sk_prot and then get rid of sk->sk_slab.
As for the protocols that don't use per socket slabcaches its just a matter
of defaulting to sk_cachep at sk_free time if sk->sk_prot is NULL.
B44 driver was using unsigned long as an io memory address.
Recent changes caused this to be a warning. This patch fixes that
and makes the readl/writel wrapper into inline's instead of macros
with magic variable side effect (yuck).
Signed-off-by: David S. Miller <davem@davemloft.net>