[IPV6]: NDISC: ensure responding to NS without link-layer information
When sending NA in response to NS, we may not know the
link-layer address for the destination of the NA
since unicast NS is not required to include its link-layer information.
In this case, we first need to resolve the link-layer address.
(RFC 2461 7.2.4)
We now create neighbour entry for the destination
and link-layer information will be automatically solved
in the output path.
Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> Signed-off-by: David S. Miller <davem@davemloft.net>
Petr Vandrovec [Thu, 16 Sep 2004 05:11:21 +0000 (22:11 -0700)]
[PATCH] matroxfb update + sparse annotations
This change switches matroxfb on x86 and x86_64 from dereferencing
pointers to {read,write}[bwl], as __pa() are gone from them, and so gcc
does not need an additional register for preserving address between
consecutive {read,write}[bwl].
Then it switches only supported architecture left (ppc/ppc64/arm) from
dereferencing pointers to __raw_{read,write}[bwl].
Third part is fixing sparse complaints: add __iomem here and there, and
switch one 1bit bitfield from int to unsigned int.
After this there should be no sparse complaints in matroxfb.
Signed-off-by: Petr Vandrovec <vandrove@vc.cvut.cz> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Thu, 16 Sep 2004 00:12:54 +0000 (17:12 -0700)]
[PATCH] back out siginfo_t.si_rusage from waitid changes
As I explained in the waitid patches, I added the si_rusage field to
siginfo_t with the idea of having the siginfo_t waitid fills in contain all
the information that wait4 or any such call could ever tell you. Nowhere
in POSIX nor anywhere else specifies this field in siginfo_t.
When Ulrich and I hashed out the system call interface we wanted, we looked
at siginfo_t and decided there was plenty of space to throw in si_rusage.
Well, it turns out we didn't check the 64-bit platforms. There struct
rusage is ridiculously large (lots of longs for things that are never in a
million years going to hit 2^32), and my changes bumped up the size of
siginfo_t. Changing that size is more trouble than it's worth.
This patch reverts the changes to the siginfo_t structure types,
and no longer provides the rusage details in SIGCHLD signal data.
Instead, I added a fifth argument to the waitid system call to fill in rusage.
waitid is the name of the POSIX function with four arguments. It might
make sense to rename the system call `waitsys' to follow SGI's system call
with the same arguments, or `wait5' in the mindless tradition. But, feh.
I just added the argument to sys_waitid, rather than worrying about
changing the name in all the tables (and choosing a new stupid name).
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Garzik [Wed, 15 Sep 2004 22:45:22 +0000 (18:45 -0400)]
[libata] add hook, and export functions needed for sata2 drivers
* add dev_select hook, and default/noop implementations
* export ata_dev_classify
* fix a couple bugs that cropped up when building with
ATA_VERBOSE_DEBUG
* export __sata_phy_reset, a variant that does not call
ata_bus_reset
[IPV6]: Missing xfrm_lookup() in icmpv6_{send,echo_reply}()
net/ipv6/icmp.c was not converted in xfrm_lookup() extraction patch.
This patch converts it; adding the missing call to xfrm_lookup in
icmpv6_{send,echo_reply}().
Signed-off-by: Kazunori Miyazawa <kazunori@miyazawa.org> Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> SIgned-off-by: David S. Miller <davem@davemloft.net>
David S. Miller [Tue, 14 Sep 2004 15:03:10 +0000 (08:03 -0700)]
[IPV4]: Make fib_semantics algorithms scale better.
A singly linked list was previously used to
do fib_info object lookup for various actions
in the routing code. This does not scale very
well with many devices and even a moderate number
of routes. This was noted by Benjamin Lahaise.
To fix all of this we use 3 hash tables, two of
which grow dynamically as the number fib_info
objects increases while the final one is fixes in
size.
The statically sized table hashes on device index.
This is used for fib_sync_down, fib_sync_up, and
ip_fib_check_default.
The first dynamically sized table is keyed on
protocol, prefsrc, and priority. This is used
by fib_create_info() to look for existing fib_info
objects matching the new one being constructed.
The last dynamically sized table is keyed on
the preferred source of the route if it has one
specified. This is used by fib_sync_down when
a local address disappears.
There are still some scalability problems for
Bens test case in fib_hash.c and I will try to
attack those next.
Signed-off-by: David S. Miller <davem@davemloft.net>
Noticed by BenH, happily harmless (nothing that uses that
code has been committed yet, and PIO seems to be pretty much
unused on at least the Apple G5 machines: all the normal
hardware is set up purely for MMIO, to the point that I
couldn't even test this thing).
Nicolas Pitre [Tue, 14 Sep 2004 10:54:04 +0000 (11:54 +0100)]
[ARM PATCH] 2094/1: don't lose the system timer after resuming from sleep on SA11x0 and
PXA2xx
Patch from Nicolas Pitre
Let's make sure OSCR doesn't end up to be restored with a value
past OSMR0 otherwise the system timer won't start ticking until
OSCR wraps around (aprox 17 min.
Also set OSCR _after_ OIER is restored to avoid matching when
corresponding match interrupt is masked out.
The final ia64 related cleanup to elf_read_implies_exec() seems to have
broken it. The effect is that the READ_IMPLIES_EXEC flag is never set
for !pt_gnu_stack binaries!
That's a bit more secure than we need to be, and might break some legacy
app that doesn't expect it.
Fix ABI in set_mempolicy() that got broken by an earlier change.
Add a check for very big input values and prevent excessive looping in the
kernel.
Cc: Paul "nyer, nyer, your mother wears combat boots!" Jackson <pj@sgi.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
As it's using the obsolete MOD_{INC,DEC}_USE_COUNT it's implicitly locked
already, but let's remove them and make it explicit so these macros can go
away completely without breaking m68k compile.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Tue, 14 Sep 2004 00:51:24 +0000 (17:51 -0700)]
[PATCH] BSD disklabel: handle more than 8 partitions
NetBSD allows 16 partitions, not just 8. This patch both ups the number,
and makes the recognition code tell you if the count in the disklabel
exceeds the number supported by the kernel.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
- I misspelled CONFIG_PREEMPT CONFIG_PREEPT as various people noticed.
But in fact that ifdef should just go, else we'll get drivers that
compile with CONFIG_PREEMPT but not without sooner or later.
- remove unused hardirq_trylock and hardirq_endlock
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] Add support for word-length UART registers
UARTS on several Intel IXP2000 systems are connected in such a way that
they can only be addressed using full word accesses instead of bytes.
Following patch adds a UPIO_MEM32 io-type to identify these UARTs.
Drop a config option which has disappeared from all archs. Btw, this
shouldn't be in the UML-specific part, but since we cannot include generic
Kconfigs to avoid problem with hardware-related configs, it's duplicated
for now.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: refer to CONFIG_USERMODE, not to CONFIG_UM
Correct one Kconfig dependency, which should refer to CONFIG_USERMODE
rather than to CONFIG_UM.
We should also figure out how to make the config process work better for
UML. We would like to make UML able to "source drivers/Kconfig" and have
the right drivers selectable (i.e. LVM, ramdisk, and so on) and the ones
for actual hardware excluded. I've been reading such a request even from
Jeff Dike at the last Kernel Summit, (in the lwn.net coverage) but without
any followup.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 14 Sep 2004 00:49:28 +0000 (17:49 -0700)]
[PATCH] uml: disable pending signals across a reboot
On reboot, all signals and signal sources are disabled so that
late-arriving signals don't show up after the reboot exec, confusing the
new image, which is not expecting signals yet.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 14 Sep 2004 00:49:04 +0000 (17:49 -0700)]
[PATCH] uml: fix scheduler race
This fixes a use-after-free bug in the context switching. A process going
out of context after exiting wakes up the next process and then kills
itself. The problem is that when it gets around to killing itself is up to
the host and can happen a long time later, including after the incoming
process has freed its stack, and that memory is possibly being used for
something else.
The fix is to have the incoming process kill the exiting process just to
make sure it can't be running at the point that its stack is freed.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ryan S. Arnold [Tue, 14 Sep 2004 00:48:26 +0000 (17:48 -0700)]
[PATCH] HVCS fix to replace yield with tty_wait_until_sent in hvcs_close
Following the same advice you gave in a recent hvc_console patch I have
modified HVCS to remove a while() { yield(); } from hvcs_close() which may
cause problems where real time scheduling is concerned and replaced it with
tty_wait_until_sent() which uses a real wait queue and is the proper method
for blocking a tty operation while waiting for data to be sent. This patch
has been tested to verify that all the paths of code that were changed were
hit during the code run and performed as expected including hotplug remove
of hvcs adapters and hangup of ttys.
- Replaced yield() in hvcs_close() with tty_wait_until_sent() to prevent
possible lockup with realtime scheduling.
- Removed hvcs_final_close() and reordered cleanup operations to prevent
discarding of pending data during an hvcs_close() call.
- Removed spinlock protection of hvcs_struct data members in
hvcs_write_room() and hvcs_chars_in_buffer() because they aren't needed.
Signed-off-by: Ryan S. Arnold <rsa@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
max_hw_sectors_kb is the maximum that the driver can handle and is
readonly. max_sectors_kb is the current max_sectors value and can be tuned
by root. PAGE_SIZE granularity is enforced.
It's all locking-safe and all affected layered drivers have been updated as
well. The patch has been in testing for a couple of weeks already as part
of the voluntary-preempt patches and it works just fine - people use it to
reduce IDE IRQ handling latencies.
This patch adds a prctl to modify current->comm as shown in /proc. This
feature was requested by KDE developers. In KDE most programs are started by
forking from a kdeinit program that already has the libraries loaded and some
other state.
Problem is to give these forked programs the proper name. It already writes
the command line in the environment (as seen in ps), but top uses a different
field in /proc/pid/status that reports current->comm. And that was always
"kdeinit" instead of the real command name. So you ended up with lots of
kdeinits in your top listing, which was not very useful.
This patch adds a new prctl PR_SET_NAME to allow a program to change its comm
field.
I considered the potential security issues of a program obscuring itself with
this interface, but I don't think it matters much because a program can
already obscure itself when the admin uses ps instead of top. In case of a
KDE desktop calling everything kdeinit is much more obfuscation than the
alternative.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] __copy_to_user() check in cdrom_read_cdda_old()
akpm: really, reads are supposed to return the number-of-bytes-read on faults,
or -EFAULT of no bytes were read. This patch returns either zero or -EFAULT,
ignoring any successfully transferred data. But the user interface (whcih is
an ioctl()) was never set up to do that.
I _think_ shmem_file_setup is protected against negative loff_t size by the
TASK_SIZE in each arch, but prefer the security of an explicit test. Wipe
those parentheses off its return(file), and update our Copyright.
Very minor adjustments to shmem_getpage return path: I now prefer it to return
NULL and let do_shmem_file_read use ZERO_PAGE(0) in that case; and we don't
need a local majmin variable, do_no_page initializes *type to VM_FAULT_MINOR
already.
If we're thinking about shmem scalability... isn't it silly that each shmem
object is added to the shmem_inodes list on creation, and removed on deletion,
yet the only use for that list is in swapoff (shmem_unuse)?
Call it shmem_swaplist; shmem_writepage add inode to swaplist when first swap
allocated (usually never); shmem_delete_inode remove inode from the list after
truncating (if called before, inode could be re-added to it).
Inode can remain on the swaplist after all its pages are swapped back in, just
be lazy about it; but if shmem_unuse finds swapped count now 0, save itself
time by then removing that inode from the swaplist.
Some might want a tmpfs mount with the improved scalability afforded by
omitting shmem superblock accounting; or some might just want to test it in an
externally-visible tmpfs mount instance.
Adopt the convention that mount option -o nr_blocks=0,nr_inodes=0 means
without resource limits, and hence no shmem_sb_info. Not recommended for
general use, but no worse than ramfs.
Disallow remounting from unlimited to limited (no accounting has been done so
far, so no idea whether it's permissible), and from limited to unlimited
(because we'd need then to free the sbinfo, and visit each inode to reset its
i_blocks to 0: why bother?).
SGI investigations have shown a dramatic contrast in scalability between
anonymous memory and shmem objects. Processes building distinct shmem objects
in parallel hit heavy contention on shmem superblock stat_lock. Across 256
cpus an intensive test runs 300 times slower than anonymous.
Jack Steiner has observed that all the shmem superblock free_blocks and
free_inodes accounting is redundant in the case of the internal mount used for
SysV shared memory and for shared writable /dev/zero objects (the cases which
most concern them): it specifically declines to limit.
Based upon Brent Casavant's SHMEM_NOSBINFO patch, this instead just removes
the shmem_sb_info structure from the internal kernel mount, testing where
necessary for null sbinfo pointer. shmem_set_size moved within CONFIG_TMPFS,
its arg named "sbinfo" as elsewhere.
This brings shmem object scalability up to that of anonymous memory, in the
case where distinct processes are building (faulting to allocate) distinct
objects. It significantly improves parallel building of a shared shmem object
(that test runs 14 times faster across 256 cpus), but other issues remain in
that case: to be addressed in later patches.
Keith Mannthey's Bugzilla #3268 drew attention to how tmpfs inodes and
dentries and long names and radix-tree nodes pin lowmem. Assuming about 1k of
lowmem per inode, we need to lower the default nr_inodes limit on machines
with significant highmem.
Be conservative, but more generous than in the original patch to Keith: limit
to number of lowmem pages, which works out around 200,000 on i386. Easily
overridden by giving the nr_inodes= mount option: those who want to sail
closer to the rocks should be allowed to do so.
Notice how tmpfs dentries cannot be reclaimed in the way that disk-based
dentries can: so even hard links need to be costed. They are cheaper than
inodes, but easier all round to charge the same. This way, the limit for hard
links is equally visible through "df -i": but expect occasional bugreports
that tmpfs links are being treated like this.
Would have been simpler just to move the free_inodes accounting from
shmem_delete_inode to shmem_unlink; but that would lose the charge on unlinked
but open files.
this patch fix a pnpbios problem with independant
resource(http://bugzilla.kernel.org/show_bug.cgi?id=3295) :
the old code assume that they are given at the beggining (before any
SMALL_TAG_STARTDEP entry), but in some case there are found after
SMALL_TAG_ENDDEP entry.
tag : 6 SMALL_TAG_STARTDEP
tag : 8 SMALL_TAG_PORT
tag : 6 SMALL_TAG_STARTDEP
tag : 8 SMALL_TAG_PORT
tag : 7 SMALL_TAG_ENDDEP
tag : 4 SMALL_TAG_IRQ <-- independant resource
tag : f SMALL_TAG_END
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Tue, 14 Sep 2004 00:44:21 +0000 (17:44 -0700)]
[PATCH] ppc64: improved VSID allocation algorithm
This patch has been tested both on SLB and segment table machines. This
new approach is far from the final word in VSID/context allocation, but
it's a noticeable improvement on the old method.
Replace the VSID allocation algorithm. The new algorithm first generates a
36-bit "proto-VSID" (with 0xfffffffff reserved). For kernel addresses this
is equal to the ESID (address >> 28), for user addresses it is:
(context << 15) | (esid & 0x7fff)
These are distinguishable from kernel proto-VSIDs because the top bit is
clear. Proto-VSIDs with the top two bits equal to 0b10 are reserved for
now.
The proto-VSIDs are then scrambled into real VSIDs with the multiplicative
hash:
This scramble is 1:1, because VSID_MULTIPLIER and VSID_MODULUS are co-prime
since VSID_MULTIPLIER is prime (the largest 28-bit prime, in fact).
This scheme has a number of advantages over the old one:
- We now have VSIDs for every kernel address (i.e. everything above
0xC000000000000000), except the very top segment. That simplifies a
number of things.
- We allow for 15 significant bits of ESID for user addresses with 20
bits of context. i.e. 8T (43 bits) of address space for up to 1M
contexts, significantly more than the old method (although we will need
changes in the hash path and context allocation to take advantage of
this).
- Because we use a real multiplicative hash function, we have better and
more robust hash scattering with this VSID algorithm (at least based on
some initial results).
Because the MODULUS is 2^n-1 we can use a trick to compute it efficiently
without a divide or extra multiply. This makes the new algorithm barely
slower than the old one.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>