As it's using the obsolete MOD_{INC,DEC}_USE_COUNT it's implicitly locked
already, but let's remove them and make it explicit so these macros can go
away completely without breaking m68k compile.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Tue, 14 Sep 2004 00:51:24 +0000 (17:51 -0700)]
[PATCH] BSD disklabel: handle more than 8 partitions
NetBSD allows 16 partitions, not just 8. This patch both ups the number,
and makes the recognition code tell you if the count in the disklabel
exceeds the number supported by the kernel.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
- I misspelled CONFIG_PREEMPT CONFIG_PREEPT as various people noticed.
But in fact that ifdef should just go, else we'll get drivers that
compile with CONFIG_PREEMPT but not without sooner or later.
- remove unused hardirq_trylock and hardirq_endlock
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] Add support for word-length UART registers
UARTS on several Intel IXP2000 systems are connected in such a way that
they can only be addressed using full word accesses instead of bytes.
Following patch adds a UPIO_MEM32 io-type to identify these UARTs.
Drop a config option which has disappeared from all archs. Btw, this
shouldn't be in the UML-specific part, but since we cannot include generic
Kconfigs to avoid problem with hardware-related configs, it's duplicated
for now.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] uml: refer to CONFIG_USERMODE, not to CONFIG_UM
Correct one Kconfig dependency, which should refer to CONFIG_USERMODE
rather than to CONFIG_UM.
We should also figure out how to make the config process work better for
UML. We would like to make UML able to "source drivers/Kconfig" and have
the right drivers selectable (i.e. LVM, ramdisk, and so on) and the ones
for actual hardware excluded. I've been reading such a request even from
Jeff Dike at the last Kernel Summit, (in the lwn.net coverage) but without
any followup.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 14 Sep 2004 00:49:28 +0000 (17:49 -0700)]
[PATCH] uml: disable pending signals across a reboot
On reboot, all signals and signal sources are disabled so that
late-arriving signals don't show up after the reboot exec, confusing the
new image, which is not expecting signals yet.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Dike [Tue, 14 Sep 2004 00:49:04 +0000 (17:49 -0700)]
[PATCH] uml: fix scheduler race
This fixes a use-after-free bug in the context switching. A process going
out of context after exiting wakes up the next process and then kills
itself. The problem is that when it gets around to killing itself is up to
the host and can happen a long time later, including after the incoming
process has freed its stack, and that memory is possibly being used for
something else.
The fix is to have the incoming process kill the exiting process just to
make sure it can't be running at the point that its stack is freed.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ryan S. Arnold [Tue, 14 Sep 2004 00:48:26 +0000 (17:48 -0700)]
[PATCH] HVCS fix to replace yield with tty_wait_until_sent in hvcs_close
Following the same advice you gave in a recent hvc_console patch I have
modified HVCS to remove a while() { yield(); } from hvcs_close() which may
cause problems where real time scheduling is concerned and replaced it with
tty_wait_until_sent() which uses a real wait queue and is the proper method
for blocking a tty operation while waiting for data to be sent. This patch
has been tested to verify that all the paths of code that were changed were
hit during the code run and performed as expected including hotplug remove
of hvcs adapters and hangup of ttys.
- Replaced yield() in hvcs_close() with tty_wait_until_sent() to prevent
possible lockup with realtime scheduling.
- Removed hvcs_final_close() and reordered cleanup operations to prevent
discarding of pending data during an hvcs_close() call.
- Removed spinlock protection of hvcs_struct data members in
hvcs_write_room() and hvcs_chars_in_buffer() because they aren't needed.
Signed-off-by: Ryan S. Arnold <rsa@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
max_hw_sectors_kb is the maximum that the driver can handle and is
readonly. max_sectors_kb is the current max_sectors value and can be tuned
by root. PAGE_SIZE granularity is enforced.
It's all locking-safe and all affected layered drivers have been updated as
well. The patch has been in testing for a couple of weeks already as part
of the voluntary-preempt patches and it works just fine - people use it to
reduce IDE IRQ handling latencies.
This patch adds a prctl to modify current->comm as shown in /proc. This
feature was requested by KDE developers. In KDE most programs are started by
forking from a kdeinit program that already has the libraries loaded and some
other state.
Problem is to give these forked programs the proper name. It already writes
the command line in the environment (as seen in ps), but top uses a different
field in /proc/pid/status that reports current->comm. And that was always
"kdeinit" instead of the real command name. So you ended up with lots of
kdeinits in your top listing, which was not very useful.
This patch adds a new prctl PR_SET_NAME to allow a program to change its comm
field.
I considered the potential security issues of a program obscuring itself with
this interface, but I don't think it matters much because a program can
already obscure itself when the admin uses ps instead of top. In case of a
KDE desktop calling everything kdeinit is much more obfuscation than the
alternative.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] __copy_to_user() check in cdrom_read_cdda_old()
akpm: really, reads are supposed to return the number-of-bytes-read on faults,
or -EFAULT of no bytes were read. This patch returns either zero or -EFAULT,
ignoring any successfully transferred data. But the user interface (whcih is
an ioctl()) was never set up to do that.
I _think_ shmem_file_setup is protected against negative loff_t size by the
TASK_SIZE in each arch, but prefer the security of an explicit test. Wipe
those parentheses off its return(file), and update our Copyright.
Very minor adjustments to shmem_getpage return path: I now prefer it to return
NULL and let do_shmem_file_read use ZERO_PAGE(0) in that case; and we don't
need a local majmin variable, do_no_page initializes *type to VM_FAULT_MINOR
already.
If we're thinking about shmem scalability... isn't it silly that each shmem
object is added to the shmem_inodes list on creation, and removed on deletion,
yet the only use for that list is in swapoff (shmem_unuse)?
Call it shmem_swaplist; shmem_writepage add inode to swaplist when first swap
allocated (usually never); shmem_delete_inode remove inode from the list after
truncating (if called before, inode could be re-added to it).
Inode can remain on the swaplist after all its pages are swapped back in, just
be lazy about it; but if shmem_unuse finds swapped count now 0, save itself
time by then removing that inode from the swaplist.
Some might want a tmpfs mount with the improved scalability afforded by
omitting shmem superblock accounting; or some might just want to test it in an
externally-visible tmpfs mount instance.
Adopt the convention that mount option -o nr_blocks=0,nr_inodes=0 means
without resource limits, and hence no shmem_sb_info. Not recommended for
general use, but no worse than ramfs.
Disallow remounting from unlimited to limited (no accounting has been done so
far, so no idea whether it's permissible), and from limited to unlimited
(because we'd need then to free the sbinfo, and visit each inode to reset its
i_blocks to 0: why bother?).
SGI investigations have shown a dramatic contrast in scalability between
anonymous memory and shmem objects. Processes building distinct shmem objects
in parallel hit heavy contention on shmem superblock stat_lock. Across 256
cpus an intensive test runs 300 times slower than anonymous.
Jack Steiner has observed that all the shmem superblock free_blocks and
free_inodes accounting is redundant in the case of the internal mount used for
SysV shared memory and for shared writable /dev/zero objects (the cases which
most concern them): it specifically declines to limit.
Based upon Brent Casavant's SHMEM_NOSBINFO patch, this instead just removes
the shmem_sb_info structure from the internal kernel mount, testing where
necessary for null sbinfo pointer. shmem_set_size moved within CONFIG_TMPFS,
its arg named "sbinfo" as elsewhere.
This brings shmem object scalability up to that of anonymous memory, in the
case where distinct processes are building (faulting to allocate) distinct
objects. It significantly improves parallel building of a shared shmem object
(that test runs 14 times faster across 256 cpus), but other issues remain in
that case: to be addressed in later patches.
Keith Mannthey's Bugzilla #3268 drew attention to how tmpfs inodes and
dentries and long names and radix-tree nodes pin lowmem. Assuming about 1k of
lowmem per inode, we need to lower the default nr_inodes limit on machines
with significant highmem.
Be conservative, but more generous than in the original patch to Keith: limit
to number of lowmem pages, which works out around 200,000 on i386. Easily
overridden by giving the nr_inodes= mount option: those who want to sail
closer to the rocks should be allowed to do so.
Notice how tmpfs dentries cannot be reclaimed in the way that disk-based
dentries can: so even hard links need to be costed. They are cheaper than
inodes, but easier all round to charge the same. This way, the limit for hard
links is equally visible through "df -i": but expect occasional bugreports
that tmpfs links are being treated like this.
Would have been simpler just to move the free_inodes accounting from
shmem_delete_inode to shmem_unlink; but that would lose the charge on unlinked
but open files.
this patch fix a pnpbios problem with independant
resource(http://bugzilla.kernel.org/show_bug.cgi?id=3295) :
the old code assume that they are given at the beggining (before any
SMALL_TAG_STARTDEP entry), but in some case there are found after
SMALL_TAG_ENDDEP entry.
tag : 6 SMALL_TAG_STARTDEP
tag : 8 SMALL_TAG_PORT
tag : 6 SMALL_TAG_STARTDEP
tag : 8 SMALL_TAG_PORT
tag : 7 SMALL_TAG_ENDDEP
tag : 4 SMALL_TAG_IRQ <-- independant resource
tag : f SMALL_TAG_END
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Tue, 14 Sep 2004 00:44:21 +0000 (17:44 -0700)]
[PATCH] ppc64: improved VSID allocation algorithm
This patch has been tested both on SLB and segment table machines. This
new approach is far from the final word in VSID/context allocation, but
it's a noticeable improvement on the old method.
Replace the VSID allocation algorithm. The new algorithm first generates a
36-bit "proto-VSID" (with 0xfffffffff reserved). For kernel addresses this
is equal to the ESID (address >> 28), for user addresses it is:
(context << 15) | (esid & 0x7fff)
These are distinguishable from kernel proto-VSIDs because the top bit is
clear. Proto-VSIDs with the top two bits equal to 0b10 are reserved for
now.
The proto-VSIDs are then scrambled into real VSIDs with the multiplicative
hash:
This scramble is 1:1, because VSID_MULTIPLIER and VSID_MODULUS are co-prime
since VSID_MULTIPLIER is prime (the largest 28-bit prime, in fact).
This scheme has a number of advantages over the old one:
- We now have VSIDs for every kernel address (i.e. everything above
0xC000000000000000), except the very top segment. That simplifies a
number of things.
- We allow for 15 significant bits of ESID for user addresses with 20
bits of context. i.e. 8T (43 bits) of address space for up to 1M
contexts, significantly more than the old method (although we will need
changes in the hash path and context allocation to take advantage of
this).
- Because we use a real multiplicative hash function, we have better and
more robust hash scattering with this VSID algorithm (at least based on
some initial results).
Because the MODULUS is 2^n-1 we can use a trick to compute it efficiently
without a divide or extra multiply. This makes the new algorithm barely
slower than the old one.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:43:18 +0000 (17:43 -0700)]
[PATCH] ppc64: clean up idle loop code
Clean up our idle loop code:
- Remove a bunch of useless includes and make most functions static
- There were places where we werent disabling interrupts before checking
need_resched then calling the hypervisor to sleep our thread. We might
race with an IPI and end up missing a reschedule. Disable interrupts
around these regions to make them safe.
- We forgot to turn off the polling flag when exiting the dedicated_idle
idle loop. This could have resulted in all manner problems as other
cpus would avoid sending IPIs to force reschedules.
- Add a missing check for cpu_is_offline in the shared cpu idle loop.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:43:06 +0000 (17:43 -0700)]
[PATCH] ppc64: enable POWER5 low power mode in idle loop
Now that we understand (and have fixed) the problem with using low power mode
in the idle loop, lets enable it. It should save a fair amount of power.
(The problem was that our exceptions were inheriting the low power mode and so
were executing at a fraction of the normal cpu issue rate. We fixed it by
always bumping our priority to medium at the start of every exception).
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:54 +0000 (17:42 -0700)]
[PATCH] ppc64: restore smt-enabled=off kernel command line option
Restore the smt-enabled=off kernel command line functionality:
- Remove the SMT_DYNAMIC state now that smt_snooze_delay allows for the
same thing.
- Remove the early prom.c parsing for the option, put it into an
early_param instead.
- In setup_cpu_maps honour the smt-enabled setting
Note to Nathan: In order to allow cpu hotplug add of secondary threads after
booting with smt-enabled=off, I had to initialise cpu_present_map to
cpu_online_map in smp_cpus_done. Im not sure how you want to handle this but
it seems our present map currently does not allow cpus to be added into the
partition that werent there at boot (but were in the possible map).
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:28 +0000 (17:42 -0700)]
[PATCH] ppc64: remove EEH command line device matching code
We have had reports of people attempting to disable EEH on POWER5 boxes. This
is not supported, and the device will most likely not respond to config space
reads/writes. Remove the IBM location matching code that was being used to
disable devices as well as the global option.
We already have the ability to ignore EEH erros via the panic_on_oops sysctl
option, advanced users should make use of that instead.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:04 +0000 (17:42 -0700)]
[PATCH] ppc64: clean up kernel command line code
Clean up some of our command line code:
- We were copying the command line out of the device tree twice, but the
first time we forgot to add CONFIG_CMDLINE. Fix this and remove the
second copy.
- The command line birec code ran after we had done some command line
parsing in prom.c. This had the opportunity to really confuse the
user, with some options being parsed out of the device tree and the
other out of birecs. Luckily we could find no user of the command
line birecs, so remove them.
- remove duplicate printing of kernel command line;
- clean up iseries inits and create an iSeries_parse_cmdline.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:41:51 +0000 (17:41 -0700)]
[PATCH] ppc64: use nm --synthetic where available
On new toolchains we need to use nm --synthetic or we miss code symbols. Sam,
I'm not thrilled about this patch but Im not sure of an easier way. Any ideas?
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Will Schmidt [Tue, 14 Sep 2004 00:40:39 +0000 (17:40 -0700)]
[PATCH] ppc64: lparcfg fixes for processor counts
This patch corrects how the lparcfg interface was presenting the number of
active and potential processors. (As reported in LTC bugzilla number 10889).
- Correct output for partition_potential_processors and
system_active_processors.
- suppress pool related values in scenarios where they do not make
sense. (non-shared processor configurations)
- Display pool_capacity as a percentage, to match the behavior from
iSeries code.
Signed-off-by: Will Schmidt <willschm@us.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jason Davis [Tue, 14 Sep 2004 00:40:15 +0000 (17:40 -0700)]
[PATCH] ES7000 subarch update
The patch below implements an algorithm to determine an unique GSI override
for mapping GSIs to IO-APIC pins correctly. GSI overrides are required in
order for ES7000 machines to function properly since IRQ to pin mappings
are NOT all one-to-one. This patch applies only to the Unisys specific
ES7000 machines and has been tested thoroughly on several models of the
ES7000 line.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
In sched_exec, schedstat_inc will dereference a null pointer if no domain
is found with the SD_BALANCE_EXEC flag set. This was exposed during
testing of the previous patches where cpus are temporarily attached to a
dummy domain without SD_BALANCE_EXEC set.
Russell King [Tue, 14 Sep 2004 00:13:03 +0000 (01:13 +0100)]
[ARM] Convert suspend to a state machine.
The original version had issues when two suspend events came
in at around the same time, causing APM to get confused:
threads became stuck in APM_IOC_SUSPEND and suspends_pending
incremented on each apm --suspend call.
Now, we only add a suspend event to a users queue and increment
suspends_pending if the user isn't already in the middle of
handling a suspend event.
Nicolas Pitre [Mon, 13 Sep 2004 08:11:27 +0000 (01:11 -0700)]
[PATCH] linux/dma-mapping.h needs linux/device.h
It seems that most architectures already include linux/device.h in their
own asm/dma-mapping.h. Most but not all, and some drivers fail to
compile on those architectures that don't. Since everybody needs it
let's include device.h from one place only and fix compilation for
everybody.
Anton Blanchard [Mon, 13 Sep 2004 07:05:30 +0000 (00:05 -0700)]
[PATCH] Backward compatibility for compat sched_getaffinity
The follow patch special cases the NR_CPUS <= BITS_PER_COMPAT_LONG case.
Without this patch, a 32bit task would be required to have a 64bit
cpumask no matter what value of NR_CPUS are used.
With this patch a compat long sized bitmask is allowed if NR_CPUS is
small enough to fit within it.
Of course applications should be using the glibc wrappers that use an
opaque cpu_mask_t type, but there could be older applications using the
syscalls directly.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Mon, 13 Sep 2004 07:05:18 +0000 (00:05 -0700)]
[PATCH] Clean up compat sched affinity syscalls
Remove the set_fs hack in the compat affinity calls. Create
sched_getaffinity and sched_setaffinity helper functions that both the
native and compat affinity syscalls use.
Also make the compat functions match what the native ones are doing now,
setaffinity calls succeed no matter what length the bitmask is, but
getaffinity calls must pass in bitmasks at least as long as the kernel
type.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[IA64] Makefile: fix for the PTRACE_SYSCALL corruption bug
Thanks to David for his help in tracking it down.
compile the kernel with sibling call optimization
turned off. There is a problem with all functions
using the optimization and the asmlinkage attribute.
The compiler should not perform the optimization on
these functions because it cannot preserve the syscall
parameters in the callee. This caused SIGSEGV on programs
traced with PTRACE_SYSCALL, for instance.
signed-off-by: stephane eranian <eranian@hpl.hp.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
This patch kills the bogus radeonfb_read/write routines. In order to do so,
it adds a new member to fb_info, along with screen_base, which is screen_size,
indicating the mapped area. The default fb_read/write will now use that instead
of fix->smem_len if it is non-0, and radeonfb now sets it to the mapped size
of the framebuffer.
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] ppc64:Fix missing register in altivec context switch
This is a resend of a patch sent in July and that got lost somewhat,
the "VSCR" register wasn't restored properly from the context on
load_up_altivec (typo), please apply the fix:
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Rusty Russell [Sun, 12 Sep 2004 10:00:47 +0000 (03:00 -0700)]
[NETFILTER]: Fix conntrack seq_file handling.
Am travelling, but this passed simple tests here. If this isn't going
in, the current seqfile stuff should be ripped out; it's a mess.
/proc/net/ip_conntrack was changed over to seq_file. However,
seq_file isn't a great fit (a linked list which is changing is not a
good candidate for seq file), and the conversion was done badly.
1) Don't do allocation: simply hand the pointer head of the correct chain.
2) Actually output the original tuple.
3) Lock only when actually traversing hash chain.
Signed-off-by: Rusty Russell <rusty@rustcorp.com.au> Signed-off-by: David S. Miller <davem@davemloft.net>