Some might want a tmpfs mount with the improved scalability afforded by
omitting shmem superblock accounting; or some might just want to test it in an
externally-visible tmpfs mount instance.
Adopt the convention that mount option -o nr_blocks=0,nr_inodes=0 means
without resource limits, and hence no shmem_sb_info. Not recommended for
general use, but no worse than ramfs.
Disallow remounting from unlimited to limited (no accounting has been done so
far, so no idea whether it's permissible), and from limited to unlimited
(because we'd need then to free the sbinfo, and visit each inode to reset its
i_blocks to 0: why bother?).
SGI investigations have shown a dramatic contrast in scalability between
anonymous memory and shmem objects. Processes building distinct shmem objects
in parallel hit heavy contention on shmem superblock stat_lock. Across 256
cpus an intensive test runs 300 times slower than anonymous.
Jack Steiner has observed that all the shmem superblock free_blocks and
free_inodes accounting is redundant in the case of the internal mount used for
SysV shared memory and for shared writable /dev/zero objects (the cases which
most concern them): it specifically declines to limit.
Based upon Brent Casavant's SHMEM_NOSBINFO patch, this instead just removes
the shmem_sb_info structure from the internal kernel mount, testing where
necessary for null sbinfo pointer. shmem_set_size moved within CONFIG_TMPFS,
its arg named "sbinfo" as elsewhere.
This brings shmem object scalability up to that of anonymous memory, in the
case where distinct processes are building (faulting to allocate) distinct
objects. It significantly improves parallel building of a shared shmem object
(that test runs 14 times faster across 256 cpus), but other issues remain in
that case: to be addressed in later patches.
Keith Mannthey's Bugzilla #3268 drew attention to how tmpfs inodes and
dentries and long names and radix-tree nodes pin lowmem. Assuming about 1k of
lowmem per inode, we need to lower the default nr_inodes limit on machines
with significant highmem.
Be conservative, but more generous than in the original patch to Keith: limit
to number of lowmem pages, which works out around 200,000 on i386. Easily
overridden by giving the nr_inodes= mount option: those who want to sail
closer to the rocks should be allowed to do so.
Notice how tmpfs dentries cannot be reclaimed in the way that disk-based
dentries can: so even hard links need to be costed. They are cheaper than
inodes, but easier all round to charge the same. This way, the limit for hard
links is equally visible through "df -i": but expect occasional bugreports
that tmpfs links are being treated like this.
Would have been simpler just to move the free_inodes accounting from
shmem_delete_inode to shmem_unlink; but that would lose the charge on unlinked
but open files.
this patch fix a pnpbios problem with independant
resource(http://bugzilla.kernel.org/show_bug.cgi?id=3295) :
the old code assume that they are given at the beggining (before any
SMALL_TAG_STARTDEP entry), but in some case there are found after
SMALL_TAG_ENDDEP entry.
tag : 6 SMALL_TAG_STARTDEP
tag : 8 SMALL_TAG_PORT
tag : 6 SMALL_TAG_STARTDEP
tag : 8 SMALL_TAG_PORT
tag : 7 SMALL_TAG_ENDDEP
tag : 4 SMALL_TAG_IRQ <-- independant resource
tag : f SMALL_TAG_END
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Tue, 14 Sep 2004 00:44:21 +0000 (17:44 -0700)]
[PATCH] ppc64: improved VSID allocation algorithm
This patch has been tested both on SLB and segment table machines. This
new approach is far from the final word in VSID/context allocation, but
it's a noticeable improvement on the old method.
Replace the VSID allocation algorithm. The new algorithm first generates a
36-bit "proto-VSID" (with 0xfffffffff reserved). For kernel addresses this
is equal to the ESID (address >> 28), for user addresses it is:
(context << 15) | (esid & 0x7fff)
These are distinguishable from kernel proto-VSIDs because the top bit is
clear. Proto-VSIDs with the top two bits equal to 0b10 are reserved for
now.
The proto-VSIDs are then scrambled into real VSIDs with the multiplicative
hash:
This scramble is 1:1, because VSID_MULTIPLIER and VSID_MODULUS are co-prime
since VSID_MULTIPLIER is prime (the largest 28-bit prime, in fact).
This scheme has a number of advantages over the old one:
- We now have VSIDs for every kernel address (i.e. everything above
0xC000000000000000), except the very top segment. That simplifies a
number of things.
- We allow for 15 significant bits of ESID for user addresses with 20
bits of context. i.e. 8T (43 bits) of address space for up to 1M
contexts, significantly more than the old method (although we will need
changes in the hash path and context allocation to take advantage of
this).
- Because we use a real multiplicative hash function, we have better and
more robust hash scattering with this VSID algorithm (at least based on
some initial results).
Because the MODULUS is 2^n-1 we can use a trick to compute it efficiently
without a divide or extra multiply. This makes the new algorithm barely
slower than the old one.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:43:18 +0000 (17:43 -0700)]
[PATCH] ppc64: clean up idle loop code
Clean up our idle loop code:
- Remove a bunch of useless includes and make most functions static
- There were places where we werent disabling interrupts before checking
need_resched then calling the hypervisor to sleep our thread. We might
race with an IPI and end up missing a reschedule. Disable interrupts
around these regions to make them safe.
- We forgot to turn off the polling flag when exiting the dedicated_idle
idle loop. This could have resulted in all manner problems as other
cpus would avoid sending IPIs to force reschedules.
- Add a missing check for cpu_is_offline in the shared cpu idle loop.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:43:06 +0000 (17:43 -0700)]
[PATCH] ppc64: enable POWER5 low power mode in idle loop
Now that we understand (and have fixed) the problem with using low power mode
in the idle loop, lets enable it. It should save a fair amount of power.
(The problem was that our exceptions were inheriting the low power mode and so
were executing at a fraction of the normal cpu issue rate. We fixed it by
always bumping our priority to medium at the start of every exception).
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:54 +0000 (17:42 -0700)]
[PATCH] ppc64: restore smt-enabled=off kernel command line option
Restore the smt-enabled=off kernel command line functionality:
- Remove the SMT_DYNAMIC state now that smt_snooze_delay allows for the
same thing.
- Remove the early prom.c parsing for the option, put it into an
early_param instead.
- In setup_cpu_maps honour the smt-enabled setting
Note to Nathan: In order to allow cpu hotplug add of secondary threads after
booting with smt-enabled=off, I had to initialise cpu_present_map to
cpu_online_map in smp_cpus_done. Im not sure how you want to handle this but
it seems our present map currently does not allow cpus to be added into the
partition that werent there at boot (but were in the possible map).
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:28 +0000 (17:42 -0700)]
[PATCH] ppc64: remove EEH command line device matching code
We have had reports of people attempting to disable EEH on POWER5 boxes. This
is not supported, and the device will most likely not respond to config space
reads/writes. Remove the IBM location matching code that was being used to
disable devices as well as the global option.
We already have the ability to ignore EEH erros via the panic_on_oops sysctl
option, advanced users should make use of that instead.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:42:04 +0000 (17:42 -0700)]
[PATCH] ppc64: clean up kernel command line code
Clean up some of our command line code:
- We were copying the command line out of the device tree twice, but the
first time we forgot to add CONFIG_CMDLINE. Fix this and remove the
second copy.
- The command line birec code ran after we had done some command line
parsing in prom.c. This had the opportunity to really confuse the
user, with some options being parsed out of the device tree and the
other out of birecs. Luckily we could find no user of the command
line birecs, so remove them.
- remove duplicate printing of kernel command line;
- clean up iseries inits and create an iSeries_parse_cmdline.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 14 Sep 2004 00:41:51 +0000 (17:41 -0700)]
[PATCH] ppc64: use nm --synthetic where available
On new toolchains we need to use nm --synthetic or we miss code symbols. Sam,
I'm not thrilled about this patch but Im not sure of an easier way. Any ideas?
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Will Schmidt [Tue, 14 Sep 2004 00:40:39 +0000 (17:40 -0700)]
[PATCH] ppc64: lparcfg fixes for processor counts
This patch corrects how the lparcfg interface was presenting the number of
active and potential processors. (As reported in LTC bugzilla number 10889).
- Correct output for partition_potential_processors and
system_active_processors.
- suppress pool related values in scenarios where they do not make
sense. (non-shared processor configurations)
- Display pool_capacity as a percentage, to match the behavior from
iSeries code.
Signed-off-by: Will Schmidt <willschm@us.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jason Davis [Tue, 14 Sep 2004 00:40:15 +0000 (17:40 -0700)]
[PATCH] ES7000 subarch update
The patch below implements an algorithm to determine an unique GSI override
for mapping GSIs to IO-APIC pins correctly. GSI overrides are required in
order for ES7000 machines to function properly since IRQ to pin mappings
are NOT all one-to-one. This patch applies only to the Unisys specific
ES7000 machines and has been tested thoroughly on several models of the
ES7000 line.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
In sched_exec, schedstat_inc will dereference a null pointer if no domain
is found with the SD_BALANCE_EXEC flag set. This was exposed during
testing of the previous patches where cpus are temporarily attached to a
dummy domain without SD_BALANCE_EXEC set.
Russell King [Tue, 14 Sep 2004 00:13:03 +0000 (01:13 +0100)]
[ARM] Convert suspend to a state machine.
The original version had issues when two suspend events came
in at around the same time, causing APM to get confused:
threads became stuck in APM_IOC_SUSPEND and suspends_pending
incremented on each apm --suspend call.
Now, we only add a suspend event to a users queue and increment
suspends_pending if the user isn't already in the middle of
handling a suspend event.
Nicolas Pitre [Mon, 13 Sep 2004 08:11:27 +0000 (01:11 -0700)]
[PATCH] linux/dma-mapping.h needs linux/device.h
It seems that most architectures already include linux/device.h in their
own asm/dma-mapping.h. Most but not all, and some drivers fail to
compile on those architectures that don't. Since everybody needs it
let's include device.h from one place only and fix compilation for
everybody.
Anton Blanchard [Mon, 13 Sep 2004 07:05:30 +0000 (00:05 -0700)]
[PATCH] Backward compatibility for compat sched_getaffinity
The follow patch special cases the NR_CPUS <= BITS_PER_COMPAT_LONG case.
Without this patch, a 32bit task would be required to have a 64bit
cpumask no matter what value of NR_CPUS are used.
With this patch a compat long sized bitmask is allowed if NR_CPUS is
small enough to fit within it.
Of course applications should be using the glibc wrappers that use an
opaque cpu_mask_t type, but there could be older applications using the
syscalls directly.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Mon, 13 Sep 2004 07:05:18 +0000 (00:05 -0700)]
[PATCH] Clean up compat sched affinity syscalls
Remove the set_fs hack in the compat affinity calls. Create
sched_getaffinity and sched_setaffinity helper functions that both the
native and compat affinity syscalls use.
Also make the compat functions match what the native ones are doing now,
setaffinity calls succeed no matter what length the bitmask is, but
getaffinity calls must pass in bitmasks at least as long as the kernel
type.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[IA64] Makefile: fix for the PTRACE_SYSCALL corruption bug
Thanks to David for his help in tracking it down.
compile the kernel with sibling call optimization
turned off. There is a problem with all functions
using the optimization and the asmlinkage attribute.
The compiler should not perform the optimization on
these functions because it cannot preserve the syscall
parameters in the callee. This caused SIGSEGV on programs
traced with PTRACE_SYSCALL, for instance.
signed-off-by: stephane eranian <eranian@hpl.hp.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
This patch kills the bogus radeonfb_read/write routines. In order to do so,
it adds a new member to fb_info, along with screen_base, which is screen_size,
indicating the mapped area. The default fb_read/write will now use that instead
of fix->smem_len if it is non-0, and radeonfb now sets it to the mapped size
of the framebuffer.
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] ppc64:Fix missing register in altivec context switch
This is a resend of a patch sent in July and that got lost somewhat,
the "VSCR" register wasn't restored properly from the context on
load_up_altivec (typo), please apply the fix:
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Rusty Russell [Sun, 12 Sep 2004 10:00:47 +0000 (03:00 -0700)]
[NETFILTER]: Fix conntrack seq_file handling.
Am travelling, but this passed simple tests here. If this isn't going
in, the current seqfile stuff should be ripped out; it's a mess.
/proc/net/ip_conntrack was changed over to seq_file. However,
seq_file isn't a great fit (a linked list which is changing is not a
good candidate for seq file), and the conversion was done badly.
1) Don't do allocation: simply hand the pointer head of the correct chain.
2) Actually output the original tuple.
3) Lock only when actually traversing hash chain.
Signed-off-by: Rusty Russell <rusty@rustcorp.com.au> Signed-off-by: David S. Miller <davem@davemloft.net>
[IPVS]: Do not use skb_checksum_help(), create and use nf_reset_debug()
Appended is a 2nd version that uses nf_reset_debug.
- do not use skb_checksum_help in input path as ipvs can handle
incoming CHECKSUM_HW packets
- do not use skb_checksum_help in forwarding path
- claim that checksum is valid (CHECKSUM_NONE) when entering output
path for out->in packets
- do not reset/destroy the nfct in IP_VS_XMIT, the intention is to
reset the debugging field just to avoid log floods from nf_debug_ip_*
functions, it is known that the ipvs packets traverse other
hooks, eg. LOCAL_IN->LOCAL_OUT. Use nf_reset_debug instead of nf_reset.
Signed-off-by: Julian Anastasov <ja@ssi.bg> Signed-off-by: David S. Miller <davem@davemloft.net>
Well, rt->rt6i_idev is always set if it is dynamically allocated.
However, when we hit ip6_null_entry here, its rt6i_idev is NULL.
This patch is minimum fix to avoid the oops for now.
Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> Signed-off-by: David S. Miller <davem@davemloft.net>
1. If the fake 5513 id bit is not set by the BIOS we must have the 5518
id in the device table.
2. If the register remapping is not set by the BIOS then the enable bit
check in ide_pci_setup_ports will fail. It's safe to switch to the
remapping mode here. Keeping the not remapped mode would need quite big
changes AFAICS.
This driver caused a _lot_ of warnings due to tons
of explicit casts to "uclong". Making all the types
sane not only removed the warnings, but got rid of
a lot of silly casting, since the types are now much
more natural to what the driver wanted to do in the
first place.
Roland McGrath [Fri, 10 Sep 2004 16:20:19 +0000 (09:20 -0700)]
[PATCH] Fix PTRACE_CONT after single-step into signal delivery
The previous single-step patch ("make single-step into signal delivery
stop in handler") took things a little too far.
It left TF set in the sigcontext on the stack, so a PTRACE_CONT after
stopping at the handler entry will step instead of resume. This
additional patch fixes it.
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Linus Torvalds <torvalds@osdl.org>