New version of the NETIF_F_LLTX for network devices patch.
This allows network drivers to set the NETIF_F_LLTX flag
and then do their own locking in start_queue_xmit.
This lowers locking overhead in this critical path.
The drivers can use try lock if they want and return -1
when the lock wasn't grabbed. In this case the packet
will be requeued. For better compatibility this is only
done for drivers with LLTX set, others don't give a special
meaning to -1.
Most of the modern drivers who have a lock around hard_start_xmit
can just set this flag. It may be a good idea to convert the spin
lock there to a try lock. The only thing that should be audited
is that they do enough locking in the set_multicast_list function
too, and not also rely on xmit_lock here.
Now doesn't move any code around and does things with gotos instead.
The loop printk is also still there even for NETIF_F_LLTX
For drivers that don't set the new flag nothing changes.
Signed-off-by: David S. Miller <davem@davemloft.net>
In this case, we know we need more fragment(s).
So, let's fill up to maxfraglen (instead of mtu)
to avoid needless copy in the next loop.
Signed-off-by: Hideaki YOSHIFUJI <yoshfuji@linux-ipv6.org> Signed-off-by: Herbert Xu <herbert@gondor.apana.org.au> Signed-off-by: David S. Miller <davem@davemloft.net>
[PATCH] uml: Avoid forcing use of the no-op scheduler
Avoid forcing use of the no-op scheduler for UBD; this may uncover some
bugs in the UBD driver, and in fact uml-ubd-no-empty-queue.patch is needed
to make this sure. But as of now, no other bugs have been discovered, so
this should be safe.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Avoid using, in the UBD driver, the elv_queue_empty function. It's for the
block layer only; in fact, the Anticipatory Scheduler can return NULL with
elv_next_request() even if the queue is not empty, because it waits for the
process to send another request before seeking on the disk.
In fact, if (with uml-ubd-any-elevator) we let UBD use any scheduler,
elevator=as will make the UBD driver Oops, if we don't have this patch.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Wed, 8 Sep 2004 00:57:38 +0000 (17:57 -0700)]
[PATCH] Speed up oprofile buffer drain code
I noticed a large machine was doing about 400,000 context switches per
second when oprofile was enabled. Upon closer inspection it looks like we
were rearming the buffer sync timer without modifying the expire time.
Now that we have schedule_delayed_work_on I believe we can remove the timer
completely. Each cpu should be offset by 1 jiffy so they dont all fire at
the same time. I bumped DEFAULT_TIMER_EXPIRE from 2 to 10 times a second
to be sure we reap cpu buffers.
With the following patch the same large machine gets about 4000 context
switches per second.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Wed, 8 Sep 2004 00:57:26 +0000 (17:57 -0700)]
[PATCH] fix oprofile vfree warning on error
On error we can call __free_cpu_buffers with only some buffers allocated.
I was getting a bunch of vfree warnings when I hit it, we should check
before calling vfree.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The problem here is, finished_one_bio() shouldn't call aio_complete() since
no work has been done. I have a fix for this - can you verify this ? I am
not really comfortable with this "tweaking". (I am not really sure about
IO errors like EIO etc. - if they can lead to calling aio_complete()
twice)
Fix is to call aio_complete() ONLY if there is something to report. Note
the we don't update dio->result with any error codes from get_user_pages(),
they just passed as "ret" value from do_direct_IO().
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Dave Jones [Wed, 8 Sep 2004 00:56:19 +0000 (17:56 -0700)]
[PATCH] Remove bogus memset from cpqfc driver
Not that this driver compiles, but coverity picked up this nonsense. If
the pci_alloc_consistent fails, we go boom. Amusingly, after the ==NULL
check, is an identical memset.
Signed-off-by: Dave Jones <davej@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] fbdev: Add module_init() and fb_get_options() per driver
This adds module_init(xxxfb_init) in all drivers. For drivers with
xxxfb_setup(), this patch also adds a
'xxxfb_setup(fb_get_options("xxxfb"))' prior to initialization.
Signed-off-by: Antonino Daplas <adaplas@pol.net> Signed-off-by: Adrian Bunk <bunk@fs.tum.de> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] fbdev: Clean up framebuffer initialization
This patch probably deserves discussion among developers.
Currently, the framebuffer system is initialized in a roundabout manner.
First, drivers/char/mem.c calls fbmem_init(). fbmem_init() will then
iterate over an array of individual drivers' xxxfb_init(), then each driver
registers its presence back to fbmem. During console_init(),
drivers/char/vt.c will call fb_console_init(). fbcon will check for
registered drivers, and if any are present, will call take_over_console()
in drivers/char/vt.c.
This patch changes the initialization sequence so it proceeds in this
manner: Each driver has its own module_init(). Each driver calls
register_framebuffer() in fbmem.c. fbmem.c will then notify fbcon of the
driver registration. Upon notification, fbcon calls take_over_console() in
vt.c.
The following are the changes brought about by this patch:
- Each subsystem (fbcon, fbmem, xxxfb) will have their own module_init.
Thus, explicit calls to each subsystem's init functions are eliminated.
- The struct fb_drivers array in fbmem.c can be removed. This slashes
around 400 lines in fbmem.c
- Parsing of kernel boot options were done by fbmem.c calling each
driver's xxxfb_setup() function. Because this is not possible with this
patch, drivers can choose to either:
- have their own __setup() routine
- call fb_get_options("xxxfb") and pass the return value to
xxxfb_setup(). This is to maintain compatibility with the
'video=xxxfb:<options>' semantics.
- Getting a framebuffer console will occur a bit late during the boot
process since the initialization sequence will depend upon the link
order. So, 'video/' is moved up in drivers/Makefile, shortly after
'pci/'
- Because driver initialization will be dependent on the link order,
hardware that depends on other subsystems (agpgart, usb, serial, etc) may
choose to initialize after the subsystems they depend on.
[PATCH] fbcon: take over console on driver registration
- This fixes another regression from 2.4. If fbcon is compiled
statically, and the framebuffer driver is compiled as a module, doing a
'modprobe xxxfb' does nothing to the console. This has generated
numerous bug reports from users.
With this patch, fbmem will notify fbcon upon driver registration
allowing fbcon to take over the console.
- This also fixes con2fbmap not working if fbcon is compiled as a module
using the same mechanism as described above.
This patch speeds up scrolling of tdfxfb by maximizing var->yres_virtual so
tdfxfb uses SCROLL_PAN_MOVE instead of SCROLL_REDRAW. This is true whether
CONFIG_FB_3DFX_ACCEL is set or not. This problem was reported by Paolo
Ornati <ornati@fastwebnet.it> who also made substantial contributions to
solve this problem.
This patch also fixes compile errors when CONFIG_FB_3DFX_ACCEL is false.
Tom Rini [Wed, 8 Sep 2004 00:54:56 +0000 (17:54 -0700)]
[PATCH] ppc32: Switch arch/ppc/boot to lib/zlib_inflate
The following patch switches arch/ppc/boot over from using its own version
of zlib to the code found under lib/zlib_inflate. In conjunction with the
previous two patches, the size of the resulting images isn't noticably
different. But this does have the advantage of removing another copy of
zlib from the kernel, and I believe this allows for lib/inflate.c to go
away (as that's basically what ppc used to use).
Signed-off-by: Tom Rini <trini@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Tom Rini [Wed, 8 Sep 2004 00:54:44 +0000 (17:54 -0700)]
[PATCH] zlib_inflate: Make zlib_inflate_trees_fixed(...) generate the table
The following changes zlib_inflate_trees_fixed(...) from using a statically
defined table, to generating this table. This cuts out 4-8kB from
inftrees.o (4kB on IBM 440GP, 8kB on PPC 74xx).
Signed-off-by: Tom Rini <trini@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The different stack frame layout for packed stacks broke cpu hotplug.
Recreate the initial stack frame of the idle thread for offline cpus coming
back online. Reenable interrupts after loading the initial registers. In
addition this patch contains two more bug fixes: a typo for 64 bit
(__SMALL_STACK_SIZE vs. __SMALL_STACK) and show_trace didn't show a trace
if for task == NULL.
Signed-off-by: Martin Schwidefsky <schwidefsky@de.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Removes unnecessary min/max macros and changes calls to use kernel.h macros
instead.
Signed-off-by: Michael Veeck <michael.veeck@gmx.net> Signed-off-by: Maximilian Attems <janitor@sternwelten.at> Signed-off-by: Martin Schwidefsky <schwidefsky@de.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
This patch removes the default stubs for init_module and cleanup_module,
and checks for NULL instead. It changes modpost to only create references
to those functions if they actually exist.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Wed, 8 Sep 2004 00:53:11 +0000 (17:53 -0700)]
[PATCH] knfsd: remove redundant initialization in nfsd4_lockt
No need to set fl_owner and fl_pid to 0, since that's already been done by
locks_init_lock.
Signed-off-by: Andy Adamson <andros@umich.edu> Signed-off-by: J. Bruce Fields <bfields@citi.umich.edu> Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Wed, 8 Sep 2004 00:53:00 +0000 (17:53 -0700)]
[PATCH] knfsd: nfsd4: store current->tgid instead of lockowner hash in fl_pid
The file_lock.fl_pid is no longer used in posix_same_owner() tests. So just
set it to current->tgid for informational purposes.
Signed-off-by: Andy Adamson <andros@umich.edu> Signed-off-by: J. Bruce Fields <bfields@citi.umich.edu> Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Wed, 8 Sep 2004 00:52:48 +0000 (17:52 -0700)]
[PATCH] knfsd: nfsd4: postpone release of stateowner on CLOSE
Postpone the release of a stateowner on CLOSE for lease time to enable the
CLOSE replay cache. Place stateowner on the close_lru list to be reaped by
the laundromat service.
Signed-off-by: Andy Adamson <andros@umich.edu> Signed-off-by: J. Bruce Fields <bfields@citi.umich.edu> Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Wed, 8 Sep 2004 00:52:37 +0000 (17:52 -0700)]
[PATCH] knfsd: nfsd4 could leak a stateid in an error path
nfsd4 could leak a stateid in a case of kmalloc failure; fix.
Signed-off-by: Andy Adamson <andros@umich.edu> Signed-off-by: J. Bruce Fields <bfields@citi.umich.edu> Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Wed, 8 Sep 2004 00:52:26 +0000 (17:52 -0700)]
[PATCH] knfsd: trivial cleanup of nfs4state.c
Whitespace cleanup, fix one dprintk, remove superfluous casts of NULL.
Signed-off-by: J. Bruce Fields <bfields@citi.umich.edu> Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Wed, 8 Sep 2004 00:52:14 +0000 (17:52 -0700)]
[PATCH] knfsd: nfsd4: Support acl_support attribute
The nfs4 attributes supported_attrs and aclsupport should not be static; they
need to depend on the exported filesystem's acl support. Test the latter by
attempting to get an acl, and adjust the returned attributes apropriately.
Signed-off-by: J. Bruce Fields <bfields@citi.umich.edu> Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Neil Brown [Wed, 8 Sep 2004 00:52:03 +0000 (17:52 -0700)]
[PATCH] knfsd: fix incorrect indentation in fh_verify
fix incorrect indentation in fh_verify
Signed-off-by: J. Bruce Fields <bfields@citi.umich.edu> Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] small wait_on_page_writeback_range() optimization
filemap_fdatawait() calls wait_on_page_writeback_range() with -1 as "end"
parameter. This is not needed since we know the EOF from the inode. Use
that instead.
Roland Dreier [Wed, 8 Sep 2004 00:51:05 +0000 (17:51 -0700)]
[PATCH] fs/compat.c: rwsem instead of BKL around ioctl32_hash_table
Currently the BKL is used to synchronize access to ioctl32_hash_table in
fs/compat.c. It seems that an rwsem would be more appropriate, since this
would allow multiple lookups to occur in parallel (and also serve the
general good of minimizing use of the BKL).
I added lock_kernel()/unlock_kernel() around the call to t->handler when a
compatibility handler is found in compat_sys_ioctl() to preserve the
expectation that the BKL will be held during driver ioctl operations. It
should be safe to do lock_kernel() while holding ioctl32_sem because of the
magic BKL sleep semantics.
Signed-off-by: Roland Dreier <roland@topspin.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] fix f_version optimization for get_tgid_list
The kernel contains an optimization that skips the linked list walk in
get_tgid_list for the common case of sequential accesses. Unfortunately
the optimization is buggy (missing NULL pointer check for the result of
find_task_by_pid) and broken (actually - broken twice: the tgid value that
is stored in f_version is always 0 because tgid is overwritten when the
string is created and additionally the common case is not filldir < 0, it's
running out of nr_tgids).
The attached patch fixes these bugs.
Roger Luethi <rl@hellgate.ch> ran a benchmark:
test: top -d 0 -b -n 10 > /dev/null
==> 2.6.8 <==
real 0m19.092s
user 0m5.013s
sys 0m12.622s
==> 2.6.8 + patch-tgid-bugfixes <==
real 0m10.062s
user 0m5.042s
sys 0m4.111s
Ram Pai [Wed, 8 Sep 2004 00:50:22 +0000 (17:50 -0700)]
[PATCH] filemap read() fix
Fix the do_generic_file_read()-reads-one-page-too-many-bug for the fifth
time.
This patch combines the best features from Nick's patch and also makes
index and end_index consistent. (i.e index 'n' covers n*PAGE_SIZE to
((n+1)PAGE_SIZE)-1. I did not feel comfortable with the way index and
end_index represented different ranges. It was like comparing apples with
oranges.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Wed, 8 Sep 2004 00:49:47 +0000 (17:49 -0700)]
[PATCH] make single-step into signal delivery stop in handler
On x86 and x86-64, setting up to run a signal handler clears the
single-step bit (TF) in the processor flags before starting the handler.
This makes sense when a process is handling its own SIGTRAPs.
But when TF is set because PTRACE_SINGLESTEP was used, and that call
specified a handled signal so the handler setup is happening, it doesn't
make so much sense. When the debugger stops to show you a signal about to
be delivered, and that signal should be handled, and then you do step or
stepi, you expect to see the signal handler code. In fact, the signal
handler runs to completion and then you see the single-step trap at the
resumed code instead of seeing the handler.
This patch changes signal handler setup so that when TF is set and the
thread is under ptrace control, it synthesizes a single-step trap after
setting up the PC and registers to start the handler. This makes that
PTRACE_SINGLESTEP not strictly a "step", since it actually runs no user
instructions at all. But it is definitely what a debugger user wants, so
that single-stepping always stops and shows each and every instruction
before it gets executed.
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Wed, 8 Sep 2004 00:49:36 +0000 (17:49 -0700)]
[PATCH] i386 syscall tracing of bogus system calls
In 2.4, strace will show you all bogus system calls a process tries. In
2.6, it only shows you stubs < __NR_syscalls, and there is no tracing stop
for large bogus system call numbers. I can't see why this was changed, so
I am assuming it was accidental.
This patch restores the expected behavior that syscall tracing shows every
bogus syscall attempt.
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Wed, 8 Sep 2004 00:49:24 +0000 (17:49 -0700)]
[PATCH] Remove RUSAGE_GROUP
After my cleanup of the rusage semantics was so quickly taken in by Andrew
and Linus without comment, I wonder if I should not have tried to be so
accommodating of potential objections as I was. :-)
In my original posting, I solicited comment on whether introducing
RUSAGE_GROUP as distinct from RUSAGE_SELF was warranted. Note that we've
now changed the behavior of the times system call when using CLONE_THREAD,
so changing getrusage RUSAGE_SELF to match would be consistent. I think
that changing the meaning of the old RUSAGE_SELF value is preferable to
introducing the new value for the proper POSIX getrusage behavior. This
patch against Linus's current tree dumps RUSAGE_GROUP and makes RUSAGE_SELF
have the fixed behavior.
If there is interest in having a new explicit interface to sample a single
thread's stats alone, then I think that would be better done by introducing
a new value for RUSAGE_THREAD. This is trivial to implement, but I won't
offer patches bloating the interface if noone is actually interested in
using it.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Wed, 8 Sep 2004 00:49:13 +0000 (17:49 -0700)]
[PATCH] ptrace userspace API preservation
This makes any ptrace operation that finds the target in TASK_STOPPED state
morph it into TASK_TRACED state before doing anything. This necessitates
reverting the last_siginfo accesses to check instead of assume last_siginfo
is set, since it's no longer impossible to be in TASK_TRACED without being
stopped in ptrace_stop (though there are no associated races to worry
about).
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Wed, 8 Sep 2004 00:48:58 +0000 (17:48 -0700)]
[PATCH] cleanup ptrace stops and remove notify_parent
This adds a new state TASK_TRACED that is used in place of TASK_STOPPED
when a thread stops because it is ptraced. Now ptrace operations are only
permitted when the target is in TASK_TRACED state, not in TASK_STOPPED.
This means that if a process is stopped normally by a job control signal
and then you PTRACE_ATTACH to it, you will have to send it a SIGCONT before
you can do any ptrace operations on it. (The SIGCONT will be reported to
ptrace and then you can discard it instead of passing it through when you
call PTRACE_CONT et al.)
If a traced child gets orphaned while in TASK_TRACED state, it morphs into
TASK_STOPPED state. This makes it again possible to resume or destroy the
process with SIGCONT or SIGKILL.
All non-signal tracing stops should now be done via ptrace_notify. I've
updated the syscall tracing code in several architectures to do this
instead of replicating the work by hand. I also fixed several that were
unnecessarily repeating some of the checks in ptrace_check_attach. Calling
ptrace_check_attach alone is sufficient, and the old checks repeated before
are now incorrect, not just superfluous.
I've closed a race in ptrace_check_attach. With this, we should have a
robust guarantee that when ptrace starts operating, the task will be in
TASK_TRACED state and won't come out of it. This is because the only way
to resume from TASK_TRACED is via ptrace operations, and only the one
parent thread attached as the tracer can do those.
This patch also cleans up the do_notify_parent and do_notify_parent_cldstop
code so that the dead and stopped cases are completely disjoint. The
notify_parent function is gone.
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Since the irq handling rework in 2.5 lots of code in the individual
<asm/hardirq.h> files is the same. This patch moves that common code
to <linux/hardirq.h>. The following differences existed:
- alpha, m68k, m68knommu and v850 were missing the ~PREEMPT_ACTIVE
masking in the CONFIG_PREEMPT case of in_atomic(). These
architectures don't support CONFIG_PREEMPT else this would have been
an easily-spottbale bug
- S390 didn't provide synchronize_irq as it doesn't fit into their
I/O model. They now get a spurious prototype/macro
- ppc added a new preemptible() macro that is provided for all
architectures now.
Most drivers were using <linux/interrupt.h> as they should, but a few
drivers and lots of architecture code has been updated to use
<linux/hardirq.h> instead of <asm/hardirq.h>
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] Cleanup & fix lost ticks handling on x86-64
This cleans up the x86-64 lost tick handling and fixes some issues:
First it moves that code into an own function.
The newer could would become very noisy when the machine loses timer ticks
regularly. This happens often on some laptops etc. during the acpi ec
access (nothing much can be really done about it) This patch prints the
warnings only once.
It also fixes the logic on when to ask cpufreq for a new estimate.
And it implements timer fallback to HPET when there are really lots of lost
ticks. This is following i386. PIT fallback isn't implemented right now
though, but I hope we don't need this.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Wed, 8 Sep 2004 00:48:18 +0000 (17:48 -0700)]
[PATCH] ppc64: handle SLB misses in realmode
Tested on pSeries and iSeries. Some future plans for VSID allocation may
mean we have to take this out again, but that's a while off yet, and in the
meantime it's a significant speedup.
This patch makes the PPC64 SLB miss handler run in real mode (i.e. MMU
off) for it's whole duration, on pSeries machines. Avoiding the rfid used
to turn relocation on saves some 70-80 cycles on Power4 and Power5. Not
having to save and restore SRR0 and SRR1 saves a few more, and means we
don't need an extra save slot for r3. Overall there's around a 27% speedup
on Power4.
Signed-off-by: David Gibson <david@gibson.dropbear.id.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Gibson [Wed, 8 Sep 2004 00:48:06 +0000 (17:48 -0700)]
[PATCH] ppc64: fix declaration order in asm-ppc64/tlb.h
In asm-ppc64/tlb.h, tlb_flush() is defined as inline after the #include of
asm-generic/tlb.h which uses it, defeating the inline directive. gcc-3.4
exposes this problem, causing a compile failure. This patch reorders the
file to fix the problem.
Signed-off-by: David Gibson <dwg@au1.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Wed, 8 Sep 2004 00:47:55 +0000 (17:47 -0700)]
[PATCH] ppc64: fix compat NUMA API on big endian 64bit
Switch the NUMA API to use compat_get_bitmap/compat_put_bitmap. In order
to use compat_alloc_userspace instead of set_fs tricks, we have to do a few
copies.
This is what we are currently using on ppc64 but are willing to entertain
the idea of going to a 32bit bitmap, especially considering how much hoops
we have to go through to get it right in this patch.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Wed, 8 Sep 2004 00:47:32 +0000 (17:47 -0700)]
[PATCH] ppc64: fix compat cpu affinity on big endian 64bit
Add compat sched affinity code. We can argue about how
USE_COMPAT_ULONG_CPUMASK works now that the non compat interface has
changed.
The old non compat behaviour was to require a bitmap long enough in both
setaffinity and getaffinity, now its only required in getaffinity. I could
do the same for the 32bit interfaces.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Wed, 8 Sep 2004 00:47:10 +0000 (17:47 -0700)]
[PATCH] ppc64: cut down paca footprint
The paca currently contains an iseries only structure which is quite large
(~1kB). The following patch removes this overhead on pseries and g5
kernels.
Since the paca is no longer required to be page aligned, remove it from the
page aligned section.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Wed, 8 Sep 2004 00:46:47 +0000 (17:46 -0700)]
[PATCH] ppc64: fix __rw_yield prototype
From: Nathan Lynch <nathanl@austin.ibm.com>
Hit this in latest bk:
include/asm/spinlock.h: In function `_raw_read_lock':
include/asm/spinlock.h:198: warning: passing arg 1 of `__rw_yield' from incompatible pointer type
include/asm/spinlock.h: In function `_raw_write_lock':
include/asm/spinlock.h:255: warning: passing arg 1 of `__rw_yield' from incompatible pointer type
This seems to have been broken by the out-of-line spinlocks patch.
You won't hit it unless you've enabled CONFIG_PPC_SPLPAR. Use the
rwlock_t for the argument type, and move the definition of rwlock_t up
next to that of spinlock_t.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Wed, 8 Sep 2004 00:46:36 +0000 (17:46 -0700)]
[PATCH] ppc64: fix hang on oprofile shutdown
We had a problem in our dummy perfmon handler where we wouldnt reset the
PMAO bit. If the bit ended up set and oprofile shutdown and removed its
handler then we would end up in a hard loop taking perfmon exceptions.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] Time interpolator: Scalability enhancements and high resolution time for IA64
This has been in the ia64 (and hence -mm) trees for a couple of months.
Changelog:
* Affects only architectures which define CONFIG_TIME_INTERPOLATION
(currently only IA64)
* Genericize time interpolation, make time interpolators easily usable
and provide instructions on how to use the interpolator for other
architectures.
* Provide nanosecond resolution for clock_gettime and an accuracy
up to the time interpolator time base.
* clock_getres() reports resolution of underlying time basis which
is typically <50ns and may be 1ns on some systems.
* Make time interpolator self-tuning to limit time jumps
and to make the interpolators work correctly on systems with
broken time base specifications.
* SMP scalability: Make clock_gettime and gettimeofday scale O(1)
by removing the cmpxchg for most clocks (tested for up to 512 CPUs)
* IA64: provide asm fastcall that doubles the performance
of gettimeofday and clock_gettime on SGI and other IA64 systems
(asm fastcalls scale O(1) together with the scalability fixes).
* IA64: provide nojitter kernel option so that IA64 systems with
correctly synchronized ITC counters may also enjoy the
scalability enhancements.
Performance measurements for single calls (ITC cycles):
A. 4 way Intel IA64 SMP system (kmart)
ITC offsets:
kmart:/usr/src/noship-tests # dmesg|grep synchr
CPU 1: synchronized ITC with CPU 0 (last diff 1 cycles, maxerr 417 cycles)
CPU 2: synchronized ITC with CPU 0 (last diff 2 cycles, maxerr 417 cycles)
CPU 3: synchronized ITC with CPU 0 (last diff 1 cycles, maxerr 417 cycles)
hid-core calls hiddev_disconnect() when the underlying device goes away
(hot unplug or system shutdown). Normally, hiddev_disconnect() will clean
up nicely and return to hid-core who then frees the hid structure.
However, if the corresponding hiddev node is open at disconnect time,
hiddev delays the majority of disconnect work until the device is closed
via hiddev_release(). hiddev_release() calls hiddev_cleanup() which
proceeds to dereference the hid struct which hid-core freed back when the
hardware was disconnected. Oops.
To solve this, we change hiddev_disconnect() to deregister the hiddev minor
and invalidate its table entry immediately and delay only the freeing of
the hiddev structure itself. We're protected against future operations on
the fd since the major fops check hiddev->exists.
Signed-off-by: Adam Kropelin <akropel1@rochester.rr.com> Signed-off-by: Vojtech Pavlik <vojtech@suse.cz> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>