Jens Axboe [Tue, 19 Oct 2004 01:01:41 +0000 (18:01 -0700)]
[PATCH] cfq-v2 I/O scheduler update
Here is the next incarnation of the CFQ io scheduler, so far known as
CFQ v2 locally. It attempts to address some of the limitations of the
original CFQ io scheduler (hence forth known as CFQ v1). Some of the
problems with CFQ v1 are:
- It does accounting for the lifetime of the cfq_queue, which is setup
and torn down for the time when a process has io in flight. For a fork
heavy work load (such as a kernel compile, for instance), new
processes can effectively starve io of running processes. This is in
part due to the fact that CFQ v1 gives preference to a new processes
to get better latency numbers. Removing that heuristic is not an
option exactly because of that.
- It makes no attempts to address inter-cfq_queue fairness.
- It makes no attempt to limit upper latency bound of a single request.
- It only provides per-tgid grouping. You need to change the source to
group on a different criteria.
- It uses a mempool for the cfq_queues. Theoretically this could
deadlock if io bound processes never exit.
- The may_queue() logic can be unfair since it fluctuates quickly, thus
leaving processes sleeping while new processes are allowed to allocate
a request.
CFQ v2 attempts to fix these issues. It uses the process io_context
logic to maintain a cfq_queue lifetime of the duration of the process
(and its io). This means we can now be a lot more clever in deciding
which process is allowed to queue or dispatch io to the device. The
cfq_io_context is per-process per-queue, this is an extension to what AS
currently does in that we truly do have a unique per-process identifier
for io grouping. Busy queues are sorted by service time used, sub sorted
by in_flight requests. Queues that have no io in flight are also
preferred at dispatch time.
Accounting is done on completion time of a request, or with a fixed cost
for tagged command queueing. Requests are fifo'ed like with deadline, to
make sure that a single request doesn't stay in the io scheduler for
ages.
Process grouping is selectable at runtime. I provide 4 grouping
criterias: process group, thread group id, user id, and group id.
As usual, settings are sysfs tweakable in /sys/block/<dev>/queue/iosched
back_seek_max
back_seek_penalty:
Useful logic stolen from AS that allow small backwards seeks in
the io stream if we deem them useful. CFQ uses a strict
ascending elevator otherwise. _max controls the maximum allowed
backwards seek, defaulting to 16MiB. _penalty denotes how
expensive we account a backwards seek compared to a forward
seek. Default is 2, meaning it's twice as expensive.
clear_elapsed:
Really a debug switch, will go away in the future. It clears the
maximum values for completion and dispatch time, shown in
show_status.
fifo_batch_expire
fifo_batch_async
fifo_batch_sync:
The settings for the expiry fifo. batch_expire is how often we
allow the fifo expire to control which request to select.
Default is 125ms. _async is the deadline for async requests
(typically writes), _sync is the deadline for sync requests
(reads and sync writes). Defaults are, respectively, 5 seconds
and 0.5 seconds.
key_type:
The grouping key. Can be set to pgid, tgid, uid, or gid. The
current value is shown bracketed:
Default is tgid. To set, simply echo any of the 4 words into the
file.
quantum:
The amount of requests we select for dispatch when the driver
asks for work to do and the current pending list is empty.
Default is 4.
queued:
The minimum amount of requests a group is allowed to queue.
Default is 8.
show_status:
Debug output showing the current state of the queues.
tagged:
Set this to 1 if the device is using tagged command queueing.
This cannot be reliably detected by CFQ yet, since most drivers
don't use the block layer (well it could, by looking at number
of requests being between dispatch and completion. but not
completely reliably). Default is 0.
The patch is a little big, but works reliably here on my laptop. There
are a number of other changes and fixes in there (like converting to
hlist for hashes). The code is commented a lot better, CFQ v1 has
basically no comments (reflecting that it was writting in one go, no
touched or tuned much since then). This is of course only done to
increase the AAF, akpm acceptance factor. Since I'm on the road, I
cannot provide any really good numbers of CFQ v1 compared to v2, maybe
someone will help me out there.
Jens Axboe [Tue, 19 Oct 2004 01:01:28 +0000 (18:01 -0700)]
[PATCH] switchable and modular io schedulers
This patch modularizes the io schedulers completely, allowing them to be
modular. Additionally it enables online switching of io schedulers. See
also http://lwn.net/Articles/102593/ .
There's a scheduler file in the sysfs directory for the block device
queue:
axboe@router:/sys/block/hda/queue> ls
iosched max_sectors_kb read_ahead_kb
max_hw_sectors_kb nr_requests scheduler
If you list the contents of the file, it will show available schedulers
and the active one:
Andrew Morton [Tue, 19 Oct 2004 01:01:03 +0000 (18:01 -0700)]
[PATCH] jbd wakeup fix
Processes can sleep in do_get_write_access(), waiting for buffers to be
removed from the BJ_Shadow state. We did this by doing a wake_up_buffer() in
the commit path and sleeping on the buffer in do_get_write_access().
With the filtered bit-level wakeup code this doesn't work properly any more -
the wake_up_buffer() accidentally wakes up tasks which are sleeping in
lock_buffer() as well. Those tasks now implicitly assume that the buffer came
unlocked. Net effect: Bogus I/O errors when reading journal blocks, because
the buffer isn't up to date yet. Hence the recently spate of journal_bmap()
failure reports.
The patch creates a new jbd-private BH flag purely for this wakeup function.
So a wake_up_bit(..., BH_Unshadow) doesn't wake up someone who is waiting for
a wake_up_bit(BH_Lock).
JBD was the only user of wake_up_buffer(), so remove it altogether.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] reduce number of parameters to __wait_on_bit() and __wait_on_bit_lock()
Some of the parameters to __wait_on_bit() and __wait_on_bit_lock() are
redundant, as the wait_bit_queue parameter holds the flags word and the bit
number. This patch updates __wait_on_bit() and __wait_on_bit_lock() to
fetch that information from the wait_bit_queue passed to them and so reduce
the number of parameters so that -mregparm may be more effective.
Incremental atop the complete out-of-lining of the contention cases and the
fastcall and wait_on_bit_lock()/test_and_set_bit() fixes.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] move wait ops' contention case completely out of line
Move the slow paths of wait_on_bit() and wait_on_bit_lock() out of line.
Also uninline wake_up_bit() to reduce the number of callsites generated,
and adjust loop startup in __wait_on_bit_lock() to properly reflect its
usage in the contention case.
Incremental atop the fastcall and wait_on_bit_lock()/test_and_set_bit()
fixes. Successfully tested on x86-64.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Eliminate specialized page and bh waitqueue hashing structures in favor of
a standardized structure, using wake_up_bit() to wake waiters using the
standardized wait_bit_key structure.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The following patch series consolidates the various instances of waitqueue
hashing to use a uniform structure and share the per-zone hashtable among all
waitqueue hashers. This is expected to increase the number of hashtable
buckets available for waiting on bh's and inodes and eliminate statically
allocated kernel data structures for greater node locality and reduced kernel
image size. Some attempt was made to look similar to Oleg Nesterov's
suggested API in order to provide some kind of credit for independent
invention of something very similar (the original versions of these patches
predated my public postings on the subject of filtered waitqueues).
These patches have the further benefit and intention of enabling aio to use
filtered wakeups by standardizing the data structure passed to wake functions
so that embedded waitqueue elements in aio structures may be succesfully
passed to the filtered wakeup wake functions, though this patch series doesn't
implement that particular functionality.
Successfully stress-tested on x86-64, and ia64 in recent prior versions.
This patch:
Move waitqueue -related functions not needing static functions in sched.c
to kernel/wait.c
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paulo Marques [Tue, 19 Oct 2004 00:59:03 +0000 (17:59 -0700)]
[PATCH] kallsyms data size reduction / lookup speedup
This patch is an improvement over my first kallsyms speedup patch posted about
2 weeks ago.
It changes scripts/kallsyms as to produce a different format for
kallsyms_names and extra data to speedup lookups. The compression algorithm
is quite simple: it uses all the char codes not actually used in symbols to
build a lookup table that translates these codes into small strings. For
instance, in my test runs the code 0xFE was being translated into "acpi_"
giving a 4 byte save on every translation.
The advantage of this algorithm is that to translate a symbol we only require
information that is stored on that symbol position, and never need to go back
on the compressed stream to get information from other symbols.
To give an idea about the benefits of this algorithm here are some benchmark
results on a P4 2.8GHz with a symbol table with 10000 entries:
kallsyms_lookup average time:
vanilla 1346.0 us
speedup 14.4 us
with this patch 0.5 us
total data produced by scripts/kallsyms:
uncompressed 169 Kb
vanilla 134 Kb
with this patch 91 Kb
(speedup was my latest patch, that only changed the way kallsyms_lookup worked
and not the data format)
I removed a cond_resched() from the proc/kallsyms handling code path, because
using stem compression, if the current position went backwards, the hole
stream would be uncompressed up to the current position. It seemed that by
removing this loop it would be safe to remove the conditional reschedule
altogether.
There is just one catch with this patch: the time it takes to compile the
kernel goes up just a bit (about 0.8s on a P4 2.8GHz with defconfig). If this
delay is not acceptable I can change the compression algorithm so that it can
use the previous table (calculating a new table is what consumes most of the
time, and not doing the actual compression) and check to see if it obtains a
similar compression ratio. If it does, then this is a sign that the symbol
patterns haven't changed that much and this table is still good to use. This
would not only cut the time down to half on any compilation (because of the 2
pass symbol build method), but in frequent cases where a developer is
compiling a single file and linking everything over and over again, the table
optimization process would never run.
I'm CC'ing Brent Casavant on this email, because last june he sent a patch
trying a different approach that used a 32 entry symbol cache, because there
was a problem with the time "top" took to read "proc/<pid>/wchan". I was
hopping he would be willing to test this patch and comment on the results.
Signed-off-by: Paulo Marques <pmarques@grupopie.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
- Key attributes:
- Key type
- Description (by which a key of a particular type can be selected)
- Payload
- UID, GID and permissions mask
- Expiry time
- Keyrings (just a type of key that holds links to other keys)
- User-defined keys
- Key revokation
- Access controls
- Per user key-count and key-memory consumption quota
- Three std keyrings per task: per-thread, per-process, session
- Two std keyrings per user: per-user and default-user-session
- prctl() functions for key and keyring creation and management
- Kernel interfaces for filesystem, blockdev, net stack access
- JIT key creation by usermode helper
David Howells [Tue, 19 Oct 2004 00:58:38 +0000 (17:58 -0700)]
[PATCH] keys: new error codes for Alpha, MIPS, PA-RISC, Sparc & Sparc64
The attached patch adds the new error codes I added for key-related errors to
those archs that don't make use of <asm-generic/errno.h>, including Alpha,
MIPS, PA-RISC, Sparc and Sparc64. This is required to compile with
CONFIG_KEYS on those platforms.
Signed-Off-By: David Howells <dhowells@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Matthew Dobson [Tue, 19 Oct 2004 00:58:00 +0000 (17:58 -0700)]
[PATCH] Create nodemask_t
The idea behind this patch is to create a nodemask_t as a node analog of
cpumask_t. As NUMA machines become more common, the need for a standard,
cross-platform bitmap of both online & possible nodes becomes more
apparent. We believe we've worked out most of the kinks of the variable
length bitmap types with the recent cpumask_t patches. Nodemasks are also
currently far less widespread than cpumasks. Further, inclusion at this
point in the kernel would mean consistency in node handling between 2.6 and
2.7.
Future goals would be to get rid of the 'numnodes' variable used to count
the number of online nodes, and replace with node_online_map. This would
allow arbitrary node numbering and facilitate node hotplugging.
(Nothing actually uses this yet, but several projects need it, and it does
model a well-defined physical grouping).
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Peter Osterlund [Tue, 19 Oct 2004 00:57:46 +0000 (17:57 -0700)]
[PATCH] cdrom: buffer sizing fix
The problem is that some drives fail the "GET CONFIGURATION" command when
asked to only return 8 bytes. This happens for example on my drive, which
is identified as:
Peter Osterlund [Tue, 19 Oct 2004 00:57:21 +0000 (17:57 -0700)]
[PATCH] packet-writing: add credits
Nigel pointed out that the earlier patches contained attributions that
are not present in this patch. The 2.4 patch contains:
Nov 5 2001, Aug 8 2002. Modified by Andy Polyakov
<appro@fy.chalmers.se> to support MMC-3 complaint DVD+RW units.
and Nigel changed it to this in his 2.6 patch:
Modified by Nigel Kukard <nkukard@lbsd.net> - support DVD+RW
2.4.x patch by Andy Polyakov <appro@fy.chalmers.se>
The patch I sent you deleted most of the earlier work and moved the
rest to cdrom.c, but the comments were not moved over, since the
earlier authors didn't modify cdrom.c.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ingo Molnar [Mon, 18 Oct 2004 16:12:06 +0000 (09:12 -0700)]
[PATCH] fix & clean up zombie/dead task handling & preemption
This patch fixes all the preempt-after-task->state-is-TASK_DEAD problems we
had. Right now, the moment procfs does a down() that sleeps in
proc_pid_flush() [it could] our TASK_DEAD state is zapped and we might be
back to TASK_RUNNING to and we trigger this assert:
schedule();
BUG();
/* Avoid "noreturn function does return". */
for (;;) ;
I have split out TASK_ZOMBIE and TASK_DEAD into a separate p->exit_state
field, to allow the detaching of exit-signal/parent/wait-handling from
descheduling a dead task. Dead-task freeing is done via PF_DEAD.
Tested the patch on x86 SMP and UP, but all architectures should work
fine.
Ingo Molnar [Mon, 18 Oct 2004 16:11:52 +0000 (09:11 -0700)]
[PATCH] sched: fix SCHED_SMT & numa=fake=2 lockup
This patch fixes an interaction between the numa=fake=<domains> feature,
the domain setup code and cpu_siblings_map[]. The bug leads to a bootup
crash when using numa=fake=2 on a 2-way/4-way SMP+HT box.
When SCHED_SMT is turned on the domains-setup code relies on siblings not
spanning multiple domains (which makes perfect sense). But numa=fake=2
creates an assymetric 1101/0010 splitup between CPUs, which results in two
siblings being on different nodes.
The patch adds a check_siblings_map() function that checks the sibling maps
and fixes them up if they violate this rule. (it also prints a warning in
that case.)
The patch also turns SCHED_DOMAIN_DEBUG back on - had this been enabled
we'd have noticed this bug much earlier.
From: Badari Pulavarty <pbadari@us.ibm.com>
arch/x86_64/mm/numa.c: In function `numa_setup':
arch/x86_64/mm/numa.c:332: error: `numa_fake' undeclared (first use in this function)
arch/x86_64/mm/numa.c:332: error: (Each undeclared identifier is reported only once
arch/x86_64/mm/numa.c:332: error: for each function it appears in.)
Matthew Dobson [Mon, 18 Oct 2004 16:11:27 +0000 (09:11 -0700)]
[PATCH] sched_domains: Make SD_NODE_INIT per-arch #2
Here's yet another version of a patch to implement per-arch SD_*_INITs.
This follows the same basic idea of my last patch, but
1) defines an arch-specific SD_NODE_INIT for the 4 NUMA arches (i386,
x86_64, IA64 & PPC64),
2) defines *default* SD_CPU_INIT & SD_SIBLING_INIT for *all* arches,
with the possibility of them being overridden by simply defining an
arch-specific version in include/asm/topology.h.
The motivation behind the third version of this patch is that Martin feels
that there should be no "default" NUMA initializer because NUMA
characteristics are *very* arch/platform specific, and hence a "default"
NUMA initializer can only lead to confusion. I agree with most of that,
but don't quite see as much harm in having a default as he does.
Nevertheless, to keep him quiet, I've run up this version of the patch.
Martin, please run this through your magic test suite and make sure I
didn't break anything trivial.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Peter Williams [Mon, 18 Oct 2004 16:11:14 +0000 (09:11 -0700)]
[PATCH] CPU Scheduler: fix potential error in runqueue nr_uninterruptible count
Problem:
In the function try_to_wake_up(), when the runqueue's nr_uninterruptible
field is decremented it's possible (on SMP systems) that the pointer no
longer points to the runqueue that the task being woken was on when it went
to sleep. This would cause the wrong runqueue's field to be decremented
and the correct one tp remain unchanged.
Fix:
Save a pointer to the old runqueue at the beginning of the function and use
it when decrementing nr_uninterruptible.
Signed-off-by: Peter Williams <pwil3058@bigpond.net.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Mon, 18 Oct 2004 16:10:50 +0000 (09:10 -0700)]
[PATCH] sched: fixes for ia64 domain setup
Still having some trouble with ia64 domain setup on the Altixes. Jesse
hasn't had much time to look into it, and I'm lacking an Altix, so I'm not
sure if this is right or not...
Anyway, it again does the right thing on the NUMAQ, and fixes some real
bugs, so can you include it please?
* Increase SD_NODES_PER_DOMAIN to 6 from 4 to better match Altix's
topology. A setting of 4 will include this node, the other one
in the brick, and the 2 nodes in the next closest brick, while 6
will catch 2 other bricks. Probably it could be increased even
more.
* Work correctly with sparse and not completely full node maps.
* Nasty typo fixed in find_next_best_node:
- val = node_distance(node, i);
+ val = node_distance(node, n);
* Ensure all nodes are themselves a member of their numa balancing
domain. This is more a precaution against creative implementations
of node_distance.. but it makes the setup easier to verify without
having to look at a table of node_distance's, which is possibly
generated at runtime.
So again, I'm not too sure if this will fix the Altix setup or not. But if
you do a release, it will surely be less broken than it was before.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Mon, 18 Oct 2004 16:09:48 +0000 (09:09 -0700)]
[PATCH] sched: IA64 add disjoint NUMA domain support
Implement disjoint NUMA domain setup for IA64 architecture. Most of the code
was what was ripped out of kernel/sched.c, which was written by Jesse Barnes
<jbarnes@sgi.com>. I fixed up the tricky NUMA groups initialistion.
Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Mon, 18 Oct 2004 16:09:35 +0000 (09:09 -0700)]
[PATCH] sched: make domain setup overridable
Allow sched domain setup to be overridden by arch code. This functionality
is needed again.
From: Paul Jackson <pj@sgi.com>
Builds of 2.6.9-rc1-mm5 ia64 NUMA configs fail, with many complaints that
SD_NODE_INIT is defined twice, in asm/processor.h and linux/sched.h.
I guess that the preprocessor conditionals were wrong when Nick added the
per-arch override ability again of SD_NODE_INIT were wrong. At least this
change lets me rebuild ia64 again.
Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Paul Jackson <pj@sgi.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Mon, 18 Oct 2004 16:09:10 +0000 (09:09 -0700)]
[PATCH] sched: sched add load balance flag
Introduce SD_LOAD_BALANCE flag for domains where we don't want to do load
balancing (so we don't have to set up meaningless spans and groups). Use this
for the initial dummy domain, and just leave isolated CPUs on the dummy
domain.
Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Mon, 18 Oct 2004 16:08:46 +0000 (09:08 -0700)]
[PATCH] sched: integrate cpu hotplug and sched domains
Register a cpu hotplug notifier which reinitializes the scheduler domains
hierarchy. The notifier temporarily attaches all running cpus to a "dummy"
domain (like we currently do during boot) to avoid balancing. It then calls
arch_init_sched_domains which rebuilds the "real" domains and reattaches the
cpus to them.
Also change __init attributes to __devinit where necessary.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com>
Alterations from Nick Piggin:
* Detach all domains in CPU_UP|DOWN_PREPARE notifiers. Reinitialise and
reattach in CPU_ONLINE|DEAD|UP_CANCELED. This ensures the domains as
seen from the scheduler won't become out of synch with the cpu_online_map.
* This allows us to remove runtime cpu_online verifications. Do that.
* Dummy domains are __devinitdata.
* Remove the hackery in arch_init_sched_domains to work around the fact that
the domains used to work with cpu_possible maps, but node_to_cpumask returned
a cpu_online map.
Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Mon, 18 Oct 2004 16:08:34 +0000 (09:08 -0700)]
[PATCH] sched: add CPU_DOWN_PREPARE notifier
Add a CPU_DOWN_PREPARE hotplug CPU notifier. This is needed so we can dettach
all sched-domains before a CPU goes down, thus we can build domains from
online cpumasks, and not have to check for the possibility of a CPU coming up
or going down.
Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Mon, 18 Oct 2004 16:08:22 +0000 (09:08 -0700)]
[PATCH] sched: trivial sched changes
The following patches properly intergrate sched domains and cpu hotplug (using
Nathan's code), by having sched-domains *always* only represent online CPUs,
and having hotplug notifier to keep them up to date.
Then tackle Jesse's domain setup problem: the disjoint top-level domains were
completely broken. The group-list builder thingy simply can't handle distinct
sets of groups containing the same CPUs. The code is ugly and specific enough
that I'm re-introducing the arch overridable domains.
I doubt we'll get a proliferation of implementations, because the current
generic code can do the job for everyone but SGI. I'd rather take a look at
it again down the track if we need to rather than try to shoehorn this into
the generic code.
Nathan and I have tested the hotplug work. He's happy with it.
I've tested the disjoint domain stuff (copied it to i386 for the test), and it
does the right thing on the NUMAQ. I've asked Jesse to test it as well, but
it should be fine - maybe just help me out and run a test compile on ia64 ;)
This really gets sched domains into much better shape. Without further ado,
the patches.
This patch:
Make a definition static and slightly sanitize ifdefs.
Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The xtime value may become incorrect when the update_wall_time(ticks)
function is called with "ticks" > 1. In such a case, the xtime variable is
updated multiple times inside the loop but it is normalized only once
outside of the loop.
This bug was reported at:
http://bugme.osdl.org/show_bug.cgi?id=3403
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Mahoney [Mon, 18 Oct 2004 16:07:45 +0000 (09:07 -0700)]
[PATCH] ReiserFS: Add I/O error handling to journal operations
This patch allows ReiserFS to handle I/O errors in the journal (or journal
flush) where it would have previously panicked. The new behavior is to
mark the filesystem read-only, disallow new transactions to be started, and
to allow existing transactions to complete (though not to commit). The
resultant filesystem can be safely umounted, and checked via normal
mechanisms. As it is a journaling filesystem, the filesystem itself will
be in a similar state to the power being cut to the machine, once umounted.
Signed-off-by: Jeff Mahoney <jeffm@novell.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Mahoney [Mon, 18 Oct 2004 16:07:32 +0000 (09:07 -0700)]
[PATCH] ReiserFS: Cleanup access of journal (cosmetic)
This patch cleans up fs/reiserfs/journal.c such that repeated uses of
SB_JOURNAL(p_s_sb) are removed in favor of a local journal variable. The
compiler won't care, and it makes the code much easier to read.
Signed-off-by: Jeff Mahoney <jeffm@novell.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Mahoney [Mon, 18 Oct 2004 16:07:20 +0000 (09:07 -0700)]
[PATCH] ReiserFS: Cleanup internal use of bh macros
This patch cleans up ReiserFS's use of buffer head flags. All direct
access of BH_* are made into macro calls, and all reiserfs-specific BH_*
macro implementations have been removed and replaced with the BUFFER_FNS
implementations found in linux/buffer_head.h
Signed-off-by: Jeff Mahoney <jeffm@novell.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andreas Herrmann [Mon, 18 Oct 2004 16:06:06 +0000 (09:06 -0700)]
[PATCH] s390: zfcp host adapter
zfcp host adapter change:
- Return -EIO if wait_event_interruptible_timeout was interrupted.
- Reduce stack uage of zfcp_cfdc_dev_ioctl.
- Make zfcp_sg_list_[alloc,free] more consistent.
- Store driver version to zfcp_data structure.
- Add missing FSF states and make corresponding log messages consistent.
- Always wait for completion in zfcp_scsi_command_sync.
- Add Andreas to authors list.
- Add timeout for cfdc upload/download.
- Add support for temporary units (units not registered to the scsi stack).
- Allow sending of ELS commands to ports by their d_id.
- Increase port refcount while link test is running.
Signed-off-by: Martin Schwidefsky <schwidefsky@de.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Since people are used to doing "make linux ARCH=um" and to use "linux" as
the kernel image, make it be an hard link to vmlinux. This should hurt the
less possible the users (actually nothing) while not slowing down the
build.
Acked-by: Jeff Dike <jdike@addtoit.com> Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Hirokazu Takata [Mon, 18 Oct 2004 16:05:18 +0000 (09:05 -0700)]
[PATCH] m32r: fix sys_tas system call for m32r
This patch fixes a sys_tas system call for m32r.
- This patch fixes an Oops at sys_tas() in case CONFIG_SMP && CONFIG_PREEMPT.
> Unable to handle kernel paging request at virtual address XXXXXXXX
It is because a page fault happens at the spin_locked region in sys_tas()
and in_atomic() checks preempt_count, but spin_lock() already counts up
the preemt_count.
arch/mm/fault.c:
137 /*
138 * If we're in an interrupt or have no user context or are runni
ng in an
139 * atomic region then we must not take the fault..
140 */
141 if (in_atomic() || !mm)
142 goto bad_area_nosemaphore;
- sys_tas() is used for user-level mutual exclusion for the m32r,
which is prepared to implement a linuxthreads library.
The above problem may be happened in a program, which uses
pthread_mutex_lock(), calls sys_tas().
The current m32r instruction set has no user-level locking
functions for mutual exclusion.
# I hope it will be fixed in the future...
- This patch fixes the problem by using _raw_spin_lock() instead of
spin_lock(). spin_lock() increments up preemt_count, on the contrary,
_raw_sping_lock() does not.
# I think this fix is just a temporary work around, and
# it is preferable to be rewrite to make it simpler by using
# asm() function or something...
* arch/m32r/kernel/sys_m32r.c:
- Fix sys_tas() for CONFIG_SMP && CONFIG_PREEMPT.
Hirokazu Takata [Mon, 18 Oct 2004 16:05:06 +0000 (09:05 -0700)]
[PATCH] m32r: SIO driver
Here is a patch to support the M32R SIO (serial IO) driver.
This driver supports the M32R serial ports.
- Supports two types M32R serial interfaces; M32R_SIO and M32R_PLDSIO.
- With SMP safeness.
Currently the M32R_PLDSIO serial interface, which is implemented on a PLD
on the M3T-M32700UT evaluation board, has slightly different specification
from the integrated peripheral SIO (M32R_SIO). Now we can select them by
CONFIG_ option.
It is a serial-core based driver, based on drivers/serial/8250.c. Any
comments or suggestions will be appreciated.
Hirokazu Takata [Mon, 18 Oct 2004 16:04:53 +0000 (09:04 -0700)]
[PATCH] m32r: AR camera driver
Here is a patch for the Renesas AR camera driver for m32r.
- AR (artificial retina) camera is newly supported.
AR camera module: Renesas M64278E-800, VGA(640x480 pixcels)
http://www.renesas.com/avs/resource/japan/jpn/pdf/assp/rjj01f0005_psmobile.pdf
This patch is required for S3 suspend-resume on noexec capable systems. On
these systems, we need to save and restore MSR_EFER during S3
suspend-resume.
Pavel Machek [Mon, 18 Oct 2004 16:03:14 +0000 (09:03 -0700)]
[PATCH] swsusp: add comments at critical places
apm.c needs save_processor_state and friends. Add a comment to keep people
from removing it. Describe a way to make swsusp work on non-PSE machines.
Document purpose of acpi_restore_state.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Randy Dunlap [Mon, 18 Oct 2004 16:02:50 +0000 (09:02 -0700)]
[PATCH] i386/io_apic init section fixups
Code section errors in i386/io_apic.c found by scripts/reference_init.pl.
Looks like they could cause problems for a few drivers or in a real hotplug
environment.
Error: ./arch/i386/kernel/io_apic.o .text refers to 000018ff R_386_PC32 .init.text
Error: ./arch/i386/kernel/io_apic.o .text refers to 00001967 R_386_PC32 .init.text
(as above thru {A}, then:)
IO_APIC_irq_trigger
irq_trigger
MPBIOS_trigger >> removing __init from this led to
needing to remove __init from
EISA_ELCR also.
Signed-off-by: Randy Dunlap <rddunlap@osdl.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Fix interaction between nosmp and pcibios_fixup_irqs().
When we boot with nosmp we dont have all the mptable info, so
IO_APIC_get_PCI_irq_vector() doesnt work and devices just end up getting a
wrong interrupt.
Suresh B. Siddha [Mon, 18 Oct 2004 16:02:26 +0000 (09:02 -0700)]
[PATCH] Disable SW irqbalance/irqaffinity for E7520/E7320/E7525 v2
As part of the workaround for the "Interrupt message re-ordering across hub
interface" errata (page #16 in
http://developer.intel.com/design/chipsets/specupdt/30288402.pdf), BIOS may
enable hardware IRQ balancing for E7520/E7320/E7525(revision ID 0x9 and
below) based platforms.
Add pci quirks to disable SW irqbalance/affinity on those platforms. Move
balanced_irq_init() to late_initcall so that kirqd will be started after
pci quirks.
Oleg Nesterov [Mon, 18 Oct 2004 16:02:14 +0000 (09:02 -0700)]
[PATCH] Fix show_trace() in irq context with CONFIG_4KSTACKS
- valid_stack_ptr() erroneously assumes that stack always lives in
task_struct->thread_info.
- the main loop in show_trace() does not recalc ebp after stack
switching. With CONFIG_FRAME_POINTER every call to print_context_stack()
will produce the same output.
With this patch, show_trace() does not use task argument in the main loop.
Instead, it converts stack to thread_info* context, and passes it to
print_context_stack() and (implicitly) to valid_stack_ptr().
valid_stack_ptr() now does bounds checking against proper context.
Some cache descriptors are missing from x86_64 table. So instead of
copying from i386 code, here is a patch to share the table between i386 and
x86_64.
Tom Rini [Mon, 18 Oct 2004 16:01:49 +0000 (09:01 -0700)]
[PATCH] sh: fix EMBEDDED_RAMDISK with O=
The following fixes EMBEDDED_RAMDISK to work with O=. The problem was that
we couldn't find the linker script, since we needed to specify the patch to
the source tree for it. I've tested this with the ramdisk set to both
'ramdisk.gz' and '../ramdisk.gz'.
Signed-off-by: Tom Rini <trini@kernel.crashing.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mundt [Mon, 18 Oct 2004 16:00:48 +0000 (09:00 -0700)]
[PATCH] sh: Broken-out CPU subtype probing
Previously we could do subtype parsing and cache configuration in the same
location.. but with the introduction of things like the SH7705 where we use
SH-3 style probing with SH-4 style caches, this is no longer the case. As
such, we move the probe code to a saner place.
Signed-off-by: Paul Mundt <paul.mundt@nokia.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mundt [Mon, 18 Oct 2004 16:00:10 +0000 (09:00 -0700)]
[PATCH] sh: cleanup + merge
This adds other random bits of sh cleanup. This includes Kconfig updates,
some exported symbols to satisfy module builds, cleanup of some whitespace
damage, some compile fixes, and some general header and mach-type cleanup.
Signed-off-by: Tom Rini <trini@kernel.crashing.org> Signed-off-by: Paul Mundt <paul.mundt@nokia.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mundt [Mon, 18 Oct 2004 15:59:19 +0000 (08:59 -0700)]
[PATCH] sh: SCBRR calculation fixes for early printk()
The early printk() code was using a fixed PCLK value that was only sane in the
SH7750 case. This updates the SCBRR value calculation to use
CONFIG_SH_PCLK_FREQ instead and thus works on other subtypes as well (tested
on SH4-202).
Signed-off-by: Paul Mundt <paul.mundt@nokia.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mundt [Mon, 18 Oct 2004 15:59:07 +0000 (08:59 -0700)]
[PATCH] sh: DMA API updates
This updates some of the sh DMA drivers and core API. Previously modules had
to register for the channels they were interested in, but now it's dealt with
transparently by the API with only the number of physical channels needing to
be specified by each module.
Signed-off-by: Paul Mundt <paul.mundt@nokia.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mundt [Mon, 18 Oct 2004 15:58:43 +0000 (08:58 -0700)]
[PATCH] sh: consistent API cleanup
This gets rid of the hardcoded workarounds for the Dreamcast in the
dma-mapping code, and now wraps into the common consistent_alloc() and
consistent_free() routines if the ones in the machvec aren't interested in
handling it.
Signed-off-by: Paul Mundt <paul.mundt@nokia.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mundt [Mon, 18 Oct 2004 15:58:17 +0000 (08:58 -0700)]
[PATCH] sh: SH7705 subtype cleanup + 32k cache support
This fixes up the existing SH7705 support and enables the 32k cache mode for
the processor.
Signed-off-by: Alex Song <songqf9@yahoo.ca> Signed-off-by: Paul Mundt <paul.mundt@nokia.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>