Roland Dreier [Thu, 21 Oct 2004 03:56:14 +0000 (20:56 -0700)]
[PATCH] ppc: fix build with O=$(output_dir)
Recent changes to arch/ppc/boot/lib/Makefile cause
CC arch/ppc/boot/lib/../../../../lib/zlib_inflate/infblock.o
Assembler messages:
FATAL: can't create arch/ppc/boot/lib/../../../../lib/zlib_inflate/infblock.o: No such file or directory
when building a ppc kernel using O=$(output_dir) with CONFIG_ZLIB_INFLATE=n,
because the $(output_dir)/lib/zlib_inflate directory doesn't get created.
This patch, which makes arch/ppc/boot/lib/Makefile create the
directory if needed, is one fix for the problem.
Signed-off-by: Roland Dreier <roland@topspin.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jesse Barnes [Wed, 20 Oct 2004 22:53:23 +0000 (22:53 +0000)]
[IA64-SGI] more sparse I/O accessor fixes
I forgot to add 'const volatile' to the I/O read/write functions in the last
patch, and also forgot to update the _relaxed variants. This patch fixes
that by adding 'const volatile' to the sn2 specific read/write routines as
well as the ia64 machine vector wrappers.
Signed-off-by: Jesse Barnes <jbarnes@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
Jesse Barnes [Wed, 20 Oct 2004 20:40:02 +0000 (20:40 +0000)]
[IA64-SGI] sparse cleanups & misc fixes for sn2
This is a big patch mostly because I trimmed shub_mmr.h down from 17M to 11k
or so. It fixes a number of things sparse discovered and removes some dead
code, fixes up some prototypes, etc. Of note:
o sn_proc_fs.c was directly dereferencing user pointers, fixed
o sn_hwperf.c was missing an include and was using asm-ia64 directly
o the I/O routines were all missing proper sparse annotations
o dead code in prominfo_proc.c has been removed
o fix generic build by putting numionodes into asm/sn/io.h
With this patch applied, the check build is pretty clean. The sn_console bit
depends on some of the other changes, so it's included here.
Signed-off-by: Jesse Barnes <jbarnes@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
John Hawkes [Wed, 20 Oct 2004 18:23:39 +0000 (18:23 +0000)]
[IA64] top level scheduler domain for ia64
Some have noticed that the overlapping sched domains code doesn't quite work
as intended (it results in disjoint domains on some machines), and that a top
level, machine spanning domain is needed. This patch from John Hawkes adds
it to the ia64 code. This allows processes to run on all CPUs in large
systems, though balancing is limited. It should go to Linus soon now
otherwise large systems will only have ~16p (depending on topology) usable by
the scheduler. I sanity checked it on a small system after rediffing John's
original, and he's done some testing on very large systems.
Nick, can you buy off on the sched.c change? Alternatively, do you want to
send that fix separately John? Nick did indeed ACK this change, but it isn't
dependent on this ia64 specific part ... so it's going to be submitted
separately.
Signed-off-by: John Hawkes <hawkes@sgi.com> Signed-off-by: Jesse Barnes <jbarnes@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
Use PIO code from ide-taskfile.c in ide-disk.c so:
* drive status is checked after PIO read
* request is failed if invalid data phase
is detected during PIO write
Russell King [Wed, 20 Oct 2004 21:34:23 +0000 (22:34 +0100)]
[ARM] Add seqlocking to timers.
Sometimes, it's useful to have locking. Especially when we're
talking about time keeping.
It would appear that shemminger's patch of 5th February 2003
completely missed updating _ANY_ ARM timer implementations and,
because linux-arch didn't exist at the time, there appears to
have been no notification to any architecture developer that
maybe, just maybe, some work was required.
One wonders how many other changes are in the kernel which
architecture maintainers have missed.
Russell King [Wed, 20 Oct 2004 15:47:09 +0000 (16:47 +0100)]
[ARM] Cleanup some quirks.
- Ensure FIQs are enabled when cpu_idle() is called.
- Remove unused members of irq_cpustat_t.
- Remove unnecessary #ifndef CONFIG_SMP...#endif around irq_exit()
macro.
- Rename __stf/__clf such that it stresses that they affect only
local state (as per local_irq_xxx).
- Move THREAD_SIZE such that it can be used in current_thread_info()
Suresh B. Siddha [Wed, 20 Oct 2004 06:43:58 +0000 (06:43 +0000)]
[IA64] fallback to swiotlb for consistent DMA mappings
Patch supplied by Suresh Siddha
This is mainly needed for EM64T platforms and makes sense for ia64 too.
Need of this was broughtup sometime(long time?) back on lkml.
http://www.ussg.iu.edu/hypermail/linux/kernel/0406.3/0112.html
Keith Owens [Wed, 20 Oct 2004 06:39:59 +0000 (06:39 +0000)]
[IA64] Avoid a rare deadlock during unwind
There is a rare deadlock condition during unwind script creation. If
build_script() is interrupted in the middle of creating the script, it
holds the script write lock. If the interrupt handler needs to call
unwind for some failure condition, unwind will try to read the
incomplete script and will deadlock on the script lock.
The fix is to disable interrupts while building the script, so
interrupt handlers never see partial scripts.
Promoting spin_lock_irqsave() from script_new() to find_save_locs()
changes the indentation, so the patch looks bigger than it really is.
Signed-off-by: Keith Owens <kaos@sgi.com> Signed-off-by: Tony Luck <tony.luck@intel.com>
The orphan list holds inodes that need to be truncated on recovery. In the
O_DIRECT case, it's used if we extend the inode --- the truncate on recovery
means we'll recover the newly-allocated disk blocks if we crash after the IO
starts but before i_size is updated on disk.
Now, the orphan list is *also* used to delete inodes that are unlinked but
still-open. Those get truncated but also deleted on recovery.
The orphan list is held both in memory and on disk. So the rules are that the
inode can't be reclaimed while on the orphan list. There are only two cases
--- either the inode is actively being written(O_DIRECT) or truncated (in
which case the inode is by definition not going to be reused), or it's
unlinked but still open (again, non-reclaimable).
But in the case where you're truncating or write(O_DIRECT)ing a file that is
*ALSO* unlinked, there's a problem --- the final unlink would put the inode on
the orphan list, but the write/truncate would try to add/remove it. End
result is that the inode disappears from the orphan list while it's still
unlinked-but-in-use.
That's just a leak-on-crash, it's not going to be detectable in normal use.
But it's still a bug, and the way we fix it is for direct-IO and truncate not
to do the ext3_orphan_del if the file is unlinked (ie. i_nlink==0).
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] nfs4 lease: separate the lease processsing code
nfsd will not have a file descriptor, nor an owner on the filp. nfsd also
will not use signals. Seperate the lease processsing coe from
fcntl_setlease() into a __setlease() call.
Signed-off-by: Andy Adamson <andros@citi.umich.edu> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
[PATCH] nfs4 lease: add a lock manager copy lock callback
The following patches provide an interface to the lease subsystem in the
current VFS locking code. NFSv4 delegations and Samba op-locks share most
architecture features. The version 4 NFS server delegation implementation
should use leases to co-ordinate behavior between local, Samba, and NFS
access.
The main design points are
- Seperate the fcntl interface from the file_lock FL_LEASE processing in
fcntl_setlease, creating __setlease() called by fcntl_setlease()
- Add new lock_manager callbacks to enable lease properties to be set,
leases to be broken, and leases to be cleaned up: with default callbacks
preserving the current fcntl_setlease properties.
- Add a new interface, setlease() which also calls __setlease(), and
remove_lease() for kernel lease managers (e.g. the v4 NFS server)
This patch:
Add a lock manager copy lock callback to locks_copy_lock() so that nfsd can
set lease properties.
Signed-off-by: Andy Adamson <andros@citi.umich.edu> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Mahoney [Wed, 20 Oct 2004 01:43:11 +0000 (18:43 -0700)]
[PATCH] reiserfs: allow user_xattr and acl options to be ignored, with warning
This patch uses the REISERFS_UNSUPPORTED_OPT flag to denote -o(no)acl, and
-o(no)user_xattr as unsupported, but allowable, when support isn't built
into the kernel.
Signed-off-by: Jeff Mahoney <jeffm@novell.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Mahoney [Wed, 20 Oct 2004 01:42:56 +0000 (18:42 -0700)]
[PATCH] reiserfs: support for REISERFS_UNSUPPORTED_OPT notation
This patch adds a REISERFS_UNSUPPORTED_OPT flag to denote when a mount
option is allowable, but is unsupported in the running configuration. This
allows the potential for the set of mount options to be consistent,
regardless of what features the kernel is compiled with.
Rather than failing the mount, a warning is issued and the mount succeeds.
Signed-off-by: Jeff Mahoney <jeffm@novell.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Kenneth W. Chen [Wed, 20 Oct 2004 01:42:41 +0000 (18:42 -0700)]
[PATCH] Enable config_schedstats for all arches
Config option CONFIG_SCHEDSTATS is currently enabled via arch specific
Kconfig.debug. Only x86 and ppc arches has code to turn it on. Why not
put it in generic lib/Kconfig.debug so it is done once to enable everyone?
Signed-off-by: Ken Chen <kenneth.w.chen@intel.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Dean Gaudet [Wed, 20 Oct 2004 01:41:28 +0000 (18:41 -0700)]
[PATCH] transmeta efficeon support and cpuid update
This patch adds efficeon as a cpu option, and makes a small update to the
transmeta cpuid code. (i wasn't sure if the various doc files are UTF-8...
if they are, then the e should be a U-275 ;)
The compile options may not be ideal, but they're probably close. i used
-march=pentium3, but -march=pentium4 would have been good enough too.
The cpuid update teaches transmeta.c about the extended processor revision
present in cpuid level 0x80860002... the external documentation does not
indicate how to break apart this field, and instructs only that the 32-bit
value should be printed in hex (alas).
Signed-off-by: dean gaudet <dean@arctic.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
I've noticed that under specific circumstances the "console=" kernel
parameter is ignored. This happens when EARLY_PRINTK is enabled and the
serial console is the only available. In this case unregister_console()
when called for the early console sets preferred_console back to -1
replacing the value that was recorded by console_setup() -- the order of
calls is as follows:
1. register_console() -- for the early console,
2. console_setup() -- recording the console index for the real console,
3. unregister_console() -- for the early console, erasing the console
index recorded above,
4. register_console() -- for the real console, picking up the first device
available, instead of the selected one.
I've observed this problem with a DECstation system using ttyS3 -- its
default console device from the firmware's point of view.
The solution is to restore the setting of "console=" upon
unregister_console(). This made a snapshot of 2.4.26 work for me. I
wasn't able to test the changes with 2.6 because DECstation drivers don't
support it yet, but the code responsible for console selection appears
functionally the same. So I've concluded it needs the same change. Here's
a patch.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ed Schouten [Wed, 20 Oct 2004 01:40:45 +0000 (18:40 -0700)]
[PATCH] lockd: remove hardcoded maximum NLM cookie length
At the moment, the NLM cookie length is fixed to 8 bytes, while 1024 is the
theoretical maximum. FreeBSD uses 16 bytes, Mac OS X uses 20 bytes.
Therefore we need to make the length dynamic (which I set to 32 bytes).
This patch is based on an old patch for Linux 2.4.23-pre9, which I changed
to patch properly (also added some stylish NIPQUAD fixes).
From: Neil Brown <neilb@cse.unsw.edu.au>
Further lockd tidyups.
- NIPQUAD everywhere that is appropriate
- use XDR_QUADLEN in more places as appropriate
- discard QUADLEN which is a lockd-specific version of XDR_QUADLEN
Signed-off-by: Neil Brown <neilb@cse.unsw.edu.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Manfred Spraul [Wed, 20 Oct 2004 01:40:18 +0000 (18:40 -0700)]
[PATCH] slab: reduce fragmentation due to kmem_cache_alloc_node
Attached is a patch that fixes the fragmentation that Badri noticed with
kmem_cache_alloc_node.
kmem_cache_alloc_node tries to allocate memory from a given node. The
current implementation contains two bugs:
- the node aware code was used even for !CONFIG_NUMA systems. Fix:
inline function that redefines kmem_cache_alloc_node as kmem_cache_alloc
for !CONFIG_NUMA.
- the code always allocated a new slab for each new allocation. This
caused severe fragmentation. Fix: walk the slabp lists and search for a
matching page instead of allocating a new page.
- the patch also adds a new statistics field for node-local allocs. They
should be rare - the codepath is quite slow, especially compared to the
normal kmem_cache_alloc.
Marcelo Tosatti [Wed, 20 Oct 2004 01:40:03 +0000 (18:40 -0700)]
[PATCH] Remove redundant AND from swp_type()
There is a useless AND in swp_type() function.
We just shifted right SWP_TYPE_SHIFT() bits the value from the swp_entry_t,
and then we AND it with "(1 << 5) - 1" (which is a mask corresponding to
the number of bits used by "type").
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
/proc shows the wrong PID as parent in the following case
Process A creates Threads 1 & 2 (using pthread_create) Thread 2 then forks
and execs process B getppid() for Process B shows Process A (rightly) as
parent, however /proc/B/status shows Thread 3 as PPid (incorrect).
Use, in the rb_entry definition, the container_of macro instead of
reinventing the wheel; compared to using offset_of() as I did in the prev.
version, it has type safety checking.
Signed-off-by: Paolo 'Blaisorblade' Giarrusso <blaisorblade_spam@yahoo.it> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Dell's Virtual Floppy (system management presents to the local system an
IDE floppy device, which is actually a floppy device in a remote system
connected over an IP link) exhibits this also, when connecting to a remote
floppy drive with no media present.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Hideo Aoki [Wed, 20 Oct 2004 01:36:51 +0000 (18:36 -0700)]
[PATCH] proc.txt cleanup
In Documentation/filesystems/proc.txt, explanation of /proc/meminfo is
described in section 1.3 (IDE devices in /proc/ide). I think that it
should be described in section 1.2 (Kernel data).
Hideo Aoki [Wed, 20 Oct 2004 01:36:36 +0000 (18:36 -0700)]
[PATCH] vm thrashing control tuning
This patch adds "swap_token_timeout" parameter in /proc/sys/vm. The
parameter means expired time of token. Unit of the value is HZ, and the
default value is the same as current SWAP_TOKEN_TIMEOUT (i.e. HZ * 300).
Matt Domsch [Wed, 20 Oct 2004 01:36:22 +0000 (18:36 -0700)]
[PATCH] EDD: use EXTENDED READ command, add CONFIG_EDD_SKIP_MBR
Some controller BIOSes have problems with the legacy int13 fn02 READ
SECTORS command. int13 fn42 EXTENDED READ is used in preference by most
boot loaders today, so lets use that. If EXTENDED READ fails or isn't
supported, fall back to READ SECTORS.
This hopefully resolves the three reports of BIOSes which would either
long-pause (30+ seconds) or hang completely on the legacy READ SECTORS
command.
This also adds CONFIG_EDD_SKIP_MBR to eliminate reading the MBR on each
BIOS-presented disk, in case there are further problems in this area.
Signed-off-by: Matt Domsch <Matt_Domsch@dell.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ryan S. Arnold [Wed, 20 Oct 2004 01:36:07 +0000 (18:36 -0700)]
[PATCH] hvc_console fix to prevent oops and late hangup and write operations
This patch prevents execution of hvc_write() and hvc_hangup() after the tty
layer has executed a final hvc_close() against a device. This patch
provides a better method than was previously used. tty->driver_data is no
longer invalidated so we'll no longer get oopses when the tty layer allows
late hangup() and write() operations.
- Removed silly tty->driver_data = NULL; from hvc_close which prevents
possible oops in hvc_write() and hvc_hangup() due to improperly acting
ldisc close ordering.
- Added hp->count <= 0 check to hvc_write() and hvc_hangup() to prevent
execution of these function after hvc_close() has been invoked by the tty
layer. Same tty ldisc issues as above are the reason.
- Added some comments to clarify the situation.
- Awaiting a forth coming patch from Alan Cox which should clean up the
close ordering and prevent the late hangup and write ops from happening.
Signed-off-by: Ryan S. Arnold <rsa@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>