Signed-off-by: Domen Puncer <domen@coderock.org> Signed-off-by: Maximilian Attems <janitor@sternwelten.at> Signed-off-by: David S. Miller <davem@davemloft.net>
Use "ifdef" rather than "if" to test for __KERNEL__
Both work, but the latter can cause warnings in user space
from compilers that don't like using undefined identifiers
in preprocessor expressions (quite reasonable).
David S. Miller [Wed, 1 Sep 2004 08:20:12 +0000 (01:20 -0700)]
[SPARC64]: Zap pci_controller_lock.
It is only taken during boot time bus probe, thus
protects nothing at run time and causes bogus bug
messages when PREEMPT is enabled. When we support
PCI controller hot plug we will add a suitable locking
mechanism.
Signed-off-by: David S. Miller <davem@davemloft.net>
Signed-off-by: Domen Puncer <domen@coderock.org> Signed-off-by: Maximilian Attems <janitor@sternwelten.at> Signed-off-by: David S. Miller <davem@davemloft.net>
Tom Rini [Wed, 1 Sep 2004 05:26:04 +0000 (22:26 -0700)]
[PATCH] ppc32: fix the 'checkbin' target
The checkbin target on PPC32 isn't quite right.
First, one of the tests (to ensure that some instructions are known to
gas) is never actually invoked because 'checkbin' doesn't know about
stuff set in .config, so we always have the 'else' case run. This
changes to always running the test and telling the user to upgrade to at
least binutils 2.12.1.
The next problem is that we were doing $(AS) -o /dev/null ... in both
that test, as well as another. The problem here is that the checkbin
target is run on the install targets, meaning that /dev/null will get
unlinked when the test passes. To get around this we use .tmp_gas_check
as the output file instead.
Acked by Sam.
Signed-off-by: Tom Rini <trini@kernel.crashing.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Alexander Viro [Wed, 1 Sep 2004 05:18:31 +0000 (22:18 -0700)]
[PATCH] nfs ->follow_link() switched to new scheme
NFS takes some thought to switch to the new symlink scheme, because we
can't rely on the pagecache lookup to find the symlink page when freeing
it - the cache might have been invalidated in the meantime.
So we hide the page information in the symlink data area itself,
by stealing the last pointer in the page used for the cache. That
way nfs_put_link() can just look up the page directly.
Signed-off-by: Al Viro <viro@parcelfarce.linux.org.uk> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Alexander Viro [Tue, 31 Aug 2004 14:46:25 +0000 (07:46 -0700)]
[PATCH] reduce stack use in altroot handling
Massaged altroot handling to avoid on-stack struct nameidata instance (and
got it faster, actually). We are in the middle of do_follow_link() recursion
here, so the stack footprint is critical.
Signed-off-by: Al Viro <viro@parcelfarce.linux.org.uk> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Alexander Viro [Tue, 31 Aug 2004 09:41:25 +0000 (02:41 -0700)]
[PATCH] usx2y cleanups and fixes
Sigh...
a) mixing of userland and kernel pointers is bad
b) so's not checking result of kmalloc()
c) so's not checking result of copy_from_user()
d) use of do { .... break; ... break; ... } while(0); is *highly*
unidiomatic. Do not confuse kernel with IOCCC, please. And if you have
religious aversion to multiple return statements in a function, at least
learn the reasons why it is frowned upon in many situations. Hint: they
all apply to use of break in that manner.
e) 0 instead of NULL
Signed-off-by: Al Viro <viro@parcelfarce.linux.org.uk> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Alexander Viro [Tue, 31 Aug 2004 09:41:13 +0000 (02:41 -0700)]
[PATCH] alpha warning fixes
pci_dma_sync_single_for_device() had wrong prototype [who TF had come up
with that name, anyway?]
->cpu in thread_info was long; it should be unsigned int.
Signed-off-by: Al Viro <viro@parcelfarce.linux.org.uk> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
zfcp host adapater changes:
- Add ability to enqueue other WKA ports besides the nameserver port.
- Document and cleanup sg_list functions.
- Add get_port_by_did/get_adapater_by_busid functions.
- Improve documentation of some functions and structures.
- Fix error handling for nameserver requests.
- Correct size check in zfcp_sg_list_copy_to_user.
- Correct parameter description for loglevel parameter.
- Remove unsused code, types and definitions.
- Add support for exchange_port_data command.
- Add infrastructure to set timers for ELS and SCSI commands.
- Avoid adapter shutdown after receiving FSF_SQ_ULP_PROGRAMMING_ERROR.
Signed-off-by: Martin Schwidefsky <schwidefsky@de.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
This adds support for the new compiler options -mkernel-backchain,
-mstack-size, -mstack-guard, -mwarn-dynamicstack and -mwarn-framesize.
The option -mkernel-backchain enables the use of modified layout for the
stack frames of kernel functions. This breaks the ABI, modules compiled
with the option won't work on a kernel compiled with the option and vice
versa. The positive effect of the option is a drastic reduction of kernel
stack use. The trick is that the new frame layout allows to overlap the 96
(31 bit)/160 (64 bit) byte bias areas of the functions on the call chain.
This lowers the minimal stack usage of a function from 96 bytes to 16 bytes
(31 bit) and 160 bytes to 24 bytes (64 bit). The kernel stack use is
decreased to a point where it is possible to use 4K (31 bit) / 8K (64 bit)
stacks. The split into process stack and interrupt stack is already in
place.
The options -mstack-size and -mstack-guard are used to detect kernel stack
overflows. The compiler adds code to the prolog of every function that
causes an illegal operation if the kernel stack is about to overflow.
The options -mwarn-dynamicstack and -mwarn-framesize cause the compiler to
emit warnings if a function uses dynamic stack allocation or if the
function frame size is bigger then a specified limit.
To play safe all the new options are configurable.
Signed-off-by: Martin Schwidefsky <schwidefsky@de.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
s390 core changes:
- Fix a race condition between kernel thread creation and preemption.
- Fix idal_is_needed for the border case 0x7ffff000.
- Get rid of compiler warnings in compat_signal.c and profile.c.
- Regenerate default configuration.
Signed-off-by: Martin Schwidefsky <schwidefsky@de.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jacek Poplawski [Tue, 31 Aug 2004 03:42:33 +0000 (20:42 -0700)]
[PATCH] stv0299 device naming fix
Name of device has been changed in 2.6.9-rc1 to "SkyStar2", but module stv0299
still compares name with "Technisat SkyStar2 driver", strings are different,
and result is that stv0299 detects invalid tuner type.
Cc: Johannes Stezenbach <js@linuxtv.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Gerd Knorr [Tue, 31 Aug 2004 03:42:07 +0000 (20:42 -0700)]
[PATCH] v4l/bttv: add sanity check (bug #3309)
Missing sanity check, overlay is supported for packed pixel formats only.
Patch below. It's not API related btw, the bug can be triggered using the
v4l2 API as well.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Antonino Daplas [Tue, 31 Aug 2004 03:41:55 +0000 (20:41 -0700)]
[PATCH] fbdev: fix copy_to/from_user in fbmem.c:fb_read/write
This patch fixes a problem reported by David S. Miller <davem@redhat.com>
"I just noticed that fb_{read,write}() uses copy_*_user() with
the kernel buffer being the frame buffer. It needs to use
the proper device address accessor functions."
The patch will do an intermediate copy of the contents to a page-sized,
kmalloc'ed buffer.
Antonino Daplas [Tue, 31 Aug 2004 03:41:42 +0000 (20:41 -0700)]
[PATCH] fbdev: fix kernel panic from FBIO_CURSOR ioctl
1. This fixes a kernel oops when issuing an FBIO_CURSOR ioctl if struct
fb_cursor_user is filled with zero/NULLs. Reported by Yuval Kogman
<nothingmuch@woobling.org>.
2. This also fixes the cursor corruption in soft_cursor when
sprite.scan_align != 1.
It's already marked BROKEN_ON_SMP, but even a UP compile yields tons of
errors. While those aren't deeply complicated to fix having them for over
a year now is a pretty good indicator no one cares.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
attached patch changes the subfrequency carrier value in the adv7175
video output driver which is part of the zr36067 driver package. The
practical consequence is that the picture will be more stable on
non-passthrough video mode in NTSC. It does not affect PAL/SECAM. Patch
originally submitted by Douglas Fraser <ds-fraser@comcast.net> (8/21).
Signed-off-by: Ronald Bultje <rbultje@ronald.bitfreak.net> Signed-off-by: Douglas Fraser <ds-fraser@comcast.net> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ronald Bultje [Tue, 31 Aug 2004 03:40:19 +0000 (20:40 -0700)]
[PATCH] zr36067 driver - use msleep() instead of schedule_timeout()
attached patch makes the zr36067 driver use msleep() instead of
schedule_timeout() with uninterruptible state. Patch originally
submitted by Nishanth Aravamudan <nacc@us.ibm.com> (7/26).
Signed-off-by: Ronald Bultje <rbultje@ronald.bitfreak.net> Signed-off-by: Nishanth Aravamudan <nacc@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ronald Bultje [Tue, 31 Aug 2004 03:40:07 +0000 (20:40 -0700)]
[PATCH] zr36067 driver - correct i2c-algo-bit dependency in Kconfig
attached patch correctly makes the zr36067 driver depend on i2c-ago-bit in
the kernel config. Bug reported and patch sent to me by Adrian Bunk
<bunk@fs.tum.de> (6/21). It wasn't signed off.
Signed-off-by: Ronald Bultje <rbultje@ronald.bitfreak.net> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Ryan S. Arnold [Tue, 31 Aug 2004 03:39:20 +0000 (20:39 -0700)]
[PATCH] interrupt driven hvc_console as vio device
This is an hvc_console patch which provides driver and ppc64 architecture
fixes to enable the hvc_console driver to register itself as a vio device
with the vio bus, provide hotplug add/remove for vty adapters, and act as
an interrupt driven driver on Power-5 hardware or remain as a polling
driver on Power-4 hardware.
- Changed hvc_get_chars() and hvc_put_chars() api to take vtermno rather
than index number.
- Added hvc_find_vtys() function which walks the bus looking for
vterm/vty devices to callback to the hvc_console driver. This provides
console output functionality prior to early console init (pre mem init
and pre device probe).
- Switch khvcd from kernel_threads to kthreads which got rid of
deprecated daemonize().
- Added module exit clause to be thorough (not terribly necessary with a
console driver of course)
- Added early discovery of vterm/vty adapters by doing a bus walk on
early console init which results in hvc_instantiate() callback and
addition of the vtermno into a static array of vtermnos supported as
console adapters (meaning the console api's work against these vtermnos
prior to full console initialization).
- This driver is now registered as a vio driver which means that vty
adapters are now managed via probe/remove. This means hvc_console
supports hotplug vty adapters.
- Driver now requests more device nodes than what was found on the
initial bus walk when registered as a tty driver to make room for hotplug
vty adapters. These secondary vty adapters provide a tty tunnel between
partitions.
- Removed static hvc_struct array and replaced with a linux list that has
elements (hvc_struct instances) added/removed on probe/remove AFTER early
console init. This is important because kmalloc can't be done at early
console init.
- Driver now either runs in interrupt driven mode or in polling mode on
older hardware. The khvcd is smart enough to not 'schedule()' when there
are no interrupts.
- kobjects are now used for ref counting on the hvc_struct instances.
- This driver puts the tty layer to sleep on hvc_close() if there are
pending data writes being blocked by firmware.
- Removed useless spinlocks in hvc_chars_in_buffer() and hvc_write_room.
Signed-off-by: Benjamin Herrenschmidt <benh@kernel.crashing.org> Signed-off-by: Ryan S. Arnold <rsa@us.ibm.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
David Howells [Tue, 31 Aug 2004 03:38:57 +0000 (20:38 -0700)]
[PATCH] Fix a NULL pointer bug in do_generic_file_read()
The attached patch fixes a bug introduced into do_generic_mapping_read() by
which a file pointer becomes required. I'd arranged things so that the
file pointer was optional so that I could call the function directly on an
inode.
Signed-Off-By: David Howells <dhowells@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
H. Peter Anvin [Tue, 31 Aug 2004 03:38:44 +0000 (20:38 -0700)]
[PATCH] Make i386 signal delivery work with -mregparm
This patch allows i386 signal delivery to work correctly when userspace is
compiled with -mregparm. This is somewhat hacky: it passes the arguments
*both* on the stack and in registers, but it works because there are only
one or three (depending on SA_SIGINFO) official arguments. If you're
relying on the unofficial arguments then you're doing something nonportable
anyway and can put in the __attribute__((cdecl,regparm(0))) in the correct
place.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Jeff Moyer [Tue, 31 Aug 2004 03:38:32 +0000 (20:38 -0700)]
[PATCH] netpoll: fix up trapped logic
This patch contains the updates necessary to fix the hangs in netconsole.
This includes the changing of trapped to an atomic_t, and the addition of a
netpoll_poll_lock. It also turns dev->netpoll_rx into a bitfield which is
used to keep from running the networking code from the netpoll_poll call path.
Signed-off-by: Jeff Moyer <jmoyer@redhat.com> Signed-off-by: Matt Mackall <mpm@selenic.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Matt Mackall [Tue, 31 Aug 2004 03:37:55 +0000 (20:37 -0700)]
[PATCH] netpoll: revert queue stopped change
Here's the first of the broken out patch set. This puts the check for
netif_queue_stopped back into netpoll_send_skb. Network drivers are not
designed to have their hard_start_xmit routines called when the queue is
stopped.
Signed-off-by: Jeff Moyer <jmoyer@redhat.com> Signed-off-by: Matt Mackall <mpm@selenic.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andy Whitcroft [Tue, 31 Aug 2004 03:37:17 +0000 (20:37 -0700)]
[PATCH] use page_to_nid
There are a couple of places where we seem to go round the houses to get
the numa node id from a page. We have a macro for this so it seems
sensible to use that.
Both lookup_node and enqueue_huge_page use page_zone() to locate the zone,
that to locate node pgdat_t and that to get the node_id. Its more
efficient to use page_to_nid() which gets the nid from the page flags,
especially if we are not using the zone for anything else it. Change these
to use page_to_nid().
Signed-off-by: Andy Whitcroft <apw@shadowen.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Andy Whitcroft [Tue, 31 Aug 2004 03:37:03 +0000 (20:37 -0700)]
[PATCH] i386 bootmem restrictions
(Comment changes only)
The bootmem allocator is initialised before the kernel virtual address
space has been fully established. As a result, any allocations which are
made before paging_init() has completed may point to invalid kernel
addresses. This patch notes this limitation and indicates where the
allocator is fully available.
Signed-off-by: Andy Whitcroft <apw@shadowen.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
The get_node function does a lookup on /dev/pts/<number> and returns the
dentry, taking a reference. We should dput the dentry after extracting the
tty pointer.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Joanne Dow [Tue, 31 Aug 2004 03:36:26 +0000 (20:36 -0700)]
[PATCH] Amiga partition reading fix
I have a large archive of files stored on Amiga volumes. Many of these
volumes are on Fujitsu magneto-optical disks with 2k sector size. The
existing partitioning code cannot properly read them since it appears the
OS automatically deblocks the large sectors into logical 512 byte sectors,
something AmigaDOS never did. I arranged the partitioning code to handle
this situation.
Second I have some rather strange test case disks, including my largest
storage partition, that have somewhat unusual partition values. As such I
needed additional information in addition to the first and last block
number information. AmigaDOS reserves N blocks, with N greater than or
equal to 1 and less than the size of the partition, for some boot time
information and signatures. I have some partitions that use other than the
usual value of 2.
There is one more "fix" that could be put in if someone needs it. Another
value in the "Rigid Disk Blocks" description of a partition is a "PreAlloc"
value. It defines a number of blocks at the end of the disk that are not
considered to be a real part of the partition. This was "important" in the
days of 20 meg and 40 meg hard disks. It is hardly important and not used
on modern drives without special user intervention.
This partitioning information is known correct. I wrote the low level
portion of the hard disk partitioning code for AmigaDOS 3.5 and 3.9. I am
also responsible for one of the more frequently used partitioning tools,
RDPrepX, before that.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Tue, 31 Aug 2004 03:36:01 +0000 (20:36 -0700)]
[PATCH] use hlist for pid hash
Use hlists for the PID hashes. This halves the memory footprint of these
hashes. No benchmarks, but I think this is a worthy improvement because
the hashes are something that would be likely to have significant portions
loaded into the cache of every CPU on some workloads.
This comes at the "expense" of
1. reintroducing the memory prefetch into the hash traversal loop;
2. adding new pids to the head of the list instead of the tail. I
suspect that if this was a big problem then the hash isn't sized
well or could benefit from moving hot entries to the head.
Also, account for all the pid hashes when reporting hash memory usage.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Nick Piggin [Tue, 31 Aug 2004 03:35:49 +0000 (20:35 -0700)]
[PATCH] fix PID hash sizing
A 4GB, 4-way Opteron would create the smallest size table (16 entries) because
pidhash_init is called before mem_init which is where x86-64 sets up max_pfn.
nr_kernel_pages is setup by paging_init, called from setup_arch, which is also
where i386 sets up max_pfn.
So export nr_kernel_pages, nr_all_pages. Use nr_kernel_pages when sizing the
PID hash. This fixes the problem.
This also makes the pid hash dependant on the size of ZONE_NORMAL instead of
total size of memory.
Signed-off-by: Nick Piggin <nickpiggin@yahoo.com.au> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Tue, 31 Aug 2004 03:35:38 +0000 (20:35 -0700)]
[PATCH] fix rusage semantics
This patch changes the rusage bookkeeping and the semantics of the
getrusage and times calls in a couple of ways.
The first change is in the c* fields counting dead child processes. POSIX
requires that children that have died be counted in these fields when they
are reaped by a wait* call, and that if they are never reaped (e.g.
because of ignoring SIGCHLD, or exitting yourself first) then they are
never counted. These were counted in release_task for all threads. I've
changed it so they are counted in wait_task_zombie, i.e. exactly when
being reaped.
POSIX also specifies for RUSAGE_CHILDREN that the report include the reaped
child processes of the calling process, i.e. whole thread group in Linux,
not just ones forked by the calling thread. POSIX specifies tms_c[us]time
fields in the times call the same way. I've moved the c* fields that
contain this information into signal_struct, where the single set of
counters accumulates data from any thread in the group that calls wait*.
Finally, POSIX specifies getrusage and times as returning cumulative totals
for the whole process (aka thread group), not just the calling thread.
I've added fields in signal_struct to accumulate the stats of detached
threads as they die. The process stats are the sums of these records plus
the stats of remaining each live/zombie thread. The times and getrusage
calls, and the internal uses for filling in wait4 results and siginfo_t,
now iterate over the threads in the thread group and sum up their stats
along with the stats recorded for threads already dead and gone.
I added a new value RUSAGE_GROUP (-3) for the getrusage system call rather
than changing the behavior of the old RUSAGE_SELF (0). POSIX specifies
RUSAGE_SELF to mean all threads, so the glibc getrusage call will just
translate it to RUSAGE_GROUP for new kernels. I did this thinking that
someone somewhere might want the old behavior with an old glibc and a new
kernel (it is only different if they are using CLONE_THREAD anyway).
However, I've changed the times system call to conform to POSIX as well and
did not provide any backward compatibility there. In that case there is
nothing easy like a parameter value to use, it would have to be a new
system call number. That seems pretty pointless. Given that, I wonder if
it is worth bothering to preserve the compatible RUSAGE_SELF behavior by
introducing RUSAGE_GROUP instead of just changing RUSAGE_SELF's meaning.
Comments?
I've done some basic testing on x86 and x86-64, and all the numbers come
out right after these fixes. (I have a test program that shows a few
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Roland McGrath [Tue, 31 Aug 2004 03:35:25 +0000 (20:35 -0700)]
[PATCH] waitid system call
This patch adds a new system call `waitid'. This is a new POSIX call that
subsumes the rest of the wait* family and can do some things the older
calls cannot. A minor addition is the ability to select what kinds of
status to check for with a mask of independent bits, so you can wait for
just stops and not terminations, for example. A more significant
improvement is the WNOWAIT flag, which allows for polling child status
without reaping. This interface fills in a siginfo_t with the same details
that a SIGCHLD for the status change has; some of that info (e.g. si_uid)
is not available via wait4 or other calls.
I've added a new system call that has the parameter conventions of the
POSIX function because that seems like the cleanest thing. This patch
includes the actual system call table additions for i386 and x86-64; other
architectures will need to assign the system call number, and 64-bit ones
may need to implement 32-bit compat support for it as I did for x86-64.
The new features could instead be provided by some new kludge inventions in
the wait4 system call interface (that's what BSD did). If kludges are
preferable to adding a system call, I can work up something different.
I added a struct rusage field si_rusage to siginfo_t in the SIGCHLD case
(this does not affect the size or layout of the struct). This is not part
of the POSIX interface, but it makes it so that `waitid' subsumes all the
functionality of `wait4'. Future kernel ABIs (new arch's or whatnot) can
have only the `waitid' system call and the rest of the wait* family
including wait3 and wait4 can be implemented in user space using waitid.
There is nothing in user space as yet that would make use of the new field.
Most of the new functionality is implemented purely in the waitid system
call itself. POSIX also provides for the WCONTINUED flag to report when a
child process had been stopped by job control and then resumed with
SIGCONT. Corresponding to this, a SIGCHLD is now generated when a child
resumes (unless SA_NOCLDSTOP is set), with the value CLD_CONTINUED in
siginfo_t.si_code. To implement this, some additional bookkeeping is
required in the signal code handling job control stops.
The motivation for this work is to make it possible to implement the POSIX
semantics of the `waitid' function in glibc completely and correctly. If
changing either the system call interface used to accomplish that, or any
details of the kernel implementation work, would improve the chances of
getting this incorporated, I am more than happy to work through any issues.
Signed-off-by: Roland McGrath <roland@redhat.com> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Anton Blanchard [Tue, 31 Aug 2004 03:35:11 +0000 (20:35 -0700)]
[PATCH] Using get_cycles for add_timer_randomness
I tested how long it took to do a dd from /dev/random on ppc64 before and
after this patch, while doing a ping flood from another machine.
before:
# /usr/bin/time dd if=/dev/random of=/dev/zero count=1k
0+51 records in
Command terminated by signal 2
0.00user 0.00system 19:18.46elapsed 0%CPU (0avgtext+0avgdata 0maxresident)k
I gave up after 19 minutes.
after:
# /usr/bin/time dd if=/dev/random of=/dev/zero count=1k
0+1024 records in
0.00user 0.00system 0:33.38elapsed 0%CPU (0avgtext+0avgdata 0maxresident)k
Just over 33 seconds. Better.
From: Arnd Bergmann <arnd@arndb.de>
I noticed that only i386 and x86-64 are currently using a high resolution
timer source when adding randomness. Since many architectures have a
working get_cycles() implementation, it seems rather straightforward to use
that.
Has this been discussed before, or can anyone comment on the implementation
below?
This patch attempts to take into account the size of cycles_t, which is
either 32 or 64 bits wide but independent of the architecture's word size.
The behavior should be nearly identical to the old one on i386, x86-64 and
all architectures without a time stamp counter, while finding more entropy
on the other architectures.
Signed-off-by: Anton Blanchard <anton@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Alan Cox [Tue, 31 Aug 2004 03:34:34 +0000 (20:34 -0700)]
[PATCH] VLAN support for 3c59x/3c90x
This adds VLAN support to the 3c59x/90x series hardware.
Stefan de Konink ported this code from the 2.4 VLAN patches and tested it
extensively. I cleaned up the ifdefs and fixed a problem with bracketing
that made older cards fail.
--
Developer's Certificate of Origin 1.0
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I have the
right to submit it under the open source license indicated in the file; or
(b) The contribution is based upon previous work that, to the best of my
knowledge, is covered under an appropriate open source license and I have
the right under that license to submit that work with modifications,
whether created in whole or in part by me, under the same open source
license (unless I am permitted to submit under a different license), as
indicated in the file; or
(c) The contribution was provided directly to me by some other person who
certified (a), (b) or (c) and I have not modified it.
I, Stefan de Konink, certify that:
The contribution is based upon previous work that, again is based on GPL
code and I have the right under that license to submit that work with
modifications, whether created in whole or in part by me, under the same
open source license.
I, Alan Cox, certify likewise.
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Oleg Nesterov [Tue, 31 Aug 2004 03:34:10 +0000 (20:34 -0700)]
[PATCH] hugetlbfs private mappings
Hugetlbfs silently coerce private mappings of hugetlb files into shared
ones. So private writable mapping has MAP_SHARED semantics. I think, such
mappings should be disallowed.
First, such behavior allows open hugetlbfs file O_RDONLY, and overwrite it
via mmap(PROT_READ|PROT_WRITE, MAP_PRIVATE), so it is security bug.
Second, private writable mmap() should fail just because kernel does not
support this.
I belisve, it is ok to allow private readonly hugetlb mappings,
sys_mprotect() does not work with hugetlb vmas.
There is another problem. Hugetlb mapping is always prefaulted, pages
allocated at mmap() time. So even readonly mapping allows to enlarge the
size of the hugetlbfs file, and steal huge pages without appropriative
permissions.
Oleg Nesterov [Tue, 31 Aug 2004 03:33:57 +0000 (20:33 -0700)]
[PATCH] /dev/zero vs hugetlb mappings.
Hugetlbfs mmap with MAP_PRIVATE becomes MAP_SHARED silently, but
vma->vm_flags have no VM_SHARED bit. Reading from /dev/zero into hugetlb
area will do:
read_zero()
read_zero_pagealigned()
if (vma->vm_flags & VM_SHARED)
break; // fallback to clear_user()
zap_page_range();
zeromap_page_range();
It will hit BUG_ON() in unmap_hugepage_range() if region is not huge page
aligned, or silently convert it into the private anonymous mapping.
Paul Mackerras [Tue, 31 Aug 2004 03:33:11 +0000 (20:33 -0700)]
[PATCH] ppc64: rework PPC64 cpu map setup
Move all cpu map initializations to one place (except for the online map --
cpus mark themselves online as they come up). This sets up
cpu_possible_map early enough that we can use num_possible_cpus for
allocating irqstacks instead of NR_CPUS. Hopefully this should also help
set the stage for kexec.
Signed-off-by: Nathan Lynch <nathanl@austin.ibm.com> Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Paul Mackerras [Tue, 31 Aug 2004 03:32:58 +0000 (20:32 -0700)]
[PATCH] Update PPC MAINTAINERS & CREDITS
David Engebretsen has moved on to other things and is no longer maintaining
ppc64. This patch adds an entry in CREDITS to note his contribution in
leading the team that did the PPC64 port originally and updates various
PPC-related MAINTAINERS entries.
Signed-off-by: Paul Mackerras <paulus@samba.org> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>
Takashi Iwai [Tue, 31 Aug 2004 03:32:46 +0000 (20:32 -0700)]
[PATCH] Fix the unnecessary entropy call in the irq handler
Currently add_interrupt_randomness() is called at each interrupt when one
of the handlers has SA_SAMPLE_RANDOM flag, regardless whether the interrupt
is processed by that handler or not. This results in the higher latency
and perfomance loss.
The patch fixes this behavior to avoid the unnecessary call by checking the
return value from each handler.
[PATCH] Jumper Probes to provide function arguments
A special kprobe type which can be placed on function entry points, and
employs a simple mirroring principle to allow seamless access to the
arguments of a function being probed. The probe handler routine should
have the same prototype as the function being probed. Currently
implemented for x86.
The way it works is that when the probe is hit, the breakpoint handler
simply irets to the probe handler's eip while retaining register and stack
state corresponding to the function entry. After it is done, the probe
handler calls jprobe_return() which traps again to restore processor state
and switch back to the probed function. Linus noted correctly at KS that
we need to be careful as gcc assumes that the callee owns arguments. We
save and restore enough stack bytes to cover argument space.
Sample Usage:
static int jip_queue_xmit(struct sk_buff *skb, int ipfragok)
{
... whatever ...
jprobe_return();
return 0;
}
This patch helps developers to trap at almost any kernel code address,
specifying a handler routine to be invoked when the breakpoint is hit.
Useful for analysing the Linux kernel by collecting debugging information
non-disruptively. Employs single-stepping out-of-line to avoid probe
misses on SMP and may be especially useful in aiding debugging elusive
races and problems on live systems. More elaborate dynamic tracing tools
such as DProbes can be built over the kprobes interface.
Helps developers to trap at almost any kernel code address, specifying a
handler routine to be invoked when the breakpoint is hit. Useful for
analysing the Linux kernel by collecting debugging information
non-disruptively. Employs single-stepping out-of-line to avoid probe
misses on SMP and may be especially useful in aiding debugging elusive
races and problems on live systems. More elaborate dynamic tracing tools
such as DProbes can be built over the kprobes interface.
Sample usage:
To place a probe on __blockdev_direct_IO:
static int probe_handler(struct kprobe *p, struct pt_regs *)
{
... whatever ...
}
struct kprobe kp = {
.addr = __blockdev_direct_IO,
.pre_handler = probe_handler
};
register_kprobe(&kp);
Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>