op-kernel-dev - Development kernel branch for OpenPOWER systems

	Commit message (Collapse)	Author	Age	Files	Lines
*	KVM: change KVM to use IOMMU API	Joerg Roedel	2009-01-03	2	-2/+3
\| \| \| \|	Signed-off-by: Joerg Roedel <joerg.roedel@amd.com>
*	KVM: rename vtd.c to iommu.c	Joerg Roedel	2009-01-03	1	-1/+1
\| \| \| \| \| \| \| \| \|	Impact: file renamed The code in the vtd.c file can be reused for other IOMMUs as well. So rename it to make it clear that it handle more than VT-d. Signed-off-by: Joerg Roedel <joerg.roedel@amd.com>
*	KVM: MMU: handle large host sptes on invlpg/resync	Marcelo Tosatti	2008-12-31	2	-3/+8
\| \| \| \| \| \| \| \| \| \| \| \|	The invlpg and sync walkers lack knowledge of large host sptes, descending to non-existant pagetable level. Stop at directory level in such case. Fixes SMP Windows XP with hugepages. Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: Add locking to virtual i8259 interrupt controller	Avi Kivity	2008-12-31	2	-4/+53
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	While most accesses to the i8259 are with the kvm mutex taken, the call to kvm_pic_read_irq() is not. We can't easily take the kvm mutex there since the function is called with interrupts disabled. Fix by adding a spinlock to the virtual interrupt controller. Since we can't send an IPI under the spinlock (we also take the same spinlock in an irq disabled context), we defer the IPI until the spinlock is released. Similarly, we defer irq ack notifications until after spinlock release to avoid lock recursion. Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: Don't treat a global pte as such if cr4.pge is cleared	Avi Kivity	2008-12-31	1	-0/+2
\| \| \| \| \| \| \| \| \| \|	The pte.g bit is meaningless if global pages are disabled; deferring mmu page synchronization on these ptes will lead to the guest using stale shadow ptes. Fixes Vista x86 smp bootloader failure. Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86: Rework user space NMI injection as KVM_CAP_USER_NMI	Jan Kiszka	2008-12-31	2	-48/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	There is no point in doing the ready_for_nmi_injection/ request_nmi_window dance with user space. First, we don't do this for in-kernel irqchip anyway, while the code path is the same as for user space irqchip mode. And second, there is nothing to loose if a pending NMI is overwritten by another one (in contrast to IRQs where we have to save the number). Actually, there is even the risk of raising spurious NMIs this way because the reason for the held-back NMI might already be handled while processing the first one. Therefore this patch creates a simplified user space NMI injection interface, exporting it under KVM_CAP_USER_NMI and dropping the old KVM_CAP_NMI capability. And this time we also take care to provide the interface only on archs supporting NMIs via KVM (right now only x86). Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: VMX: Fix pending NMI-vs.-IRQ race for user space irqchip	Jan Kiszka	2008-12-31	1	-1/+3
\| \| \| \| \| \| \| \|	As with the kernel irqchip, don't allow an NMI to stomp over an already injected IRQ; instead wait for the IRQ injection to be completed. Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: check for present pdptr shadow page in walk_shadow	Marcelo Tosatti	2008-12-31	1	-0/+2
\| \| \| \| \| \| \| \| \| \|	walk_shadow assumes the caller verified validity of the pdptr pointer in question, which is not the case for the invlpg handler. Fixes oops during Solaris 10 install. Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: Consolidate userspace memory capability reporting into common code	Avi Kivity	2008-12-31	1	-1/+0
\| \| \| \|	Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: prepopulate the shadow on invlpg	Marcelo Tosatti	2008-12-31	3	-14/+38
\| \| \| \| \| \| \| \| \| \| \| \|	If the guest executes invlpg, peek into the pagetable and attempt to prepopulate the shadow entry. Also stop dirty fault updates from interfering with the fork detector. 2% improvement on RHEL3/AIM7. Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: skip global pgtables on sync due to cr3 switch	Marcelo Tosatti	2008-12-31	3	-10/+57
\| \| \| \| \| \| \| \| \|	Skip syncing global pages on cr3 switch (but not on cr4/cr0). This is important for Linux 32-bit guests with PAE, where the kmap page is marked as global. Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: collapse remote TLB flushes on root sync	Marcelo Tosatti	2008-12-31	1	-5/+14
\| \| \| \| \| \| \| \| \| \|	Collapse remote TLB flushes on root sync. kernbench is 2.7% faster on 4-way guest. Improvements have been seen with other loads such as AIM7. Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: use page array in unsync walk	Marcelo Tosatti	2008-12-31	1	-55/+140
\| \| \| \| \| \| \| \| \| \|	Instead of invoking the handler directly collect pages into an array so the caller can work with it. Simplifies TLB flush collapsing. Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: Fix handling of VMMCALL instruction	Amit Shah	2008-12-31	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \|	The VMMCALL instruction doesn't get recognised and isn't processed by the emulator. This is seen on an Intel host that tries to execute the VMMCALL instruction after a guest live migrates from an AMD host. Signed-off-by: Amit Shah <amit.shah@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: add the emulation of shld and shrd instructions	Guillaume Thouvenin	2008-12-31	1	-2/+15
\| \| \| \| \| \| \|	Add emulation of shld and shrd instructions Signed-off-by: Guillaume Thouvenin <guillaume.thouvenin@ext.bull.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: add the assembler code for three operands	Guillaume Thouvenin	2008-12-31	1	-0/+39
\| \| \| \| \| \| \| \|	Add the assembler code for instruction with three operands and one operand is stored in ECX register Signed-off-by: Guillaume Thouvenin <guillaume.thouvenin@ext.bull.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: add a new "implied 1" Src decode type	Guillaume Thouvenin	2008-12-31	1	-0/+5
\| \| \| \| \| \| \| \|	Add SrcOne operand type when we need to decode an implied '1' like with regular shift instruction Signed-off-by: Guillaume Thouvenin <guillaume.thouvenin@ext.bull.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: add Src2 decode set	Guillaume Thouvenin	2008-12-31	1	-0/+29
\| \| \| \| \| \| \| \| \|	Instruction like shld has three operands, so we need to add a Src2 decode set. We start with Src2None, Src2CL, and Src2ImmByte, Src2One to support shld/shrd and we will expand it later. Signed-off-by: Guillaume Thouvenin <guillaume.thouvenin@ext.bull.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: Extend the opcode descriptor	Guillaume Thouvenin	2008-12-31	1	-4/+4
\| \| \| \| \| \| \| \|	Extend the opcode descriptor to 32 bits. This is needed by the introduction of a new Src2 operand type. Signed-off-by: Guillaume Thouvenin <guillaume.thouvenin@ext.bull.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: VMX: fix sparse warning	Hannes Eder	2008-12-31	1	-1/+1
\| \| \| \| \| \| \| \| \|	Impact: make global function static arch/x86/kvm/vmx.c:134:3: warning: symbol 'vmx_capability' was not declared. Should it be static? Signed-off-by: Hannes Eder <hannes@hanneseder.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: Remove extraneous semicolon after do/while	Avi Kivity	2008-12-31	1	-1/+1
\| \| \| \| \| \|	Notices by Guillaume Thouvenin. Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: fix popf emulation	Avi Kivity	2008-12-31	1	-0/+2
\| \| \| \| \| \|	Set operand type and size to get correct writeback behavior. Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: fix ret emulation	Avi Kivity	2008-12-31	1	-0/+2
\| \| \| \| \| \| \|	'ret' did not set the operand type or size for the destination, so writeback ignored it. Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: switch 'pop reg' instruction to emulate_pop()	Avi Kivity	2008-12-31	1	-7/+4
\| \| \| \|	Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: allow pop from mmio	Avi Kivity	2008-12-31	1	-3/+3
\| \| \| \|	Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: Extract 'pop' sequence into a function	Avi Kivity	2008-12-31	1	-4/+17
\| \| \| \| \| \|	Switch 'pop r/m' instruction to use the new function. Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: consolidate emulation of two operand instructions	Avi Kivity	2008-12-31	1	-51/+28
\| \| \| \| \| \|	No need to repeat the same assembly block over and over. Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: reduce duplication in one operand emulation thunks	Avi Kivity	2008-12-31	1	-43/+23
\| \| \| \|	Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: optimize set_spte for page sync	Marcelo Tosatti	2008-12-31	1	-0/+9
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The write protect verification in set_spte is unnecessary for page sync. Its guaranteed that, if the unsync spte was writable, the target page does not have a write protected shadow (if it had, the spte would have been write protected under mmu_lock by rmap_write_protect before). Same reasoning applies to mark_page_dirty: the gfn has been marked as dirty via the pagefault path. The cost of hash table and memslot lookups are quite significant if the workload is pagetable write intensive resulting in increased mmu_lock contention. Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: VMX: Conditionally request interrupt window after injecting irq	Avi Kivity	2008-12-31	1	-0/+2
\| \| \| \| \| \| \| \| \| \| \|	If we're injecting an interrupt, and another one is pending, request an interrupt window notification so we don't have excess latency on the second interrupt. This shouldn't happen in practice since an EOI will be issued, giving a second chance to request an interrupt window, but... Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: SVM: move svm_hardware_disable() code to asm/virtext.h	Eduardo Habkost	2008-12-31	1	-5/+1
\| \| \| \| \| \| \|	Create cpu_svm_disable() function. Signed-off-by: Eduardo Habkost <ehabkost@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: SVM: move has_svm() code to asm/virtext.h	Eduardo Habkost	2008-12-31	1	-14/+5
\| \| \| \| \| \| \| \| \|	Use a trick to keep the printk()s on has_svm() working as before. gcc will take care of not generating code for the 'msg' stuff when the function is called with a NULL msg argument. Signed-off-by: Eduardo Habkost <ehabkost@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: VMX: extract kvm_cpu_vmxoff() from hardware_disable()	Eduardo Habkost	2008-12-31	1	-2/+11
\| \| \| \| \| \| \| \|	Along with some comments on why it is different from the core cpu_vmxoff() function. Signed-off-by: Eduardo Habkost <ehabkost@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: VMX: move cpu_has_kvm_support() to an inline on asm/virtext.h	Eduardo Habkost	2008-12-31	1	-2/+2
\| \| \| \| \| \| \| \|	It will be used by core code on kdump and reboot, to disable vmx if needed. Signed-off-by: Eduardo Habkost <ehabkost@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: SVM: move svm.h to include/asm	Eduardo Habkost	2008-12-31	2	-329/+1
\| \| \| \| \| \| \| \|	svm.h will be used by core code that is independent of KVM, so I am moving it outside the arch/x86/kvm directory. Signed-off-by: Eduardo Habkost <ehabkost@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: VMX: move vmx.h to include/asm	Eduardo Habkost	2008-12-31	3	-369/+2
\| \| \| \| \| \| \| \|	vmx.h will be used by core code that is independent of KVM, so I am moving it outside the arch/x86/kvm directory. Signed-off-by: Eduardo Habkost <ehabkost@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: Fix cpuid iteration on multiple leaves per eac	Nitin A Kamble	2008-12-31	1	-1/+2
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The code to traverse the cpuid data array list for counting type of leaves is currently broken. This patches fixes the 2 things in it. 1. Set the 1st counting entry's flag KVM_CPUID_FLAG_STATE_READ_NEXT. Without it the code will never find a valid entry. 2. Also the stop condition in the for loop while looking for the next unflaged entry is broken. It needs to stop when it find one matching entry; and in the case of count of 1, it will be the same entry found in this iteration. Signed-Off-By: Nitin A Kamble <nitin.a.kamble@intel.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: Fix cpuid leaf 0xb loop termination	Nitin A Kamble	2008-12-31	1	-1/+1
\| \| \| \| \| \| \| \| \|	For cpuid leaf 0xb the bits 8-15 in ECX register define the end of counting leaf. The previous code was using bits 0-7 for this purpose, which is a bug. Signed-off-by: Nitin A Kamble <nitin.a.kamble@intel.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: Fix aliased gfns treated as unaliased	Izik Eidus	2008-12-31	1	-4/+10
\| \| \| \| \| \| \| \| \| \| \|	Some areas of kvm x86 mmu are using gfn offset inside a slot without unaliasing the gfn first. This patch makes sure that the gfn will be unaliased and add gfn_to_memslot_unaliased() to save the calculating of the gfn unaliasing in case we have it unaliased already. Signed-off-by: Izik Eidus <ieidus@redhat.com> Acked-by: Marcelo Tosatti <mtosatti@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: Enable Function Level Reset for assigned device	Sheng Yang	2008-12-31	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Ideally, every assigned device should in a clear condition before and after assignment, so that the former state of device won't affect later work. Some devices provide a mechanism named Function Level Reset, which is defined in PCI/PCI-e document. We should execute it before and after device assignment. (But sadly, the feature is new, and most device on the market now don't support it. We are considering using D0/D3hot transmit to emulate it later, but not that elegant and reliable as FLR itself.) [Update: Reminded by Xiantao, execute FLR after we ensure that the device can be assigned to the guest.] Signed-off-by: Sheng Yang <sheng@linux.intel.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: VMX: Handle mmio emulation when guest state is invalid	Guillaume Thouvenin	2008-12-31	1	-12/+15
\| \| \| \| \| \| \| \| \| \| \|	If emulate_invalid_guest_state is enabled, the emulator is called when guest state is invalid. Until now, we reported an mmio failure when emulate_instruction() returned EMULATE_DO_MMIO. This patch adds the case where emulate_instruction() failed and an MMIO emulation is needed. Signed-off-by: Guillaume Thouvenin <guillaume.thouvenin@ext.bull.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: allow emulator to adjust rip for emulated pio instructions	Guillaume Thouvenin	2008-12-31	4	-3/+3
\| \| \| \| \| \| \| \| \| \| \| \| \|	If we call the emulator we shouldn't call skip_emulated_instruction() in the first place, since the emulator already computes the next rip for us. Thus we move ->skip_emulated_instruction() out of kvm_emulate_pio() and into handle_io() (and the svm equivalent). We also replaced "return 0" by "break" in the "do_io:" case because now the shadow register state needs to be committed. Otherwise eip will never be updated. Signed-off-by: Guillaume Thouvenin <guillaume.thouvenin@ext.bull.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: SVM: Set the 'busy' flag of the TR selector	Amit Shah	2008-12-31	1	-0/+7
\| \| \| \| \| \| \| \|	The busy flag of the TR selector is not set by the hardware. This breaks migration from amd hosts to intel hosts. Signed-off-by: Amit Shah <amit.shah@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: SVM: Set the 'g' bit of the cs selector for cross-vendor migration	Amit Shah	2008-12-31	1	-0/+9
\| \| \| \| \| \| \| \| \|	The hardware does not set the 'g' bit of the cs selector and this breaks migration from amd hosts to intel hosts. Set this bit if the segment limit is beyond 1 MB. Signed-off-by: Amit Shah <amit.shah@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86: Fix typo in function name	Amit Shah	2008-12-31	1	-5/+5
\| \| \| \| \| \| \|	get_segment_descritptor_dtable() contains an obvious type. Signed-off-by: Amit Shah <amit.shah@redhat.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86: Optimize NMI watchdog delivery	Jan Kiszka	2008-12-31	2	-8/+26
\| \| \| \| \| \| \| \| \| \|	As suggested by Avi, this patch introduces a counter of VCPUs that have LVT0 set to NMI mode. Only if the counter > 0, we push the PIT ticks via all LAPIC LVT0 lines to enable NMI watchdog support. Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com> Acked-by: Sheng Yang <sheng@linux.intel.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86: Fix and refactor NMI watchdog emulation	Jan Kiszka	2008-12-31	3	-17/+20
\| \| \| \| \| \| \| \| \| \| \| \| \|	This patch refactors the NMI watchdog delivery patch, consolidating tests and providing a proper API for delivering watchdog events. An included micro-optimization is to check only for apic_hw_enabled in kvm_apic_local_deliver (the test for LVT mask is covering the soft-disabled case already). Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com> Acked-by: Sheng Yang <sheng@linux.intel.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: x86 emulator: Add decode entries for 0x04 and 0x05 opcodes (add acc, imm)	Guillaume Thouvenin	2008-12-31	1	-1/+1
\| \| \| \| \| \| \| \|	Add decode entries for 0x04 and 0x05 (ADD) opcodes, execution is already implemented. Signed-off-by: Guillaume Thouvenin <guillaume.thouvenin@ext.bull.net> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: VMX: Move private memory slot position	Sheng Yang	2008-12-31	2	-3/+4
\| \| \| \| \| \| \| \| \| \| \|	PCI device assignment would map guest MMIO spaces as separate slot, so it is possible that the device has more than 2 MMIO spaces and overwrite current private memslot. The patch move private memory slot to the top of userspace visible memory slots. Signed-off-by: Sheng Yang <sheng@linux.intel.com> Signed-off-by: Avi Kivity <avi@redhat.com>
*	KVM: MMU: Extend kvm_mmu_page->slot_bitmap size	Sheng Yang	2008-12-31	1	-3/+3
\| \| \| \| \| \| \| \|	Otherwise set_bit() for private memory slot(above KVM_MEMORY_SLOTS) would corrupted memory in 32bit host. Signed-off-by: Sheng Yang <sheng@linux.intel.com> Signed-off-by: Avi Kivity <avi@redhat.com>