op-kernel-dev - Development kernel branch for OpenPOWER systems

	Commit message (Collapse)	Author	Age	Files	Lines
*	bcache: remove nested function usage	John Sheu	2014-03-18	2	-72/+76
\| \| \| \| \| \| \| \| \| \| \|	Uninlined nested functions can cause crashes when using ftrace, as they don't follow the normal calling convention and confuse the ftrace function graph tracer as it examines the stack. Also, nested functions are supported as a gcc extension, but may fail on other compilers (e.g. llvm). Signed-off-by: John Sheu <john.sheu@gmail.com>
*	bcache: Kill bucket->gc_gen	Kent Overstreet	2014-03-18	4	-11/+9
\| \| \| \| \| \| \| \|	gc_gen was a temporary used to recalculate last_gc, but since we only need bucket->last_gc when gc isn't running (gc_mark_valid = 1), we can just update last_gc directly. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Kill unused freelist	Kent Overstreet	2014-03-18	5	-125/+110
\| \| \| \| \| \| \| \|	This was originally added as at optimization that for various reasons isn't needed anymore, but it does add a lot of nasty corner cases (and it was responsible for some recently fixed bugs). Just get rid of it now. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Rework btree cache reserve handling	Kent Overstreet	2014-03-18	6	-139/+145
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This changes the bucket allocation reserves to use _real_ reserves - separate freelists - instead of watermarks, which if nothing else makes the current code saner to reason about and is going to be important in the future when we add support for multiple btrees. It also adds btree_check_reserve(), which checks (and locks) the reserves for both bucket allocation and memory allocation for btree nodes; the old code just kinda sorta assumed that since (e.g. for btree node splits) it had the root locked and that meant no other threads could try to make use of the same reserve; this technically should have been ok for memory allocation (we should always have a reserve for memory allocation (the btree node cache is used as a reserve and we preallocate it)), but multiple btrees will mean that locking the root won't be sufficient anymore, and for the bucket allocation reserve it was technically possible for the old code to deadlock. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Kill btree_io_wq	Kent Overstreet	2014-03-18	3	-24/+2
\| \| \| \| \| \| \| \|	With the locking rework in the last patch, this shouldn't be needed anymore - btree_node_write_work() only takes b->write_lock which is never held for very long. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: btree locking rework	Kent Overstreet	2014-03-18	4	-52/+133
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Add a new lock, b->write_lock, which is required to actually modify - or write - a btree node; this lock is only held for short durations. This means we can write out a btree node without taking b->lock, which _is_ held for long durations - solving a deadlock when btree_flush_write() (from the journalling code) is called with a btree node locked. Right now just occurs in bch_btree_set_root(), but with an upcoming journalling rework is going to happen a lot more. This also turns b->lock is now more of a read/intent lock instead of a read/write lock - but not completely, since it still blocks readers. May turn it into a real intent lock at some point in the future. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix a race when freeing btree nodes	Kent Overstreet	2014-03-18	1	-33/+20
\| \| \| \| \| \| \| \| \| \| \| \|	This isn't a bulletproof fix; btree_node_free() -> bch_bucket_free() puts the bucket on the unused freelist, where it can be reused right away without any ordering requirements. It would be better to wait on at least a journal write to go down before reusing the bucket. bch_btree_set_root() does this, and inserting into non leaf nodes is completely synchronous so we should be ok, but future patches are just going to get rid of the unused freelist - it was needed in the past for various reasons but shouldn't be anymore. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Add a real GC_MARK_RECLAIMABLE	Kent Overstreet	2014-03-18	4	-14/+21
\| \| \| \| \| \| \|	This means the garbage collection code can better check for data and metadata pointers to the same buckets. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Add bch_keylist_init_single()	Kent Overstreet	2014-03-18	2	-4/+7
\| \| \| \| \| \| \|	This will potentially save us an allocation when we've got inode/dirent bkeys that don't fit in the keylist's inline keys. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Improve priority_stats	Kent Overstreet	2014-03-18	1	-6/+20
\| \| \| \| \| \|	Break down data into clean data/dirty data/metadata. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Better alloc tracepoints	Kent Overstreet	2014-03-18	2	-5/+12
\| \| \| \| \| \| \|	Change the invalidate tracepoint to indicate how much data we're invalidating, and change the alloc tracepoints to indicate what offset they're for. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Kill dead cgroup code	Kent Overstreet	2014-03-18	5	-202/+0
\| \| \| \| \| \|	This hasn't been used or even enabled in ages. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: stop moving_gc marking buckets that can't be moved.	Nicholas Swenson	2014-03-18	1	-1/+4
\| \| \| \|	Signed-off-by: Nicholas Swenson <nks@daterainc.com>
*	bcache: Fix moving_pred()	Kent Overstreet	2014-03-18	1	-5/+3
\| \| \| \| \| \|	Avoid a potential null pointer deref (e.g. from check keys for cache misses) Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix moving_gc deadlocking with a foreground write	Nicholas Swenson	2014-03-18	5	-8/+16
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Deadlock happened because a foreground write slept, waiting for a bucket to be allocated. Normally the gc would mark buckets available for invalidation. But the moving_gc was stuck waiting for outstanding writes to complete. These writes used the bcache_wq, the same queue foreground writes used. This fix gives moving_gc its own work queue, so it was still finish moving even if foreground writes are stuck waiting for allocation. It also makes work queue a parameter to the data_insert path, so moving_gc can use its workqueue for writes. Signed-off-by: Nicholas Swenson <nks@daterainc.com> Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix discard granularity	Kent Overstreet	2014-03-18	1	-0/+1
\| \| \| \| \| \|	blk_stack_limits() doesn't like a discard granularity of 0. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix another bug recovering from unclean shutdown	Kent Overstreet	2014-03-18	3	-65/+36
\| \| \| \| \| \| \| \| \| \| \|	The on disk bucket gens are allowed to be out of date, when we reuse buckets that didn't have any live data in them. To deal with this, the initial gc has to update the bucket gen when we find a pointer gen newer than the bucket's gen. Unfortunately we weren't doing this for pointers in the journal that we're about to replay. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix a bug recovering from unclean shutdown	Kent Overstreet	2014-03-18	1	-2/+2
\| \| \| \| \| \| \|	The code to fixup incorrect bucket prios incorrectly did not skip btree node freeing keys Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix a journalling reclaim after recovery bug	Kent Overstreet	2014-03-18	1	-2/+8
\| \| \| \| \| \| \| \|	On recovery we weren't correctly keeping track of what journal buckets had open journal entries, thus it was possible for them to be overwritten until we'd written all new journal entries. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix a null ptr deref in journal replay	Kent Overstreet	2014-03-17	1	-1/+5
\| \| \| \|	Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix a lockdep splat in an error path	Kent Overstreet	2014-03-17	1	-3/+5
\| \| \| \|	Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix a shutdown bug	Kent Overstreet	2014-02-25	3	-2/+12
\| \| \| \| \| \|	Shutdown wasn't cancelling/waiting on journal_write_work() Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix flash_dev_cache_miss() for real this time	Kent Overstreet	2014-02-25	1	-14/+5
\| \| \| \| \| \| \|	The code was using sectors to count the number of sectors it was zeroing... but then it passed it to bio_advance()... after it had been set to 0. Amusing... Signed-off-by: Kent Overstreet <kmo@daterainc.com>
*	bcache: Fix another compiler warning on m68k	Kent Overstreet	2014-02-18	1	-2/+2
\| \| \| \| \| \| \|	Use a bigger hammer this time Signed-off-by: Kent Overstreet <kmo@daterainc.com> Cc: linux-stable <stable@vger.kernel.org>
*	Merge branch 'bcache-for-3.14' of git://evilpiepirate.org/~kent/linux-bcache ↵	Jens Axboe	2014-01-30	6	-10/+15
\|\ \| \| \| \| \| \|	into for-linus
\| *	bcache: bugfix - gc thread now gets woken when cache is full	Nicholas Swenson	2014-01-29	1	-3/+3
\| \| \| \| \| \| \| \|	Signed-off-by: Nicholas Swenson <nks@daterainc.com>
\| *	bcache: Minor fixes from kbuild robot	Kent Overstreet	2014-01-29	4	-5/+8
\| \| \| \| \| \| \| \|	Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: fix BUG_ON due to integer overflow with GC_SECTORS_USED	Darrick J. Wong	2014-01-29	2	-2/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The BUG_ON at the end of __bch_btree_mark_key can be triggered due to an integer overflow error: BITMASK(GC_SECTORS_USED, struct bucket, gc_mark, 2, 13); ... SET_GC_SECTORS_USED(g, min_t(unsigned, GC_SECTORS_USED(g) + KEY_SIZE(k), (1 << 14) - 1)); BUG_ON(!GC_SECTORS_USED(g)); In bcache.h, the SECTORS_USED bitfield is defined to be 13 bits wide. While the SET_ code tries to ensure that the field doesn't overflow by clamping it to (1<<14)-1 == 16383, this is incorrect because 16383 requires 14 bits. Therefore, if GC_SECTORS_USED() + KEY_SIZE() = 8192, the SET_ statement tries to store 8192 into a 13-bit field. In a 13-bit field, 8192 becomes zero, thus triggering the BUG_ON. Therefore, create a field width constant and a max value constant, and use those to create the bitfield and check the inputs to SET_GC_SECTORS_USED. Arguably the BITMASK() template ought to have BUG_ON checks for too-large values, but that's a separate patch. Signed-off-by: Darrick J. Wong <darrick.wong@oracle.com>
* \|	Merge branch 'for-3.14/drivers' of git://git.kernel.dk/linux-block	Linus Torvalds	2014-01-30	21	-1755/+2212
\|\ \ \| \|/ \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Pull block IO driver changes from Jens Axboe: - bcache update from Kent Overstreet. - two bcache fixes from Nicholas Swenson. - cciss pci init error fix from Andrew. - underflow fix in the parallel IDE pg_write code from Dan Carpenter. I'm sure the 1 (or 0) users of that are now happy. - two PCI related fixes for sx8 from Jingoo Han. - floppy init fix for first block read from Jiri Kosina. - pktcdvd error return miss fix from Julia Lawall. - removal of IRQF_SHARED from the SEGA Dreamcast CD-ROM code from Michael Opdenacker. - comment typo fix for the loop driver from Olaf Hering. - potential oops fix for null_blk from Raghavendra K T. - two fixes from Sam Bradshaw (Micron) for the mtip32xx driver, fixing an OOM problem and a problem with handling security locked conditions * 'for-3.14/drivers' of git://git.kernel.dk/linux-block: (47 commits) mg_disk: Spelling s/finised/finished/ null_blk: Null pointer deference problem in alloc_page_buffers mtip32xx: Correctly handle security locked condition mtip32xx: Make SGL container per-command to eliminate high order dma allocation drivers/block/loop.c: fix comment typo in loop_config_discard drivers/block/cciss.c:cciss_init_one(): use proper errnos drivers/block/paride/pg.c: underflow bug in pg_write() drivers/block/sx8.c: remove unnecessary pci_set_drvdata() drivers/block/sx8.c: use module_pci_driver() floppy: bail out in open() if drive is not responding to block0 read bcache: Fix auxiliary search trees for key size > cacheline size bcache: Don't return -EINTR when insert finished bcache: Improve bucket_prio() calculation bcache: Add bch_bkey_equal_header() bcache: update bch_bkey_try_merge bcache: Move insert_fixup() to btree_keys_ops bcache: Convert sorting to btree_keys bcache: Convert debug code to btree_keys bcache: Convert btree_iter to struct btree_keys bcache: Refactor bset_tree sysfs stats ...
\| *	bcache: Fix auxiliary search trees for key size > cacheline size	Kent Overstreet	2014-01-08	1	-14/+14
\| \| \| \| \| \| \| \|	Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Don't return -EINTR when insert finished	Kent Overstreet	2014-01-08	1	-2/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	We need to return -EINTR after a split because we invalidated iterators (and freed the btree node) - but if we were finished inserting, we don't want to redo the traversal. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Improve bucket_prio() calculation	Kent Overstreet	2014-01-08	2	-3/+16
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When deciding what order to reuse buckets we take into account both the bucket's priority (which indicates lru order) and also the amount of live data in that bucket. The way they were scaled together wasn't as correct as it could be... this patch improves and documents it. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Add bch_bkey_equal_header()	Nicholas Swenson	2014-01-08	3	-8/+11
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Checks if two keys have equivalent header fields. (good enough for replacement or merging) Used in bch_bkey_try_merge, and replacing a key in the btree. Signed-off-by: Nicholas Swenson <nks@daterainc.com> Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: update bch_bkey_try_merge	Nicholas Swenson	2014-01-08	3	-16/+28
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Added generic header checks to bch_bkey_try_merge, which then calls the bkey specific function Removed extraneous checks from bch_extent_merge Signed-off-by: Nicholas Swenson <nks@daterainc.com>
\| *	bcache: Move insert_fixup() to btree_keys_ops	Kent Overstreet	2014-01-08	4	-229/+257
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Now handling overlapping extents/keys is a method that's specific to what the btree node contains. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Convert sorting to btree_keys	Kent Overstreet	2014-01-08	3	-36/+33
\| \| \| \| \| \| \| \| \| \| \| \|	More work to disentangle various code from struct btree Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Convert debug code to btree_keys	Kent Overstreet	2014-01-08	9	-217/+264
\| \| \| \| \| \| \| \| \| \| \| \|	More work to disentangle various code from struct btree Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Convert btree_iter to struct btree_keys	Kent Overstreet	2014-01-08	6	-38/+41
\| \| \| \| \| \| \| \| \| \| \| \|	More work to disentangle bset.c from struct btree Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Refactor bset_tree sysfs stats	Kent Overstreet	2014-01-08	3	-47/+54
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	We're in the process of turning bset.c into library code, so none of the code in that file should know about struct cache_set or struct btree - so, move the btree traversal part of the stats code to sysfs.c. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Add bch_btree_keys_u64s_remaining()	Kent Overstreet	2014-01-08	2	-13/+30
\| \| \| \| \| \| \| \| \| \| \| \|	Helper function to explicitly check how much space is free in a btree node Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Add struct btree_keys	Kent Overstreet	2014-01-08	8	-263/+322
\| \| \| \| \| \| \| \| \| \| \| \|	Soon, bset.c won't need to depend on struct btree. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Abstract out stuff needed for sorting	Kent Overstreet	2014-01-08	9	-289/+423
\| \| \| \| \| \| \| \|	Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Rename/shuffle various code around	Kent Overstreet	2014-01-08	8	-276/+341
\| \| \| \| \| \| \| \| \| \| \| \|	More work to disentangle bset.c from the rest of the code: Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Add struct bset_sort_state	Kent Overstreet	2014-01-08	6	-49/+87
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	More disentangling bset.c from the rest of the bcache code - soon, the sorting routines won't have any dependencies on any outside structs. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Split out sort_extent_cmp()	Kent Overstreet	2014-01-08	4	-32/+73
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Only use extent comparison for comparing extents, so we're not using START_KEY() on other key types (i.e. btree pointers) Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Bkey indexing renaming	Kent Overstreet	2014-01-08	6	-52/+62
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	More refactoring: node() -> bset_bkey_idx() end() -> bset_bkey_last() Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Make bch_keylist_realloc() take u64s, not nptrs	Kent Overstreet	2014-01-08	4	-16/+26
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Getting away from KEY_PTRS and moving toward KEY_U64s - and getting rid of magic 2s Also - split out the part that checks against journal entry size so as to avoid a dependancy on struct cache_set in bset.c Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Remove/fix some header dependencies	Kent Overstreet	2014-01-08	3	-24/+26
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	In the process of disentagling/libraryizing bset.c from the rest of the bcache code. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Use a mempool for mergesort temporary space	Kent Overstreet	2014-01-08	3	-16/+8
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	It was a single element mempool before, it's slightly cleaner to just use a real mempool. Signed-off-by: Kent Overstreet <kmo@daterainc.com>
\| *	bcache: Btree verify code improvements	Kent Overstreet	2014-01-08	6	-40/+83
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Used this fixed code to find and fix the bug fixed by a4d885097b0ac0cd1337f171f2d4b83e946094d4. Signed-off-by: Kent Overstreet <kmo@daterainc.com>