bleve

Author	SHA1	Message	Date
Steve Yen	596d990eb9	scorch zap optimize when zero hits Instead of allocating brand-new empty postingsList/Iterator instances, reuse some empty singletons.	2018-03-27 15:39:33 -07:00
Sreekanth Sivasankaran	6c6c1419b5	Merge pull request #855 from blevesearch/tfr_advance TermFieldReader Advance optimisation	2018-03-27 22:49:48 +05:30
Sreekanth Sivasankaran	72ac352961	TermFieldReader Advance optimization skips to the target segment and avoid un necesary read of freq,loc,norm details	2018-03-27 20:18:16 +05:30
Steve Yen	1cab701f85	scorch zap postingsIter skips freq/norm/locs parsing if allowed In this optimization, the zap PostingsIterator skips the parsing of freq/norm/locs chunks based on the includeFreq\|Norm\|Locs flags. In bleve-query microbenchmark on dev macbookpro, with 50K en-wiki docs, on a medium frequency term search that does not ask for term vectors, throughput was ~750 q/sec before the change and went to ~1400 q/sec after the change.	2018-03-26 09:49:44 -07:00
Steve Yen	192621f402	scorch includeFreq/Norm/Locs params for postingsList.Iterator API This commit adds boolean flag params to the scorch PostingsList.Iterator() method, so that the caller can specify whether freq/norm/locs information is needed or not. Future changes can leverage these params for optimizations.	2018-03-26 09:49:44 -07:00
Steve Yen	fc7584f5a0	scorch zap prealloc extra locs for future growth	2018-03-26 09:49:44 -07:00
Steve Yen	3f4b161850	scorch zap postingsIter reuses array positions slice	2018-03-26 09:49:44 -07:00
Steve Yen	db792717a6	scorch zap postingsIter reuses nextLocs/nextSegmentLocs The previous code would inefficiently throw away the nextLocs and would also throw away the []segment.Location slice if there were no locations, such as if it was a 1-hit postings list. This change tries to reuse the nextLocs/nextSegmentLocs for all cases.	2018-03-26 09:49:44 -07:00
Steve Yen	ba644f3893	scorch zap fix postingsIter.nextBytes() when 1-bit encoded The previous commit's optimization that replaced the locsBitmap was incorrectly handling the case when there was a 1-bit encoding optimization in the postingsIterator.nextBytes() method, incorrectly generating the freq-norm bytes. Also as part of this change, more unused locsBitmap's were removed.	2018-03-26 09:19:00 -07:00
Steve Yen	7a19e6fd7e	scorch zap replace locsBitmap w/ 1 bit from freq-norm varint encoding This is attempt #2 of the optimization that replaces the locsBitmap, without any changes from the original commit attempt. A commit that follows this one contains the actual fix. See also... - commit `621b58dd83` (the 1st attempt) - commit `49a4ee60ba` (the revert) ------------- The original commit message body from 621b58 was... NOTE: this is a zap file format change. The separate "postings locations" roaring Bitmap that encoded whether a posting has locations info is now replaced by the least significant bit in the freq varint encoded in the freq-norm chunkedIntCoder. encode/decodeFreqHasLocs() are added as helper functions.	2018-03-23 12:50:24 -07:00
Steve Yen	49a4ee60ba	Revert "scorch zap replace locsBitmap w/ 1 bit from freq-norm varint encoding" Testing with the cbft application led to cbft process exits... AsyncError exit()... error reading location field: EOF -- main.initBleveOptions.func1() at init_bleve.go:85 This reverts commit `621b58dd83`.	2018-03-23 10:01:30 -07:00
Steve Yen	621b58dd83	scorch zap replace locsBitmap w/ 1 bit from freq-norm varint encoding NOTE: this is a zap file format change. The separate "postings locations" roaring Bitmap that encoded whether a posting has locations info is now replaced by the least significant bit in the freq varint encoded in the freq-norm chunkedIntCoder. encode/decodeFreqHasLocs() are added as helper functions.	2018-03-22 17:43:07 -07:00
Steve Yen	b506fae4f7	scorch zap postingsItr remove unused offset/locoffset fields	2018-03-21 18:00:14 -07:00
Steve Yen	d1e2b55c72	scorch zap postingsItr.nextDocNum() maintains allNChunk correctly When PostingsIterator.nextDocNum() moves the 'all' roaring bitmap iterator forwards, it was incorrectly not keeping the allNChunk value aligned.	2018-03-21 17:57:54 -07:00
Steve Yen	b411e65234	scorch zap optimize postingsIterator reuse of freq/locChunkOffsets	2018-03-16 11:22:50 -07:00
Sreekanth Sivasankaran	23cebae5a8	Merge pull request #815 from blevesearch/loadchunk_minor minor optimisation to loadChunk method	2018-03-16 08:15:37 +05:30
Sreekanth Sivasankaran	d1155c223a	zap version bump, changed the offset slice format ,UTs	2018-03-15 23:25:53 +05:30
Sreekanth Sivasankaran	1775602958	posting iterator array positions clean up, max segment size limit adjustment for hit-1 optimisation	2018-03-15 14:40:00 +05:30
Sreekanth Sivasankaran	5271b582bb	Merge branch 'master' of https://github.com/blevesearch/bleve into loadchunk_minor	2018-03-13 11:59:29 +05:30
Steve Yen	b1f3969521	scorch zap reuse roaring Bitmap in postings lists	2018-03-12 09:18:11 -07:00
Steve Yen	2a20a36e15	scorch zap optimimze to avoid bitmaps for 1-hit posting lists This commit avoids creating roaring.Bitmap's (which would have just a single entry) when a postings list/iterator represents a single "1-hit" encoding.	2018-03-10 06:33:09 -08:00
Sreekanth Sivasankaran	d6522e7e17	minor optimisation to loadChunk method	2018-03-09 16:10:39 +05:30
Steve Yen	eac9808990	scorch zap optimize FST val encoding for terms with 1 hit NOTE: this is a scorch zap file format change / bump to version 4. In this optimization, the uint64 val stored in the vellum FST (term dictionary) now may either be a uint64 postingsOffset (same as before this change) or a uint64 encoding of the docNum + norm (in the case where a term appears in just a single doc).	2018-03-08 09:19:54 -08:00
abhinavdangeti	5c721226cf	Fixing the scorch search request memory estimate Do not re-account for certain referenced data in the zap structures. New estimates: ESTIMATE BENCHMEM TermQuery 11396 12437 MatchQuery 12244 12951 DisjunctionQuery (Term queries) 20644 20709	2018-03-06 16:03:10 -08:00
abhinavdangeti	38b6c522b0	Address build breakage after rebase Removed attribute: iterator of type Posting	2018-03-06 14:00:54 -08:00
abhinavdangeti	7e36109b3c	MB-28162: Provide API to estimate memory needed to run a search query This API (unexported) will estimate the amount of memory needed to execute a search query over an index before the collector begins data collection. Sample estimates for certain queries: {Size: 10, BenchmarkUpsidedownSearchOverhead} ESTIMATE BENCHMEM TermQuery 4616 4796 MatchQuery 5210 5405 DisjunctionQuery (Match queries) 7700 8447 DisjunctionQuery (Term queries) 6514 6591 ConjunctionQuery (Match queries) 7524 8175 Nested disjunction query (disjunction of disjunctions) 10306 10708 …	2018-03-06 13:53:42 -08:00
Steve Yen	5b86da85f3	scorch zap optimize postings itr with tf/loc reader/decoder reuse	2018-03-06 13:30:59 -08:00
Steve Yen	530a3d24cf	scorch zap optimize merge by byte copying freq/norm/loc's This change adds a zap PostingsIterator.nextBytes() method, which is similar to Next(), but instead of returning a Posting instance, nextBytes() returns the encoded freq/norm and location byte slices. The zap merge code then provides those byte slices directly to the intCoder's via a new method, intCoder.AddBytes(), thereby avoiding having to encode many uvarint's.	2018-03-06 13:30:59 -08:00
Steve Yen	655268bec8	scorch zap postings iterator nextDocNum() helper method Refactored out a nextDocNum() helper method from Next() that future optimizations can use.	2018-03-06 07:55:26 -08:00
Steve Yen	502e64c256	scorch zap Posting doesn't use iterator field	2018-03-05 16:33:13 -08:00
Steve Yen	dd7d93ee5e	scorch zap loadChunk reuses Location slices	2018-02-27 18:01:48 -08:00
Steve Yen	4dbb4b1495	scorch zap posting reuses freqNorm & loc reader and decoder	2018-02-27 18:01:48 -08:00
Steve Yen	f177f07613	scorch zap segment merging reuses prealloc'ed PostingsIterator During zap segment merging, a new zap PostingsIterator was allocated for every field X segment X term. This change optimizes by reusing a single PostingsIterator instance per persistMergedRest() invocation. And, also unused fields are removed from the PostingsIterator.	2018-02-08 17:24:30 -08:00
Steve Yen	684ee3c0e7	scorch zap DictIterator term count fixed and more merge unit tests The zap DictionaryIterator Next() was incorrectly returning the postingsList offset as the term count. As part of this, refactored out a PostingsList.read() helper method. Also added more merge unit test scenarios, including merging a segment for a few rounds to see if there are differences before/after merging.	2018-01-30 21:22:06 -08:00
Steve Yen	10dd5489c2	scorch zap Dict.postingsList() takes []byte for more mem control This allows callers that already have a []byte term to avoid string'ification garbage.	2018-01-27 11:35:10 -08:00
Steve Yen	5a035dc9aa	scorch zap in-memory segment representation (SegmentBase) The zap SegmentBase struct is a refactoring of the zap Segment into the subset of fields that are needed for read-only ops, without any persistence related info. This allows us to use zap's optimized data encoding as scorch's in-memory segments. The zap Segment struct now embeds a zap SegmentBase struct, and layers on persistence. Both the zap Segment and zap SegmentBase implement scorch's Segment interface.	2018-01-27 11:35:10 -08:00
Steve Yen	8f8333e01b	scorch optimize zap Count() This proposed approach avoids building a temporary AndNot() bitmap, following the same kind of optimization used by mem segments.	2017-12-19 18:02:27 -08:00
Steve Yen	d0e4f85026	scorch avoid extra clone by using roaring.AndNot(x, y)	2017-12-19 13:37:04 -08:00
Steve Yen	730d906a50	scorch reuses Posting instance in PostingsIterator.Next() With this change, there are no more memory allocations in the calls to PostingsIterator.Next() in the micro benchmarks of bleve-query. On a dev macbook, on an index of 50K wikipedia docs, using high frequency search of "text:date"... 400 qps - upsidedown/moss 565 qps - scorch before 680 qps - scorch after	2017-12-18 16:15:38 -08:00
Marty Schoch	927216df8c	fix postings list count impl	2017-12-12 08:42:13 -05:00
Marty Schoch	74b2eeb14d	refactor where we do some work so we can return error	2017-12-11 15:59:36 -05:00
Marty Schoch	f13b786609	fix up issues to get all bleve unit tests passing for scorch make scorch default	2017-12-11 15:47:41 -05:00
Marty Schoch	9781d9b089	add initial version of zap file format	2017-12-09 14:28:33 -05:00

43 Commits