bleve

Author	SHA1	Message	Date
Marty Schoch	da794d3762	fix bug introduced by reuse of TermFrequencyRow values in a recent commit, we changed the code to reuse TermFrequencyRow objects intsead of constantly allocating new ones. unfortunately, one of the original methods was not coded with this reuse in mind, and a lazy initialization cause us to leak data from previous uses of the same object. in particular this caused term vector information from previous hits to still be applied to subsequent hits. eventually this causes the highlighter to try and highlight invalid regions of a slice. fixes #404	2016-08-05 08:33:04 -04:00
Steve Yen	3c82086805	optimize upside_down reader & 64-bit struct alignments The UpsideDownCouchTermFieldReader.Next() only needs the doc ID from the key, so this change provides a specialized parseKDoc() method for that optimization. Additionally, fields in various structs are more 64-bit aligned, in an attempt to reduce the invocations of runtime.typedmemmove() and runtime.heapBitsBulkBarrier(), which the go compiler seems to automatically insert to transparently handle misaligned data.	2016-07-23 10:37:40 -07:00
Steve Yen	2498ccc913	optimize upside_down reader Next() to reuse TermFrequencyRow Before this change, upside down's reader would alloc a new TermFrequencyRow on every Next(), which would be immediately transformed into an index.TermFieldDoc{}. This change reuses a pre-allocated TermFrequencyRow that's a field in the reader.	2016-07-21 11:10:49 -07:00
Steve Yen	a29dd25a48	upside_down dict row value size accounts for large uvarint's This is somewhat unlikely, but if a term is (incredibly) popular, its uvarint count value representation might go beyond 8 bytes. Some KVStore implementations (like forestdb) provide a BatchEx cgo optimization that depends on proper preallocated counting, so this change provides a proper worst-case estimate based on the max-unvarint of 10 bytes instead of the previously incorrect 8 bytes.	2016-02-22 11:52:51 -08:00
Silvan Jegen	d326898f7b	Remove unneeded brackets	2016-01-14 16:41:41 +01:00
Steve Yen	82b8b3468e	upside_down analysis converts to docIDBytes once	2016-01-06 23:38:02 -08:00
Patrick Mezard	e85c9c542e	row: expose TermFrequencyRow term and freq fields Rows content is an implementation detail of bleve index and may change in the future. That said, they also contains information valuable to assess the quality of the index or understand its performances. So, as long as we agree that type asserting rows should only be done if you know what you are doing and are ready to deal with future changes, I see no reason to hide the row fields from external packages. Fix #268	2015-11-17 17:21:26 +01:00
Marty Schoch	71cbb13e07	modify code to reuse buffer for kv generation	2015-10-05 17:49:50 -04:00
Marty Schoch	d06b526cbf	more refactoring	2015-09-28 16:50:27 -04:00
Marty Schoch	9db850a53e	Merge branch 'fix/MaxVarintLen64' of https://github.com/tukdesk/bleve into tukdesk-fix/MaxVarintLen64	2015-07-31 15:16:16 -04:00
Marty Schoch	c1c4941dde	Merge branch 'feature/term_vector' of https://github.com/tukdesk/bleve into tukdesk-feature/term_vector	2015-07-29 14:31:15 -04:00
Marty Schoch	2768c2da3c	fix previous sloppy fix which hadn't been adequately tested	2015-05-27 19:15:55 -07:00
Marty Schoch	201fb91171	fix up to correctly trim off separator even though it should never be present	2015-05-27 19:10:12 -07:00
Marty Schoch	a58592ceff	fix case where NewBackIndexRowKV returns nil, nil the logic for reading the docID from the keys in this row relies on the keys NEVER containing the byte separator character (0xff), this is OK as we require that all keys be valid utf-8 however, it turns out that in the case where this rule was violated, we would panic, because we return nil, nil and later try to print the doc id	2015-05-27 19:04:57 -07:00
dtynn	59c97ae577	use binary.MaxVarintLen64	2015-05-26 15:35:31 +08:00
dtynn	89dc2c22bc	update TermVector	2015-05-17 13:07:14 +08:00
Marty Schoch	30a0ba1f9b	fix bug, dictionary row encoding buffer too small we incorrectly created a []byte of length 8 but the max for a uvarint is 10 closes #197	2015-05-06 10:04:02 -04:00
Marty Schoch	867110e03b	major improvements to index row encoding improvements uncovered some issues with how k/v data was copied or not. to address this, kv abstraction layer now lets impl specify if the bytes returned are safe to use after a reader (or writer since writers are also readers) are closed See index/store/KVReader - BytesSafeAfterClose() bool false is the safe value if you're not sure it will cause index impls to copy the data Some kv impls already have created a copy a the C-api barrier in which case they can safely return true. Overall this yields ~25% speedup for searches with leveldb. It yields ~10% speedup for boltdb. Returning stored fields is now slower with boltdb, as previously we were returning unsafe bytes.	2015-04-03 16:50:48 -04:00
Marty Schoch	a44a7c01af	rewrite to used fixed size []byte instead of buffer removes unchecked errors in calls to buffer.Write and also benchmarks considerably faster	2015-03-11 15:12:13 -04:00
Marty Schoch	522f9d5cc7	significant change to index format, support dictionary rows this introduces disk format v4 now the summary rows for a term are stored in their own "dictionary row" format, previously the same information was stored in special term frequency rows this now allows us to easily iterate all the terms for a field in sorted order (useful for many other fuzzy data structures) at the top-level of bleve you can now browse terms within a field using the following api on the Index interface: FieldDict(field string) (index.FieldDict, error) FieldDictRange(field string, startTerm []byte, endTerm []byte) (index.FieldDict, error) FieldDictPrefix(field string, termPrefix []byte) (index.FieldDict, error) fixes #127	2015-03-10 16:22:19 -04:00
Marty Schoch	300ec79c96	first pass at checking errors that were ignored part of #169	2015-03-06 14:46:29 -05:00
Marty Schoch	a2ad7634f2	update term freq rows to use varint where possible benchmark old ns/op new ns/op delta BenchmarkLevelDBIndexing1Workers 1138292 657901 -42.20% BenchmarkLevelDBIndexing2Workers 1619323 647628 -60.01% BenchmarkLevelDBIndexing4Workers 1172845 636478 -45.73% BenchmarkLevelDBIndexing1Workers10Batch 465556545 448153394 -3.74% BenchmarkLevelDBIndexing2Workers10Batch 504203911 449657355 -10.82% BenchmarkLevelDBIndexing4Workers10Batch 510766435 439839335 -13.89% BenchmarkLevelDBIndexing1Workers100Batch 307657846 268976464 -12.57% BenchmarkLevelDBIndexing2Workers100Batch 302257400 269110215 -10.97% BenchmarkLevelDBIndexing4Workers100Batch 305320485 259084902 -15.14% BenchmarkLevelDBIndexing1Workers1000Batch 301320576 258070231 -14.35% BenchmarkLevelDBIndexing2Workers1000Batch 334174454 261175641 -21.84% BenchmarkLevelDBIndexing4Workers1000Batch 267732436 261461739 -2.34% closes #165	2015-03-06 13:00:53 -05:00
Sergey Avseyev	a8351be5a6	Update protobuf imports	2014-12-10 01:24:59 +03:00
Marty Schoch	198ca1ad4d	major refactor of kvstore/index internals, see below In the index/store package introduce KVReader creates snapshot all read operations consistent from this snapshot must close to release introduce KVWriter only one writer active access to all operations allows for consisten read-modify-write must close to release introduce AssociativeMerge operation on batch allows efficient read-modify-write for associative operations used to consolidate updates to the term summary rows saves 1 set and 1 get op per shared instance of term in field In the index package introduced an IndexReader exposes a consisten snapshot of the index for searching At top level All searches now operate on a consisten snapshot of the index	2014-09-12 17:21:35 -04:00
Marty Schoch	d534b0836b	converted ALL_CAPS constants to CamelCase	2014-09-03 17:48:40 -04:00
Marty Schoch	7a7eb2e94c	add newline between license and package this avoids cluttering godocs with the license	2014-09-02 10:54:50 -04:00
Marty Schoch	082a5b0b03	major change to fields now can track array positions for field values stored fields now include this in the key and the back index now uses protobufs to simplify serialization closes #73	2014-08-19 08:58:26 -04:00
Marty Schoch	c526a38369	major refactor of analysis files, now wired up to registry ultimately this is make it more convenient for us to wire up different elements of the analysis pipeline, without having to preload everything into memory before we need it separately the index layer now has a mechanism for storing internal key/value pairs. this is expected to be used to store the mapping, and possibly other pieces of data by the top layer, but not exposed to the user at the top.	2014-08-13 21:14:47 -04:00
Marty Schoch	292af78b9e	implemented prefix search closes #4	2014-08-07 13:45:39 -04:00
Marty Schoch	b16c1d7f79	changed term row encoding previously we used the format: 't' <utf-8 term> <byte separator> <16-bit field id> <utf-8 docID> <byte separator> now we have moved the field before the term, resulting in: 't' <16-bit field id> <utf-8 term> <byte separator> <utf-8 docID> <byte separator> this means now instead of all fields with the same term being grouped together all terms within the same field are grouped together this allows us to enumerate the terms used with a field this allows us to implement prefix search, and possibly improve numeric range queries	2014-08-07 09:39:04 -04:00
Marty Schoch	41d4f67ee2	fix storing/retrieving numeric and date fields also includes new ability to request stored fields be returned with results closes #55 and closes #56 and closes #58	2014-08-06 13:52:20 -04:00
Marty Schoch	fda861d4e7	add formatted printing of stored rows fix critcal bug in prefix matching on stored row keys	2014-07-03 14:51:06 -04:00
Marty Schoch	9bebbec267	added support for stored fields and highlighting results	2014-06-26 11:43:13 -04:00
Marty Schoch	4af76f539d	fewer allocations building byte array encodings	2014-05-19 11:02:15 -04:00
Marty Schoch	1f1ac3e4a8	added some negative tests to row	2014-04-18 22:31:13 -04:00
Marty Schoch	15726437eb	fix issue identified by go vet	2014-04-18 21:11:32 -04:00
Marty Schoch	f92f274665	refactored to remove panics, return errors, and fewer type assertions	2014-04-18 21:07:41 -04:00
Marty Schoch	a3e04d8697	rewrote to not handle errors which cannot occur	2014-04-18 16:36:03 -04:00
Marty Schoch	bb2f66be92	Revert "refactor to use less panics, return more errors" This reverts commit `dec37fed07`.	2014-04-18 16:09:34 -04:00
Marty Schoch	dec37fed07	refactor to use less panics, return more errors	2014-04-18 15:54:29 -04:00
Marty Schoch	3d842dfaf2	initial commit	2014-04-17 16:55:53 -04:00

41 Commits