kz930 opened a new issue, #6761:
URL: https://github.com/apache/texera/issues/6761

   ### What happened?
   
   # What happened
   
   `KeywordSearchOpExec` picks one of two Lucene analyzers based on the
   `isCaseSensitive` flag:
   
   
`common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/keywordSearch/KeywordSearchOpExec.scala:38`
   
   ```scala
   @transient private lazy val analyzer: Analyzer = {
     if (desc.isCaseSensitive) new CaseSensitiveAnalyzer() else new 
StandardAnalyzer()
   }
   ```
   
   - **Case-insensitive (default)** → `StandardAnalyzer`: `StandardTokenizer` 
splits
     on Unicode word boundaries, so **punctuation is stripped** and `"perfect."`
     tokenizes to `perfect`. Query `perfect` matches.
   - **Case-sensitive** → `CaseSensitiveAnalyzer`, which is a bare
     `WhitespaceTokenizer`:
   
   
`common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/keywordSearch/CaseSensitiveAnalyzer.scala:29`
   
   ```scala
   // Achieves case sensitivity by skipping the lowercasing and normalization
   // pipeline used in StandardAnalyzer.
   class CaseSensitiveAnalyzer extends Analyzer {
     override protected def createComponents(fieldName: String): 
TokenStreamComponents = {
       val tokenizer = new WhitespaceTokenizer()
       val stream: TokenStream = new StopFilter(tokenizer, 
CharArraySet.EMPTY_SET)
       new TokenStreamComponents(tokenizer, stream)
     }
   }
   ```
   
   `WhitespaceTokenizer` splits **only** on whitespace, so **punctuation stays 
glued
   to the token**: `"...absolutely perfect."` tokenizes to `[absolutely, 
perfect.]`.
   The query term `perfect` (tokenized the same way → `perfect`) does not equal 
the
   indexed token `perfect.`, so **the row is not matched**.
   
   The intent of `CaseSensitiveAnalyzer` was only to *skip lowercasing*. But 
swapping
   `StandardTokenizer` for `WhitespaceTokenizer` also dropped Unicode 
word-boundary
   tokenization. So toggling "case sensitive" silently changes **punctuation
   handling** too — a side effect the option name does not imply. A user who 
turns on
   case sensitivity to match `Trump` exactly will find that `USA.` no longer 
matches
   `USA`, `day!` no longer matches `day`, etc.
   
   **Impact:** any keyword search over real text (which contains commas, 
periods,
   `!`, `?`) behaves inconsistently between the two modes. Words adjacent to
   punctuation become unmatchable in case-sensitive mode.
   
   ** Proposed Fix:** make `CaseSensitiveAnalyzer` mirror `StandardAnalyzer` 
*minus* the
   lowercasing step — i.e. use `StandardTokenizer` (Unicode word-boundary
   tokenization, strips punctuation) and simply omit the `LowerCaseFilter`, 
instead
   of falling back to `WhitespaceTokenizer`:
   
   ```scala
   override protected def createComponents(fieldName: String): 
TokenStreamComponents = {
     val tokenizer = new StandardTokenizer()
     // StandardAnalyzer pipeline without LowerCaseFilter: keep case, still 
strip
     // punctuation on Unicode word boundaries.
     val stream: TokenStream = new StopFilter(tokenizer, CharArraySet.EMPTY_SET)
     new TokenStreamComponents(tokenizer, stream)
   }
   ```
   
   Then "case sensitive" changes *only* case behavior, matching user 
expectation and
   staying consistent with the default mode's tokenization.
   
   Notes:
   - Present on `main`; not introduced by the standalone-translation work.
   - Related: the standalone translation 
(`KeywordSearchOpDesc.generateStandaloneCode`)
     does not implement case sensitivity at all (always case-insensitive 
`\b`-regex),
     so it already diverges from this mode. Fixing the analyzer would make the
     case-insensitive behavior of both paths line up on punctuated text; full
     case-sensitive parity in the standalone is a separate gap.
   
   ### How to reproduce?
   
   
   **Option A — user-facing (workflow):**
   
   1. Add a **Keyword Search** operator over a text column whose values contain
      trailing punctuation, e.g. a row `"everything was absolutely perfect."`.
   2. Set keyword = `perfect`, **Case Sensitive = off** → the row matches 
(expected).
   3. Flip **Case Sensitive = on**, rerun → the same row **no longer matches**, 
even
      though `perfect` obviously appears. Only the case handling was supposed to
      change.
   
   **Option B — unit test (mirrors `KeywordSearchOpExecSpec`):**
   
   ```scala
   val opDesc = new KeywordSearchOpDesc()
   opDesc.attribute = "text"
   opDesc.keyword = "perfect"
   
   val schema = Schema().add(new Attribute("text", AttributeType.STRING))
   def row(t: String) = Tuple.builder(schema).add(schema.getAttribute("text"), 
t).build()
   val data = List(row("everything was absolutely perfect."))
   
   // case-insensitive: matches
   opDesc.isCaseSensitive = false
   val a = new KeywordSearchOpExec(objectMapper.writeValueAsString(opDesc)); 
a.open()
   assert(data.exists(t => a.processTuple(t, 0).nonEmpty))   // passes
   
   // case-sensitive: SHOULD still match "perfect", but does not
   opDesc.isCaseSensitive = true
   val b = new KeywordSearchOpExec(objectMapper.writeValueAsString(opDesc)); 
b.open()
   assert(data.exists(t => b.processTuple(t, 0).nonEmpty))   // FAILS on 
current main
   ```
   
   ### Version/Branch
   
   1.3.0-incubating-SNAPSHOT (main)
   
   ### Commit Hash (Optional)
   
   _No response_
   
   ### What browsers are you seeing the problem on?
   
   _No response_
   
   ### Relevant log output
   
   ```shell
   
   ```


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to