kz930 opened a new issue, #6761:
URL: https://github.com/apache/texera/issues/6761
### What happened?
# What happened
`KeywordSearchOpExec` picks one of two Lucene analyzers based on the
`isCaseSensitive` flag:
`common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/keywordSearch/KeywordSearchOpExec.scala:38`
```scala
@transient private lazy val analyzer: Analyzer = {
if (desc.isCaseSensitive) new CaseSensitiveAnalyzer() else new
StandardAnalyzer()
}
```
- **Case-insensitive (default)** → `StandardAnalyzer`: `StandardTokenizer`
splits
on Unicode word boundaries, so **punctuation is stripped** and `"perfect."`
tokenizes to `perfect`. Query `perfect` matches.
- **Case-sensitive** → `CaseSensitiveAnalyzer`, which is a bare
`WhitespaceTokenizer`:
`common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/keywordSearch/CaseSensitiveAnalyzer.scala:29`
```scala
// Achieves case sensitivity by skipping the lowercasing and normalization
// pipeline used in StandardAnalyzer.
class CaseSensitiveAnalyzer extends Analyzer {
override protected def createComponents(fieldName: String):
TokenStreamComponents = {
val tokenizer = new WhitespaceTokenizer()
val stream: TokenStream = new StopFilter(tokenizer,
CharArraySet.EMPTY_SET)
new TokenStreamComponents(tokenizer, stream)
}
}
```
`WhitespaceTokenizer` splits **only** on whitespace, so **punctuation stays
glued
to the token**: `"...absolutely perfect."` tokenizes to `[absolutely,
perfect.]`.
The query term `perfect` (tokenized the same way → `perfect`) does not equal
the
indexed token `perfect.`, so **the row is not matched**.
The intent of `CaseSensitiveAnalyzer` was only to *skip lowercasing*. But
swapping
`StandardTokenizer` for `WhitespaceTokenizer` also dropped Unicode
word-boundary
tokenization. So toggling "case sensitive" silently changes **punctuation
handling** too — a side effect the option name does not imply. A user who
turns on
case sensitivity to match `Trump` exactly will find that `USA.` no longer
matches
`USA`, `day!` no longer matches `day`, etc.
**Impact:** any keyword search over real text (which contains commas,
periods,
`!`, `?`) behaves inconsistently between the two modes. Words adjacent to
punctuation become unmatchable in case-sensitive mode.
** Proposed Fix:** make `CaseSensitiveAnalyzer` mirror `StandardAnalyzer`
*minus* the
lowercasing step — i.e. use `StandardTokenizer` (Unicode word-boundary
tokenization, strips punctuation) and simply omit the `LowerCaseFilter`,
instead
of falling back to `WhitespaceTokenizer`:
```scala
override protected def createComponents(fieldName: String):
TokenStreamComponents = {
val tokenizer = new StandardTokenizer()
// StandardAnalyzer pipeline without LowerCaseFilter: keep case, still
strip
// punctuation on Unicode word boundaries.
val stream: TokenStream = new StopFilter(tokenizer, CharArraySet.EMPTY_SET)
new TokenStreamComponents(tokenizer, stream)
}
```
Then "case sensitive" changes *only* case behavior, matching user
expectation and
staying consistent with the default mode's tokenization.
Notes:
- Present on `main`; not introduced by the standalone-translation work.
- Related: the standalone translation
(`KeywordSearchOpDesc.generateStandaloneCode`)
does not implement case sensitivity at all (always case-insensitive
`\b`-regex),
so it already diverges from this mode. Fixing the analyzer would make the
case-insensitive behavior of both paths line up on punctuated text; full
case-sensitive parity in the standalone is a separate gap.
### How to reproduce?
**Option A — user-facing (workflow):**
1. Add a **Keyword Search** operator over a text column whose values contain
trailing punctuation, e.g. a row `"everything was absolutely perfect."`.
2. Set keyword = `perfect`, **Case Sensitive = off** → the row matches
(expected).
3. Flip **Case Sensitive = on**, rerun → the same row **no longer matches**,
even
though `perfect` obviously appears. Only the case handling was supposed to
change.
**Option B — unit test (mirrors `KeywordSearchOpExecSpec`):**
```scala
val opDesc = new KeywordSearchOpDesc()
opDesc.attribute = "text"
opDesc.keyword = "perfect"
val schema = Schema().add(new Attribute("text", AttributeType.STRING))
def row(t: String) = Tuple.builder(schema).add(schema.getAttribute("text"),
t).build()
val data = List(row("everything was absolutely perfect."))
// case-insensitive: matches
opDesc.isCaseSensitive = false
val a = new KeywordSearchOpExec(objectMapper.writeValueAsString(opDesc));
a.open()
assert(data.exists(t => a.processTuple(t, 0).nonEmpty)) // passes
// case-sensitive: SHOULD still match "perfect", but does not
opDesc.isCaseSensitive = true
val b = new KeywordSearchOpExec(objectMapper.writeValueAsString(opDesc));
b.open()
assert(data.exists(t => b.processTuple(t, 0).nonEmpty)) // FAILS on
current main
```
### Version/Branch
1.3.0-incubating-SNAPSHOT (main)
### Commit Hash (Optional)
_No response_
### What browsers are you seeing the problem on?
_No response_
### Relevant log output
```shell
```
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]