adityamparikh opened a new pull request, #235:
URL: https://github.com/apache/solr-mcp/pull/235
## What
This removes `FieldNameSanitizer`. JSON keys and Markdown front-matter keys
now reach Solr exactly as given, as CSV and XML field names already do since
#205. Nested JSON objects are still flattened with underscores (`{"user":
{"name": …}}` → `user_name`), and the indexing response still lists the indexed
field names.
## Why
The sanitizer lowercased every name, replaced non-word characters with `_`,
trimmed and collapsed underscores, and prefixed a leading digit with `field_`.
Its javadoc justified this as "Solr's field naming requirements". Solr doesn't
have such requirements: the [reference
guide](https://solr.apache.org/guide/solr/latest/indexing-guide/fields.html)
only *recommends* alphanumeric/underscore names and says this is not strictly
enforced.
I checked unsanitized JSON against Solr 9.9 with the `_default` configset:
| Key sent | Field Solr created | `q=field:value` |
|---|---|---|
| `first-name` | `first-name` | ✅ |
| `product.price` | `product.price` | ✅ |
| `Title` | `Title` | ✅ (`title:` → undefined field) |
| `User Name` | `User_Name`, renamed by Solr | — |
| `café` | `caf_`, renamed by Solr | — |
- **Lowercasing was harmful.** Solr field names are case-sensitive. On a
collection with an explicit schema, a key such as `Title` or `firstName` was
renamed to a field the schema doesn't define, and Solr rejected the document.
- **Character replacement duplicated Solr.** The `_default` schemaless
update chain already replaces characters outside `[\w.-]`.
`testSpecialCharactersInFieldNames` passes **unchanged**: `field@with@at` still
ends up as `field_with_at`, but Solr now does the renaming.
- **It was inconsistent across formats.** CSV and XML names were never
sanitized.
**Behaviour change:** a JSON or front-matter key that differs from an
existing field only by case (`ID`, `Title`) is no longer folded into
`id`/`title`.
## Changes
- `JsonDocumentCreator` and `MarkdownDocumentCreator` use keys as given.
`FieldNameSanitizer` is deleted.
- Removed the sanitization claims from the `index-json-documents` tool
description, from the indexed-field-names line in the response, and from the
javadoc, `AGENTS.md`, `dev-docs/ARCHITECTURE.md` and `docs/FAQ.md`.
- Tests that pinned the sanitized names now assert that names are kept as
given:
- `testFieldNamesAreUsedAsGiven`: real-Solr round trip of `first-name`,
`product.price` and `Title`, with case preserved.
- `testDirectFieldNamesAreUsedAsGiven`
- `indexJsonDocuments_reportsFieldNamesAsIndexed`
- `testFrontMatterFieldNamesAreUsedAsGiven`
- `JsonDocumentCreatorTest`
`./gradlew build`: 421 tests, 0 failures, 0 skipped.
Merge note: this conflicts trivially with #233 (adjacent bullet lines in
`IndexingService`'s class javadoc) and with #207 (the front-matter loop in
`MarkdownDocumentCreator`). Whichever lands second needs a one-line fix-up.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]