adityamparikh opened a new issue, #241:
URL: https://github.com/apache/solr-mcp/issues/241

   To turn a question into `q` and `fq`, a client needs the fields that 
**actually hold data** in a collection, and their types. `get-schema` answers a 
different question:
   
   - It returns the full Schema API representation (`SchemaService.java:286`). 
On the `_default` configset that is 68 field types (with analyzers) and 69 
dynamic-field *patterns*, about **34.8 KB**.
   - It never says which dynamic fields are in use. The books sample has 
`series_t`, `sequence_i` and `genre_s`; `get-schema` shows only the patterns 
`*_t`, `*_i` and `*_s`.
   
   Solr's Luke handler lists the fields that are really indexed, including 
dynamic-field instances. On the same collection its response is about **3.7 
KB**. SolrJ already parses it (`LukeRequest`, `LukeResponse.FieldInfo`), and 
the server already calls it for `get-collection-stats` 
(`CollectionService.java:537`).
   
   ## Luke is not distributed
   
   On a 2-shard SolrCloud 9.9.0 collection, each collection-level Luke call 
went to **one random shard**:
   - A field present on only one shard appeared in one call and was missing 
from the next.
   - The per-field `docs` counts were per shard.
   
   A correct listing has to ask every shard and merge the answers.
   
   ## Proposal
   
   ### Tool: `list-fields`
   
   Read-only. Returns the collection's `uniqueKey` and, per field:
   
   | Field | Source |
   |---|---|
   | `name` | Luke |
   | `type` | `FieldInfo.getType()` |
   | `dynamicBase` (e.g. `*_s`); null for explicit fields | Luke's raw 
`NamedList` (`FieldInfo` has no getter for it) |
   | `indexed`, `stored`, `docValues`, `multiValued` | 
`FieldInfo.getSchemaFlags()` (already parsed into `EnumSet<FieldFlag>`) |
   | `docs`, summed across shards; **null** for point (numeric) fields, which 
Luke can't count | `FieldInfo.getDocs()` |
   
   How it works:
   1. A `CLUSTERSTATUS` request for the collection 
(`CollectionAdminRequest.getClusterStatus()`) returns one active replica per 
shard.
   2. Send `LukeRequest` with `numTerms=0` and `distrib=false` to each 
replica's core. The server's `HttpJdkSolrClient` addresses a core by name the 
same way it addresses a collection.
   3. Merge: union the fields, sum `docs`.
   
   The shard requests can run concurrently; virtual threads are already 
enabled. The response is a new record, so it needs a `SolrNativeHints` entry 
like the other tool responses.
   
   ### Resource: `solr://{collection}/fields`
   
   The server already pairs reference-data tools with resources 
(`solr://collections`, `solr://{collection}/schema`). Add the same pairing:
   - a `solr://{collection}/fields` resource backed by the same method, 
returning JSON like `getSchemaResource` does (including its `{"error": ...}` 
shape on failure);
   - an `@McpComplete(uri = "solr://{collection}/fields")` entry next to the 
existing schema one (`CollectionService.java:338`), so `{collection}` 
autocompletes.
   
   A user attaching "the fields of `shows`" adds about 4 KB of context instead 
of about 35 KB. The tool stays the primary path, because many clients don't let 
the model read resources on its own.
   
   ### Descriptions split by purpose
   
   Two field-related tools only route well if their descriptions name different 
jobs:
   
   - `list-fields`: "List the fields that hold data in a collection, with their 
types and how many documents have each. Use this before `search` to decide 
which fields go in `q` and `fq`."
   - `get-schema`: "Full schema definition: field types, analyzers and 
dynamic-field patterns. Use this before `add-fields` or `add-field-types` to 
see what already exists."
   
   Then point the other guidance at the right tool:
   - the server instructions' "use list-collections and get-schema to see what 
exists before searching or indexing" → `list-fields` before searching, 
`get-schema` before changing the schema;
   - step 1 of the `search-collection` prompt → `list-fields`.
   
   ## Tests
   
   - [ ] Integration, 2 shards: a field indexed on only one shard is listed, 
and its summed `docs` equals `numFound` for `fq=field:[* TO *]`. This is 
exactly what a single Luke call gets wrong.
   - [ ] Integration: dynamic-field instances report their `dynamicBase`; 
explicit fields report null.
   - [ ] Integration: numeric point fields report `docs` as null, not 0.
   - [ ] Integration via `McpClientIntegrationTestBase`: the tool and the 
resource return the same fields for `shows`; `{collection}` completion works 
for the new URI; `listToolsReturnsExpectedTools` and `toolsExposeBehaviorHints` 
include `list-fields` as read-only.
   - [ ] Passes in the Solr compatibility matrix (8.11 to 10).
   
   ## Notes
   
   - #231 also changes the Luke call in `get-collection-stats`; whichever lands 
second rebases onto it.
   - Related: #183 (collections sharing `_default` see each other's schemaless 
fields). Listing a collection's fields makes that visible; it doesn't fix it.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to