[ 
https://issues.apache.org/jira/browse/SOLR-18324?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121228#comment-18121228
 ] 

Marc Byrd commented on SOLR-18324:
----------------------------------

I hit a closely related NPE while testing the real Solr 10.1.0 RC1 build via 
the V2 admin API.  Same root symptom as this ticket:  a request fails before 
V2HttpCall attaches a SolrQueryRequest to the Jersey context, and downstream 
code unconditionally dereferences the null req, but a different specific call 
site than the ones covered by PR #4676.

Request: POST /api/collections/\{collection}/shards/\{shard}/replicas 
(ADDREPLICA), immediately followed by what looks like an Admin-UI-triggered 
collection reload, both on the same thread:
First NPE (covered by existing fix): 

{noformat}
java.lang.NullPointerException: Cannot invoke 
"org.apache.solr.request.SolrQueryRequest.getCore()" because "req" is null at 
org.apache.solr.handler.RequestHandlerBase.processReceivedException(RequestHandlerBase.java:326)
 at 
org.apache.solr.jersey.CatchAllExceptionMapper.processAndRespondToException(CatchAllExceptionMapper.java:93)
 at 
org.apache.solr.jersey.CatchAllExceptionMapper.toResponse(CatchAllExceptionMapper.java:84)
 ...
Second NPE (new call site): java.lang.NullPointerException: Cannot invoke 
"org.apache.solr.request.SolrQueryRequest.getContext()" because "req" is null 
at 
org.apache.solr.cloud.api.collections.AdminCmdContext.(AdminCmdContext.java:46) 
at 
org.apache.solr.handler.admin.api.AdminAPIBase.submitRemoteMessageAndHandleAsync(AdminAPIBase.java:150)
 at 
org.apache.solr.handler.admin.api.AdminAPIBase.submitRemoteMessageAndHandleResponse(AdminAPIBase.java:168)
 at 
org.apache.solr.handler.admin.api.ReloadCollectionAPI.reloadCollection(ReloadCollectionAPI.java:67)
 ...{noformat}

The first NPE is the exact one this ticket already fixes 
(RequestHandlerBase.processReceivedException -> req.getCore()). The second is a 
different null-req dereference, in AdminCmdContext's constructor via 
AdminAPIBase.submitRemoteMessageAndHandleAsync, which isn't in PR #4676's 
current file list: * RequestHandlerBase
 * CatchAllExceptionMapper
 * PostRequestDecorationFilter
 * PostRequestLoggingFilter
 * RequestMetricHandling


AdminAPIBase.submitRemoteMessageAndHandleAsync/Response is shared by a lot of 
V2 collection-admin operations (ReloadCollectionAPI and others), so I'd guess 
this same NPE is reachable from other admin operations that go through the same 
base method whenever req ends up null in this code path and therefore worth the 
same null-guard treatment as the other classes in this fix, if the underlying 
"why is req null here" isn't itself something this PR already addresses 
upstream of AdminCmdContext.

Net effect client-side (before this ticket's fix): The client gets zero 
diagnostic information for whatever the original failure was — just an opaque 
500.

Happy to provide more repro detail if useful (this was reproduced against a 
live Solr 10.1.0 RC1 build, commit b29b61bcc163c6fb083666f3e5237d2b86d13833).

> V2 (JAX-RS) requests can throw a cascading NPE / RTimer assertion failure 
> when a `SolrQueryRequest` isn't attached to the request context
> -----------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: SOLR-18324
>                 URL: https://issues.apache.org/jira/browse/SOLR-18324
>             Project: Solr
>          Issue Type: Bug
>          Components: v2 API
>    Affects Versions: 10.0
>            Reporter: Eric Pugh
>            Assignee: Eric Pugh
>            Priority: Major
>              Labels: pull-request-available
>         Attachments: repro-get-collection-status.sh
>
>          Time Spent: 1h 20m
>  Remaining Estimate: 0h
>
> BIG CAVEAT: This was all CLAUDE...   It found it, and raised it to me.   I 
> asked claude to make me a script that would reproduce the error on main 
> branch, and it does fail, and then pass on this branch.   
> However, I am feeling a bit out of my depth on is this really a problem and 
> really need a fix, since the issue would presumably have been there for a 
> while?    Hoping smarter than I folks can look at this one...   
> [~gerlowskija]?
> Discovered while migrating `bin/solr delete` off direct ZooKeeper access 
> (SOLR-18321) onto V2 APIs. `CollectionsApi.GetCollectionStatus` (`GET 
> /api/collections/{name}`) intermittently failed under a basic-auth-secured 
> cluster with:
> {code}
> SEVERE: An exception has been thrown from an exception mapper class 
> org.apache.solr.jersey.CatchAllExceptionMapper.
> java.lang.NullPointerException: Cannot invoke 
> "org.apache.solr.request.SolrQueryRequest.getCore()" because "req" is null
>       at 
> org.apache.solr.handler.RequestHandlerBase.processReceivedException(RequestHandlerBase.java:326)
>       at 
> org.apache.solr.jersey.CatchAllExceptionMapper.processAndRespondToException(CatchAllExceptionMapper.java:93)
> {code}
> which cascaded into a second NPE 
> (`PostRequestDecorationFilter`/`PostRequestLoggingFilter` dereferencing the 
> same null `SolrQueryRequest`), which caused Jersey to re-invoke the response 
> filter chain via its "already mapped exception" fallback path, which in turn 
> caused `RequestMetricHandling.PostRequestMetricsFilter` to call `.stop()` a 
> second time on an already-stopped timer — tripping `RTimer`'s internal 
> `assert state == STARTED || state == PAUSED`. That assertion is compiled out 
> in production builds (invisible there), but fires under test/dev builds run 
> with `-ea`.
> *Root cause:* when a V2 request fails **before** `V2HttpCall` attaches a 
> `SolrQueryRequest` to the Jersey request context — the trigger appears to be 
> a rare timing issue in inter-node request handling, not yet fully root-caused 
> and not specific to any one endpoint — several filters/mappers in 
> `org.apache.solr.jersey` unconditionally dereference that possibly-null 
> request, turning a recoverable failure into a crash. 
> `MediaTypeOverridingFilter` already had the correct null-check, with a `// 
> TODO Is it valid for SQRequest to be null?` comment confirming this was a 
> known-but-unaddressed possibility elsewhere.
> *Reproduction:* confirmed via a throwaway integration test hitting 
> `GetCollectionStatus` against a `SecurityJson.SIMPLE`-secured 2-node 
> `MiniSolrCloudCluster` — roughly a 25–50% failure rate over repeated runs 
> pre-fix, 0% post-fix (20+ consecutive passes observed).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to