[ 
https://issues.apache.org/jira/browse/CAMEL-25025?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119982#comment-18119982
 ] 

Andrea Cosentino commented on CAMEL-25025:
------------------------------------------

PR opened: https://github.com/apache/camel/pull/26961

Both fixes in one PR, a commit each, since they touch the same consumer threads.

----
_Claude Code on behalf of oscerd (Andrea Cosentino)._

> camel-mongodb - a change stream event with a non-ObjectId _id loops the 
> consumer forever
> ----------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25025
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25025
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-mongodb
>            Reporter: Andrea Cosentino
>            Assignee: Andrea Cosentino
>            Priority: Major
>
> h3. Summary
> A change stream event whose {{_id}} is not an {{ObjectId}} - a string, an int 
> or a compound key, all
> ordinary in MongoDB - makes the consumer fail on that event, regenerate its 
> cursor from the same resume
> token, read the same event again and fail again, once per 
> {{cursorRegenerationDelay}} for as long as the
> route runs.
> h3. Details
> {{MongoDbChangeStreamsThread.doRun()}}:
> {code:java}
> ObjectId documentId = dbObj.getDocumentKey().getObjectId(MONGO_ID).getValue();
> {code}
> Verified against {{bson-5.9.2}} and {{mongodb-driver-core-5.9.2}}:
> * {{BsonDocument.getObjectId(key)}} is {{throwIfKeyAbsent(key); 
> get(key).asObjectId();}}, so it throws
>   {{BsonInvalidOperationException}} when {{_id}} is absent or is not an 
> {{ObjectId}}.
> * {{BsonInvalidOperationException extends BSONException extends 
> RuntimeException}}, while
>   {{MongoException extends RuntimeException}} - they are siblings, so the 
> enclosing
>   {{catch (MongoException e)}} in {{doRun()}} does not catch it.
> * {{ChangeStreamDocument.getDocumentKey()}} is annotated {{@Nullable}} (as 
> are {{getFullDocument()}},
>   {{getResumeToken()}} and {{getOperationType()}}), so events without a 
> document key - {{invalidate}},
>   {{drop}}, {{rename}}, {{dropDatabase}} - throw a {{NullPointerException}} 
> on the same line.
> Either exception reaches {{MongoAbstractConsumerThread.run()}}, which logs 
> it, closes the cursor,
> regenerates it and continues. Because the failure happens *before* the 
> exchange is created, the resume
> token is never advanced, so the regenerated cursor resumes at the same event 
> and the cycle repeats
> indefinitely at {{cursorRegenerationDelay}} (default 1000 ms).
> h3. Relation to CAMEL-16025
> CAMEL-16025 reported the {{NullPointerException}} on this line when the 
> watched collection is dropped and
> was fixed in 3.10.0 by hardening {{MongoAbstractConsumerThread.run()}} so the 
> thread survives and the
> cursor is regenerated. That addressed thread death, which was the reported 
> symptom. This issue is the
> half that remains: for a *valid, resumable* event - a document with a string 
> {{_id}} - the consumer now
> survives but can never make progress past it.
> h3. Also worth fixing here
> The hardening from CAMEL-16025 logs the stack trace only on the branch where 
> the thread is stopping:
> {code:java}
> if (keepRunning) {
>     log.warn("... Will try again on next poll.", e.getMessage());
> } else {
>     log.warn("... ConsumerThread will be stopped.", e.getMessage(), e);
> }
> {code}
> The recurring-failure branch - the one that fires once a second in the loop 
> above - gets no stack trace,
> which makes this hard to diagnose from a log. The two branches should be the 
> other way round, or both
> should carry the exception.
> h3. Proposed fix
> Read the document key defensively: skip the header rather than throw when 
> {{getDocumentKey()}} is null or
> {{_id}} is not an {{ObjectId}}, so that the event still reaches the route and 
> the resume token can
> advance. Keep the {{CamelMongoDbObjectId}} header for the {{ObjectId}} case 
> so existing routes are
> unaffected.
> ----
> _Reported by Claude Code on behalf of oscerd (Andrea Cosentino)._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to