[
https://issues.apache.org/jira/browse/CAMEL-25016?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Claus Ibsen resolved CAMEL-25016.
---------------------------------
Resolution: Fixed
Fixed by https://github.com/apache/camel/pull/26881 (merged to main for 4.23.0).
_Claude Code on behalf of davsclaus_
> Remote file consumers (FTP, FTPS, SFTP, SMB, Azure Files): with noop=true or
> idempotent=true, a file whose download fails is never retried
> ------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25016
> URL: https://issues.apache.org/jira/browse/CAMEL-25016
> Project: Camel
> Issue Type: Bug
> Components: camel-core
> Reporter: shashank
> Priority: Minor
> Fix For: 4.23.0
>
>
> In {{GenericFileConsumer.processExchange}}
> ({{GenericFileConsumer.java:469-513}}) a file is retrieved after
> {{processStrategy.begin(...)}} ({{:426}}) succeeded. When the retrieve fails,
> {{tryRetrievingFile}} ({{:516-556}}) throws, and the catch block at
> {{:499-511}} only removes the file from the in-progress repository, calls the
> exception handler and returns {{true}}. So {{processBatch}} counts the file
> as started, and {{GenericFileOnCompletion}}, which is only registered after a
> successful retrieve ({{:482}}), never runs.
> With {{idempotent=true}} and the default {{idempotentEager=true}}, which
> includes every {{noop=true}} endpoint, the idempotent key is added while
> polling ({{isValidFile}} -> {{notUnique}}, {{:669-689}}). Nothing removes it
> after the failed retrieve: {{processBatch}} only removes the key of the files
> for which {{processExchange}} returned {{false}} ({{:256-266}}, added by
> CAMEL-21947 for the read-lock-not-acquired case). The next polls skip the
> file because its key is present, so it is never retried, and only a WARN
> "Cannot retrieve file ..." is logged once. With the default memory repository
> this lasts until the application is restarted. With a persistent idempotent
> repository (JDBC, Infinispan, Hazelcast, ...) the file is never consumed
> again, not even after a restart.
> This happens on the remote file components, whose downloads fail on ordinary
> transient errors: FTP and FTPS (socket timeout, connection reset, {{426
> Transfer aborted}}), SFTP, SMB and Azure Files. The local file component is
> not affected in practice, because {{FileOperations.retrieveFile}} always
> returns true ({{FileOperations.java:247-251}}).
> With {{preMove}}, {{GenericFileRenameProcessStrategy.begin}} binds the pre
> moved file to the exchange ({{FILE_EXCHANGE_FILE}}), while the eager key was
> added for the original path, so removing the key of the exchange file would
> remove the wrong key.
> The same catch block also skips {{processStrategy.abort}}, so the exclusive
> read lock taken by {{begin}} is not released on a failed retrieve. For the
> built-in read locks of the remote components ({{none}}, {{changed}},
> {{rename}}) that has no lasting effect, but a custom
> {{exclusiveReadLockStrategy}} bean would keep its lock.
> *Reproduction*: the real {{FileConsumer}} / {{GenericFileConsumer}} / process
> strategies, with a {{FileOperations}} subclass (plugged in through the
> protected {{FileEndpoint.newFileConsumer}}) whose {{retrieveFile}} throws
> {{GenericFileOperationFailedException("Connection reset")}} once, as an FTP
> download failure does. {{FileConsumer}} shares {{GenericFileConsumer}} with
> the remote consumers, so this exercises the same code path. One file, poll
> delay 100 ms, result after 3 s:
> {noformat}
> noop=true no failure: processed=1 | 1 failure:
> processed=0 retrieveCalls=1 fileStillInDir=true keyInRepository=true
> idempotent=true (default move) no failure: processed=1 | 1 failure:
> processed=0 retrieveCalls=1 fileStillInDir=true keyInRepository=true
> {noformat}
> Expected: the next poll retrieves and processes the file. Actual: the file is
> never retried.
> A TLA+ model of the consumer (list, begin, retrieve, commit/rollback)
> violates "a listed file is eventually processed" in 4 states (List -> Begin
> -> RetrieveFail); with the fix below it holds.
> *Proposed fix:* handle a failed or ignored retrieve like a failed begin: call
> {{processStrategy.abort}} (releases the read lock, deletes a partial local
> work file, releases the retrieved resources), remove the in-progress entry
> and the eager idempotent key of the original file (by {{absoluteFileName}} or
> the {{FILE_IDEMPOTENT_KEY}} snapshot, as
> {{GenericFileOnCompletion.processStrategyRollback}} does, so it also works
> with {{preMove}}), and return {{false}}, then report the exception as today.
> Behaviour change: a poll in which no file could be downloaded now counts as
> idle ({{backoffIdleThreshold}}, {{sendEmptyMessageWhenIdle}}, {{greedy}}),
> the same as a poll in which no read lock was acquired.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)