tomatotomata commented on issue #994: URL: https://github.com/apache/stormcrawler/issues/994#issuecomment-5159094569
I was thinking about taking this on as a focused change in the redirect handling path. I would first trace how a FETCHED URL becomes REDIRECT, then add the smallest configurable retry or error transition that lets the existing deletion flow remove the stale index entry. I would keep the current behavior available by default, add regression coverage for repeated redirects, and avoid changing unrelated WARC or indexing code. Does that direction still fit the issue, and may I take it? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
