Jannik Rebmann created NIFI-16359:
-------------------------------------

             Summary: Port GitHubFlowRegistryClient API-call optimizations 
(NIFI-14837) to GitLabFlowRegistryClient
                 Key: NIFI-16359
                 URL: https://issues.apache.org/jira/browse/NIFI-16359
             Project: Apache NiFi
          Issue Type: Improvement
          Components: Flow Versioning
    Affects Versions: 2.12.0
            Reporter: Jannik Rebmann


We are using GitLabFlowRegistryClient against a self-hosted GitLab instance. 
Our infrastructure team flagged a high load on the GitLab server that is being 
generated by the NiFi flow registry client. With a larger number of versioned 
process groups, NiFi's periodic version checks and version listings translate 
into a high volume of GitLab REST API calls (repository/commits, 
repository/tree, repository/files).

Looking into it, NIFI-14837 (fixed in 2.6.0, PR #10186) already reduced GitHub 
API load in GitHubFlowRegistryClient by (a) adding an in-memory LRU cache of 
commit details keyed by SHA (max 1000), and (b) limiting commit listing to the 
first page instead of paging the full history.

These optimizations were applied only to GitHubRepositoryClient. The sibling 
GitLabRepositoryClient (nifi-gitlab-bundle) still has no caching and, in 
getCommits(path, branch), calls CommitsApi.getCommits(projectPath, branch, 
resolvedPath), which pages the
entire commit history for the path (default 100/page). This is the most likely 
driver of the load our GitLab team is seeing.

Proposal: apply the same two improvements to GitLabRepositoryClient - an 
internal commit cache keyed by commit SHA, and a bounded commit listing (e.g. 
first page / configurable or hard limit) - mirroring the approach in PR #10186.

Reference: NIFI-14837, PR https://github.com/apache/nifi/pull/10186



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to