[ 
https://issues.apache.org/jira/browse/HDDS-16631?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18121000#comment-18121000
 ] 

Adweta Ojha commented on HDDS-16631:
------------------------------------

Hi,

I'd like to work on this issue as my first Apache Ozone contribution.

I've successfully set up the Ozone development environment locally and would 
like to reproduce the bug, investigate the relevant code paths, and submit a 
fix with regression tests.

Please let me know if someone is already working on a fix so I can coordinate 
and avoid duplicating effort.

Thanks!

> NetworkTopology chooseRandom returns null early when one rack name is a 
> prefix of another
> -----------------------------------------------------------------------------------------
>
>                 Key: HDDS-16631
>                 URL: https://issues.apache.org/jira/browse/HDDS-16631
>             Project: Apache Ozone
>          Issue Type: Bug
>          Components: SCM
>            Reporter: Chu Cheng Li
>            Priority: Major
>
> When one rack name is a prefix of another, such as /rack1 and /rack10, 
> NetworkTopology#chooseRandom can return null even though datanodes are 
> available.
> InnerNodeImpl#getLeaf subtracts excluded scopes from each child with a plain 
> string check:
> {code:java}
> if (entry.getKey().startsWith(child.getNetworkFullPath())) {
> {code}
> The check has no path separator, so an excluded /rack10/dn1 is also 
> subtracted from /rack1. The per-rack counts then add up to less than the 
> number of available nodes that chooseNodeInternal gets from 
> getAvailableNodesCount. Random indexes near the top of that range fall past 
> the last rack, and getLeaf returns null. A rack whose count drops to zero can 
> never be picked. An excluded node is never returned; the effect is early 
> nulls and uneven picks.
> Example: racks /rack1 and /rack10 with 3 datanodes each, and two /rack10 
> datanodes passed as excluded scopes. chooseRandom(ROOT, excludedScopes, null, 
> null, 0) returned null in about half of 1000 calls, although 4 datanodes were 
> available.
> SCMContainerPlacementRackAware makes this call when it has no affinity node. 
> In clusters whose rack names share a prefix, it can fall back early and relax 
> the rack rule, or throw SCMException.
> HDDS-6640 fixed the same pattern in NodeImpl#isAncestor and 
> InnerNodeImpl#getLeafOnLeafParent but missed this loop. The fix is to use 
> child.isAncestor(scope), which compares with a trailing separator.
> This came up while working on a RackAware placement bug found by the SCM 
> simulation test in HDDS-16627.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to