Kindly look into this for full trace of RS. http://pastebin.com/VS17vVd8
Thanks On Wed, Jun 10, 2015 at 11:35 AM, Ted Yu <[email protected]> wrote: > Can you pastebin the complete stack trace for the region server ? > > Thanks > > > > > On Jun 9, 2015, at 10:52 PM, mukund murrali <[email protected]> > wrote: > > > > We are using HBase-1.0.0. Just before the client stalled, in RS there > were > > few handler threads that were blocked for MVCC(thread stack below) > check. > > Not sure if it could cause a problem. I don't see anything unusual in RS > > threads. Also the same client can connect to regionserver after restart. > At > > that instant what causing the problem is what we are confused. > > > > > > java.lang.Thread.State: BLOCKED (on object monitor) > > at java.lang.Object.wait(Native Method) > > at > > > org.apache.hadoop.hbase.regionserver.MultiVersionConsistencyControl.waitForPreviousTransactionsComplete(MultiVersionConsistencyControl.java:224) > > - locked <0x00000007ac0e0e88> (a java.util.LinkedList) > > at > > > org.apache.hadoop.hbase.regionserver.MultiVersionConsistencyControl.completeMemstoreInsertWithSeqNum(MultiVersionConsistencyControl.java:127) > > at > > > org.apache.hadoop.hbase.regionserver.HRegion.doMiniBatchMutation(HRegion.java:2822) > > at > > > org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:2476) > > at > > > org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:2430) > > at > > > org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:2434) > > at > > > org.apache.hadoop.hbase.regionserver.RSRpcServices.doBatchOp(RSRpcServices.java:640) > > at > > > org.apache.hadoop.hbase.regionserver.RSRpcServices.doNonAtomicRegionMutation(RSRpcServices.java:604) > > at > > > org.apache.hadoop.hbase.regionserver.RSRpcServices.multi(RSRpcServices.java:1832) > > at > > > org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:31313) > > at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2031) > > at org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:107) > > at > > > org.apache.hadoop.hbase.ipc.RpcExecutor.consumerLoop(RpcExecutor.java:130) > > at > > org.apache.hadoop.hbase.ipc.RpcExecutor$1.run(RpcExecutor.java:107) > > at java.lang.Thread.run(Thread.java:745) > > > > > > > > > >> On Tue, Jun 9, 2015 at 6:48 PM, Anoop John <[email protected]> > wrote: > >> > >> Can you see at this time, what the threads at RS doing? Handlers > mainly.. > >> which version oh hbase? > >> > >>> On Tuesday, June 9, 2015, mukund murrali <[email protected]> > wrote: > >>> Hi > >>> > >>> I wrote a sample program with default client configurations and > created a > >>> single connection. I spawn client threads > > hbase.hconnection.threads.max > >>> from my client application and each thread insert data to hbase > cluster. > >>> Once a region split happens, all the hconnection threads(core pool and > >> max > >>> pool size were kept at 256) stalled at BoundedCompletionService.take() > >>> indefinitely. Even after the split completed it never resumed. > >>> > >>> So does it mean I have to create more instances of connection object > for > >> a > >>> cluster in such scenarios (which is really not needed) ? There was no > >>> exception (I expected a RejectedExecution) also in client side. So > >> changing > >>> the hbase.hconnection.threads.max, hbase.hconnection.threads.core can > >>> create such problem? > >>> > >>> > >>> > >>> On Sat, Jun 6, 2015 at 5:02 PM, ramkrishna vasudevan < > >>> [email protected]> wrote: > >>> > >>>> Not very sure on what could be the problem when the meta update > >> happened. > >>>> I would think that when the region split happened, there was some > issue > >> on > >>>> the meta update (as you said in the later mail). The splitted regions > >> would > >>>> not have been updated properly in the META. So any client > updates/reads > >>>> happening to this region would have stalled and hence your client > >>>> application also stalled. > >>>> > >>>> As I said the logs would be important here to know what happened. > This > >>>> could be one of a case and could be identified with the logs. > >>>> > >>>> Regards > >>>> Ram > >>>> > >>>> On Sat, Jun 6, 2015 at 1:25 PM, mukund murrali < > >> [email protected]> > >>>> wrote: > >>>> > >>>>> Sorry for misleading by specifying it as meta split. It was meta > >> update > >>>>> during a user region split. This had caused the stallation probably. > >> We > >>>>> have right now reverting client configs. Till now we didn't face the > >>>> issue > >>>>> again. Those changes causing some kindof exceptions or timeout was > >> what > >>>> we > >>>>> expected, but clients stalling indefinitely is what worrying us. > >>>>> > >>>>> On Friday 5 June 2015, Vladimir Rodionov <[email protected]> > >> wrote: > >>>>> > >>>>>> I would suggest reverting client config changes back to defaults. At > >>>>> least > >>>>>> we will know if the issue is somehow related to client config > >> changes. > >>>>>> On Jun 5, 2015 6:15 AM, "ramkrishna vasudevan" < > >>>>>> [email protected] <javascript:;>> wrote: > >>>>>> > >>>>>>> Hbase:meta getting split? It may b some user region, can u check > >>>> that? > >>>>> If > >>>>>>> ur meta was splitting then there is something wrong. > >>>>>>> Can u attach the log snippets. > >>>>>>> > >>>>>>> Sent from phone. Excuse typos. > >>>>>>> On Jun 5, 2015 6:00 PM, "mukund murrali" < > >> [email protected] > >>>>>> <javascript:;>> wrote: > >>>>>>> > >>>>>>>> Hi > >>>>>>>> > >>>>>>>> In our case there at that instance when the client thread > >> stalled, > >>>>>> there > >>>>>>>> was a hbase:meta region split happening. So what went wrong? If > >>>> there > >>>>>> is > >>>>>>> a > >>>>>>>> split why should hconnection thread stall? Since we changed the > >>>>> client > >>>>>>>> configuration caused this? I am once again specifying our client > >>>>>> related > >>>>>>>> changes we did > >>>>>>>> > >>>>>>>> hbase.client.retries.number => 5 > >>>>>>>> zookeeper.recovery.retry => 0 > >>>>>>>> zookeeper.session.timeout => 1000 > >>>>>>>> zookeeper.recovery.retry. > >>>>>>>> intervalmilli => 1 > >>>>>>>> hbase.rpc.timeout => 30000. > >>>>>>>> > >>>>>>>> Is zk timeout too low? > >>>>>>>> > >>>>>>>> > >>>>>>>> > >>>>>>>> > >>>>>>>> > >>>>>>>> > >>>>>>>> On Fri, Jun 5, 2015 at 11:37 AM, ramkrishna vasudevan < > >>>>>>>> [email protected] <javascript:;>> wrote: > >>>>>>>> > >>>>>>>>> When you started your client server was the META table > >> assigned. > >>>>>> May > >>>>>>> be > >>>>>>>>> some thing happened around that time and the client app was > >> just > >>>>>>> waiting > >>>>>>>> on > >>>>>>>>> the meta table to be assigned. It would have retried - Can > >> you > >>>>> check > >>>>>>> the > >>>>>>>>> logs.? > >>>>>>>>> > >>>>>>>>> So the best part here is the stand alone client was able to be > >>>>>>>> successful - > >>>>>>>>> which means the new clients were able to talk successfully > >> with > >>>> the > >>>>>>>>> server. And hence the restart of your client has solved your > >>>>>> problem. > >>>>>>>> It > >>>>>>>>> may be difficult to trouble shoot the exact issue with the > >>>> limited > >>>>>>> info - > >>>>>>>>> but see if your client app regularly gets stalled and then it > >> is > >>>>>> better > >>>>>>>> to > >>>>>>>>> trouble shoot your app and the way it accesses the server. > >>>>>>>>> > >>>>>>>>> On Fri, Jun 5, 2015 at 11:21 AM, PRANEESH KUMAR < > >>>>>>>> [email protected] <javascript:;> > >>>>>>>>> wrote: > >>>>>>>>> > >>>>>>>>>> The client connection was in stalled state. But there was > >> only > >>>>> one > >>>>>>>>>> hconnection thread found in our thread dump, which was > >> waiting > >>>>>>>>> indefinitely > >>>>>>>>>> in BoundedCompletionService.take call. Meanwhile we ran a > >>>>>> standalone > >>>>>>>> test > >>>>>>>>>> program which was successful. > >>>>>>>>>> > >>>>>>>>>> Once we restarted the client server, the problem got > >> resolved. > >>>>>>>>>> > >>>>>>>>>> The basic doubt is, when the hconnection thread stalled, why > >>>> the > >>>>>>> HBase > >>>>>>>>>> client failed to create any more hconnections(max pool size > >> was > >>>>>> 10). > >>>>>>> In > >>>>>>>>>> case of problem with table/meta regions how come the test > >>>> program > >>>>>>>>>> succeeded. > >>>>>>>>>> > >>>>>>>>>> Regards, > >>>>>>>>>> Praneesh > >>>>>>>>>> > >>>>>>>>>> On Fri, Jun 5, 2015 at 10:21 AM, ramkrishna vasudevan < > >>>>>>>>>> [email protected] <javascript:;>> wrote: > >>>>>>>>>> > >>>>>>>>>>> Can you tell us more. Is your client not working at all > >> and > >>>> it > >>>>> is > >>>>>>>>>> stalled ? > >>>>>>>>>>> Are you seeing some results but you find it slow than you > >>>>>> expected? > >>>>>>>>>>> > >>>>>>>>>>> What type of workload are you running? All the tables are > >>>>>> healthy? > >>>>>>>>> Are > >>>>>>>>>>> you able to read or write to them individually using the > >>>> hbase > >>>>>>> shell? > >>>>>>>>>>> > >>>>>>>>>>> On Fri, Jun 5, 2015 at 10:18 AM, PRANEESH KUMAR < > >>>>>>>>>> [email protected] <javascript:;> > >>>>>>>>>>> wrote: > >>>>>>>>>>> > >>>>>>>>>>>> Hi Ram, > >>>>>>>>>>>> > >>>>>>>>>>>> The cluster ran without any problem for about 2 to 3 > >> days > >>>>> with > >>>>>>> low > >>>>>>>>>> load, > >>>>>>>>>>>> once we enabled it for high load we immediately faced > >> this > >>>>>> issue. > >>>>>>>>>>>> > >>>>>>>>>>>> > >>>>>>>>>>>> Regards, > >>>>>>>>>>>> Praneesh. > >>>>>>>>>>>> > >>>>>>>>>>>> On Thursday 4 June 2015, ramkrishna vasudevan < > >>>>>>>>>>>> [email protected] <javascript:;>> wrote: > >>>>>>>>>>>> > >>>>>>>>>>>>> Is your cluster in working condition. Can you see if > >> the > >>>>>> META > >>>>>>>> has > >>>>>>>>>> been > >>>>>>>>>>>>> assigned properly? If the META table is not > >> initialized > >>>>> and > >>>>>>>> opened > >>>>>>>>>>> then > >>>>>>>>>>>>> your client thread will hang. > >>>>>>>>>>>>> > >>>>>>>>>>>>> Regards > >>>>>>>>>>>>> Ram > >>>>>>>>>>>>> > >>>>>>>>>>>>> On Thu, Jun 4, 2015 at 9:05 PM, PRANEESH KUMAR < > >>>>>>>>>>>> [email protected] <javascript:;> > >>>>>>>>>>>>> <javascript:;>> > >>>>>>>>>>>>> wrote: > >>>>>>>>>>>>> > >>>>>>>>>>>>>> Hi, > >>>>>>>>>>>>>> > >>>>>>>>>>>>>> We are using Hbase-1.0.0. We also facing the same > >> issue > >>>>>> that > >>>>>>>>> client > >>>>>>>>>>>>>> connection thread is waiting at > >> > >> > org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegionInMeta(ConnectionManager.java:1200). > >>>>>>>>>>>>>> > >>>>>>>>>>>>>> Any help is appreciated. > >>>>>>>>>>>>>> > >>>>>>>>>>>>>> Regards, > >>>>>>>>>>>>>> Praneesh > >> >
