Thanks for the pointer! It works. btw, in pig 0.8, you need to "set pig.splitCombination false" to get all the paths.
Regards, Shawn On Tue, Mar 29, 2011 at 10:47 AM, Kim Vogt <[email protected]> wrote: > Yeah like Will said, you need to add something like this to your load func: > > https://gist.github.com/892718 > > > > On Tue, Mar 29, 2011 at 6:31 AM, Will Duckworth > <[email protected]>wrote: > >> I have done exactly this. You are going to want to override >> prepareToRead and then look at getWrappedSplit on PigSplit object. >> >> -----Original Message----- >> From: Xiaomeng Wan [mailto:[email protected]] >> Sent: Monday, March 28, 2011 4:45 PM >> To: [email protected] >> Subject: how to get input path in loader >> >> Hi, >> I am trying to write a loader which can append the input path as a >> field, something like this: >> >> hadoop fs -ls a/ >> 1.txt >> 2.txt >> 3.txt >> >> a = load 'a/*' using MyLoader() as (id, path); >> dump a; >> >> (x,a/1.txt) >> (y,a/2.txt) >> (z,a/3.txt) >> >> I tried to subclass PigStorage, and append location to the end of the >> returned tuple, but what I got is: >> >> (x,a/*) >> (y,a/*) >> (z,a/*) >> >> Where can I find the absolute path of the input file? Thanks! >> >> Shawn >> >
