Thanks for the pointer! It works.

btw, in pig 0.8, you need to "set pig.splitCombination false" to get
all the paths.

Regards,
Shawn

On Tue, Mar 29, 2011 at 10:47 AM, Kim Vogt <[email protected]> wrote:
> Yeah like Will said, you need to add something like this to your load func:
>
>   https://gist.github.com/892718
>
>
>
> On Tue, Mar 29, 2011 at 6:31 AM, Will Duckworth 
> <[email protected]>wrote:
>
>> I have done exactly this.  You are going to want to override
>> prepareToRead and then look at getWrappedSplit on PigSplit object.
>>
>> -----Original Message-----
>> From: Xiaomeng Wan [mailto:[email protected]]
>> Sent: Monday, March 28, 2011 4:45 PM
>> To: [email protected]
>> Subject: how to get input path in loader
>>
>> Hi,
>> I am trying to write a loader which can append the input path as a
>> field, something like this:
>>
>> hadoop fs -ls a/
>> 1.txt
>> 2.txt
>> 3.txt
>>
>> a = load 'a/*' using MyLoader() as (id, path);
>> dump a;
>>
>> (x,a/1.txt)
>> (y,a/2.txt)
>> (z,a/3.txt)
>>
>> I tried to subclass PigStorage, and append location to the end of the
>> returned tuple, but what I got is:
>>
>> (x,a/*)
>> (y,a/*)
>> (z,a/*)
>>
>> Where can I find the absolute path of the input file? Thanks!
>>
>> Shawn
>>
>

Reply via email to