[ 
https://issues.apache.org/jira/browse/SPARK-6527?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15259963#comment-15259963
 ] 

Steve Loughran edited comment on SPARK-6527 at 4/1/17 12:41 PM:
----------------------------------------------------------------

I've not seen a JIRA surface;

#  if anyone does, -link it to HADOOP-11694, S3a Phase II, which I'm trying to 
wrap up this week.-
# what are the characters in question?
# if it's not just when there are complex characters in a name, how many files 
in a directory tree does it take to trigger this problem.


looking into the Hadoop code, this specific error string appears if there is no 
match on a path containing a pattern, 
{code}
      Path p = dirs[i];
      FileSystem fs = p.getFileSystem(job.getConfiguration()); 
      FileStatus[] matches = fs.globStatus(p, inputFilter);
      if (matches == null) {
        errors.add(new IOException("Input path does not exist: " + p));
      } else if (matches.length == 0) {
        errors.add(new IOException("Input Pattern " + p + " matches 0 files"));
...
{code}

It might be that odd chars in filenames are confusing that pattern matching



was (Author: [email protected]):
I've not seen a JIRA surface;

#  if anyone does, link it to HADOOP-11694, S3a Phase II, which I'm trying to 
wrap up this week.
# what are the characters in question?
# if it's not just when there are complex characters in a name, how many files 
in a directory tree does it take to trigger this problem.


looking into the Hadoop code, this specific error string appears if there is no 
match on a path containing a pattern, 
{code}
      Path p = dirs[i];
      FileSystem fs = p.getFileSystem(job.getConfiguration()); 
      FileStatus[] matches = fs.globStatus(p, inputFilter);
      if (matches == null) {
        errors.add(new IOException("Input path does not exist: " + p));
      } else if (matches.length == 0) {
        errors.add(new IOException("Input Pattern " + p + " matches 0 files"));
...
{code}

It might be that odd chars in filenames are confusing that pattern matching


> sc.binaryFiles can not access files on s3
> -----------------------------------------
>
>                 Key: SPARK-6527
>                 URL: https://issues.apache.org/jira/browse/SPARK-6527
>             Project: Spark
>          Issue Type: Bug
>          Components: EC2, Input/Output
>    Affects Versions: 1.2.0, 1.3.0
>         Environment: I am running Spark on EC2
>            Reporter: Zhao Zhang
>            Priority: Minor
>
> The sc.binaryFIles() can not access the files stored on s3. It can correctly 
> list the number of files, but report "file does not exist" when processing 
> them. I also tried sc.textFile() which works fine.



--
This message was sent by Atlassian JIRA
(v6.3.15#6346)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to