Skip to content

[Bug]: GCS glob matching drops objects when the pattern contains non-ASCII characters (Go SDK) #39969

Description

@Shizuma5

What happened?

Beam version: 2.75.0 through current master (2.77.0-SNAPSHOT). SDK: Go. Filesystem: GCS (gs://).

Summary

fileio.MatchAll fails to return existing GCS objects when the glob pattern contains non-ASCII (multi-byte) characters. When empty matches are allowed, it returns no error at all; the affected objects are simply missing from the results.
This is a regression in 2.75.0 (#38099); the same patterns worked correctly in 2.74.0 and earlier.

Examples

A bucket gs://my-bucket containing:

gs://my-bucket/
├── singlebyte/
│   ├── english1.txt
│   └── english2.txt
├── multibyte/
│   ├── japanese.txt
│   └── 日本語.txt
└── multibytepath_é/
    ├── singlebyte_file1.txt
    └── singlebyte_file2.txt
patterns := beam.Create(s,
    "gs://my-bucket/singlebyte/*.txt",                      // all ASCII → OK
    "gs://my-bucket/multibyte/*.txt",                       // ASCII pattern, non-ASCII object names → OK
    "gs://my-bucket/multibyte/日本語.txt",                   // non-ASCII file name → broken
    "gs://my-bucket/multibytepath_é/*.txt",                 // non-ASCII directory → broken
    "gs://my-bucket/multibytepath_é/singlebyte_file1.txt",  // non-ASCII directory, exact path → broken
)
matched := fileio.MatchAll(s, patterns)  // PCollection<fileio.FileMetadata>
glob passed to MatchAll returned on 2.74 returned on 2.75+
.../singlebyte/*.txt english1.txt, english2.txt english1.txt, english2.txt
.../multibyte/*.txt japanese.txt, 日本語.txt japanese.txt, 日本語.txt
.../multibyte/日本語.txt 日本語.txt (none)
.../multibytepath_é/*.txt singlebyte_file1.txt, singlebyte_file2.txt (none)
.../multibytepath_é/singlebyte_file1.txt singlebyte_file1.txt (none)

When empty matches are allowed, no error is raised at all, so the missing objects are a silent data loss.

Root cause (summary)

globToRegex() in sdks/go/pkg/beam/io/filesystem/gcs/gcs.go iterates over the glob pattern byte by byte instead of rune by rune. In its default branch, it passes each byte c to regexp.QuoteMeta(string(c)).

For a multi-byte UTF-8 character, converting each byte separately with string(c) changes its encoding. As a result, globToRegex() produces a corrupted regular expression that does not match the original GCS object name.

This behavior was introduced by #38099 in version 2.75.0. Earlier versions used the rune-aware filepath.Match, so version 2.74.0 and earlier are unaffected.

Issue Priority

Priority: 2 (default / most bugs should be filed as P2)

Issue Components

  • Component: Python SDK
  • Component: Java SDK
  • Component: Go SDK
  • Component: Typescript SDK
  • Component: IO connector
  • Component: Beam YAML
  • Component: Beam examples
  • Component: Beam playground
  • Component: Beam katas
  • Component: Website
  • Component: Infrastructure
  • Component: Spark Runner
  • Component: Flink Runner
  • Component: Prism Runner
  • Component: Twister2 Runner
  • Component: Hazelcast Jet Runner
  • Component: Google Cloud Dataflow Runner

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions