This is an automated email from the ASF dual-hosted git repository.
damccorm pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/beam.git
The following commit(s) were added to refs/heads/master by this push:
new 0f106b32122 [Python] Don't skip a delimiter right after an escaped one
in ReadFromText (#40460)
0f106b32122 is described below
commit 0f106b32122b4dd07c7e112a188e9bf3ebcff7db
Author: Anish Mehta <[email protected]>
AuthorDate: Thu Oct 8 19:11:32 2026 +0530
[Python] Don't skip a delimiter right after an escaped one in ReadFromText
(#40460)
* Don't skip a delimiter right after an escaped one in TextIO
ReadFromText with escapechar merged the next record into the current one.
* Add CHANGES.md entry
---
CHANGES.md | 1 +
sdks/python/apache_beam/io/textio.py | 2 +-
sdks/python/apache_beam/io/textio_test.py | 9 +++++++++
3 files changed, 11 insertions(+), 1 deletion(-)
diff --git a/CHANGES.md b/CHANGES.md
index 9d8067a7126..5c5370006d6 100644
--- a/CHANGES.md
+++ b/CHANGES.md
@@ -92,6 +92,7 @@
* (Java) IcebergIO now writes rows containing `EnumerationType` (proto enum)
fields as strings, instead of throwing `Unsupported Beam logical type Enum`
([#40299](https://github.com/apache/beam/issues/40299)).
* (Python) Fixed stateful DoFns with side inputs sometimes taking the timer
key coder from a side input instead of the main input, which could make the
worker fail to decode timer keys with `Unknown type tag`
([#40374](https://github.com/apache/beam/issues/40374)).
* (Python) `Duration` built from float seconds now rounds to the nearest
microsecond instead of truncating, which could lose a microsecond
([#40263](https://github.com/apache/beam/issues/40263)).
+* (Python) `ReadFromText` with `escapechar` no longer skips a delimiter that
directly follows an escaped delimiter, which merged two records into one
([#40459](https://github.com/apache/beam/issues/40459)).
* Fixed X (Java/Python) ([#X](https://github.com/apache/beam/issues/X)).
## Security Fixes
diff --git a/sdks/python/apache_beam/io/textio.py
b/sdks/python/apache_beam/io/textio.py
index 5b2e6fc4736..a2af75f89d1 100644
--- a/sdks/python/apache_beam/io/textio.py
+++ b/sdks/python/apache_beam/io/textio.py
@@ -321,7 +321,7 @@ class _TextSource(filebasedsource.FileBasedSource):
if self._escapechar is not None and self._is_escaped(read_buffer,
next_delim):
# Skip an escaped delimiter.
- current_pos = next_delim + delimiter_len + 1
+ current_pos = next_delim + delimiter_len
continue
else:
# Found a delimiter. Accepting that as the next delimiter.
diff --git a/sdks/python/apache_beam/io/textio_test.py
b/sdks/python/apache_beam/io/textio_test.py
index f5f2860d3ee..b67b1b3f0e6 100644
--- a/sdks/python/apache_beam/io/textio_test.py
+++ b/sdks/python/apache_beam/io/textio_test.py
@@ -1388,6 +1388,15 @@ class TextSourceTest(unittest.TestCase):
self._run_read_test(
file_name, expected_data, delimiter=b'*|', escapechar=b'\\')
+ def test_read_escaped_delimiter_followed_by_delimiter(self):
+ with TempDir() as tempdir:
+ file_name = tempdir.create_temp_file(lines=[b'li\\\n\nne\n'])
+ self._run_read_test(file_name, ['li\\\n', 'ne'], escapechar=b'\\')
+
+ file_name = tempdir.create_temp_file(lines=[b'li\\*|*|ne*|'])
+ self._run_read_test(
+ file_name, ['li\\*|', 'ne'], delimiter=b'*|', escapechar=b'\\')
+
def test_read_escaped_lf_at_buffer_edge(self):
file_name, expected_data = write_data(3, eol=EOL.LF,
line_value=b'line\\\n')
assert len(expected_data) == 3