This is an automated email from the ASF dual-hosted git repository.

damccorm pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/beam.git


The following commit(s) were added to refs/heads/master by this push:
     new 0f106b32122 [Python] Don't skip a delimiter right after an escaped one 
in ReadFromText (#40460)
0f106b32122 is described below

commit 0f106b32122b4dd07c7e112a188e9bf3ebcff7db
Author: Anish Mehta <[email protected]>
AuthorDate: Thu Oct 8 19:11:32 2026 +0530

    [Python] Don't skip a delimiter right after an escaped one in ReadFromText 
(#40460)
    
    * Don't skip a delimiter right after an escaped one in TextIO
    
    ReadFromText with escapechar merged the next record into the current one.
    
    * Add CHANGES.md entry
---
 CHANGES.md                                | 1 +
 sdks/python/apache_beam/io/textio.py      | 2 +-
 sdks/python/apache_beam/io/textio_test.py | 9 +++++++++
 3 files changed, 11 insertions(+), 1 deletion(-)

diff --git a/CHANGES.md b/CHANGES.md
index 9d8067a7126..5c5370006d6 100644
--- a/CHANGES.md
+++ b/CHANGES.md
@@ -92,6 +92,7 @@
 * (Java) IcebergIO now writes rows containing `EnumerationType` (proto enum) 
fields as strings, instead of throwing `Unsupported Beam logical type Enum` 
([#40299](https://github.com/apache/beam/issues/40299)).
 * (Python) Fixed stateful DoFns with side inputs sometimes taking the timer 
key coder from a side input instead of the main input, which could make the 
worker fail to decode timer keys with `Unknown type tag` 
([#40374](https://github.com/apache/beam/issues/40374)).
 * (Python) `Duration` built from float seconds now rounds to the nearest 
microsecond instead of truncating, which could lose a microsecond 
([#40263](https://github.com/apache/beam/issues/40263)).
+* (Python) `ReadFromText` with `escapechar` no longer skips a delimiter that 
directly follows an escaped delimiter, which merged two records into one 
([#40459](https://github.com/apache/beam/issues/40459)).
 * Fixed X (Java/Python) ([#X](https://github.com/apache/beam/issues/X)).
 
 ## Security Fixes
diff --git a/sdks/python/apache_beam/io/textio.py 
b/sdks/python/apache_beam/io/textio.py
index 5b2e6fc4736..a2af75f89d1 100644
--- a/sdks/python/apache_beam/io/textio.py
+++ b/sdks/python/apache_beam/io/textio.py
@@ -321,7 +321,7 @@ class _TextSource(filebasedsource.FileBasedSource):
           if self._escapechar is not None and self._is_escaped(read_buffer,
                                                                next_delim):
             # Skip an escaped delimiter.
-            current_pos = next_delim + delimiter_len + 1
+            current_pos = next_delim + delimiter_len
             continue
           else:
             # Found a delimiter. Accepting that as the next delimiter.
diff --git a/sdks/python/apache_beam/io/textio_test.py 
b/sdks/python/apache_beam/io/textio_test.py
index f5f2860d3ee..b67b1b3f0e6 100644
--- a/sdks/python/apache_beam/io/textio_test.py
+++ b/sdks/python/apache_beam/io/textio_test.py
@@ -1388,6 +1388,15 @@ class TextSourceTest(unittest.TestCase):
     self._run_read_test(
         file_name, expected_data, delimiter=b'*|', escapechar=b'\\')
 
+  def test_read_escaped_delimiter_followed_by_delimiter(self):
+    with TempDir() as tempdir:
+      file_name = tempdir.create_temp_file(lines=[b'li\\\n\nne\n'])
+      self._run_read_test(file_name, ['li\\\n', 'ne'], escapechar=b'\\')
+
+      file_name = tempdir.create_temp_file(lines=[b'li\\*|*|ne*|'])
+      self._run_read_test(
+          file_name, ['li\\*|', 'ne'], delimiter=b'*|', escapechar=b'\\')
+
   def test_read_escaped_lf_at_buffer_edge(self):
     file_name, expected_data = write_data(3, eol=EOL.LF, 
line_value=b'line\\\n')
     assert len(expected_data) == 3

Reply via email to