On Wed, Sep 16, 2026 at 06:38:03PM +0300, Eli Zaretskii wrote: > > Date: Wed, 16 Sep 2026 16:47:28 +0200 > > From: [email protected] > > Cc: [email protected], [email protected] > > > > On Wed, Sep 16, 2026 at 05:07:28PM +0300, Eli Zaretskii wrote: > > > > Date: Wed, 16 Sep 2026 15:55:27 +0200 > > > > From: [email protected] > > > > Cc: Gavin Smith <[email protected]>, [email protected] > > > > It depends where. If it is some isolated sequence, you get some message > > > > that there is something wrong going on, but the result should be usable. > > > > > > OK, but what happens with the byte sequence which couldn't be > > > converted in this case? is it skipped, and thus nothing due to it will > > > appear in the produced Info file? > > > > If it happens upon reading the Texinfo input file, there is an error > > message in Perl and in C, the byte sequence is skipped in C, it will not > > appear in the tree at all, in Perl it depends on what Encode::decode > > does, I think that it is skipped too. > > I wonder whether skipping is the best we can do. How about emitting > something instead, like a string showing the un-decodable bytes in hex > or something?
The fact that we already have different output for C and for Perl indicates that it may not be very easy to change the behaviour for such input. We already output a clear warning message and so this should be enough. In short, I do not believe it is worth spending any time or effort on trying to change this. I was able to test what texi2any did with a test file which was not correctly formed UTF-8: $ cat test.texi | xxd 00000000: 5c69 6e70 7574 2074 6578 696e 666f 2e74 \input texinfo.t 00000010: 6578 0a0a 4074 6f70 2054 6573 740a 0a40 ex..@top Test..@ 00000020: 6368 6170 7465 7220 4120 4368 6170 7465 chapter A Chapte 00000030: 720a 0a41 41c3 4242 2e0a 0a40 6279 650a r..AA.BB...@bye. None the lone "c3" byte on the last line, between AA and BB. This input file is presumed by texi2any to be encoded in UTF-8 as it does not have a @documentencoding declaration. $ texi2any test.texi test.texi:7: C:encoding error at byte 0xc3 test.texi: warning: document without nodes Here the byte is omitted (just as Patrice said): $ cat test.info | xxd | grep AABB 00000060: 2a2a 2a2a 2a2a 2a2a 2a2a 0a0a 4141 4242 **********..AABB With Perl code, it is different: $ TEXINFO_XS=omit texi2any test.texi test.texi:7: encoding error at byte 0xc3 test.texi: warning: document without nodes $ cat test.info | xxd | grep AA --after-context=1 00000060: 2a2a 2a2a 2a2a 2a2a 2a2a 0a0a 4141 efbf **********..AA.. 00000070: bd42 422e 0a0a 1f0a 5461 6720 5461 626c .BB.....Tag Tabl Here the bytes "ef bf bd" are output - UTF-8 for U+FFFD, the Unicode replacement character (https://www.compart.com/en/unicode/U+FFFD). Since presumably the Perl behaviour is built into Encode::decode, there is no point trying to change what is done with the Perl code. It may be possible to output U+FFFD instead in the C code as well (although I haven't looked at the part of the C code responsible for reading such input).
