On Wed, Sep 16, 2026 at 06:38:03PM +0300, Eli Zaretskii wrote:
> > Date: Wed, 16 Sep 2026 16:47:28 +0200
> > From: [email protected]
> > Cc: [email protected], [email protected]
> > 
> > On Wed, Sep 16, 2026 at 05:07:28PM +0300, Eli Zaretskii wrote:
> > > > Date: Wed, 16 Sep 2026 15:55:27 +0200
> > > > From: [email protected]
> > > > Cc: Gavin Smith <[email protected]>, [email protected]
> > > > It depends where.  If it is some isolated sequence, you get some message
> > > > that there is something wrong going on, but the result should be usable.
> > > 
> > > OK, but what happens with the byte sequence which couldn't be
> > > converted in this case? is it skipped, and thus nothing due to it will
> > > appear in the produced Info file?
> > 
> > If it happens upon reading the Texinfo input file, there is an error
> > message in Perl and in C, the byte sequence is skipped in C, it will not
> > appear in the tree at all, in Perl it depends on what Encode::decode
> > does, I think that it is skipped too.
> 
> I wonder whether skipping is the best we can do.  How about emitting
> something instead, like a string showing the un-decodable bytes in hex
> or something?

The fact that we already have different output for C and for Perl indicates
that it may not be very easy to change the behaviour for such input.  We
already output a clear warning message and so this should be enough.  In short,
I do not believe it is worth spending any time or effort on trying to change
this.

I was able to test what texi2any did with a test file which was not correctly
formed UTF-8:

$ cat test.texi | xxd
00000000: 5c69 6e70 7574 2074 6578 696e 666f 2e74  \input texinfo.t
00000010: 6578 0a0a 4074 6f70 2054 6573 740a 0a40  ex..@top Test..@
00000020: 6368 6170 7465 7220 4120 4368 6170 7465  chapter A Chapte
00000030: 720a 0a41 41c3 4242 2e0a 0a40 6279 650a  r..AA.BB...@bye.

None the lone "c3" byte on the last line, between AA and BB.  This
input file is presumed by texi2any to be encoded in UTF-8 as it does
not have a @documentencoding declaration.

$ texi2any test.texi
test.texi:7: C:encoding error at byte 0xc3
test.texi: warning: document without nodes

Here the byte is omitted (just as Patrice said):

$ cat test.info | xxd | grep AABB
00000060: 2a2a 2a2a 2a2a 2a2a 2a2a 0a0a 4141 4242  **********..AABB

With Perl code, it is different:

$ TEXINFO_XS=omit texi2any test.texi
test.texi:7: encoding error at byte 0xc3
test.texi: warning: document without nodes

$ cat test.info | xxd | grep AA --after-context=1
00000060: 2a2a 2a2a 2a2a 2a2a 2a2a 0a0a 4141 efbf  **********..AA..
00000070: bd42 422e 0a0a 1f0a 5461 6720 5461 626c  .BB.....Tag Tabl

Here the bytes "ef bf bd" are output - UTF-8 for U+FFFD, the Unicode
replacement character (https://www.compart.com/en/unicode/U+FFFD).

Since presumably the Perl behaviour is built into Encode::decode, there
is no point trying to change what is done with the Perl code.  It may
be possible to output U+FFFD instead in the C code as well (although I
haven't looked at the part of the C code responsible for reading such
input).




Reply via email to