On Mon, Sep 14, 2026 at 01:16:15PM +0300, Eli Zaretskii wrote: > > Date: Sun, 13 Sep 2026 21:21:31 +0200 > > From: Patrice Dumas <[email protected]> > > Cc: [email protected] > > > > > > My proposal is to use UTF-8 as encoding independentely of Texinfo input > > > > file encoding, with the possibility to set the output encoding with > > > > OUTPUT_ENCODING_NAME. > > > > > > > > What do you think? > > > > > > I think this is an unusual thing to do, > > > > This is not unusual. We do the same for LaTeX, and for EPUB > > OUTPUT_ENCODING_NAME is even ignored (though this is mandated by > > the specification, not really our choice). > > That we do it in other cases (which I presume are less popular) > doesn't yet justify this one. > > > > and needs a serious justification. Do we have one? > > > > This was already discussed in the previous thread, to me there are at > > least three reasons > > * it is easier for Info reader, in particular for cross-references, which > > are problematic when the different manuals do not have the same > > encoding. Currently this is an issue with the stand-alone Info > > reader > > * if UTF-8 becomes the only encoding (or almost), it could simplify > > readers, install-info and any software that deals with the Info format > > I understand why it would be simpler, but we are making an existing > feature harder to use, so that's a disadvantage for which we should > have a good justification, IMO. > > > * there is no reason why a user would want a specific encoding, as > > interaction with Info is through software, so using one that is > > practical and can render any character is a good thing > > Using characters UTF-8 doesn't support is one (albeit rare) case. > Another use case is reading the Info file with a pager on a terminal > that cannot show UTF-8, or in general using the Info file in a > non-UTF-8 locale, where stuff like Grep might produce mojibake or fail > to find non-ASCII text entirely. > > > > Forcing users to use > > > OUTPUT_ENCODING_NAME is a nuisance. > > > > Again, I fail to see a use case where a user would want a specific > > encoding. The Info format cannot be manually edited anyway without > > messing up the tag tables, and there are escaping with control > > characters that also make manual editing very hazardous. > > See above. > > In general, removing features that were supported for a long time is > IMO harsh and should be rarely if ever done. But that's me.
If I understand correctly, the functionality still would exist, and be provided by OUTPUT_ENCODING_NAME. > > More generally, you asked for opinions, and I provided one. I hope > it's helpful in some way. > I don't fully understand your use case for non-UTF-8 encoded Info files, but it does appear to exist and it appears to be easy to continue the current behaviour of letting the input encoding determine the output encoding, so I think we should not change the behaviour. Patrice's question was to discover if users depended on or preferred the current behaviour and from your experience this appears to be the case, so this is a strong argument for not changing it. It does appear to me that this use case would entail the author of the Texinfo files also being the reader of the Info files, as they are produced and consumed in a similar locale environment. Usually, they would be independent: e.g. just because a Texinfo manual was written in a Latin-1 locale, doesn't mean that the user is reading it in a Latin-1 locale.
