swzoh wrote:
>> Does someone know the best way to reliably identify (and skip) the
>> whole UTF-8 character in such cases?
>
> I think you can use, like the one below:
> [^\x01-\x{10FFFF}]
Is a regex really the best way to identify a UTF-8 character and
determine its size?
I don't know what \x{} does so I can't say if it's realiable but there
has to be a better way... I think I'd rather study the UTF-8 rules and
implement them myself! Surely there's a standard library function that
can take care of this.
>> To reproduce the bug, replace your line:
>>> local szRes=regex.umg(zText;;+
>>> ,"[^"++esc(?"\xC3\x82",?"\")++"]+",?"\0 ")
>> with this:
>> regex.umg(zText,"([^"++esc(?"\xC3\x82",?"\")++"]+|)",?"\0 ")
>
> Is this what you had in mind?
>
> local szRes=regex.umg(zText;;+
> ,"([^"++esc(?"\xC3\x82",?"\")++?"]+|[^\x01-\x{10FFFF}])",?"\0 ")
No, what I had in mind is exactly what I wrote.
What you wrote doesn't seem to trigger any bug... do you think it does?