--- In [email protected], Julien Pierrehumbert <[EMAIL PROTECTED]>
wrote:
>
> Does someone know the best way to reliably identify (and skip) the
> whole UTF-8 character in such cases?
I think you can use, like the one below:
[^\x01-\x{10FFFF}]
> To reproduce the bug, replace your line:
> > local szRes=regex.umg(zText;;+
> > ,"[^"++esc(?"\xC3\x82",?"\")++"]+",?"\0 ")
> with this:
> regex.umg(zText,"([^"++esc(?"\xC3\x82",?"\")++"]+|)",?"\0 ")
Is this what you had in mind?
local szRes=regex.umg(zText;;+
,"([^"++esc(?"\xC3\x82",?"\")++?"]+|[^\x01-\x{10FFFF}])",?"\0 ")
> > My previous example didn't require other plugins except the file
> > plugin.
>
> You used a plugin to call MS functions. But yeah, my PP is
> shamefully old.
Oh, they were commented inside /* ... */.
I left it in case for users to want to test other cases.
Sean