--- In [email protected], Julien Pierrehumbert <[EMAIL PROTECTED]> 
wrote:
>
> Does someone know the best way to reliably identify (and skip) the
> whole UTF-8 character in such cases?

I think you can use, like the one below:
[^\x01-\x{10FFFF}]

> To reproduce the bug, replace your line:
> > local szRes=regex.umg(zText;;+
> > ,"[^"++esc(?"\xC3\x82",?"\")++"]+",?"\0 ")
> with this:
> regex.umg(zText,"([^"++esc(?"\xC3\x82",?"\")++"]+|)",?"\0 ")

Is this what you had in mind?

local szRes=regex.umg(zText;;+
,"([^"++esc(?"\xC3\x82",?"\")++?"]+|[^\x01-\x{10FFFF}])",?"\0 ")


> > My previous example didn't require other plugins except the file
> > plugin.
> 
> You used a plugin to call MS functions. But yeah, my PP is
> shamefully old.

Oh, they were commented inside /* ... */.
I left it in case for users to want to test other cases.

Sean

Reply via email to