--- In [email protected], Julien Pierrehumbert <[EMAIL PROTECTED]>
wrote:
>
> Is a regex really the best way to identify a UTF-8 character and
> determine its size?
I'm not sure what you're actually looking for. There are unicode
sequences as \p, \P, \X, but it requires --enable-unicode-properties
at compliling time.
> > local szRes=regex.umg(zText;;+
> > ,"([^"++esc(?"\xC3\x82",?"\")++?"]+|[^\x01-\x{10FFFF}])",?"\0 ")
>
> No, what I had in mind is exactly what I wrote.
>
> What you wrote doesn't seem to trigger any bug... do you think it does?
No, it was supposed to not trigger the symptom.
BTW, I can't tell for sure if it's a bug or not. The empty pattern
simply seems to work at byte level, not at character level, like the
operator \C. So, it made PCRE engine stopping further process after
its first trigger as it made the (UTF-8) following byte \x82 orphaned
by removing the first byte \xC3.
Sean