--- In [email protected], Julien Pierrehumbert <[EMAIL PROTECTED]> 
wrote:
>
> > I'm not sure what you're actually looking for.
> 
> What I'm looking for is a standard library (as in glibc) function that 
> identifies UTF-8 characters and their length. That or consice and 
> precise rules for UTF-8 that would allow me to implement it myself.
I've 
> never manipulated UTF-8 so I have no idea where to look (not that you 
> should go out of your way to find that for me).

I thought you already found it.
Here is one where the rule is clearly summarized:

http://en.wikipedia.org/wiki/UTF-8


> Exactly. It's the plugin that makes the empty pattern "work at a byte 
> level" (more precisely: skip a byte and declare it to be unmatched each 
> time there's a zero-sized match). This is someting I did because it
made 
> sense and because Luciano said other regex implementations did the same 
> thing... but I didn't think about UTF-8 or CRLF back then obviously.
> I'd say it's a bug because, if you're offering UTF-8 services, you 
> should handle those characters properly and not as multiple ASCII 
> characters.
> 
> Of course I could invalidate zero-char matches instead like Sheri 
> suggests (I don't need the library to support it) but, last I heard, 
> this is not how such cases are usually handled. I don't know if this is 
> spelled out in some reference document somewhere or not.
>

IMHO, empty match should be avoided and it's the user's job to do it.
When saying that, I think the regex engine should help the user to
track down an error in the pattern code. To conclude my opinion, if
regex plugin meets an empty match, then it should stop processing
immediately and return the results so far.

For example, consider regex.mg("abcdef",".*?","\0|").
One would expect that the engine immediately escapes the would-be
infinite loop and returns only "|". However, in current
implementation, it returns "||||||". This could be very confusing.

Sean

Reply via email to