swzoh wrote: > So, if a first byte or following byte(s) is missing or orphaned, then > PCRE engine would probably detect it and produce an error. Thus, I > reckon there would be no need to worry about it in UTF-8 case.
Sorry I sent you the wrong pseudo-expression... it's more specific cases that are problematic actually: stuff like ([^x]+|) It seems it simply stops matching in this case so no error is raised by the plugin. Oops! I should look into it really but I think I already know what the problem is so... Does someone know the best way to reliably identify (and skip) the whole UTF-8 character in such cases? To reproduce the bug, replace your line: > local szRes=regex.umg(zText;;+ > ,"[^"++esc(?"\xC3\x82",?"\")++"]+",?"\0 ") with this: regex.umg(zText,"([^"++esc(?"\xC3\x82",?"\")++"]+|)",?"\0 ") > My previous example didn't require other plugins except the file > plugin. You used a plugin to call MS functions. But yeah, my PP is shamefully old.
