How to find a string within another string using a Regular Expression instead of strpos or stripos
I was writing a piece of code in PHP the other day where I had to find a snippet of text within another longer piece of text (e.g an article) that contained a word. I then wanted to take X number of characters from that first point and return a snippet that didn't cut off the last word in the sentence.
At first I was using the PHP functions strpos and stripos but these don't allow you to use Regular Expressions as the search term (needle in the haystack as PHP.net calls the parameters) and therefore it meant that I was returning mismatches due to the search term being contained within other words.
E.G if I was looking for the word wool it would match woollen.
Therefore the answer was to use a custom function that made use of preg_match and a non greedy capture group at the beginning of a pattern that could be passed to the function (without delimiters).
The function is below
As you can see the pattern needs to be passed in without delimiters e.g instead of /\bwool\b/ or @\bwool\b@ just pass in \bwool\b.
I then add a capture group to the beginning that is non greedy so that it finds the first match from the start of the input string ^(.*?) and then if the pattern is found I can do a strlen on the matching group to get the starting position of the pattern.
If you want the pattern to be case-sensitive then you can just pass in TRUE or FALSE as the extra parameter and the ignore flag will be added to the end of the pattern.
An example of this code being used is below. The code is looping through an array of words looking for the first match within a longer string (some HTML) and then taking 250 characters of text from the starting point, ensuring the last word is a whole word match.
Also remember to wrap your word in preg_quote so that any special characters that are used by the Regular Expression engine e.g ? . + * [ ] ( ) { } etc are all characters that need to be escaped properly.
I found this function quite useful.
I was writing a piece of code in PHP the other day where I had to find a snippet of text within another longer piece of text (e.g an article) that contained a word. I then wanted to take X number of characters from that first point and return a snippet that didn't cut off the last word in the sentence.
At first I was using the PHP functions strpos and stripos but these don't allow you to use Regular Expressions as the search term (needle in the haystack as PHP.net calls the parameters) and therefore it meant that I was returning mismatches due to the search term being contained within other words.
E.G if I was looking for the word wool it would match woollen.
Therefore the answer was to use a custom function that made use of preg_match and a non greedy capture group at the beginning of a pattern that could be passed to the function (without delimiters).
The function is below
/**
* Function to find the first occurence of a regular expression pattern within a string
*
* @param string $regex
* @param string $str
* @param bool $ignorecase
* @return variant
*/
function preg_pos( $regex, $str, $ignorecase )
{
// build up the RegEx wrapping it in @ delimiters
$pattern = "@^(.*?)" . $regex . "@" . ($ignorecase===true ? "i" : "");
if( preg_match( $pattern, $str, $matches ) ) {
return strlen( $matches[ 1 ] );
}
return false;
}
As you can see the pattern needs to be passed in without delimiters e.g instead of /\bwool\b/ or @\bwool\b@ just pass in \bwool\b.
I then add a capture group to the beginning that is non greedy so that it finds the first match from the start of the input string ^(.*?) and then if the pattern is found I can do a strlen on the matching group to get the starting position of the pattern.
If you want the pattern to be case-sensitive then you can just pass in TRUE or FALSE as the extra parameter and the ignore flag will be added to the end of the pattern.
An example of this code being used is below. The code is looping through an array of words looking for the first match within a longer string (some HTML) and then taking 250 characters of text from the starting point, ensuring the last word is a whole word match.
// find first occurence of any of the terms I am looking for and then take 250 characters from the first word
// ensuring I get a whole word at the end
$a = explode(" ",$terms);
foreach($a as $w){
// skip empty or small terms
if(!empty($w) && strlen($w) > 2){
// get the position of the word ensuring its not within another word - using \b word boundary - notice no RegEx delimiters @regex@ or /regex/
// also ensure any special characters within the word are delimited to prevent a mismatch
$pos = preg_pos( "\b" . preg_quote($w) . "\b", $html, true ) ;
// if pos is false then its empty otherwise
if($pos !== false){
// found the word take 250 chars from the first occurrence
$text = substr($html, $pos, 250);
// roll back to last space before our last word to ensure we don't get partial words
$text = substr($text, 0, strrpos($text," "));
// now we have found a term exit
break;
}
}
}
Also remember to wrap your word in preg_quote so that any special characters that are used by the Regular Expression engine e.g ? . + * [ ] ( ) { } etc are all characters that need to be escaped properly.
I found this function quite useful.