HTML Tag Regex
Find opening, closing and self-closing HTML tags in a piece of text. Good for quick searches and cleanup; for real parsing, use the browser’s parser.
Link optionspattern only
The address bar holds the pattern, flags and replacement, so Copy link shares them. Your test text stays out of it unless you include it (up to 2,000 characters), because it may be private. Nothing is sent to a server.
Matches
12matches
The first is “<article class="post">”, at position 0.
- Groups
- none
- Characters matched
- 127 of 246
What the pattern meansHover or tap a part to see it in the pattern and what it matched
How this works: Method, 5 sources, Checked against 1 worked example,
How this works
Method
Your pattern runs in your own browser’s JavaScript engine, in a background worker that is stopped after 1.5 seconds, so a pattern that backtracks catastrophically can’t freeze the page. The explanation comes from our own parser of the ECMAScript pattern grammar, checked against the engine; what each part matched is found by wrapping that part in one more group and running the pattern again. Conversions to other flavours only rewrite the syntax and list what their documentation says works differently.
Sources
How it’s tested
One worked example for this page is checked by automated tests before every release: given the inputs, the tool must show the expected answer.
Changes
- First published, with a library of 19 common patterns and their test cases.
Worked example
Take <p>Hello</p>, the first case in the tester. Reading the pattern left to right, each part takes its share of the text:
<“<” matches </?Optionally “/” (matches a position, no characters)[a-z]One character: a–z (either case) matches p[a-z0-9-]*Zero or more characters, each a–z, 0–9 or “-” (either case) (matches a position, no characters)(?:\s[^<>]*)?A group (not captured), optional (zero or one time): (matches a position, no characters)/?Optionally “/” (matches a position, no characters)>“>” matches >
How the HTML tag regex works
A tag starts with <, and /? allows the slash of a closing tag. The name must begin with a letter, [a-z], followed by letters, digits or hyphens, [a-z0-9-]*, which covers custom elements such as my-widget. The i flag accepts names in any case, as HTML does.
Attributes are optional: (?:\s[^<>]*)? needs a whitespace character after the name, then takes everything up to the next angle bracket. Requiring that whitespace is what stops a name like <p2x from being split, and refusing < and > keeps one match inside one tag. Finally /?> closes the tag, with an optional slash for self-closing syntax.
Because the name must start right after the <, comparisons such as a < b and comments or doctypes, which begin with <!, are not matched. Run it with the g flag to list every tag in a document, or with replace to strip them.
- The HTML standard says a start tag begins with <, then the tag name, optional attributes, an optional / for void and foreign elements, and ends with >. Source: HTML Living Standard, section 13.1.2 (elements and start tags).
- HTML tag names use only ASCII alphanumerics and are case-insensitive in the HTML syntax. Source: HTML Living Standard, section 13.1.2 (elements and start tags).
- DOMParser parses HTML or XML source from a string into a DOM Document. Source: MDN, DOMParser.
Test cases
What the pattern finds in each line, and why.
| Text | Finds | Why |
|---|---|---|
| <p>Hello</p> | <p> · </p> | an opening and a closing tag |
| Line one<br>Line two<br/> | <br> · <br/> | void elements, with and without a slash |
| <a href="/docs" class="link">Docs</a> | <a href="/docs" class="link"> · </a> | attributes are part of the tag |
| <img src="cat.png" alt="A cat" /> | <img src="cat.png" alt="A cat" /> | a self-closing tag with a space before the slash |
| <my-widget data-id="7"></my-widget> | <my-widget data-id="7"> · </my-widget> | custom elements have hyphens in their names |
| <DIV>Shouting</DIV> | <DIV> · </DIV> | tag names are case-insensitive |
| if (a < b && c > d) | nothing | a less-than sign followed by a space is not a tag |
| <!-- a comment --> | nothing | comments start with <! and are skipped |
| <!DOCTYPE html> | nothing | the doctype is not an element |
| I <3 regex | nothing | a tag name must start with a letter |
| <p> is escaped text | nothing | escaped markup is just text |
| <a title="1 > 0">link</a> | <a title="1 > · </a> | where regex fails: it stops at the > inside the quotes |
What it doesn’t check
- This is a search pattern, not an HTML parser. A > inside a quoted attribute value is legal HTML and cuts the match short, as one of the failing cases shows.
- It also matches tag-like text inside script and style elements, comments and CDATA, where it is not markup at all.
- Do not use it to sanitise HTML for security: attackers can write markup that a regex misreads but a browser runs. Use a maintained sanitiser instead.
- In the browser, DOMParser turns a string into a real document, which is the reliable way to read tags and attributes.
The same pattern in other languages
Converted automatically from the JavaScript version; the tester above always runs JavaScript.
| Flavour | Pattern | Notes |
|---|---|---|
| Python | (?i)</?[a-z][a-z0-9\-]*(?:\s[^<>]*)?/?> | |
| PCRE | (?i)</?[a-z][a-z0-9\-]*(?:\s[^<>]*)?/?> | |
| Go | (?i)</?[a-z][a-z0-9\-]*(?:\s[^<>]*)?/?> | |
| Java | (?i)</?[a-z][a-z0-9\-]*(?:\s[^<>]*)?/?> | |
| .NET | (?i)</?[a-z][a-z0-9\-]*(?:\s[^<>]*)?/?> |
Sources
Frequently Asked Questions
Can you parse HTML with a regular expression?
Not reliably. HTML allows nesting, quoted attributes containing angle brackets, comments and scripts, which a single regex cannot follow. A regex is fine for quick searches and cleanup of simple, known markup.
How do I strip all HTML tags from a string?
Replace every match of this pattern with an empty string using the g flag. For text from users, read textContent from a parsed document instead, which also decodes entities such as & correctly.
How do I match only one kind of tag, such as links?
Replace the name part with the tag you want, for example a followed by a word boundary, and keep the attribute part. To read the link target as well, parse the HTML and read the href attribute.