URL Regex Tester
A readable pattern for web links with an http or https scheme, a dotted host name and an optional port, path, query and fragment.
One case per line; the m flag makes ^ and the end anchor work line by line.
Link optionspattern only
The address bar holds the pattern, flags and replacement, so Copy link shares them. Your test text stays out of it unless you include it (up to 2,000 characters), because it may be private. Nothing is sent to a server.
Matches
6matches
The first is “https://example.com”, at position 0.
- Groups
- none
- Characters matched
- 197 of 351
What the pattern meansHover or tap a part to see it in the pattern and what it matched
How this works: Method, 6 sources, Checked against 1 worked example,
How this works
Method
Your pattern runs in your own browser’s JavaScript engine, in a background worker that is stopped after 1.5 seconds, so a pattern that backtracks catastrophically can’t freeze the page. The explanation comes from our own parser of the ECMAScript pattern grammar, checked against the engine; what each part matched is found by wrapping that part in one more group and running the pattern again. Conversions to other flavours only rewrite the syntax and list what their documentation says works differently.
Sources
How it’s tested
One worked example for this page is checked by automated tests before every release: given the inputs, the tool must show the expected answer.
Changes
- First published, with a library of 19 common patterns and their test cases.
Worked example
Take https://example.com, the first case in the tester. Reading the pattern left to right, each part takes its share of the text:
^The start of the text (matches a position, no characters)httpThe text “http” (any case) matches https?Optionally “s” (either case) matches s://The text “://” matches ://(?:[a-z0-9](?:[a-z0-9-]*[a-z0-9])?\.)+A group (not captured), one or more times: matches example.[a-z]{2,}Two or more characters: a–z (either case) matches com(?::\d{1,5})?A group (not captured), optional (zero or one time): (matches a position, no characters)(?:[/?#]\S*)?A group (not captured), optional (zero or one time): (matches a position, no characters)$The end of the text (matches a position, no characters)
How the URL regex works
It starts with https?://: the letters http, an optional s, then the colon and two slashes. Limiting the scheme keeps out javascript: and data: links, which is usually what you want when someone pastes a link into a form.
The host is one or more labels, each [a-z0-9](?:[a-z0-9-]*[a-z0-9])? followed by a dot, so a label starts and ends with a letter or digit and may hold hyphens in between. The top-level part, [a-z]{2,}, needs two or more letters, which is why bare hosts and raw IP addresses do not match.
Everything after the host is optional. (?::\d{1,5})? allows a port of up to five digits, and (?:[/?#]\S*)? accepts a path, query or fragment as long as it starts with a slash, question mark or hash and has no spaces.
- RFC 3986 includes its own regular expression in Appendix B, a non-validating parser that splits any URI reference into scheme, authority, path, query and fragment. Source: RFC 3986, Appendix B (parsing a URI with a regular expression).
- URI schemes and host names are case-insensitive, with lower case the canonical form, which is why the pattern uses the i flag. Source: RFC 3986, URI generic syntax.
- URL.canParse() returns true or false for whether a string parses as a URL, and has worked across major browsers since December 2023. Source: MDN, URL.canParse().
Test cases
Every case runs as an automated test of this page’s pattern, so the table can’t drift from what the pattern really does.
| Text | Result | Why |
|---|---|---|
| https://example.com | Passes | scheme and host, nothing else |
| http://www.example.co.uk/path/to/page | Passes | several labels and a path |
| https://shop.example.org:8080/search?q=regex&lang=en | Passes | a port and a query string |
| https://example.com/#section-2 | Passes | a fragment |
| HTTPS://EXAMPLE.COM/Docs | Passes | schemes and hosts are case-insensitive (the i flag) |
| https://xn--bcher-kva.example/books | Passes | a punycode host |
| example.com | Fails | no scheme |
| ftp://example.com/file.txt | Fails | only http and https are accepted |
| https://exa mple.com | Fails | a space in the host |
| https://-example.com | Fails | a label may not start with a hyphen |
| http://localhost:3000 | Fails | a host without a dot |
| https:/example.com | Fails | one slash instead of two |
| https://example.com:123456 | Fails | a port has at most five digits |
What it doesn’t check
- The pattern checks shape, not reachability: a well-formed link can still point to a domain that does not exist.
- It rejects localhost, raw IP addresses such as http://192.168.0.1/ and other schemes; widen the host part if your users need them.
- It accepts any characters after the host except whitespace, so it will not catch a malformed query string or percent sign.
- In the browser, URL.canParse() or new URL() follows the URL Standard exactly and is the better test when you need a real parser rather than a quick filter.
The same pattern in other languages
Converted automatically from the JavaScript version; the tester above always runs JavaScript.
| Flavour | Pattern | Notes |
|---|---|---|
| Python | (?ai)^https?://(?:[a-z0-9](?:[a-z0-9\-]*[a-z0-9])?\.)+[a-z]{2,}(?::\d{1,5})?(?:[/?#]\S*)?\Z | |
| PCRE | (?i)^https?://(?:[a-z0-9](?:[a-z0-9\-]*[a-z0-9])?\.)+[a-z]{2,}(?::\d{1,5})?(?:[/?#]\S*)?\z | |
| Go | (?i)^https?://(?:[a-z0-9](?:[a-z0-9\-]*[a-z0-9])?\.)+[a-z]{2,}(?::\d{1,5})?(?:[/?#]\S*)?$ | |
| Java | (?i)^https?://(?:[a-z0-9](?:[a-z0-9\-]*[a-z0-9])?\.)+[a-z]{2,}(?::\d{1,5})?(?:[/?#]\S*)?\z | |
| .NET | (?i)^https?://(?:[a-z0-9](?:[a-z0-9\-]*[a-z0-9])?\.)+[a-z]{2,}(?::\d{1,5})?(?:[/?#]\S*)?\z | In .NET, \d and \w match any Unicode digit or letter. Pass RegexOptions.ECMAScript, or write [0-9] and [A-Za-z0-9_], for JavaScript’s ASCII-only meaning. |
Sources
Frequently Asked Questions
How do I match a URL with or without http?
Make the scheme optional by wrapping it in a group with a question mark, as in (?:https?://)? at the start. Be aware that this also accepts plain words with a dot in them, such as file names like report.pdf.
Why not use a regex that follows RFC 3986 exactly?
A complete RFC 3986 pattern is long, hard to read and still accepts strings browsers treat differently. For links people type, a readable pattern plus the URL constructor in code catches more real mistakes.
How do I find all URLs inside a block of text?
Remove the start and end anchors and add the g flag, then use matchAll. Expect to trim trailing punctuation, because a link at the end of a sentence is often followed by a full stop or a closing bracket.