Remove HTML tags without specifying names by JavaScript

Question

On JavaScript, it can remove all HTML tags in the text with regular expressions like this:

replace(/(<([^>]+)>)/ig, "")

In addition, I would like to keep specific tags.

ex)<h1>Text</h1><input type="text">Text</input><b>Text</b> → <h1>Text</h1>Text<b>Text</b>

I tried this code, but it doesn't work correctly.

replace(/<\/{0,1}!(font|h\d|p|hr|pre|blockquote|ol|ul|...).*?>/ig, "");

Please let me know the best formula.

There's numerous posts around explaining that parsing (i.e., analyzing) HTML will not succeed with REs. Simple tasks, such as eliminating all markup, will work, but more complex stuff won't work. This is due to the simplicity of regular languages (strings described by REs), compared to the complexity of HTML. My first attempt would be to filter the DOM. — virtualnobi
– virtualnobi, Commented Feb 6, 2014 at 12:11
I'll use strip_tags for JavaScript. I appreciate everybody's reply. — Tank2005
– Tank2005, Commented Feb 7, 2014 at 3:53

Community · Accepted Answer · 2017-05-23 12:28:40Z

1

THE PONY HE COMES

Especially in JavaScript, there is no excuse.

var div = document.createElement('div');
div.innerHTML = your_input_here;
var allowedtags = "font|h[1-6]|p|hr|...";

var rgx = new RegExp("^(?:"+allowedtags+")$","i");
var tags = div.getElementsByTagName('*');
var length = tags.length;
var i;
for( i=length-1; i>=0; i--) {
    if( !tags[i].nodeName.match(rgx)) {
        while(tags[i].firstChild) {
            tags[i].parentNode.insertBefore(tags[i].firstChild,tags[i]);
            // this will take all children and extract them
        }
        tags[i].parentNode.removeChild(tags[i]);
    }
}

var result = div.innerHTML;

edited May 23, 2017 at 12:28

CommunityBot

11 silver badge

answered Feb 6, 2014 at 12:07

Niet the Dark Absol

326k86 gold badges480 silver badges604 bronze badges

Sign up to request clarification or add additional context in comments.

Comments

dfsq · Accepted Answer · 2014-02-06 12:26:09Z

What about using such a simple function to remove unwanted tags:

function sanitize(text, allowed) {

    var tags = typeof allowed === 'string' ? allowed.split(',') : allowed;

    var a = document.createElement('div');
    a.innerHTML = text;

    for (var c = a.childNodes, i = c.length; i--;) {
        if (c[i].nodeType == 1) {
            c[i].innerHTML = sanitize(c[i].innerHTML, tags);
            if (tags.indexOf(c[i].tagName.toLowerCase()) === -1) {
                c[i].parentNode.removeChild(c[i]);
            }
        }
    }

    return a.innerHTML;
}

sanitize('<h1>This is a <script>alert(1)</script> test</h1> <input type="text"> and <b>this</b> should stay.', 'font,h1,h2,p,b,ul')

Output:

"<h1>This is a  test</h1>  and <b>this</b> should stay."

Or you can replace tag with it's text content if you use

c[i].parentNode.replaceChild(document.createTextNode(c[i].innerText);

instead of c[i].parentNode.removeChild(c[i]);

anubhava · Accepted Answer · 2014-02-07 05:23:48Z

0

You need to use negative lookahead:

replace(/<\/?(?!(font|h[1234]|p|hr|input|pre|blockquote|ol|ul))[^>]*>/ig, "");

Caution: HTML parsing and manipulation is error prone using regex like this. Better to use DOM parsers.

edited Feb 7, 2014 at 5:23

answered Feb 6, 2014 at 12:06

anubhava

790k67 gold badges603 silver badges671 bronze badges

1 Comment

Tank2005 Over a year ago

I trid this regex, but some tags will remain. <h1>Text</h1><input type="text">Text</input><b>Text</b> → <h1>Text<input type="text">Text<b>Text

Collectives™ on Stack Overflow

Remove HTML tags without specifying names by JavaScript

3 Answers 3

Comments

Comments

1 Comment

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

3 Answers 3

Comments

Comments

1 Comment

Your Answer

Sign up or log in

Post as a guest

Linked

Related