Free Online HTML, JavaScript and XML Escape and Unescape
Escape text so it cannot break out of HTML, a JavaScript string or an XML document, and unescape it again. Each language gets its own rules: HTML entities named or numeric, JavaScript backslash sequences including the script tag trap, and the five entities XML actually defines. Nothing is uploaded.
Enjoying WizTools123? Help keep our server infrastructure 100% free and open for everyone.
Why This Page Does Not Unescape With innerHTML
The one-line trick that turns an unescaper into a hole.
There is a well known shortcut for unescaping HTML entities in a browser: make an element, set its
innerHTML to the text, and read back textContent. It is three lines
and it handles every entity the browser knows, which is all of them. It is also an XSS hole
the moment the text is not yours.
The reason is that innerHTML does not decode a string, it parses HTML. Give it
<img src=x onerror=alert(1)> and the browser builds an image element, tries
to load x, fails, and runs the handler. Script tags inserted that way do not run,
which is what makes people think the trick is safe, but event handlers on other elements do,
and so do a dozen other constructions. On a page where the input is pasted by the person
using it the damage is limited to themselves; in a shared component, or anywhere the text
arrived from a URL, a database or another user, it is a real vulnerability.
So the unescaping here is done the long way: a table of names to code points, written into the
page, and a scan that replaces one reference at a time. It costs a few hundred lines and it
cannot run anything, because nothing is ever handed to the HTML parser. Use
DOMParser with text/html if you need the browser's own table in your
own code, and never innerHTML on text you did not write.
How to Escape Text for HTML, JavaScript or XML
A few steps, and nothing is uploaded.
What to Know About Escaping
Including why unescaping with innerHTML is a bad idea.
innerHTML and reading textContent, hands untrusted text to the HTML parser. <img src=x onerror=...> runs code that way, which is how a decoder becomes an XSS hole. Nothing on this page is ever parsed as HTML.
</script inside a script element whether or not it is in the middle of a string, and ends the script there. So data that contains that text, escaped perfectly for JavaScript, still breaks the page and can inject markup. Writing it with a backslash in the middle means the same thing to JavaScript and nothing to the HTML parser, which is what the switch on this page does.
in XML, nor any of the other 247 names HTML 4 defined. An XML parser meeting an unknown entity does not ignore it, it refuses the whole document, which is why one copied line can make an entire feed or configuration file unreadable. Unescaping here names the ones XML will not accept instead of silently fixing them.
Key Features & Capabilities
What this tool does, and what it deliberately does not.
About Escaping and Unescaping
Escaping is what keeps text as text. A few characters mean something to whatever is going to read the string next, and if they are passed through untouched the reader stops treating them as content and starts treating them as structure. That is the whole of how cross site scripting works, and it is also why a single apostrophe in a surname can break a page that has worked for years. The fix is never to remove the character; it is to write it in the form that particular reader understands.
Which form depends entirely on who is reading. HTML wants entities, and in practice wants five of them. A JavaScript string wants backslash escapes, and it also wants the closing script tag broken up, which has nothing to do with JavaScript at all and everything to do with the HTML parser that gets there first. XML wants entities too, but it defines only five names and refuses a document containing any other, so text that is perfect HTML is often unusable XML. One escape button cannot serve all three, which is why this page asks where the text is going.
The unescaping is the part that took care, and not because the formats are hard. The obvious way to decode HTML entities in a browser is to let the browser do it through innerHTML, and that is a genuine security hole on untrusted text, so this page builds its own table instead and scans the string by hand. It costs more code, it is limited to the names in that table, and it cannot execute anything, which is the trade worth making.
Frequently Asked Questions
The five that matter, the script tag trap, and what XML does not have.
Every Other Security Tool
10 more tools in this set. All free, all in your browser.