Free Online HTML Entity Encoder and Decoder
Turn text into HTML entities or turn entities back into text, both directions on one page. Encode only the five characters that matter, use named entities where they exist, or write everything non-ASCII as decimal or hex references. Decoding reads named, decimal and hex forms, and the legacy ones missing a semicolon.
Enjoying WizTools123? Help keep our server infrastructure 100% free and open for everyone.
Why This Page Does Not Decode With innerHTML
The one-line trick that turns a decoder into a hole.
There is a well known shortcut for decoding entities in a browser: make an element, set its
innerHTML to the text, and read back textContent. It is three lines
and it handles every entity the browser knows, which is all of them. It is also an XSS hole
the moment the text is not yours.
The reason is that innerHTML does not decode a string, it parses HTML. Give it
<img src=x onerror=alert(1)> and the browser builds an image element, tries
to load x, fails, and runs the handler. Script tags inserted that way do not run,
which is what makes people think the trick is safe, but event handlers on other elements do,
and so do a dozen other constructions. On a page where the input is pasted by the person
using it the damage is limited to themselves; in a shared component, or anywhere the text
arrived from a URL, a database or another user, it is a real vulnerability.
So the decoding here is done the long way: a table of names to code points, written into the
page, and a scan that replaces one reference at a time. It costs a few hundred lines and it
cannot run anything, because nothing is ever handed to the HTML parser. Use
DOMParser with text/html if you need the browser's own table in your
own code, and never innerHTML on text you did not write.
How to Encode and Decode HTML Entities
A few steps, and nothing is uploaded.
What to Know About HTML Entities
Including why decoding with innerHTML is a bad idea.
innerHTML and reading textContent, hands untrusted text to the HTML parser. <img src=x onerror=...> runs code that way, which is how a decoder becomes an XSS hole. Nothing on this page is ever parsed as HTML.
' failed in older documents and still fails in XHTML served as XML in some parsers. The numeric reference ' has always worked everywhere. That is why the numeric form is the default here and the named one is a switch.
Key Features & Capabilities
What this tool does, and what it deliberately does not.
About the HTML Entity Tool
Entities exist because a few characters mean something to the HTML parser and cannot be written literally in text. In practice there are five of them, and everything else is either an old habit from the days of uncertain file encodings or a way of typing a character your keyboard does not have. Both directions of the conversion are needed constantly: escaping text on the way into a page, and cleaning up text that has been escaped twice on the way out of one.
This page does both, with the amount of encoding under your control. The five-character mode is what a template engine does and what you almost always want. The named mode is for hand-written HTML where a reader may open the source. The numeric modes exist for files whose encoding is not certain, where every non-ASCII character becomes a pure ASCII reference and the file survives any transport.
The decoding is the part that took care, and not because entities are hard. The obvious way to decode them in a browser is to let the browser do it through innerHTML, and that is a genuine security hole on untrusted text, so this page builds its own table instead and scans the string by hand. It costs more code, it is limited to the names in that table, and it cannot execute anything, which is the trade worth making.
Frequently Asked Questions
The five that matter, apostrophes, numeric forms and missing semicolons.