Loading...

Security · 11 min read

What is XSS? Cross-site scripting, and how to shut it out

Cross-site scripting (XSS) is a bug that lets an attacker get their own JavaScript running in your page, in your visitors' browsers, with your site's privileges. It happens when untrusted text reaches the page as markup or code instead of as plain text. The fix is in your code: encode output for its context, sanitize the HTML you must accept, and add HttpOnly cookies and a Content-Security-Policy as the second line.

Updated

What is XSS? Cross-site scripting, and how to shut it out

What cross-site scripting is

Cross-site scripting (XSS) is a vulnerability where an attacker gets a page on your site to include JavaScript they wrote. The browser cannot tell that script apart from your own: it arrives from your origin, so it runs with everything your origin is allowed to do, including your users' sessions.

The root cause is always the same. Text that someone else controls, such as a comment, a search term, a display name or a URL fragment, is inserted into a page in a way that lets the browser read it as markup or code instead of as plain text. A comment that should appear as the literal characters <script> instead becomes a script element.

The name is historical and slightly misleading: the classic attack involved a second site, but today most XSS is simply untrusted data executed inside your own page. OWASP files it under the Injection category of the OWASP Top 10, next to its server-side cousin, SQL injection. The difference is the interpreter. SQL injection fools your database; XSS fools your users' browsers.

XSS is a bug in the application, not in the browser or the network. HTTPS does not prevent it, a firewall does not fully prevent it, and only the code that builds the page can remove it.

The three kinds: stored, reflected and DOM-based

Stored XSS (also called persistent) is the most damaging. The malicious input is saved by the server, in a comment, a profile field, a product review or a support ticket, and then served to every visitor who opens that page, often including administrators reading the back office. No one has to click a special link.

Reflected XSS is not stored. The server takes something from the request, typically a query parameter, and writes it straight back into the response: a search page that prints "Results for ...", or an error page that echoes the bad value. The attacker has to get the victim to open a crafted link, through email, chat or another site.

DOM-based XSS happens entirely in the browser. Your own JavaScript reads a value the attacker controls, such as location.hash, location.search, postMessage data or document.referrer, and writes it into a dangerous sink: innerHTML, document.write, eval, setTimeout with a string, or a javascript: URL in an href. The server may never see the payload at all, since everything after # is not sent with the request.

The three textbook examples below show the shape of each bug. In every case the fix is the same idea: treat the value as text.

Minimal vulnerable patterns (do not ship these)

<!-- Stored: a saved comment printed without escaping (PHP) -->
<p><?= $comment->body ?></p>
<!-- body saved as: <script>alert(document.domain)</script> -->

<!-- Reflected: the search term echoed from the URL -->
<h1>Results for <?= $_GET['q'] ?></h1>
<!-- link: /search?q=<script>alert(document.domain)</script> -->

// DOM-based: client code writes the URL fragment as HTML
document.getElementById('greeting').innerHTML =
  decodeURIComponent(location.hash.slice(1));
// link: /welcome#<img src=x onerror=alert(document.domain)>

What an attacker can do with it

alert(1) is how testers prove the bug; it is not what attackers run. Once their script executes in your page, it acts as the logged-in user, inside your origin.

Session theft. If the session cookie is readable from JavaScript, the script sends it to the attacker, who then logs in as the victim. Tokens kept in localStorage or sessionStorage are always readable by script, which is the main argument against storing long-lived tokens there.

Actions as the user. Even when the cookie is HttpOnly, the script can call your API with fetch; the browser attaches the cookie itself. It can read CSRF tokens from the page, change the email address on the account, create an admin user, post messages or place orders. Same-origin rules, CORS and SameSite cookies do not stop this, because the request comes from your own site.

Reading what the user sees. Personal data, messages, invoices, anything on the page or reachable through your API.

Phishing and keylogging on your domain. The script can redraw the page as a login form or record what the user types into real forms. The address bar shows your domain and a valid certificate, so the user has no reason to doubt it.

Defacement and spreading. Stored XSS can change what every visitor sees, or copy itself into each victim's own profile. The 2005 Samy worm on MySpace spread to more than a million profiles in under a day this way.

How bad a given XSS is depends on who sees the page. A bug in an admin-only screen reached through stored input, such as a support ticket subject, is often worse than one on the public homepage.

The fix: context-aware output encoding

The core defence is to encode untrusted data at the moment you write it into the page, in the way the surrounding context requires. Validating input helps (a postcode should look like a postcode), but it cannot be the main control, because the same value may be safe in one context and dangerous in another.

HTML body: escape &, <, >, " and ' into entities, so <script> is displayed rather than parsed.

HTML attribute: always quote attribute values, and escape the same characters. An unquoted attribute can be broken out of with a single space.

URL in href or src: encoding is not enough, because javascript:alert(1) contains nothing to escape. Parse the URL and allow only https: and http: (and mailto: if you need it); encode query values with encodeURIComponent.

Inside a <script> block: do not concatenate strings into JavaScript. Serialise the data as JSON with a function that also escapes <, or put it in a data- attribute and read it with element.dataset.

Styles: avoid placing user data in CSS at all.

On the client: prefer the safe APIs. textContent, setAttribute on harmless attributes and createElement never parse HTML; innerHTML, outerHTML, insertAdjacentHTML and document.write do.

Writing user data safely in the browser

// Text, never markup
el.textContent = userName;

// Links: allow only http(s)
function safeUrl(value) {
  try {
    const url = new URL(value, location.origin);
    return ['https:', 'http:'].includes(url.protocol) ? url.href : '#';
  } catch {
    return '#';
  }
}
link.href = safeUrl(profile.website);

// Data for scripts: a data attribute, not string concatenation
// <div id="app" data-user="{{ user_json }}">  (escaped by the template)
const user = JSON.parse(document.getElementById('app').dataset.user);

Framework auto-escaping, and its escape hatches

Modern frameworks already encode for the HTML context by default. {{ }} in Blade, Twig, Jinja, Django templates, Vue and Angular, <%= %> in Rails ERB and {value} in React JSX all escape what they print. That is the main reason XSS is rarer than it was in hand-built PHP pages, and the reason to keep using the framework's default way of printing values.

Almost every real XSS in a modern codebase lives in an escape hatch, a feature that deliberately turns escaping off:

- React dangerouslySetInnerHTML, Vue v-html, Svelte {@html}, Angular bypassSecurityTrustHtml. - Laravel Blade {!! $value !!}, Jinja |safe and Markup(), Django mark_safe and {% autoescape off %}, Rails raw and html_safe. - Direct DOM writes from component code: ref.current.innerHTML = ..., jQuery .html(), $(userInput).

The other gaps are contexts the auto-escaper does not understand. A template that escapes HTML still lets javascript: through in href={url}, still allows a user value inside an inline <script> or event handler to break the code, and never protects a string you pass to eval.

A practical rule: make escape hatches rare, searchable and reviewed. A code search for the list above finds most of your exposure, and a linter rule (for example ESLint react/no-danger, or a Semgrep rule for {!!) keeps new ones from slipping in unnoticed.

When you must accept HTML: use a real sanitizer

Some features need user HTML: a rich-text editor, a CMS body, a forum post with formatting, an email preview. Escaping would destroy the formatting, so the answer is sanitizing: parsing the HTML and keeping only an allowlist of safe elements and attributes.

Use a maintained library built for this, never a regular expression or a blocklist of "bad" tags. Browsers parse HTML in surprising ways, and every hand-written filter that removes <script> has been bypassed with an event-handler attribute, an SVG element, odd nesting or an encoding trick. Established choices are DOMPurify in the browser and in Node.js with jsdom, HTML Purifier for PHP, nh3 (the Rust library ammonia) for Python, and the OWASP Java HTML Sanitizer for Java.

Three details make the difference. Sanitize with an allowlist as small as the feature needs. Sanitize at output, or again at output, so that a later change in the library or in your allowlist protects old data. And do not modify the HTML after sanitizing it: re-parsing, string replacement or inserting it into a different context can turn a safe result into an unsafe one.

For Markdown, the rendered HTML needs the same treatment: most Markdown renderers allow raw HTML by default.

Limit the damage: HttpOnly, SameSite and Content-Security-Policy

Assume one XSS will eventually slip through, and make it worth less.

Cookies. Mark session cookies HttpOnly, so document.cookie cannot read them, plus Secure and SameSite=Lax or Strict. That stops the stolen-cookie scenario. It does not stop the script from acting as the user while the page is open, so it is damage control rather than a fix. Keep long-lived tokens out of localStorage for the same reason.

Content-Security-Policy. A CSP tells the browser which scripts may run. The policy that stops XSS is a strict one: a random nonce generated fresh for every response, put both in the header and on each legitimate <script> tag, plus 'strict-dynamic', object-src 'none' and base-uri 'none'. An injected <script> has no valid nonce, and inline event handlers such as onerror= are blocked, so most XSS bugs become a console error and a report. A policy that only lists domains and keeps 'unsafe-inline' gives little protection against XSS.

Roll it out as Content-Security-Policy-Report-Only first, fix what the reports show, then enforce. The HTTP security headers guide covers the report-only rollout, the other headers that belong next to it, and why X-XSS-Protection should no longer be sent.

One caching trap: a nonce must be unpredictable. If a CDN or page cache serves the same HTML to everyone, every visitor gets the same nonce, and an attacker can read it from the page. For cached pages, use script hashes ('sha256-...') instead, or keep the HTML that carries a nonce out of the cache.

Where browsers support it, require-trusted-types-for 'script' goes further: DOM sinks such as innerHTML refuse plain strings, so DOM-based XSS has to pass through code you wrote.

A strict, nonce-based policy (new nonce for every response)

Content-Security-Policy: script-src 'nonce-R4nd0mPerResponse' 'strict-dynamic'; object-src 'none'; base-uri 'none'

<script nonce="R4nd0mPerResponse" src="/js/app.js"></script>

Set-Cookie: session=...; Path=/; Secure; HttpOnly; SameSite=Lax

Testing your own application for XSS

Test only systems you own or are authorised to test. For your own app, a harmless marker finds most bugs without any attack code.

Trace every input to every output. Put a unique marker containing HTML-significant characters, for example xss7"'<b>bold</b>, into each field, query parameter, header your app displays, and file name you accept. Then visit every page where that value appears, including admin screens, emails, exports and notifications. If the word appears bold, or the page's source shows <b> unescaped or a quote that closes an attribute, the output is not encoded.

Check the DOM, not just the source. For DOM-based bugs, put the marker in location.hash and the query string and inspect the live DOM in developer tools: view-source only shows what the server sent.

Search the code. Grep for the escape hatches listed above and for innerHTML, insertAdjacentHTML, document.write, eval and new Function. Each hit should either go away or have a short comment explaining why the input is safe.

Use tools. OWASP ZAP and Burp Suite crawl and test inputs automatically and find the obvious reflected cases; static analysis such as Semgrep or CodeQL follows data flows a crawler cannot see. Neither replaces the manual trace for stored XSS in back-office screens.

Watch CSP reports. Once a report-only policy is in place, violations from unexpected inline scripts are a free early warning.

Find the escape hatches in a codebase
grep -rnE 'dangerouslySetInnerHTML|v-html|\{@html|bypassSecurityTrust|innerHTML|insertAdjacentHTML|document\.write' src/
grep -rnE '\{!!|\|safe|mark_safe|html_safe|raw\(' resources/ templates/ app/

Where a WAF fits, and what CDN.com.tr's WAF does

A web application firewall inspects requests and blocks those that look like attacks. For XSS that is useful: script tags and event handlers in a query string or form body are recognisable, so reflected attempts and automated scanners are stopped before they reach your application, and a stored payload sent through a normal form is often refused on the way in.

It is a layer, not the fix. A WAF sees requests, not your pages, so it cannot know how a value will be used later. DOM-based XSS that lives in the URL fragment never reaches it, input that arrives through an import, an API you call or a channel the WAF does not inspect slips past, and attackers tune payloads to avoid patterns. The OWASP Top 10 guide maps in detail what a firewall can and cannot catch. Encode output, sanitize HTML and set a CSP whether or not a WAF is in front.

On CDN.com.tr the WAF is ModSecurity with the OWASP Core Rule Set, running at the edge and switched on for the account from the delivery rules page. The Core Rule Set's 941 family is its cross-site scripting rules. A blocked visitor gets a branded 403 page with a Reference ID, and the WAF Logs page lists blocked events with the attack category, country, IP and rule, searchable by that ID over the last 30 days and exportable to CSV or XLSX; cdnctl waf logs returns the same list from the command line.

The usual false positive for XSS rules is a legitimate HTML post, such as an editor saving formatted content or a code sample. Exceptions for single paths are not available in the panel yet: if the WAF blocks a legitimate request, send its Reference ID to support rather than switching the WAF off.

Headers can be added at the edge too. Single-value headers go in a delivery rule's Custom Headers, one per line, and apply to 2xx and 3xx responses, cache hits included. A full Content-Security-Policy cannot go there, because the panel and the edge refuse ; in a header line; a single directive such as frame-ancestors 'self' works, and a nonce has to be generated per response anyway, so the policy belongs in your application.

Blocked XSS attempts from the command line
cdnctl waf logs --account $ACCOUNT_UUID --range 1d
cdnctl waf show $REFERENCE_ID --account $ACCOUNT_UUID

XSS FAQ

What is the difference between stored, reflected and DOM-based XSS?

Where the payload lives. Stored XSS is saved on the server and served to everyone who opens the page. Reflected XSS comes back from the same request, so the victim has to open a crafted link. DOM-based XSS is created by your own client-side JavaScript writing an attacker-controlled value, often from the URL, into the page; the server may never see it.

Does HTTPS or a valid certificate protect against XSS?

No. TLS protects data in transit. An XSS payload is delivered by your own server or your own script over the same encrypted connection, and the padlock makes a fake login form on your domain look more trustworthy, not less.

Is HttpOnly enough to stop XSS?

No. It stops the script from reading the session cookie, which removes one outcome. The script can still send requests as the user, because the browser attaches the cookie for it, and it can still read and change the page. HttpOnly is damage control; output encoding is the fix.

Can a WAF stop cross-site scripting?

It stops many reflected and automated attempts, because script payloads in requests are recognisable. It cannot see DOM-based XSS in the URL fragment, does not know how your app will use a stored value, and can be evaded by a tailored payload. Treat it as a layer in front of correct code.

Should I still send X-XSS-Protection?

No. The filter it controlled has been removed from current browsers, and in older ones it could be abused. Send nothing or X-XSS-Protection: 0, and use a Content-Security-Policy instead.

Does React or Vue make my app immune to XSS?

They make the common case safe by escaping what you print. You are still exposed through dangerouslySetInnerHTML and v-html, through user URLs in href that start with javascript:, through direct DOM writes, and through server-rendered HTML that the framework does not control.

What is self-XSS?

A script the victim is tricked into pasting into their own browser console. It needs no bug in your site, which is why browsers warn when you paste into developer tools. It is social engineering rather than a vulnerability in your code.