GuidesHTML IN THE BROWSER

Fetch HTML from another website in JavaScript

Browser JavaScript can read the HTML of another website only when that site sends Access-Control-Allow-Origin, which regular web pages never do, so a direct fetch() fails with a CORS error and mode: no-cors returns an empty opaque response. Put https://proxy.cors.dev/ in front of the page URL and the HTML comes back readable: free, no signup, no API key, HTTPS pages up to 1 MiB within 10 seconds. Parse the text with DOMParser and query it like any document; the parsed page runs no scripts, so you get exactly what the server rendered.

Why a direct fetch fails

The same-origin policy lets your page send a request anywhere, but the browser only hands the response to your code when it carries an Access-Control-Allow-Origin header that matches your origin. Web pages are made to be viewed, not read by scripts on other sites, so they send none. mode: 'no-cors' is not an escape hatch: it gives you an opaque response with status 0 and an empty body, explained in no-cors and opaque responses.

What a direct fetch of another site gives you
// Default mode: the browser blocks reading the response
await fetch('https://github.blog/');
// TypeError: Failed to fetch
// Console: No 'Access-Control-Allow-Origin' header is present on the requested resource.

// mode: 'no-cors' does not help: the response is opaque
const opaque = await fetch('https://github.blog/', { mode: 'no-cors' });
opaque.status; // 0
await opaque.text(); // '' (always empty)

Fetch through the proxy and parse with DOMParser

The proxy makes the request from its own servers, where no browser policy applies, and relays the status, Content-Type and body with Access-Control-Allow-Origin set to your origin. DOMParser with text/html turns the text into a document you can query with querySelector; it executes no scripts and loads no subresources. Relative links and image URLs in the parsed document still point at paths on the fetched site, so resolve them against the page URL (or its base element) before you use them.

fetchPage(): fetch the HTML through the proxy and parse it
const PROXY = 'https://proxy.cors.dev/';

async function fetchPage(pageUrl) {
  const response = await fetch(PROXY + pageUrl, { credentials: 'omit' });
  if (!response.ok) {
    const code = response.headers.get('X-Cors-Error'); // null when the site itself answered
    throw new Error(`HTTP ${response.status}${code ? ` (${code})` : ''} for ${pageUrl}`);
  }

  const doc = new DOMParser().parseFromString(await response.text(), 'text/html');

  // Make relative URLs usable outside the page they came from
  const base = doc.querySelector('base[href]')?.getAttribute('href') ?? pageUrl;
  for (const element of doc.querySelectorAll('a[href], img[src], link[href]')) {
    const attribute = element.hasAttribute('href') ? 'href' : 'src';
    try {
      element.setAttribute(attribute, new URL(element.getAttribute(attribute), base).href);
    } catch {
      element.removeAttribute(attribute);
    }
  }
  return doc;
}

const doc = await fetchPage('https://github.blog/');
const headings = [...doc.querySelectorAll('h2, h3')].map((heading) => heading.textContent.trim());
const links = [...doc.querySelectorAll('a[href^="https://github.blog/"]')].map((link) => link.href);
console.log(doc.title, headings.slice(0, 5), new Set(links).size);

Checked while writing this page, https://github.blog/ returned 285 KB of HTML through the proxy, https://unity.com/ 527 KB and https://www.theverge.com/ 971 KB, just under the free limit. Compressed transfer is handled by the proxy; the 1 MiB cap applies to the decoded body.

Extract text, links or data

Once parsed, extraction is ordinary DOM code: textContent for readable text, attribute reads for links and images, querySelectorAll with the selectors you would use in DevTools. JSON-LD blocks in script[type="application/ld+json"] are often the cleanest source for prices, events and articles, and JSON.parse on their text costs nothing.

Readable text from the page, no scripts or styles
function pageText(doc) {
  const copy = doc.body.cloneNode(true);
  copy.querySelectorAll('script, style, noscript, template, svg').forEach((node) => node.remove());
  return copy.textContent.replace(/\s+/g, ' ').trim();
}

const text = pageText(doc);
console.log(text.length, text.slice(0, 200));

What you get is the server-rendered page

  • No JavaScript runs: a single-page app returns its empty shell. Look for the JSON API the app itself calls, which is usually easier to consume anyway.
  • No cookies travel: you see the logged-out, default-region version of the page.
  • Status codes pass through: a 404 or a 403 from the site arrives as that status with X-Cors-Source: upstream, while proxy-side problems set X-Cors-Error (response_too_large, upstream_timeout, owner_opted_out).
  • HTTPS only: the proxy rejects http:// targets with invalid_request, and follows redirects for up to 4 hops as long as every hop is a public HTTPS host.
  • Text only on the free tier: HTML, JSON, XML, CSV, plain text and SVG. Images and other binary types need Pro, which relays any content type up to 6 MiB.

Limits and pricing

cors.dev limits for fetching pages
LimitFree, no keyPro, $5 per month
Price$0$5 per month; 7-day trial with 1,000 requests, no card
MethodsGET and HEADGET, HEAD, POST, PUT, PATCH and DELETE
DestinationsAny public HTTPS hostPublic HTTPS hosts enabled on your Connection
Response size1 MiB, text types only6 MiB, any content type
Request bodyNone1 MiB
Upstream deadline10 seconds10 seconds
RateShared pool with fair-use limits600 requests per minute, 10 concurrent per account
Monthly requestsNo quota500,000 per billing period, $5 per extra 500,000
RedirectsFollowed, up to 4 hopsFollowed when every hop host is enabled
CachingNone, no-storeOpt-in X-Cors-Cache, 1 to 300 seconds

When to move this to a server

Fetching one page on demand, when a user pastes a URL or opens a dashboard, is a browser job. Crawling many pages, rendering client-side apps with a headless browser, keeping a history, or scraping on a schedule while nobody has the tab open is a server job, and the fair-use pool will slow a browser crawl down long before a server would. Pro raises the ceiling to 600 requests per minute and 6 MiB pages for $5 per month, but it does not change where the code runs.

Good questions.

Can I fetch http:// pages?

No. The proxy accepts public HTTPS targets on port 443 only and answers 400 with X-Cors-Error: invalid_request for anything else. Most sites answer on https:// as well; try that form.

Why is the HTML different from what I see in my browser?

Your browser runs the page JavaScript, sends your cookies and sits in your region. The proxy requests the page once, with no cookies and no script execution, so you get the initial server response. Consent banners, personalised blocks and client-rendered content differ for that reason.

Does DOMParser run scripts or load images from the parsed page?

No. A document created by DOMParser is inert: scripts do not execute and images, styles and frames are not fetched. Reading from it is safe; inserting its nodes into your live page with innerHTML is not, so copy text and attributes instead.

How do I get readable text instead of markup?

Clone the body, remove script, style, noscript, template and svg elements, then read textContent and collapse whitespace. For article extraction beyond that, run a readability library on the parsed document in the browser.

Can I submit a form or POST to the site?

Anonymous requests are GET and HEAD only. A Pro Connection forwards POST, PUT, PATCH and DELETE with bodies up to 1 MiB to hosts you enabled, which covers APIs; posting HTML forms to sites you do not control is rarely what you want.

Is there a rate limit?

Anonymous requests share a fair-use pool and wait up to 2 seconds for a slot before a 429 with Retry-After. Pro gives your account 600 requests per minute and 10 concurrent requests, with 500,000 requests per billing period.