TL;DR:

Headless Chrome can be a great entry-level tool for turning dynamic JS sites into static HTML pages. Running it on a web server allows you to prerender any page with any modern JS features, making page content load faster and be indexable by crawler tools.

The technical approach in this article shows everyone how to use Puppeteer 's API to add server-side rendering (SSR) functionality to an Express web server. The best part is that the Express app itself requires almost no code changes. The heavy lifting is left to Headless Chrome. You only need to add a few short lines of code and you can SSR any page and get the final complete HTML of that page.
You can refer to the code below:
import puppeteer from 'puppeteer';

async function ssr(url) {
  const browser = await puppeteer.launch({headless: true});
  const page = await browser.newPage();
  await page.goto(url, {waitUntil: 'networkidle0'});
  const html = await page.content(); // serialized HTML of page DOM.
  await browser.close();
  return html;
}
In this article I will use ES Modules (import), which requires a Node version higher than 8.5.0 and the runtime flag enabled for —experimental-modules flag. If it bothers you, you are also free to choose to use require()。click hereto check Node's support for ES Modules.

Introduction

If SEO is done well, there are probably two reasons why you visited this article. One is that you wrote a web application, but it cannot be indexed by search engines! Your app might be a single-page application (SPA), or aprogressive web app (PWA), or written using vanilla JS, or written using some complex library or framework. To be honest, what exactly your tech stack uses is not the issue. The issue is that you spent a lot of time writing this awesome web page, but your users cannot discover it. Another reason you visited this article might also be that some articles on the internet mentioned that server-side rendering is very helpful for performance improvement. You came here to find a quick solution to reduce The time cost of JavaScript startupand improving the page'sFirst meaningful paint。
There are now some frameworks, such as Preact, that themselves provide tools that can be used to solve SSR. If the framework you are using itself provides a prerendering solution, then please use it. Mixing in other tools (headless chrome/Puppeteer) is meaningless.

Crawling modern web pages

Search engine crawlers, social sharing platforms,and even browsershave historically relied only on static HTML tags to index pages and extract content. However, modern web pages have evolved into something completely different. Programs written based on JavaScript are very popular now, which means that in many cases, our content appears empty to crawling tools.
However, some crawler tools have become smarter, such as Google Search. Google's crawler uses Chrome 41 torun the page's JavaScript and then render the final page, but this feature is still imperfect since it was only recently introduced. For example, if a page uses newer syntax features, such as ES6 classes,Modules and arrow functions, then it will cause JS to run with errors in this relatively old browser, ultimately resulting in incorrect page rendering. As for other search engines, who knows what they will do!? ¯\(ツ)/¯

Using Headless Chrome to prerender pages

All crawlers can recognize HTML. The indexing problem we need to solve is finding a tool that can run JS and then output HTML. What if I told you such a tool already exists?
  1. This tool knows how to run all modern JavaScript and then output static HTML
  2. This tool stays up to date as new features are added to the Web
  3. You can quickly use this tool in an existing app, with almost no code changes
Sounds great, right?This tool is the browser!
Headless Chrome doesn't care what library, framework, or toolchain you use. It eats JavaScript for "breakfast" and outputs static HTML before "lunch". Of course, hopefully it can be faster than that :)- Eric
If you use Node, then the simple and fast way to use headless chrome is to use Puppeteer. Its API makes it possible to take a front-end app and prerender its tags. Below is an example of prerendering with it.

1. Example JS app

We start with a dynamic page that generates HTML with JavaScript:
public/index.html
<html>
<body>
  <div id="container">
    <!-- Populated by the JS below. -->
  </div>
</body>
<script>
function renderPosts(posts, container) {
  const html = posts.reduce((html, post) => {
    return `${html}
      <li class="post">
        <h2>${post.title}</h2>
        <div class="summary">${post.summary}</div>
        <p>${post.content}</p>
      </li>`;
  }, '');

  // CAREFUL: assumes html is sanitized.
  container.innerHTML = `<ul id="posts">${html}</ul>`;
}

(async() => {
  const container = document.querySelector('#container');
  const posts = await fetch('/posts').then(resp => resp.json());
  renderPosts(posts, container);
})();
</script>
</html>

2. SSR function

Next, let's optimize the previous ssr() function:
ssr.mjs
import puppeteer from 'puppeteer';

// In-memory cache of rendered pages. Note: this will be cleared whenever the
// server process stops. If you need true persistence, use something like
// Google Cloud Storage (https://firebase.google.com/docs/storage/web/start).
const RENDER_CACHE = new Map();

async function ssr(url) {
  if (RENDER_CACHE.has(url)) {
    return {html: RENDER_CACHE.get(url), ttRenderMs: 0};
  }

  const start = Date.now();

  const browser = await puppeteer.launch();
  const page = await browser.newPage();
  try {
    // networkidle0 waits for the network to be idle (no requests for 500ms).
    // The page's JS has likely produced markup by this point, but wait longer
    // if your site lazy loads, etc.
    await page.goto(url, {waitUntil: 'networkidle0'});
    await page.waitForSelector('#posts'); // ensure #posts exists in the DOM.
  } catch (err) {
    console.err(err);
    throw new Error('page.goto/waitForSelector timed out.');
  }

  const html = await page.content(); // serialized HTML of page DOM.
  await browser.close();

  const ttRenderMs = Date.now() - start;
  console.info(`Headless rendered page in: ${ttRenderMs}ms`);

  RENDER_CACHE.set(url, html); // cache rendered page.

  return {html, ttRenderMs};
}

export {ssr as default};
The main changes are:
  1. Add caching. Caching the rendered HTML can greatly improve response times. When a page is requested repeatedly, you don't have to restart headless chrome to render the HTML again. I will discuss other optimization options later.
  2. Add simple page load timeout error handling
  3. Add a line page.waitForSelector('#post') . This ensures that the article exists in the DOM before we store the serialized page.
  4. Add response time. Record how long headless chrome takes to render the page and return the render time along with the HTML.
  5. Consolidate the code into a module called ssr.mjs module

3. Example Web Service

Finally, this is an Express server that integrates all the code together. The main route handler pre-renders http://localhost/index.html URLs (e.g. the main entry page) and returns the rendered result as the server's response. When users visit the page, they will immediately see the article, because at this point the page response already contains static HTML.
server.mjs
import express from 'express';
import ssr from './ssr.mjs';

const app = express();

app.get('/', async (req, res, next) => {
  const {html, ttRenderMs} = await ssr(`${req.protocol}://${req.get('host')}/index.html`);
  // Add Server-Timing! See https://w3c.github.io/server-timing/.
  res.set('Server-Timing', `Prerender;dur=${ttRenderMs};desc="Headless render time (ms)"`);
  return res.status(200).send(html); // Serve prerendered page as response.
});

app.listen(8080, () => console.log('Server started. Press Ctrl+C to quit'));
To run this example, you need to install the dependencies (npm i —save puppeteer express) and use a Node version higher than 8.5.0, with the —experimental-modules parameter enabled.
The response returned by this server is as follows:
<html>
<body>
  <div id="container">
    <ul id="posts">
      <li class="post">
        <h2>Title 1</h2>
        <div class="summary">Summary 1</div>
        <p>post content 1</p>
      </li>
      <li class="post">
        <h2>Title 2</h2>
        <div class="summary">Summary 2</div>
        <p>post content 2</p>
      </li>
      ...
    </ul>
  </div>
</body>
<script>
...
</script>
</html>

A Perfect Use Case for the New Server-Timing API

Server-Timing API allows you to send web server performance metrics (such as request/response time, database query time, etc.) back to the browser. Client-side code can use this information to track the overall performance of the web app.
A perfect use case for Server-Timing is reporting the time consumed by headless Chrome to pre-render a page! To implement this feature, you just need to add the Server-Timing response header to the server response:
res.set('Server-Timing', `Prerender;dur=1000;desc="Headless render time (ms)"`);
On the client side, you can use the Performance Timeline API and/or PerformaceObserver to retrieve this data:
const entry = performance.getEntriesByType('navigation').find(
    e => e.name === location.href);
console.log(entry.serverTiming[0].toJSON());
{
  "name": "Prerender",
  "duration": 3808,
  "description": "Headless render time (ms)"
}

Performance Results

So how does the performance data look? Ione of the apps(the code is here). On the server side, headless Chrome took about 1 second to render the page. Once the pre-rendered page is cached, when 3G slow simulation is enabled through DevTools,FCP The result was 8.37 seconds faster than the client-side rendering version.
NameFirst Paint(FP)First Contentful Paint(FCP)
Client-side rendering app4s11s
Server-side rendering app2.3s~2.3s
This result looks great. Users can see meaningful page content more quickly, because the server-side rendered page no longer depends on JavaScript to load and render the article.

Avoid multiple executions

Remember when I said earlier, "We don't need to modify any client-side code"? Well, I lied to you.
Our Express server accepts a request, uses Puppeteer to load the page into headless Chrome, and returns the result to the client as the response. But a problem arose in this process.
That piece ofJS code running in server-side headless Chromewill, when the user's browser finishes loading the page,run once more on the front end. We generate HTML tags in two places. #Duplicate rendering!
Let's fix it. We need to tell the page that its HTML has already been generated. The solution I came up with is to have the page's JS determine, after the page finishes loading, <ul id="posts"> whether the tag already exists in the DOM. If it does, then we know the current page was rendered on the server side, and we can avoid adding the article again. 👍
public/index.html
<html>
<body>
  <div id="container">
    <!-- Populated by JS (below) or by prerendering (server). Either way,
         #container gets populated with the posts markup:
      <ul id="posts">...</ul>
    -->
  </div>
</body>
<script>
...
(async() => {
  const container = document.querySelector('#container');

  // Posts markup is already in DOM if we're seeing a SSR'd.
  // Don't re-hydrate the posts here on the client.
  const PRE_RENDERED = container.querySelector('#posts');
  if (!PRE_RENDERED) {
    const posts = await fetch('/posts').then(resp => resp.json());
    renderPosts(posts, container);
  }
})();
</script>
</html>

Optimization

In addition to caching the rendered result, we can also make ssr() many interesting optimizations to the function. Some can yield optimization results quickly, while others may be more speculative. The performance benefits you see may ultimately depend on the type of web page you are pre-rendering and its complexity.

Cancel unnecessary requests

Currently, the entire page (including all the resources it requests) is loaded unconditionally in headless Chrome. However, we are only interested in two things:
  1. The final rendered HTML tags
  2. JS requests that produce HTML tags
Network requests that do not participate in building the DOM tree are very wasteful.Resources such as images, fonts, CSS, and media files do not participate in building the page's HTML. They add styling and supplements to the web page's structure, but do not explicitly create it. We should tell the browser to ignore these requests! Doing so can reduce the workload of headless Chrome,reduce network bandwidth requests, and mayspeed up the pre-rendering time of some larger pages。
Devtools Protocolsupports a Network interception powerful feature that canmodify these requests before the browser actually sends them. Puppeteer can enable network interception by setting page.setRequestInterception(true) to enable network interception,and listen for the page's request event. This allows us to cancel requests for specific resources and let other resources be requested normally.
ssr.mjs
async function ssr(url) {
  ...
  const page = await browser.newPage();

  // 1. Intercept network requests.
  await page.setRequestInterception(true);

  page.on('request', req => {
    // 2. Ignore requests for resources that don't produce DOM
    // (images, stylesheets, media).
    const whitelist = ['document', 'script', 'xhr', 'fetch'];
    if (!whitelist.includes(req.resourceType())) {
      return req.abort();
    }

    // 3. Pass through all other requests.
    req.continue();
  });

  await page.goto(url, {waitUntil: 'networkidle0'});
  const html = await page.content(); // serialized HTML of page DOM.
  await browser.close();

  return {html};
}

Inline important resources

When building pages, we often use different build tools (such as gulp) to process the app and inline some important CSS/JS into the page. This can speed up the first meaningful paint, because the browser can make fewer requests when initially loading the page.
Instead of using different build tools,we can also use the browser as your build tool! We can use Puppeteer to manipulate the page's DOM, inline styles, JavaScript, and anything else you want to add to the page before prerendering the page.
This example shows how to intercept responses for local stylesheet files and inline these resources as <style> tags into the page.
ssr.mjs
import urlModule from 'url';
const URL = urlModule.URL;

async function ssr(url) {
  ...
  const stylesheetContent = {};

  // 1. Stash the responses of local stylesheets.
  page.on('response', async resp => {
    const responseUrl = resp.url();
    const sameOrigin = new URL(responseUrl).origin === new URL(url).origin;
    const isStylesheet = resp.request().resourceType() === 'stylesheet';
    if (sameOrigin && isStylesheet) {
      stylesheetContent[responseUrl] = await resp.text();
    }
  });

  // 2. Load page as normal, waiting for network requests to be idle.
  await page.goto(url, {waitUntil: 'networkidle0'});

  // 3. Inline the CSS.
  // Replace stylesheets in the page with their equivalent <style>.
  await page.$$eval('link[rel="stylesheet"]', (links, content) => {
    links.forEach(link => {
      const cssText = content[link.href];
      if (cssText) {
        const style = document.createElement('style');
        style.textContent = cssText;
        link.replaceWith(style);
      }
    });
  }, stylesheetContents);

  // 4. Get updated serialized HTML of page.
  const html = await page.content();
  await browser.close();

  return {html};
}
This code does the following:
  1. Uses a page.on('response') listener to listen for network responses
  2. Caches the responses for local stylesheets
  3. Finds all <link rel="stylesheet"> tags on the DOM and replaces them with the equivalent <style> tags. You can refer to page.$$eval 's API documentation.style.textContent is set to the response for the stylesheet.

Automatically compressing resources

Through the network interception feature, you can also implement another little trick, which is to modify the response to a request.
For example, suppose you want to compress CSS in your app, but also want to keep an uncompressed version for debugging during development. Suppose you have already configured another tool to compress in advance style.css files, then you can use Request.respond() to use style.min.css to rewrite the contents of style.css the response to the file.
ssr.mjs
import fs from 'fs';

async function ssr(url) {
  ...

  // 1. Intercept network requests.
  await page.setRequestInterception(true);

  page.on('request', req => {
    // 2. If request is for styles.css, respond with the minified version.
    if (req.url().endsWith('styles.css')) {
      return req.respond({
        status: 200,
        contentType: 'text/css',
        body: fs.readFileSync('./public/styles.min.css', 'utf-8')
      });
    }
    ...

    req.continue();
  });
  ...

  const html = await page.content();
  await browser.close();

  return {html};
}

Reusing the same Chrome instance across different renders

Starting a new browser process on every prerender incurs significant overhead. Instead, you can start just one instance and reuse it while rendering multiple pages.
Puppeteer can, by calling puppeteer.connect() and passing the instance's remote debugging URL to it, reconnect to an existing Chrome instance. To be able to keep a long-running browser instance, we can move the code that launches Chrome from ssr() function into the Express server.
server.mjs
import express from 'express';
import puppeteer from 'puppeteer';
import ssr from './ssr.mjs';

let browserWSEndpoint = null;
const app = express();

app.get('/', async (req, res, next) => {
  if (!browserWSEndpoint) {
    const browser = await puppeteer.launch();
    browserWSEndpoint = await browser.wsEndpoint();
  }

  const url = `${req.protocol}://${req.get('host')}/index.html`;
  const {html} = await ssr(url, browserWSEndpoint);

  return res.status(200).send(html);
});
ssr.mjs
import puppeteer from 'puppeteer';

/**
 * @param {string} url URL to prerender.
 * @param {string} browserWSEndpoint Optional remote debugging URL. If
 *     provided, Puppeteer's reconnects to the browser instance. Otherwise,
 *     a new browser instance is launched.
 */
async function ssr(url, browserWSEndpoint) {
  ...
  console.info('Connecting to existing Chrome instance.');
  const browser = await puppeteer.connect({browserWSEndpoint});

  const page = await browser.newPage();
  ...
  await page.close(); // Close the page we opened here (not the browser).

  return {html};
}

Example: Creating a scheduled task to periodically prerender

In my App Engine dashboard app, I set up a scheduled task handler to periodically re-render the site's most important pages. This lets visitors always see fast, up-to-date pages, and avoids making them experience the startup delay caused by creating new prerenders. Creating multiple Chrome instances is very wasteful for this situation. Instead, I used a shared browser instance to render multiple pages at the same time:
import puppeteer from 'puppeteer';
import * as prerender from './ssr.mjs';
import urlModule from 'url';
const URL = urlModule.URL;

app.get('/cron/update_cache', async (req, res) => {
  if (!req.get('X-Appengine-Cron')) {
    return res.status(403).send('Sorry, cron handler can only be run as admin.');
  }

  const browser = await puppeteer.launch();
  const homepage = new URL(`${req.protocol}://${req.get('host')}`);

  // Re-render main page and a few pages back.
  prerender.clearCache();
  await prerender.ssr(homepage.href, await browser.wsEndpoint());
  await prerender.ssr(`${homepage}?year=2018`);
  await prerender.ssr(`${homepage}?year=2017`);
  await prerender.ssr(`${homepage}?year=2016`);
  await browser.close();

  res.status(200).send('Render cache updated!');
});
I also ssr.js added a method to the exported variables of the file: clearCache() method:
...
function clearCache() {
  RENDER_CACHE.clear();
}

export {ssr, clearCache};

Other advisable approaches

Add a flag to the page: "You are currently being rendered in headless mode"

When your page is rendered via headless Chrome on the server, it may be helpful for the client-side code to let the client logic know this. In my app, I use this hook to "turn off" some page logic so that it does not participate in rendering the article's HTML tags. For example, I disabled the code that lazy-loads firebase-auth.js file. Because there is simply no user who can log in (in server-side rendering).
Adding a parameter to the URL that needs to be rendered is the simplest way to add a hook to the page: ?headless parameter is the simplest way to add a hook to the page:
ssr.mjs
import urlModule from 'url';
const URL = urlModule.URL;

async function ssr(url) {
  ...
  // Add ?headless to the URL so the page has a signal
  // it's being loaded by headless Chrome.
  const renderUrl = new URL(url);
  renderUrl.searchParams.set('headless', '');
  await page.goto(renderUrl, {waitUntil: 'networkidle0'});
  ...

  return {html};
}
Then in the page, we can check whether that parameter exists:
public/index.html
<html>
<body>
  <div id="container">
    <!-- Populated by the JS below. -->
  </div>
</body>
<script>
...

(async() => {
  const params = new URL(location.href).searchParams;

  const RENDERING_IN_HEADLESS = params.has('headless');
  if (RENDERING_IN_HEADLESS) {
    // Being rendered by headless Chrome on the server.
    // e.g. shut off features, don't lazy load non-essential resources, etc.
  }

  const container = document.querySelector('#container');
  const posts = await fetch('/posts').then(resp => resp.json());
  renderPosts(posts, container);
})();
</script>
</html>
A handy tip: another useful approach is to use Page.evaluateOnNewDocument() . It lets you inject code into the page and have Puppeteer run that code before any other JavaScript code on the page runs.

Avoid double-counting page visits

If you use Analytics (or other similar products) for your site, be careful. Pre-rendering pages may cause the number of pageviews to increase. More precisely, you will see twice the pageviews: one from headless Chrome pre-rendering the page, and another from when the user's browser renders the page.
So how do you fix it? Use network interception to cancel any requests related to loading the analytics library.
page.on('request', req => {
  // Don't load Google Analytics lib requests so pageviews aren't 2x.
  const blacklist = ['www.google-analytics.com', '/gtag/js', 'ga.js', 'analytics.js'];
  if (blacklist.find(regex => req.url().match(regex))) {
    return req.abort();
  }
  ...
  req.continue();
});
Page visits will never be recorded, as long as the relevant code is not loaded.
Or, keep loading the Analytics library to gain insight into the number of pre-renders performed by the server.

Conclusion

Puppeteer makes server-side rendering of pages very simple by using headless Chrome as a companion service on your web server. My favorite "feature" of this approach is that, without major code changes, you can improve loading performance and the indexability of your pages!
A friendly tip: if you're curious to learn about which runnable apps use the technical approaches mentioned above, you can take a look atthis app andits code。

Appendix

Discussion of existing technologies

Server-side rendering of client apps is hard. How hard? You only need to look at how many npm packages people have written for this fieldhow many npm packagesto know. There are countlesspatterns,toolsas well asservicesto help server-render JS applications.
Isomorphic / Universal JavaScript
The concept of Universal JavaScript is simple: code that runs on the server also runs on the client (the browser). On both the client and the server you share the same set of code, and everyone seems to feel a moment of Zen in that instant.
In practice, I've found that Universal JavaScript struggles to stand out. A story of my own:
I recently created anew project, and wanted to try using lit-html . Lit is a great library that lets you use JS template strings to write HTML <template>, and then render those templates to the DOM very efficiently. The problem is that its core functionality (using the <template> element) doesn't work outside the browser. This means it can't run on a Node server. My hopes of sharing code between Node and the frontend for SSR were cast to the horizon.
Eventually I realized I could server-render this app by using headless Chrome. Whether Chrome is running in the user's hands or automatically on the server makes no difference. Chrome is perfectly happy to run any JS you give it. No questions asked.
Headless Chrome makes isomorphic JS between the client and server possible. If the library you're using doesn't support running on the server (Node), then it's a great choice.
Prerendering tools
The Node community has created a great many tools to solve server-side rendering of JS apps. This doesn't surprise us! Personally, I find these tools vary from person to person, so be sure to do your homework before using a particular tool. For example, some SSR tools are too old and don't use headless Chrome (or any other headless browser). Instead, they use PhantomJS (that is, the familiar old Safari), which also means that if your page uses new syntax features, it won't be rendered correctly.
Prerender is a notable special case.Prerender What's interesting about it is that it uses headless Chrome and it comes with Express middleware:
const prerender = require('prerender');
const server = prerender();
server.use(prerender.removeScriptTags());
server.use(prerender.blockResources());
server.start();
It's also worth noting that Prerender glosses over the details of downloading and installing Chrome on different platforms. Many times, implementing this step correctlyis very troublesome, which is why Puppeteer does this step for you. Some of my apps have also run into problems when usingonline serviceswhen using them:
the browser-rendered page
the same site using prerender.io rendering