Blog Froxy | News, Helpful Articles About Using Proxies

Cheerio vs Puppeteer: Which to Use for Web Scraping

Written by Team Froxy | Sep 30, 2026, 3:00:00 AM

Both Cheerio and Puppeteer are designed to work in the Node.js environment. This backend is written in JavaScript and can run on virtually any modern operating system, including servers and virtual containers. However, comparing Cheerio vs Puppeteer is a bit like trying to compare something warm with something soft.

Both libraries can be used for website scraping, but they work in fundamentally different ways. Cheerio processes HTML content only and cannot render dynamic website pages, while Puppeteer allows you to control headless browsers such as Google Chrome or Mozilla Firefox through dedicated protocols, CDP or BiDi.

Because of this, the choice of a specific solution should depend on the initial scraping requirements: what types of websites you need to work with, what level of performance you need, what content extraction capabilities are required, how proxy connections should be organized, and so on. The same considerations apply when building a cheerio web scraping workflow.

Below, we will look at the details and compare Cheerio vs Puppeteer: what these libraries are, when each one is most effective, how to work with them, and what their main differences and specific features are.

What Is Cheerio?

Cheerio is a solution for traversing the DOM structure of a resulting page containing HTML markup. In other words, it takes HTML code as input and allows you to access specific elements and extract their content.

The library works in much the same way as the Beautiful Soup parser for Python. The main difference is that Cheerio syntax is closer to jQuery syntax. Developers already familiar with that framework will find it easier to understand how element searches work.

At the same time, you cannot build a complete cheerio scraper without additional components because Cheerio cannot directly request websites and cannot render pages.

To work around this limitation, the following combinations are most commonly used:

  • The Fetch HTTP client or an axios cheerio setup.
  • The Playwright web driver or a puppeteer cheerio combination.

Technically, Cheerio has the following method:

.fromURL(url, options?): Promise<CheerioAPI>

So, if you really want to, you can retrieve HTML or XML documents directly from URLs. HTTPS connections are also supported through undici. You can change the User-Agent string and other HTTP headers.

However, the built-in HTTP client cannot replace a full browser and comes with many technical limitations. It works only with HTML and XML responses due to a strict Content-Type filter, cannot download content larger than 500 MB, processes no more than five redirects, and so on. JavaScript on the page is not executed.

The key advantages of Cheerio are execution speed, simple syntax, and flexibility. This is why cheerio scraping works particularly well when the required data are already present in the HTML.

Installation command:

npm install cheerio

Importing the library into your scraping script:

import * as cheerio from 'cheerio';

or:

const cheerio = require('cheerio');

Suppose the target HTML looks like this:

<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>My First Page</title>
</head>
<body>
<h2 class="main-title">List of My Favorite Activities</h2>
<ul id="hobby-list">
<li class="my-class-for-list-elements">Reading Books</li>
<li class="my-class-for-list-elements">Programming</li>
<li class="my-class-for-list-elements">Traveling</li>
<li class="my-class-for-list-elements">Sports</li>
<li class="my-class-for-list-elements">Music</li>
</ul>
</body>
</html>

Then extracting the content of the Title tag would look like this:

const $ = cheerio.load('HTML CODE GOES HERE');
// Extract the Title element
const title = $('title').text();
// Extract the H2 element
const h2 = $('h2.main-title').text();
// Second list item
const secondItem = $('li').eq(1).text();

All list elements:

const { listLi } = $.extract({ listLi: ['.my-class-for-list-elements'] });

General summary for clarity:

Cheerio is a tool for quickly and conveniently extracting data from HTML. In practical terms, cheerio web scraping is primarily about parsing content that has already been retrieved.

Residential Proxies

Perfect proxies for accessing valuable data from around the world.

Try With Trial $1.99, 100Mb

What Is Puppeteer?

Puppeteer is a library that provides control over headless browsers: Chrome and Chromium through the Chrome DevTools Protocol (CDP), and Firefox through WebDriver BiDi. WebMCP API support is also available.

Read the full guide on web scraping with Puppeteer.

In other words, Puppeteer is primarily designed for browser automation. In fact, the library was originally created to demonstrate the capabilities of the CDP protocol.

From this alone, you can understand that Puppeteer supports much more than simply starting and stopping browser instances. It can execute virtually all commands related to user actions: keyboard input, cursor movement, page scrolling, and so on. It can also search for elements within the DOM structure using the resulting HTML after rendering or while rendering is still in progress.

For this reason, installing and configuring Puppeteer is noticeably more complicated. At the same time, its feature set is much broader than Cheerio's.

Because Puppeteer is directly connected to a browser, it can work with virtually any type of web content: PDFs, screenshots and images, video, XML, documents, and more. JavaScript rendering and web applications are also supported.

Installation:

npm install puppeteer

If you do not need headless browsers to be installed automatically, you can install the puppeteer-core package instead.

Importing it into your scripts:

import puppeteer from 'puppeteer';

or:

const puppeteer = require('puppeteer');

Here is an example of opening a website and finding specific layout elements:

const puppeteer = require('puppeteer');

(async () => {
  // Launch the browser
  const browser = await puppeteer.launch({
    headless: true
  });

  // Open a new tab
  const page = await browser.newPage();

  // Navigate to the website
  await page.goto('https://quotes.toscrape.com/');

  // 1. Get the <title>
  const title = await page.title();
  console.log('Title:', title);

  // 2. Find an element using "tag + CSS class"
  const firstQuote = await page.$eval(
    'span.text',
    element => element.textContent.trim()
  );
  console.log('First quote:', firstQuote);

  // 3. Get the second element
  const secondQuote = await page.$eval(
    'span.text:nth-of-type(2)',
    element => element.textContent.trim()
  );
  console.log('Second quote:', secondQuote);

  // 4. Get a list of all elements
  const quotes = await page.$$eval(
    'span.text',
    elements => elements.map(element => element.textContent.trim())
  );
  console.log('All quotes:', quotes);

  // 5. Get elements with a specific attribute
  const links = await page.$$eval(
    'a[href]',
    elements => elements.map(element => ({
      text: element.textContent.trim(),
      href: element.getAttribute('href')
    }))
  );
  console.log('Links:', links);

  // Close the browser
  await browser.close();
})();

As you can see, the code required to work with Puppeteer is noticeably more complex. This is because you need to:

  • Configure the browser parameters.
  • Launch a new browser instance and open a new tab.
  • Navigate to a specific website and wait for it to load.
  • Only after that can you work with the page's HTML code.

Puppeteer supports direct interaction with both JavaScript and HTML and also allows you to export the resulting code to other processes and libraries. For example, you can pass the final HTML to Cheerio for processing.

Key Differences: Cheerio vs Puppeteer and Comparison Table

As a reminder, in a Cheerio vs Puppeteer comparison, we are comparing two fundamentally different tools: Cheerio is a markup parser, while Puppeteer is a browser automation tool. All of their major differences follow from this distinction.

Criterion

Cheerio

Puppeteer

Tool type

Specialized HTML/XML parser

Web driver for automating headless browsers

JavaScript execution + dynamic content (SPA, infinite scroll, etc.)

No

Yes

Page interaction (clicks, form completion, login)

No

Yes

Speed

Very high

Slower because the browser needs to be launched

Resources (CPU/RAM)

Low

High

Additional capabilities

Direct HTTP/HTTPS requests without browser functionality

Screenshots, PDFs, device emulation, and user behavior emulation

Ideal use case

Parsing ready-made HTML, working with XML, scraping static websites without anti-bot protection

Working with complex dynamic websites and web applications, scraping content after logging into user accounts, and similar tasks

In general, Cheerio is mainly useful for scraping simple static websites or XML documents when working through an API. Puppeteer is the "heavy artillery" for complex dynamic websites that modify or deliver their content through JavaScript and may also be protected by advanced anti-bot mechanisms such as web application firewalls, CAPTCHAs, CSRF tokens, and so on.

Although Puppeteer has its own built-in equivalent of an HTML parser, its syntax is relatively cumbersome. There is nothing preventing you from using both tools together: Puppeteer retrieves the final HTML, while Cheerio parses it. A puppeteer cheerio stack can therefore be more convenient than relying on Puppeteer's DOM extraction syntax alone.

Performance and Ease of Use

Because Cheerio does not launch a browser, it works several times faster than Puppeteer. Comparing their raw performance is therefore somewhat pointless: Cheerio will always be faster. This is one of the main reasons a cheerio scraper is so efficient for static pages.

However, to make the Cheerio vs Puppeteer comparison as practical as possible, we conducted a real test. To avoid distortions caused by architectural differences or different website types, we used the same static resource for both tools.

For reference and to better illustrate the difference, we sent requests to https://quotes.toscrape.com/, specifically its homepage. Processing took:

  • Requesting the website and parsing the response with Cheerio: 914–1158 ms.
  • Puppeteer: 1963–2351 ms.

Measurements were made using performance.now().

The difference is almost twofold: Cheerio is faster. That is exactly what we expected to demonstrate.

You should keep in mind that script execution speed is affected by many factors: internet connection quality, PC hardware, architecture, availability and response speed of the target website, and so on.

For example, if we had accessed a dynamic website, Cheerio might not have received valid HTML at all, while Puppeteer might have had to wait 20–60 seconds for the page to finish loading.

But even repeated measurements would not fundamentally change the situation. Puppeteer is a "heavy monster" that consumes considerably more computing resources and also requires more time to execute each request. In contrast, cheerio scraping avoids the overhead of launching and controlling a browser.

Why One Library Is Not Enough in Production

Neither Cheerio nor Puppeteer was originally created specifically for scraping. Therefore, neither library includes built-in mechanisms for IP rotation, CAPTCHA bypassing, quick proxy integration, and similar tasks.

It is important to remember that no modern business-grade scraper can operate effectively without proxies. It needs to be able to quickly change connection addresses, browser profiles, digital fingerprints, User-Agent settings, and other parameters. The same applies to large-scale cheerio web scraping.

The libraries' built-in capabilities are limited:

  • When using the fromURL function, Cheerio allows you to configure a custom User-Agent and some other HTTP headers, much like basic HTTP clients for Python.
  • Puppeteer supports the HTTP_PROXY and HTTPS_PROXY environment variables. Its configuration also allows you to override the User-Agent and other digital fingerprint parameters. However, there is still no flexible built-in proxy rotation mechanism. You need to implement the logic manually or use additional rotation solutions.

How to Scale Scraping Without Getting Banned Using Proxies

As soon as your scraper starts sending a large number of requests from your IP address, the risk of being blocked increases. Using proxies and quickly rotating outgoing IP addresses reduces the risk of bans, helps balance or distribute the load, and hides your real address.

We have previously covered the main types of proxies:

For business use cases, rotating residential or mobile proxies are usually the best option.

How to Use a Proxy with Cheerio

Cheerio does not provide a dedicated proxy API. To route requests through alternative IP addresses, you need to use the functionality of an HTTP client available in Node.js, such as undici.

In this case, HTTP request settings can be passed through requestOptions:

const cheerio = require('cheerio');
const { ProxyAgent } = require('undici');

(async () => {
  const proxyAgent = new ProxyAgent('http://127.0.0.1:8080');

  const $ = await cheerio.fromURL('https://example.com', {
    requestOptions: {
      dispatcher: proxyAgent
    }
  });

  console.log($('title').text());

  await proxyAgent.close();
})();

Proxies requiring authentication can be connected through ProxyAgent:

const proxyAgent = new ProxyAgent(
  'http://username:password@127.0.0.1:8080'
);

Instead of undici, you can use other HTTP clients such as Fetch, Axios, and others. An axios cheerio configuration is often convenient when the rest of the project already relies on Axios.

For a cheerio scraper that needs frequent IP changes, the proxy rotation logic should be handled by the HTTP layer rather than by Cheerio itself.

How to Connect Puppeteer Through a Proxy

In Puppeteer, a proxy is configured when launching the Chromium browser instance through the --proxy-server parameter.

If you need to define proxies at the individual tab level, you will need to install additional browser extensions or use third-party libraries.

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({
    headless: true,
    args: [
      '--proxy-server=http://127.0.0.1:8080'
    ]
  });

  const page = await browser.newPage();
  await page.goto('https://example.com');

  console.log(await page.title());

  await browser.close();
})();

For authenticated proxies, Puppeteer allows you to provide the username and password separately:

await page.authenticate({
  username: 'username',
  password: 'password'
});

What to Keep in Mind When Scraping to Avoid Bans

  1. IP rotation is the most important factor. Rotation can be tied to sessions, time intervals, or every new request. The appropriate strategy depends on how sophisticated the target website's protection is.
  2. Follow the rules specified in the Terms of Service and the directives in robots.txt. Read more about the legality of web scraping.
  3. Use the correct User-Agent and HTTP headers.
  4. Maintain high-quality digital fingerprints. These are your browser profiles: installed fonts, extensions, cookies from previously visited websites, screen resolution, and other parameters.
  5. Handle CAPTCHAs and WAF protection.
  6. Hide signs that popular headless browsers or HTTP clients are being used.
  7. Process tokens and traps correctly.
  8. Imitate user behavior realistically: use natural delays, scrolling, mouse pointer movements, and similar actions.

More details: Guide to Successful Web Scraping Without Blocks.

Conclusion

Cheerio offers maximum speed and simple syntax, but the library can work only with static HTML. That is why cheerio scraping is best suited to pages where the required content is already present in the response.

Puppeteer is a browser automation tool, so it can work with dynamic websites and imitate user behavior. In the Cheerio vs Puppeteer choice, Puppeteer is therefore the more capable option when rendering or browser interaction is required.

In practice, the best solution is often not one library but a combination of both. Puppeteer can retrieve the final rendered code when a page requires JavaScript, while Cheerio extracts the required data from the HTML. This puppeteer cheerio approach combines browser rendering with simpler parsing syntax, which is why Cheerio vs Puppeteer does not always have to be an either/or decision.

FAQ

Can Cheerio and Puppeteer Be Used Together?

Yes! Sometimes they should be, especially if you find Puppeteer's built-in syntax for accessing DOM elements through the DevTools Protocol difficult to work with.

Cheerio syntax is much simpler and more intuitive. It will also be familiar to many developers who have previously worked with jQuery. In practice, Cheerio vs Puppeteer is often better treated as a question of how to combine the two tools rather than which one must completely replace the other.

Which Is Faster, Cheerio or Puppeteer?

Cheerio, without question. But speed may not be the most important factor because Cheerio cannot work with dynamic pages. It is limited to static content and cannot handle protection mechanisms on its own.

Can Cheerio Scrape JavaScript Websites?

No. To render such pages, you will need Puppeteer or another library that works with headless browsers.

Are Proxies Mandatory for Cheerio / Puppeteer?

For one-off and small scraping tasks, you can work without proxies. For large-scale scraping, proxies become essential. The more requests you send to the target website, the more likely your IP address is to be blocked.

If you are using an axios cheerio stack for large-scale collection, proxy configuration and rotation need to be implemented in Axios or another networking layer.

Puppeteer or Playwright?

This is largely a matter of habit and personal preference.

Puppeteer was originally written for Node.js, although Playwright also has a Node.js version. Puppeteer once supported only Google Chrome, but its browser support has since expanded thanks to the BiDi protocol.

For a more informed choice, you need to look more closely at the configuration options and available features.

More details in the Playwright vs Puppeteer comparison.

Does Cheerio Replace jQuery?

That is a very good question.

If you retrieve HTML and execute it inside a browser, you can manually inject jQuery files into the page and then use the familiar syntax to find layout elements.

Cheerio allows you to parse HTML without additional workarounds and without launching a browser. So, yes, it can replace jQuery for this type of task.