Antidetect Browser + Crawlee: Resilient Node.js Crawlers
The run that made me put an antidetect browser behind Crawlee looked fine from the console. Crawlee's stats table said requestsFinished: 1204, requestsFailed: 3. The dataset had eleven rows.
Eleven. Out of twelve hundred pages that all returned HTTP 200, rendered a full DOM, fired no errors, and tripped none of Crawlee's retry logic. The crawler thought it had a great night. What it actually had was 1,193 pages of polite empty shell — the version of a site you get served once you've been made, where the markup is intact and the data is gone.
I lost most of a Saturday to that one. Rewrote the selector twice. Added waits. Blamed hydration.
Here's the part that took me embarrassingly long to accept: Crawlee wasn't the problem, and neither were my selectors. Crawlee's fingerprint injection is genuinely good — it's better than raw Playwright by a wide margin, and the header generator alone will get you past a lot of mid-tier defenses. But it's JavaScript, patching a stock Chromium from the outside, after the page context already exists. The checks that matter now run below that line. Canvas rasterization at the driver level. Audio DSP timing. Font metrics that come from the actual text shaper, not from a navigator property somebody overwrote.
So don't pick between them. Keep Crawlee for what it's excellent at — the request queue, the autoscaling, the retry semantics, the dataset — and hand the browser identity to a native antidetect browser engine over CDP.
That's the whole integration. It's about forty lines, and one of them is the one everybody gets wrong.
What We're Building
A Crawlee PlaywrightCrawler that:
- Leases a JustBrowser profile from a pool instead of launching a local Chromium
- Connects to it over CDP through a custom launcher
- Keeps Crawlee's
RequestQueue,Dataset, autoscaling, and retry handling untouched - Turns Crawlee's own fingerprint layer off, deliberately
- Returns the profile to the pool when the browser retires, without orphaning it
The architecture point worth stating plainly: Crawlee orchestrates, the antidetect browser is the identity. Two layers, one job each. Every failure mode below comes from blurring that boundary.
Unpopular take while I'm here: most scraping stacks don't need an antidetect browser at all, and reaching for one first is how you end up with a slow crawler that still gets blocked. Fix your request rate, your headers, your retry behaviour. Then come back to canvas hashes.
Prerequisites
- Node 20+. Crawlee 3.x and Playwright as a peer dep:
npm i crawlee playwright. - JustBrowser — $9.99/mo for unlimited profiles, or $99.99/year. The REST API and CDP launch are in the 7-day trial too, so the launcher below can be built and tested before the first charge.
- A host that stays up. Every active profile is a real Chromium process, not a tab. My 8GB box handles six concurrent profiles and starts swapping at eight. Size accordingly and leave headroom.
- Windows x64 or macOS on Apple Silicon for the machine the profiles run on — there's no Linux build, so a bare VPS isn't the host for this half of the stack. The crawler process can sit elsewhere and reach the API over a tunnel.
- Proxies. JustBrowser routes them per profile over HTTP, HTTPS, or SOCKS5. Bring your own, or buy from the in-app proxy marketplace — they're billed separately from the subscription either way.
Step 1: Turn Crawlee's Fingerprint Layer Off
Start here, because everything else is wasted if you skip it.
const crawler = new PlaywrightCrawler({
browserPoolOptions: {
useFingerprints: false, // <- the line everyone forgets
maxOpenPagesPerBrowser: 1,
},
});
Leave useFingerprints on and Crawlee generates a plausible fingerprint, then injects it into every new context. Meanwhile the engine underneath has already made its own decisions about canvas, WebGL, audio, fonts, screen, timezone — 40+ parameters, patched in C++ before any script runs.
Now a detector asks the same question twice and gets two answers.
That's worse than no spoofing at all. An unusual-but-internally-consistent browser looks like someone running a niche Linux distro. A browser whose JavaScript-visible navigator disagrees with its own rendering output looks like exactly what it is. I shipped this misconfiguration for about a week and watched profiles that passed every isolated fingerprint test still collect blocks, which is a special kind of maddening.
One engine owns identity. Pick which one. (It should be the one that isn't a script.)
Step 2: The Antidetect Browser Launcher
Crawlee's PlaywrightPlugin calls launcher.launch(opts) and expects a Playwright Browser back. It does not care where that browser came from — so we hand it a launcher that connects instead of spawns.
import { PlaywrightCrawler, Dataset } from 'crawlee';
import { chromium } from 'playwright';
const JB = process.env.JB_BASE_URL; // e.g. http://127.0.0.1:36542
const TOKEN = process.env.JB_API_TOKEN;
async function jb(path, init = {}) {
const res = await fetch(`${JB}/api/v1${path}`, {
...init,
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${TOKEN}`,
...(init.headers ?? {}),
},
});
if (!res.ok) throw new Error(`JustBrowser ${path} → ${res.status} ${await res.text()}`);
return res.json();
}
const POOL = ['prof_a1', 'prof_a2', 'prof_a3', 'prof_a4'];
const idle = new Set(POOL);
const jbLauncher = {
name: () => 'chromium',
async launch() {
const profileId = [...idle][0];
if (!profileId) throw new Error('no idle profile — lower maxConcurrency');
idle.delete(profileId);
// set the proxy on the profile (POST /profiles/batch/proxy) rather than at launch
const { data: { cdp_url: wsEndpoint } } = await jb(`/profiles/${profileId}/start`, {
method: 'POST',
body: JSON.stringify({ headless: false }),
});
const browser = await chromium.connectOverCDP(wsEndpoint);
browser.on('disconnected', async () => {
await jb(`/profiles/${profileId}/stop`, { method: 'POST' }).catch(() => {});
idle.add(profileId);
});
return browser;
},
};
Two things I want to flag honestly.
headless: false is on purpose. Run the profiles headed on a real desktop session — headless Chromium still gives up a handful of signals that no flag combination closes, and we went through the specifics in the Playwright and Puppeteer stealth breakdown.
And the launcher shim leans on Crawlee's plugin contract rather than a documented public API for swapping browsers. It's stable in 3.x, it's how a lot of people do remote-browser work, but pin your Crawlee version in package.json and don't let a minor bump ride into production unwatched. I'd rather tell you that than pretend this is blessed.
Step 3: Wire It Into the Crawler
const crawler = new PlaywrightCrawler({
launchContext: { launcher: jbLauncher },
browserPoolOptions: {
useFingerprints: false,
maxOpenPagesPerBrowser: 1,
retireBrowserAfterPageCount: 40,
closeInactiveBrowserAfterSecs: 300,
},
maxConcurrency: POOL.length, // never exceed the pool
maxRequestRetries: 3,
navigationTimeoutSecs: 90,
sessionPoolOptions: {
maxPoolSize: POOL.length,
sessionOptions: { maxUsageCount: 200 },
},
async requestHandler({ page, request, enqueueLinks, log }) {
await page.waitForSelector('.product-card', { timeout: 20_000 });
const rows = await page.$$eval('.product-card', (cards) =>
cards.map((c) => ({
title: c.querySelector('.title')?.textContent?.trim() ?? null,
price: c.querySelector('.price')?.textContent?.trim() ?? null,
})),
);
if (!rows.length) {
log.warning(`empty grid at ${request.url} — likely served a shell`);
throw new Error('empty result set'); // let Crawlee retry on another profile
}
await Dataset.pushData(rows);
await enqueueLinks({ selector: 'a.next-page' });
},
failedRequestHandler({ request, log }) {
log.error(`gave up on ${request.url}`);
},
});
await crawler.run(['https://example.com/catalog']);
maxConcurrency: POOL.length is not a performance tuning knob here, it's a correctness constraint. Crawlee will happily ask for a seventh browser when you have six profiles, and the start endpoint returns an error for a profile that is already running.
The throw on an empty result set is the fix for my eleven-row Saturday. A 200 with no data isn't success — it's a soft block, and unless you turn it into an exception Crawlee has no way to know. Throwing sends the request back through the queue, and because the browser gets retired on repeat failures, the retry lands on a different profile with a different IP. That one change took a crawl from 11 usable rows to roughly 1,150.
maxUsageCount: 200 is deliberately high. Crawlee's default assumes sessions are cheap — an IP and a cookie jar, discard freely. But a profile that's been through cookie warm-up across 50+ sites and has weeks of profile aging behind it is an asset, and rotating it every 50 requests throws that away. If you're not warming profiles before they crawl, start; the cookie warm-up automation walkthrough covers the scripted version.
Step 4: Proxies Go On the Profile
Notice the proxy isn't passed to Crawlee's proxyConfiguration, and it isn't passed at start time either — it's set on the profile itself (POST /profiles/batch/proxy) before the crawl runs.
Small detail. Big blast radius.
Set the proxy in Crawlee and Playwright applies it at the context level — which sounds equivalent right up until you remember that the engine's WebRTC protection and its per-profile DNS-over-HTTPS were both configured around the profile's own network path, a path that now disagrees with where the packets are actually going. You get a browser whose network stack and identity stack tell different stories. Detectors read both.
Match the exit geo to the profile, too. A São Paulo residential IP behind a profile set to Europe/Warsaw is a free flag, and it's the most common self-inflicted wound in this whole category — more on the mechanics in the timezone and geolocation mismatch breakdown. I've done it twice. Second time on a profile I'd spent a week warming.
Common Errors and How to Fix Them
connectOverCDP: connect ECONNREFUSED
The API is up but the CDP endpoint isn't, or the start call returned a cdp_url bound to localhost while your crawler runs elsewhere. Check curl -H "Authorization: Bearer $JB_API_TOKEN" $JB_BASE_URL/api/v1/profiles first. If that returns and the connect still fails, you're crossing a host boundary and need a tunnel — not a firewall rule, and definitely not the API port (36542 by default) on a public IP — it is bound to 127.0.0.1 for a reason.
profile already running
Concurrency exceeded the pool, or a previous crash left a profile up with nobody holding the handle. The disconnected listener in Step 2 covers the clean path. For the unclean one, list profiles at startup and stop anything already running before crawler.run() — three lines, saves an afternoon.
Every profile gets blocked at roughly the same page count
That's not fingerprinting, that's behavior. Identical timing, identical scroll patterns, identical request ordering across four profiles reads as one actor wearing four hats. Add jitter to requestHandler, vary maxRequestsPerCrawl per profile, stagger the start times.
This one annoys me more than it should, because it's the failure that looks most like a fingerprint problem and isn't. You'll spend a day on canvas hashes. The canvas was fine the whole time.
Launch calls piling up
Launch calls are heavier than a normal REST call, and a crawler that retires browsers aggressively will hammer them. Serialize them or back off between starts rather than firing them in a tight loop — same concurrency and backoff patterns that apply to any profile-heavy workload.
Fingerprint tests pass, blocks continue
Check consistency, not scores. Run the profile against CreepJS and Pixelscan and look at whether the values agree with each other rather than whether the page shows green. A perfect fingerprint on a browser with no history is still a browser with no history. Green is not a receipt.
Next Steps
- Lease profiles by target, not round-robin. One profile per site builds site-specific cookie history instead of spreading thin identity across everything. This is the biggest payoff on the list and almost nobody does it, mostly because round-robin is what every pool example on the internet shows you.
- Poll profile health. Hit
GET /profiles/{id}/fingerprint/validateor the health score between crawls instead of relying on a fixed counter — there are no webhooks, so schedule the check yourself. - Re-test after Chromium updates. The engine tracks Chromium; your proxies and profile state drift on their own schedule. Monthly is enough.
- Watch what the crawl costs you. If scraped data feeds a dashboard, JustAnalytics handles the metrics side without reporting back to the platforms you're measuring, and if you're buying traffic next to the scraping infrastructure, ClickzProtect is the traffic-quality half of the same problem.
Honest caveat to end on: this setup trades Crawlee's zero-config convenience for a moving part you own. The pool can leak. The host can die. You're now on the hook for both.
Worth it, in my opinion. But I'd rather you decide that before 3 a.m. on a Saturday than after.
Frequently Asked Questions
Does Crawlee's built-in fingerprint generator conflict with a native antidetect engine?
Yes, and it's the single most common way this integration goes wrong. Crawlee injects generated fingerprint values through JavaScript at context creation. JustBrowser patches the same surfaces in the C++ engine underneath. Run both and a detector reads two different answers for the same question — a canvas hash from the engine, a navigator block from the injection, and no consistency between them. Set browserPoolOptions.useFingerprints to false and let the engine own identity.
Can I use CheerioCrawler instead, or do I need PlaywrightCrawler?
If the target renders server-side and doesn't check TLS or browser signals, CheerioCrawler is faster and cheaper — no browser at all. The moment the site fingerprints the client, plain HTTP requests have no fingerprint to spoof, so there's nothing an antidetect browser can help with in that path. Use PlaywrightCrawler with a CDP connection for the pages that need a real browser identity, and keep Cheerio for the pagination and sitemap crawling that doesn't.
How do I stop Crawlee's session rotation from burning warm profiles?
Raise sessionOptions.maxUsageCount and lower maxPoolSize so sessions live longer and map cleanly onto your profile pool. Crawlee's defaults assume sessions are disposable — a session is just an IP plus a cookie jar to it. A warmed profile with weeks of aging is not disposable. Pin one session to one profile, retire on real block signals rather than usage counters, and never let the pool size exceed the number of profiles you actually have running.
Can I build this on the trial, or do I need to subscribe first?
The trial covers it. Seven days, card taken at checkout, and nothing is held back — the REST API and the CDP launch endpoint are there from the first hour, so you can wire the launcher, run a real crawl, and cancel inside the week if it isn't for you. After that it's one plan: $9.99/mo for unlimited profiles, or $99.99/year. There's no API tier to upgrade into and no profile cap to design around, which removes the usual question of when to switch.
Try JustBrowser
Native Chromium antidetect browser — not extension-based. Real C++ engine patches at the canvas / WebGL / audio / font / screen layer, so 40+ identity parameters are genuine, not faked. REST API for Playwright, Puppeteer, Selenium. $9.99/month or $99.99/year. 7-day free trial, card required — cancel any time in the seven days and you are not charged. Unlimited profiles.
Get started → · How it differs from Multilogin / GoLogin / AdsPower
Related Posts
Ready to manage multiple accounts?
Seven days free, then $9.99/month — one plan, everything included.