DataDome bot detection no longer works like a perimeter filter that maps an IP address to a rule and a block. In 2026 it behaves like a real-time classification pipeline: network, browser, behavioral and reputation signals are fused into features, scored by machine-learning models and turned into a decision while the request is still in flight. This guide explains how that architecture fits together, what each signal layer contributes, and what it means for teams that run authorized automation and web-data pipelines.
The short version
DataDome does not rely on one fingerprint. It correlates many weak signals. Server-side data (HTTP, TLS, IP and ASN) is combined with client-side telemetry (browser, device, interaction events), session behavior and historical reputation. Models turn that evidence into a risk and intent assessment, and the response is proportional: allow, verify invisibly, challenge or mitigate.
About the diagrams and code in this guide. The architecture diagrams and Python snippets are conceptual illustrations of how a modern bot-management system can be built. They are not a description of DataDome’s private implementation. Statements about DataDome’s product are drawn from its public documentation.
Key takeaways
- The core engineering idea is signal fusion. No single fingerprint, header or IP decides the outcome.
- Signals are features, not verdicts. Weak observations become strong evidence when they occur together or contradict each other.
- “Is this a bot?” is the wrong question. The system has to work out who is acting, whether the identity can be trusted and what the actor is trying to do.
- Mitigation is proportional to confidence: allow, verify invisibly with Device Check, fall back to a CAPTCHA, or block.
- AI agents mean automation is no longer the same thing as abuse. Identity, provenance, behavior and intent are converging.
- Network infrastructure is one layer of the stack. Changing where a connection originates does not change browser telemetry, session behavior or intent signals.
From IP Rules to Distributed Intelligence
Why the perimeter model stopped workingBot detection used to be mostly a perimeter problem. The old model was straightforward:
IP โ Rule โ Block
Traditional bot controls assumed that automation would reveal itself through relatively stable infrastructure: suspicious IP ranges, abnormal request rates, malformed headers or known user agents. Those signals still have value. They simply cannot carry the entire decision any more.
Several things weakened them:
- Residential proxy networks weaken the relationship between IP reputation and identity.
- Browser automation can produce traffic that superficially resembles normal browser traffic.
- Distributed systems spread requests across addresses and sessions.
- AI agents add a new problem, because some automated traffic is explicitly desirable.
The modern model looks much more like a real-time classification pipeline:
Request
โ
Network signals
โ
Browser / device signals
โ
Behavior
โ
Reputation
โ
ML models
โ
Risk / intent
โ
Decision
That shift matters for e-commerce, marketplaces, travel platforms, financial services and any organization operating high-value APIs. An IP address is no longer a reliable representation of a user, and a CAPTCHA is too expensive in conversion, accessibility and friction to be the default answer to every ambiguous request.
DataDome’s own documentation illustrates the transition. Its current materials describe analysis across hundreds of signals, including device fingerprints, behavioral patterns, IP reputation, session dynamics, and HTTP and TLS information, with machine-learning models making real-time decisions.
A contemporary bot-management system therefore has to answer a more complicated question:
What does all available evidence tell us about this request, the actor behind it, and its likely intent?
Answering it means collecting evidence at the edge, browser, device, session, network and historical levels, and correlating it quickly enough to affect the request currently in flight.
What DataDome Actually Has to Determine
Intent, not just automation“Is this a bot?” is increasingly the wrong abstraction. Traffic arriving at an application might represent any of the following:
| Actor | Automated? | Typical stance |
|---|---|---|
| Legitimate human | No | Should pass without friction |
| Verified search crawler | Yes | Often commercially valuable |
| Uptime or monitoring system | Yes | Can be operationally essential |
| Approved commercial automation | Yes | Expected and permitted |
| AI crawler or agent | Yes | Depends on identity and authorization; a shopping agent may act for a customer |
| Abusive scraping automation | Yes | Extracting an entire product catalog may be undesirable |
| Credential-stuffing or account-takeover infrastructure | Yes | Harmful |
| Fraud automation targeting checkout, inventory, payments or promotions | Yes | Actively harmful, for example automated checkout acquiring scarce inventory |
Two requests can both be automated and still require opposite decisions. That is why intent classification is becoming more important than binary bot classification.
DataDome’s 2026 material increasingly uses this framing. Its intent-based protection documentation describes combining server-side signals such as request headers, IP reputation and traffic patterns with browser and device telemetry, while its Agent Trust product evaluates dimensions including intent, provenance and continuity.
The security question is moving from human or machine? toward what actor is this, can its claimed identity be trusted, and what is it attempting to accomplish?
The DataDome Detection Architecture
A multi-stage signal pipelineConceptually, the architecture can be represented as a pipeline that takes two families of signals, turns them into features and feeds several kinds of analysis into the models that produce a decision.
Incoming Request
โ
โโโโโโโโโโโโโโดโโโโโโโโโโโโโ
โ โ
Server-Side Signals Client-Side Signals
โ โ
โโโโโโโโผโโโโโโโ โโโโโโโโโผโโโโโโโ
โ โ โ โ โ โ
HTTP TLS IP/ASN Browser Device Events
โ โ
โโโโโโโโโโโโโโฌโโโโโโโโโโโโโ
โ
Feature Pipeline
โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโ
โ โ โ
Signature Behavior Reputation
โ โ โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโ
โผ
ML Models
โ
โผ
Risk / Intent
โ
โผ
Mitigation Decision
Read this as a conceptual detection architecture rather than a claim about DataDome’s private implementation. The important property is that signals are not independent verdicts. They become features.
A network characteristic may be weak evidence by itself. A browser anomaly may also be weak evidence. A navigation sequence may look unusual but legitimate. When those observations occur together, and contradict historical reputation or the claimed client identity, the combined evidence can become substantially stronger.
DataDome’s Protection API documentation reinforces the server-side half of this picture: integrations provide request metadata including HTTP headers and client IP information, with JA3 and JA4 TLS fingerprints recommended where available.
Server-Side Signals: HTTP, TLS, Network and Reputation
Evidence that exists before any JavaScript runsServer-side telemetry is valuable because it is available before client-side JavaScript necessarily executes. Three groups of signals matter most.
| Signal group | What can be examined | What makes it useful |
|---|---|---|
| HTTP | Header structure, protocol behavior, request consistency, endpoint usage | Consistency between characteristics that should normally agree |
| TLS | Handshake characteristics, JA3 and JA4 fingerprints | Whether transport behavior makes sense for the client the request claims to be |
| Network | IP reputation, ASN context, geography, traffic history, request velocity | Context that adjusts confidence in the other signals |
HTTP characteristics
HTTP provides considerably more information than a user-agent string. Detection systems can examine header structure, protocol behavior, request consistency, endpoint usage and the relationships among characteristics that should normally agree.
The useful feature is often not “header X has value Y.” It is consistency between signals.
TLS metadata
TLS adds another observational layer. Handshake characteristics can provide information about the software stack establishing a connection, and DataDome explicitly recommends collecting JA3 and JA4 fingerprints for its custom Protection API integrations where possible.
The objective is not to treat a TLS fingerprint as identity. The useful information comes from correlation: does the transport behavior make sense given what the application-layer client claims to be?
Network context
Network-derived features can include IP reputation, ASN context, geography, traffic history, request velocity and TLS fingerprint metadata when available. A defensive telemetry layer might normalize those observations like this:
def build_network_features(request):
return {
"method": request.method,
"asn": request.asn,
"country": request.country,
"request_rate": request.request_rate,
"tls_family": request.tls_family,
"known_network": int(request.network_reputation == "known"),
}
Mind the encoding. Production systems typically transform high-cardinality fields before feeding them into statistical models. ASN numbers and country codes should not accidentally become ordinal numeric variables.
The objective is feature engineering, not one brittle rule such as “this ASN equals bot.”
Client-Side Detection: Why the Browser Became a Sensor
The execution environment is also a telemetry sourceThe browser has become one of the most important telemetry surfaces in modern bot detection.
DataDome’s JavaScript Tag documentation describes collecting behavioral information such as mouse movements and keystrokes, along with information about the operating system, the browser, the GPU and the consistency of built-in browser functionality. Separately, Device Check executes client-side code that collects hundreds of device and environment signals and runs automated checks when a request requires additional verification.
The browser is no longer just the execution environment. It is also a telemetry source.
Client-side instrumentation can provide evidence about:
- browser state and JavaScript execution;
- device characteristics and hardware-visible properties;
- rendering capabilities;
- interaction events.
No individual property should be read as a secret password proving humanity. What matters is consistency. A genuine browser is a complicated collection of APIs, runtime behavior, rendering characteristics, hardware capabilities, event patterns, network properties and application state. Detecting automation becomes a matter of deciding whether those observations form a coherent system. Our explainer on why modern web automation runs on Chromium covers the browser side of that picture in more depth.
This is also why server-side and client-side detection complement each other. The server sees the request from the infrastructure perspective. The browser gives another view of the execution environment. Disagreement between those two perspectives is itself useful evidence.
Behavioral Analysis: From Individual Requests to Time Series
A request is an event, a session is a sequenceA request is an event. A session is a sequence. That distinction dramatically expands what a detection system can learn.
Session
โ
โโโ Homepage
โโโ Search
โโโ Category
โโโ Product
โโโ Product
โโโ Cart
โโโ Checkout
Once traffic is modeled over time, features can describe request frequency, session duration, endpoint transitions, repeated sequences, navigation entropy, failed actions and deviations from normal workflows.
Start with simple session features
A simple defensive feature extractor might begin with counts and path frequencies:
from collections import Counter
def session_features(events):
paths = [event["path"] for event in events]
return {
"event_count": len(events),
"unique_paths": len(set(paths)),
"path_frequency": Counter(paths),
}
Add timing statistics
A production feature pipeline would go further and derive statistics such as inter-event-time distributions and transition frequencies:
def timing_features(events):
timestamps = [event["timestamp"] for event in events]
gaps = [
timestamps[i] - timestamps[i - 1]
for i in range(1, len(timestamps))
]
if not gaps:
return {"mean_gap": 0, "gap_variance": 0}
mean = sum(gaps) / len(gaps)
variance = sum((x - mean) ** 2 for x in gaps) / len(gaps)
return {
"mean_gap": mean,
"gap_variance": variance,
}
The goal is not to assume that “fast equals bot.” Humans can be fast and software can intentionally be slow. The model instead asks whether the observed sequence resembles expected behavior for this application and context.
DataDome says its broader detection engine incorporates behavioral patterns, time-series anomalies, aggregate traffic behavior, supervised learning, signatures and anomaly detection.
Reputation Systems: Why Historical Context Matters
Real-time classification with a memoryReal-time classification becomes substantially more useful when the system can remember. Reputation can exist at several scopes, and each scope answers a different question.
| Scope | What it tells the system |
|---|---|
| IP | Infrastructure history for a single address |
| ASN | Broader network context |
| Session | A short-lived behavioral identity that connects requests |
| Account | Security decisions tied to authenticated application activity |
| Device | Continuity when network addresses change |
These signals also need different time windows. A five-minute request spike and a 30-day history of suspicious account behavior are fundamentally different evidence. In a robust design, reputation is treated as time-bounded, continuously updated context rather than a permanent “good” or “bad” label.
That matters because legitimate and malicious infrastructure overlap. Residential networks, cloud platforms, carrier NAT, corporate proxies and shared devices make static reputation increasingly fragile.
Historical context should modify confidence, not replace current evidence.
Machine Learning and Risk Scoring
Many weak signals, one proportional decisionA simplistic rule engine produces discrete decisions:
if suspicious:
block()
A richer architecture first creates a risk representation. For illustration:
def calculate_risk(features, weights):
score = 0.0
for name, weight in weights.items():
score += features.get(name, 0) * weight
return max(0.0, min(score, 1.0))
This is intentionally simpler than a production classifier, but it shows the architecture: many weak signals contribute to one decision, and the response scales with the score.
| Risk level | Response | What it means in practice |
|---|---|---|
| Low | Allow | Known legitimate traffic proceeds without unnecessary friction |
| Medium | Additional verification | Ambiguous traffic is asked for more evidence |
| High | Mitigate | Unambiguous malicious automation is stopped immediately |
Real systems can use ensembles, supervised classifiers, anomaly detection, signature models and specialized models for particular attack classes. The architectural advantage is that mitigation becomes proportional to confidence.
A risk score does not have to be read as a literal “bot probability” either. In a mature architecture, scoring can combine several dimensions such as confidence, reputation, identity consistency, behavioral anomaly and intent.
Device Check and Invisible Verification
Resolving uncertainty without asking the userOne of the more important developments is verification that does not immediately require human interaction.
DataDome documents Device Check for suspicious requests, including cases where bot evidence is not strong enough to justify an immediate block or CAPTCHA. Client-side JavaScript collects hundreds of device and environment signals and performs automated checks. If more information is still required, DataDome can then present a CAPTCHA.
The result is a progressive decision tree:
Request
โ
โผ
Initial Classification
โ
โโโ Low risk โโโโโโโโโโโโโโโโบ Allow
โ
โโโ Uncertain โโโโโโโโโโโโโโโบ Invisible verification
โ โ
โ โโโ Pass โ Allow
โ โโโ Uncertain โ CAPTCHA
โ
โโโ High-confidence threat โโบ Mitigate
That architecture matters commercially. CAPTCHAs impose latency and interaction cost. They can affect accessibility and conversion, and they force legitimate users to pay for uncertainty in the security model. Invisible verification moves some of that uncertainty into machine-readable evidence instead.
Why Modern Bot Detection Is a Streaming Systems Problem
Security software that looks like data infrastructureAt sufficient scale, bot detection stops looking exclusively like cybersecurity software and starts looking like streaming data infrastructure. A generalized architecture might resemble this:
Edge
โ
โผ
Event Collector
โ
โผ
Kafka
โ
โโโโโโโโโโดโโโโโโโโโ
โผ โผ
Feature Processing Reputation Store
โ โ
โโโโโโโโโโฌโโโโโโโโโ
โผ
ML Models
โ
โผ
Decision API
Reference architecture only. This is not a description of DataDome’s private backend. The engineering constraints are nevertheless representative of the problem.
The synchronous path must remain extremely fast because security processing sits in front of application workloads. At the same time, behavioral detection needs events to be aggregated across requests, sessions, accounts and time windows. Typical building blocks include:
- A streaming platform such as Kafka to distribute telemetry to feature processors.
- Low-latency stores such as Redis to maintain rolling counters or recent reputation state.
- Feature stores to keep model training and inference consistent.
- Model-serving infrastructure with predictable latency, versioning, monitoring and rollback.
The result is effectively two systems operating together:
| Path | Flow | Constraint |
|---|---|---|
| Synchronous | Request โ Features โ Model โ Decision | Must be fast enough to sit in front of the application |
| Asynchronous | Events โ Aggregation โ Reputation โ Model updates | Must aggregate across requests, sessions, accounts and time windows |
DataDome’s intent-based protection documentation explicitly distinguishes synchronous real-time detection from asynchronous behavioral analysis across sessions and users, with a third layer feeding signals and behavioral data back into ML model training.
False Positives: The Metric Security Teams Can’t Ignore
Blocking everything is not a strategyA security system that blocks everything has excellent recall and catastrophic business performance. Two basic metrics illustrate the problem:
precision = TP / (TP + FP)
recall = TP / (TP + FN)
High recall means detecting a large proportion of malicious traffic. High precision means that traffic classified as malicious usually is malicious. Optimizing one without the other can be dangerous:
- For an e-commerce platform, a false positive can mean rejecting a paying customer during checkout.
- For an authentication system, excessive verification can increase abandonment.
- For APIs, aggressive controls can disrupt legitimate integrations.
Bot-management performance therefore sits inside a multi-objective optimization problem that balances security, conversion, latency, user friction and revenue.
Observability should reflect that. Security teams should correlate detection outcomes with application metrics: authentication failures, checkout conversion, challenge completion, cart abandonment, API errors, latency percentiles and customer-support incidents.
DataDome itself describes using customer-side business metrics, including login denial rates, cart abandonment, bounce rates and traffic anomalies, as feedback for evaluating detection quality.
DataDome in the Age of AI Agents
Automation is no longer a synonym for abuseAI agents break the cleanest remaining assumption in traditional bot management: that automation is inherently undesirable. The spectrum now runs from a human, to a bot, to an AI agent, to an AI agent acting on behalf of a human.
An agent might research products, compare prices, book travel, interact with APIs or eventually complete transactions with explicit user authorization. Blocking it simply because it is automated may be incorrect. Trusting it because it claims to be a recognized AI service is also dangerous. For background on how these systems reach the web in the first place, see our guide to how LLMs, RAG pipelines and agents collect web data.
DataDome reported on July 16, 2026 that its network processed 17.7 billion AI-agent requests during Q2 2026, up 45% from Q1. Its research also highlights impersonation as a significant problem: an HTTP user-agent string claiming to represent a known AI service is not, by itself, proof of identity.
That changes the core security model. The old question was simply human or bot? The new questions form a chain:
- Who is acting?
- Can that identity be verified?
- On whose behalf are they acting?
- What are they trying to accomplish?
- Is that action permitted?
DataDome now describes its platform in terms of bot and agent trust management, including policies for distinguishing legitimate from malicious agents and mechanisms for authenticating approved bots and AI agents. Identity, provenance, behavior and intent are converging.
Network Infrastructure Is Still Part of the Architecture
One layer, in its proper contextNone of this makes network infrastructure irrelevant. It puts networking in its proper context, as the first of several layers a detection system evaluates:
Network โ Browser โ Session โ Behavior โ Application โ Intent
For legitimate enterprise data operations, proxy infrastructure remains useful for global QA, localization testing, market intelligence, regional content validation, AI data acquisition where authorized, and geographically distributed data systems. Providers such as ProxyEmpire occupy one layer of that larger architecture, whether the workload calls for rotating residential proxies or another network type.
The mistake is treating the network layer as a complete solution to modern bot detection. Changing where a connection originates does not erase browser telemetry, session behavior, application semantics, historical reputation or intent signals. Conversely, security systems cannot infer everything about an actor from an IP address.
For enterprise teams, the relevant design principle is separation of concerns:
Modern data infrastructure has to engineer all of them.
What This Means for CTOs Building Data Platforms
Don’t optimize only the scraperFor teams operating authorized web-data pipelines, the strategic mistake is optimizing only the scraper. The production system is larger, and every layer of it requires observability.
| Layer | What to measure |
|---|---|
| Network infrastructure | Request success rates, regional availability |
| Browser / runtime | Runtime failures |
| Session management | Session stability |
| Acquisition and extraction | Extraction completeness, schema drift |
| Validation | Duplicate rates, freshness, downstream validation failures |
| Data products | Cost per usable record |
Security and data engineering are converging here for the same reason: both sides increasingly depend on telemetry. The defender needs enough telemetry to determine whether an interaction is trusted. The data platform needs enough telemetry to determine whether its acquisition pipeline is functioning correctly and operating within authorized boundaries.
For CTOs, this is ultimately an architecture problem rather than a scraping-library decision. If you are working through the runtime layer of that stack, our guide to designing a browser cluster for 10,000 concurrent Chromium sessions covers capacity planning, scheduling and observability in detail.
Conclusion: Bot Detection Is Becoming Intent Detection
The system is the correlation, not the signalDataDome bot detection in 2026 illustrates a larger transformation in web security. IP reputation still matters. So do HTTP characteristics, TLS metadata, browser and device telemetry, behavioral sequences and historical reputation. But none of them is the system.
The system is the machinery that collects those observations, turns them into features, correlates them across time and identity scopes, applies statistical and machine-learning models, and produces a decision quickly enough to protect an application without unnecessarily disrupting legitimate users.
AI agents make that architecture even more important because automation is no longer synonymous with abuse. Some machines should be blocked. Others should be authenticated. Some should be constrained by policy. Others may eventually deserve access comparable to a human acting through a trusted client.
The question is changing from “is this traffic automated?” to “who or what is acting, can it be trusted, and is its behavior acceptable in this context?”
For teams building authorized data pipelines, the practical takeaway is the same one defenders have reached: treat the network, runtime, session and data layers as one observable system, and engineer each of them deliberately.
Frequently Asked Questions
DataDome bot detection, in shortHow does DataDome detect bots?
By correlating many signals rather than relying on one. Server-side data such as HTTP headers, TLS fingerprints, IP reputation and ASN context is combined with client-side telemetry from the browser and device, session behavior over time and historical reputation. Machine-learning models turn that evidence into a real-time decision.
Does DataDome only look at IP addresses?
No. IP reputation is one input among many. Residential networks, cloud platforms, carrier NAT, corporate proxies and shared devices mean legitimate and malicious infrastructure overlap, so an IP address on its own is not a reliable representation of a user.
What is DataDome Device Check?
Device Check is a verification step for suspicious requests where the evidence is not strong enough to justify an immediate block or CAPTCHA. Client-side JavaScript collects hundreds of device and environment signals and runs automated checks without prompting the user. If more information is still needed, DataDome can then present a CAPTCHA.
Why does DataDome use JA3 and JA4 TLS fingerprints?
TLS handshake characteristics give information about the software stack opening a connection. DataDome recommends collecting JA3 and JA4 fingerprints in custom Protection API integrations where possible. The value comes from correlation: whether the transport behavior is consistent with the client the request claims to be.
How does DataDome treat AI agents?
Not as automatically good or bad. DataDome describes its platform in terms of bot and agent trust management, with policies for distinguishing legitimate from malicious agents and mechanisms for authenticating approved bots and AI agents. A user-agent string that claims to be a known AI service is not proof of identity by itself.
Does a proxy change how a request is classified?
It changes one layer. A proxy determines where a connection originates, which is useful for legitimate work such as global QA, localization testing and regional content validation. It does not alter browser telemetry, session behavior, application semantics, historical reputation or intent signals, which are evaluated alongside the network layer.
Build the network layer of your data platform on ProxyEmpire
Residential, mobile and datacenter proxies with sticky or rotating sessions and precise geographic targeting, for authorized QA, localization testing and web-data pipelines.














