BLOG

eCommerce Insights

Traffic Is Growing, but Customer Activity Is Not: How to Separate Bots from Real Shopify Visitors

Why native filtering is not enough, which signals are useful, and how to restore the decision value of your reports

A sharp rise in traffic usually looks like good news. If sessions multiply while orders, enquiries, and add-to-cart activity remain flat, however, I first examine data quality rather than marketing performance.

Bot traffic may be the cause: search crawlers, price-comparison services, monitoring tools, product-data scrapers, AI agents, or automated attacks. These visitors are not equivalent. Some support indexing and content distribution, some create operational load, and others distort reporting or imitate customer actions.

The answer is therefore not a single “block bots” button. We first need to understand which sessions entered analytics, which ones affect decisions, and which controls will avoid harming search visibility, advertising, and legitimate customers.

What the problem looks like in real data

In one publicly documented case, the store's normal volume was about 40,000 sessions per month. It received more than 500,000 automated sessions in five days. In a sample of 234,790 sessions, every session used the same 1366 × 1366 screen resolution, Chrome browser, and Macintosh operating system. Location resolved to Singapore with no city or region, while average engagement stayed within a narrow 20-to-28-second range.

That last detail matters. A longer session does not prove that a visitor is human. A modern headless browser can execute JavaScript, wait for a page to load, and send events, making its behavior look plausible in a high-level report.

The same activity expanded 466 real base paths into 99,984 unique URLs. More than 54,000 represented filtered collection pages. The crawler combined filter parameters and produced a combinatorial stream of addresses. Robots.txt rules did not stop it: compliant search engines observe these instructions, while an abusive crawler can simply ignore them.

In another store, 52,843 sessions were recorded from August 1 through August 21. Of those, 46,768, or 88.5%, were classified as bots. The published bot segment had a zero conversion rate, a 96.37% bounce rate, a 99.1% desktop share, and 96.36% direct traffic with no referrer. The store's genuine audience was primarily mobile, and its home market did not appear among the main countries generating the automated traffic.

These are merchant-reported data from individual stores, not Shopify-wide benchmarks. They are still valuable because they show why a single metric cannot be interpreted in isolation. Suspicion becomes meaningful when several anomalies form one persistent profile.

What bots distort

The first consequence is a misleading conversion rate. If purchases remain stable while automated sessions inflate the denominator, the conversion rate falls. A manager may conclude that product pages, advertising, or checkout have deteriorated even though real customer behavior has not changed.

The second is false product and collection interest. A bot can open product pages in bulk, cycle through filters, and generate thousands of URL variants. Reports then show page popularity that carts and sales do not support.

The third is polluted advertising signals. If an automated browser executes an advertising pixel, page views and other upper-funnel events can reach audiences and attribution systems. Advertising platforms also filter invalid activity, but different systems should not be expected to classify every session in the same way.

The fourth is higher cost in tools priced by sessions, profiles, events, or requests. Shopify may absorb the traffic while a connected service exhausts its allowance faster or stores a large amount of unusable data.

The fifth is lost team attention. Alerts about a traffic surge, a conversion decline, or unexpected countries begin competing with genuine store issues. This is why early Shopify store warning signs are useful only when the reliability of the underlying data is understood.

Which signals can identify probable bot traffic

Geography, bounce rate, and missing referrer information do not prove automation on their own. A campaign, corporate network, new audience, or tracking error can create similar patterns. I use a combination of signals and compare it with the normal profile of the specific store.

Signal Why it matters Check before concluding
Sudden session surge Traffic rises without comparable order or campaign growth Campaign launches, press, seasonality, and tracking errors
Unexpected geography Countries do not match sales markets; city or region may be missing Shipping markets, international campaigns, travel, and corporate networks
Direct traffic dominates A high share arrives without a source and lands directly on products Lost tags, apps, browser privacy, and links shared in messages
Identical technical profile Browser, operating system, resolution, and duration repeat Real audience patterns and the accuracy of captured parameters
URL-count explosion Automation cycles through collection filters and parameters Search behavior, apps, canonical URLs, and campaign tags
No commercial actions Sessions create no carts, checkouts, or purchases Funnel length, cart defects, and the quality of a new traffic source

 

A practical first-pass rule is to mark a segment as probable bot traffic when three or four independent signals coincide, then investigate it separately. This is a working diagnostic heuristic, not a Shopify or Google standard. It does not justify permanently deleting data or blocking an entire country without validation.

Automated traffic types also need to be separated. A search crawler, social preview service, price scraper, spam generator, and card-testing program pursue different goals and leave different traces. Grouping them as “bots” is convenient for a summary report but insufficient for selecting a control.

Shopify's native filter is the right starting point

In Shopify Analytics > Reports, a merchant can add the Human or bot session dimension and filter a report to human sessions. I consider this the first mandatory step: it requires no theme change and quickly shows how automated traffic affects sessions, visitors, and conversion.

Shopify's official guidance describes the classification as conservative. The system prefers to miss some bots rather than misclassify a real customer as automated activity. Sophisticated traffic may not be identified immediately. The filter applies only to session-related metrics, covers new data from October 7, 2025 onward, and is not available for Headless and Hydrogen storefronts.

In one of the cases reviewed, hundreds of suspicious sessions per day remained after the filter was enabled. This is consistent with the published limitation: native classification is useful, but it is not final truth.

My assessment is that the Shopify filter should be used, but a management report should not rely on one label alone. The classification needs to be compared with geography, device mix, source, landing URL, and confirmed commercial events.

Google Analytics automatic filtering does not solve everything either

Google Analytics 4 automatically excludes known bots and spiders. This is a useful baseline, and there is no separate standard switch to enable. The limitation is contained in the word “known.” Automation that imitates a normal Chrome browser, executes JavaScript, and changes its network origin can escape that classification.

I therefore do not treat a difference between Shopify Analytics and Google Analytics 4 as proof that one system is wrong. They use different event sources, identification rules, processing windows, and consent constraints. The first step is to align periods and segments, then compare sessions with orders, checkout starts, purchase events, and Google Search Console clicks.

My assessment is that Google Analytics 4 is valuable as an independent verification layer, but not as a universal bot-traffic cleaner.

Popular recommendations that help only partially

Block countries

If automated traffic originates in a region where a store neither sells nor ships, a geographic restriction may reduce noise. Country is still a weak identifier: proxies and virtual private networks change origin, while a genuine customer may be travelling or using a corporate network.

Some apps also block a visitor only after the page begins loading. The request has already reached the store, and analytics or advertising events may already have fired.

My assessment is that geographic restrictions are justified only by a clear business rule and after testing when the block occurs. They should not be the primary method for identifying bots.

Add a CAPTCHA challenge

CAPTCHA is useful when a store needs to protect an action: a form, account creation, login, or a selected checkout step. It is much less effective against a program that only reads product pages and cycles through collection URLs.

My assessment is that CAPTCHA is a local safeguard for sensitive actions, not a filter for all visits. Excessive challenges also damage the legitimate customer journey.

Disallow crawling in robots.txt

Robots.txt guides compliant search crawlers and helps prevent unnecessary crawl paths. It is not a network-level prohibition. In the documented case, a significant share of automated requests targeted URL combinations already disallowed by Shopify's standard rules.

My assessment is that a correct robots.txt file matters for search optimization, but it is not a security control. The relationship between indexing, URL structure, and data quality is covered in our Shopify SEO and GEO optimization resources.

Hide URL parameters in reporting

Normalizing addresses and combining variants with unnecessary parameters makes a report easier to read. Instead of tens of thousands of rows, the analyst can see a base collection or product. The bot session remains, however, and the advertising pixel can continue receiving events.

My assessment is that this is useful reporting hygiene, not removal of the cause. Raw URLs should remain available in a technical report because their structure helps identify the crawl mechanism.

Install a blocking app or add theme code

Such a control may stop a known pattern by country, browser agent, or URL. A theme does not see every network signal, however, and a bot may access platform interfaces without rendering the page. A block that runs after loading begins does not prevent the request itself.

My assessment is that a narrow rule is reasonable after a pattern is confirmed and false positives are tested. Installing an app without understanding where its decision occurs creates a false sense of protection.

Move filtering to the server side

Server-side event delivery can keep some suspicious page views out of advertising systems and provide better control over event rules. This is one of the strongest recommendations when the main loss involves attribution, remarketing audiences, and paid analytics.

The configuration cleans measurement flows; it does not necessarily block requests to the storefront. A poor rule can remove genuine conversions, so it requires logs, a comparison period, and order reconciliation. Consent and privacy requirements also remain applicable.

My assessment is that server-side filtering is a strong method for protecting advertising signal quality, but only when the rules are evidence-based and the outcome is monitored.

Use Cloudflare in an orange-to-orange setup

The discussion includes one merchant's positive experience with Cloudflare O2O. At the time of verification, Cloudflare's official documentation states that O2O can be used with any Cloudflare zone plan, not only Shopify Plus. That primary documentation carries more weight than unconfirmed comments.

O2O changes the infrastructure layer. Domain configuration, certificates, caching, security rules, checkout, and app compatibility must be tested. Some Cloudflare capabilities are restricted on checkout paths. The setup should be deployed as a controlled technical change with a rollback plan. Our Cloudflare for Shopify page explains the broader capabilities and constraints.

My assessment is that O2O can be a powerful pre-Shopify traffic-control layer, but it is not a default switch or the first action for every store.

Separating reporting cleanup from traffic blocking

I separate this work into two outcomes.

The first is trustworthy reporting. The store needs a segment that represents customer behavior with reasonable confidence while preserving all traffic for technical observation. In some cases, that alone restores the ability to make conversion and advertising decisions.

The second is reducing unwanted activity. This may involve narrow app or theme rules, form protection, changes to advertising event delivery, an edge layer, or a different architecture. These controls are not always required and should be judged by cost, false-positive risk, and demonstrated business impact.

Combining the goals is dangerous. A team can clean reporting without reducing requests. It can block some traffic while historical data remains distorted. It can also protect the store too aggressively and lose useful search crawlers or genuine customers.

How an individual analysis is built

Within ongoing support, a Shopify store diagnostic, or optimization work, we normally handle this task as part of a broader package because it requires store-specific data and observation over time.

1. Preserve the raw layer

Before filtering, we record source reports, periods, tracking settings, and the time the anomaly began. Historical data is not deleted. This makes it possible to review the rule if a segment proves incorrect.

2. Establish the store's normal profile

We compare the anomaly with several normal periods across geography, devices, sources, landing pages, engagement, bounce, carts, checkouts, and purchases. A store running international campaigns will have a different acceptable profile from a local store without active advertising.

3. Build segments from combined signals

We separately examine identical technical characteristics, direct traffic without a source, geographic outliers, unusual URLs, and the absence of commercial actions. Doubtful sessions are labeled probable rather than declared bots without sufficient evidence.

4. Create two reporting layers

The management report shows the cleaned human segment and supports decisions about conversion, pages, and campaigns. The technical report preserves all traffic, the share of known and probable bots, new patterns, and the volume excluded from management reporting.

5. Validate against commercial events

Every rule is reconciled with orders, checkout activity, refunds, and actual revenue. If a rule removes paid orders or established customers, it is too broad. Advertising should prioritize confirmed purchase and checkout events rather than relying only on page views.

6. Protect downstream systems

After validation, we decide which events should be withheld from advertising and paid analytics tools. If a server-side layer is needed, it begins in comparison mode. A permanent rule is applied only after measuring data loss and false exclusions.

7. Monitor and adjust

Automated traffic profiles change. Criteria that worked in August cannot be treated as permanent. Alert thresholds, regular reconciliation, and a clear process owner are required. This is what turns a one-time analysis into part of store support.

Why complete bot removal cannot be promised

Shopify has strengthened restrictions for bots and agents. Since May 2026, unsigned automated requests to hosted storefront pages and the Storefront API receive the strictest rate limits, while Web Bot Auth supports identification. This is a useful platform change, but it does not mean that every automated request will be named, stopped, and removed from every connected analytics system.

Eliminating all automated traffic without side effects is not realistic. Search engines, social platforms, and monitoring services use bots as well. The specialist's task is not to reach an attractive “zero bots” number. It is to make data reliable enough for decisions and reduce the harm caused by specific automation types.

My conclusion

Bot traffic can be isolated statistically when the analysis uses a persistent combination of geography, devices, source, behavior, URL structure, and commercial events rather than one indicator. The result still applies to a specific store, period, and dataset.

I recommend beginning with two reports: a cleaned management view and a complete technical view. The next step is to determine where suspicious events are sent, what they cost the business, and whether blocking is justified. Only then should the team select an app, server-side filtering, Cloudflare, or another technical control.

This analysis fits naturally within a support or optimization package because its value comes from recurring data-quality review, not a one-time formula. No store is protected from a new automation pattern forever. It can, however, be prepared to detect one quickly, avoid mistaking false numbers for customer behavior, and avoid spending budget on fixing a problem that does not exist.