A collection page on a large catalogue can stop showing filters without anyone touching a line of code. Shopify’s own documentation states that “collections that contain more than 5,000 products don’t display filters”, so a store that grows past that threshold hands its shoppers an unfiltered list of several thousand items. Nobody designed that experience. It arrived with the catalogue.
That is the kind of problem an ecommerce UX audit exists to surface, and it is also the kind that tends to get lost. Most audits produce a ranked list of usability violations ordered by how badly each one breaches a design principle. Revenue is not organised that way. A usability problem is worth fixing in proportion to how many shoppers meet it, where in the funnel they meet it, and what share of them could plausibly be recovered. An ecommerce UX audit that does not carry those three figures for every finding usually produces a document that reads well and changes very little.
What an Ecommerce UX Audit Examines, and What It Does Not
HubSpot describes a UX audit as “a systematic approach to assessing a product for gaps, challenges, and opportunities in user experience (UX)”, and notes that for a commerce site “your audit plan generally follows the customer’s path from the product page to the payment page”. That framing sets a useful boundary: the audit follows a journey, not a sitemap, and it stops where the journey stops mattering commercially. On an established store the scope usually covers category navigation, internal search and filtering, the product page, the cart, the checkout, account and returning-customer flows, and the post-purchase pages where support tickets originate.
It is worth separating this from two neighbouring exercises often bought under the same name. A conversion audit, of the kind described in our guide to what a professional ecommerce CRO audit includes, starts from commercial performance and works outwards into whatever is causing it, which may be pricing, merchandising or traffic quality as easily as usability. A technical audit starts from the platform and asks whether the store is correctly built, indexed and integrated. A usability review sits between them and asks a narrower question: where does the interface itself make a willing buyer’s task harder than it needs to be?
The narrowness matters. A store that commissions a usability review hoping it will explain a revenue decline is often disappointed, because a decline driven by a shift in paid traffic mix will not surface in a heuristic review of the product page.
Why Usability Problems Rarely Show Up in Standard Reporting
Standard analytics reports the outcome of a behaviour rather than the behaviour itself. A high exit rate on a collection page is compatible with a shopper who found nothing relevant, one who found what they wanted and opened it in a new tab, and one whose filter panel never rendered. The number is identical in all three cases, so the report says where to look but rarely what to look at.
Aggregation compounds it. Site-wide averages absorb the segments where an interface fails hardest, so a checkout step that works acceptably on desktop and poorly on older Android devices may show a modest overall completion rate that nobody flags. There is also the problem that broken things and disliked things look alike from a distance: when shoppers repeatedly fail to use a control, the data does not say whether it is defective, mislabelled or simply unwanted, and that distinction can change the cost of the fix by an order of magnitude. Several signals suggest a store has accumulated usability debt its reporting is not surfacing:
- Support answers the same questions repeatedly even though the site states the answer somewhere, which often means the answer sits where shoppers do not reach it at the moment they need it.
- Internal search terms include category names and navigation labels, suggesting people are searching for things the menu was supposed to provide.
- Add-to-cart rates differ sharply between product templates nobody intentionally designed differently, which points to template drift rather than product appeal.
- Design disagreements are settled by seniority, because no shared evidence exists to settle them any other way.
- Nobody can say what proportion of mobile sessions reach the payment step, because the funnel has never been built for that segment specifically.
Four Kinds of Evidence, and What Each One Cannot Tell You
A usability finding is only as credible as the evidence behind it, and the four methods below answer genuinely different questions. Combining them is what allows a finding to survive the meeting in which someone asks how you know.
Expert review against usability heuristics
A heuristic evaluation is, in HubSpot’s description, “a usability test where experts assess an interface based on a set of principles, called heuristics”. The same source recommends “at least two evaluators to avoid bias” while cautioning that “any more than ten can make the data harder to handle”. Its strengths are speed and coverage, including the low-traffic templates behavioural tools will never gather enough sessions to judge. Its weakness is that it produces expert opinion rather than user behaviour, since a reviewer fluent in ecommerce conventions does not experience a first-time buyer’s confusion. Used alone, it generates long lists of small findings with no basis for ranking them.
Behavioural instrumentation on live traffic
Session recordings, heatmaps and behavioural flags observe what shoppers do without asking them anything. Microsoft’s Clarity documentation defines several of these signals precisely: a rage click is recorded when “the user clicks multiple times in a clustered area in rapid succession”, a dead click when “a user clicks on an element but gets no feedback in a reasonable amount of time”, and excessive scrolling flags sessions with “higher amounts of vertical scrolling than the expected average”.
These signals are close to unambiguous, and a cluster of dead clicks on a size selector is not a matter of taste. The limitation is that they describe symptoms without motive and favour high-traffic pages. A defect on a template receiving forty sessions a week may matter commercially for a high-value product and still produce almost no recordings worth watching.
Quantitative funnel analysis
Funnel analysis converts a usability observation into a number of lost sessions. Google’s documentation explains that funnel exploration “lets you visualize the steps your users take to complete a task and quickly see how well they are succeeding or failing at each step”, and distinguishes open funnels, where users may enter at any step, from closed funnels, where users must begin at the first step to be counted. A closed funnel measures the journey you designed; an open funnel measures the journey people actually take.
Breakdowns are where most of the value sits, since a funnel split by device or market usually exposes a gap the blended figure conceals. One caveat belongs in the analysis: Google notes that when a breakdown is applied, “users are only attributed to the first instance of the breakdown value that applies to them”. Someone who browses on a phone and buys on a laptop will not appear in both columns, so cross-device behaviour is easy to misread as a mobile conversion failure.
Moderated testing with real participants
Where heuristic review gives expert judgement and behavioural tools give observed actions, moderated testing gives reasoning. HubSpot draws the contrast directly, noting that usability testing “puts real users in the driver’s seat” while evaluators work through the interface themselves. Watching a participant try to establish whether an item will arrive before a deadline, and hearing what they conclude when the page does not say, is usually more actionable than any amount of scroll-depth data. The trade-off is cost and sample size, so it works better as a diagnostic instrument than a measurement one, and it earns that cost when the other methods agree something is wrong at a step and disagree about why.
How to Turn a Usability Finding into a Revenue Estimate
This is the step most audits skip, and skipping it is why so many usability reports are read once and shelved. The sequence below is deliberately conservative, because an estimate that later proves optimistic damages the credibility of every recommendation behind it.
- Attach the finding to a funnel step you already measure. A finding that cannot be located on a step your reporting contains cannot be sized, and belongs in a hypothesis log rather than on the roadmap.
- Count the affected population, not the total audience. A defect confined to one browser, market or template affects a subset, and that subset should be weighted by commercial value rather than volume. Internal search illustrates the point: Salesforce reports that “although only 16% of shoppers use the search function, they generate 55% of all online revenue”. Search findings may therefore deserve more attention than their traffic share implies, though the gap reflects the intent of people who search as much as the quality of the search itself.
- Estimate the recoverable share, and be pessimistic about it. Not every shopper who meets an obstacle abandons because of it, and not every one who abandons would have bought. A defensible estimate assumes only a modest fraction is recoverable and states that fraction explicitly, so it can be argued with.
- Convert to margin rather than revenue. Multiplying recovered orders by average order value overstates the case for any store with meaningful cost of goods, and makes low-margin categories look more attractive than they are.
- Divide by the cost to implement. Ranking by estimated value alone favours large, slow projects. Ranking value against development effort tends to surface a different and more useful order, in which several small defects outrank one redesign.
- Decide whether the change needs a test or simply needs shipping. Testing whether a broken control should be repaired wastes traffic. Genuine design choices where the direction is uncertain belong in an experiment.
These figures are estimates and belong in the report as ranges, not forecasts. Their purpose is to order work rather than predict revenue, and an ordering built on transparent assumptions survives scrutiny far better than a severity score whose basis nobody can reconstruct three months later.
Key takeaway: the value of a usability finding is set by the number of shoppers who meet it, the commercial weight of the step where they meet it, and the share who can realistically be recovered. A finding carrying none of those three numbers is an observation, not a priority, and it will lose every argument against a feature request that has a business case attached.
Which Findings Should Be Tested and Which Should Simply Be Fixed
Treating every usability finding as a test candidate is a common and expensive habit. Experimentation consumes traffic, and traffic is finite, so it is better reserved for questions where the answer is genuinely unknown.
| Finding type | Typical evidence | Test or fix directly | Risk if handled wrongly |
|---|---|---|---|
| Functional defect | Dead clicks, click errors, failure reproducible on a device or browser | Fix directly, since there is no hypothesis to test | Testing delays a repair that has no downside |
| Layout instability | Poor layout shift scores, misdirected taps near a primary action | Fix directly, then verify the metric moved | Treating it as taste leaves a mechanical cause unaddressed |
| Missing decision information | Support tickets, search queries, participants asking for it aloud | Usually fix directly, though placement may be worth testing | Withholding delivery cost until checkout moves abandonment later rather than removing it |
| Navigation and filtering structure | Search terms matching category names, deep exits from collections | Research first, then test if the change is substantial | Restructuring on intuition can disturb organic landing pages |
| Layout, wording and hierarchy | Heuristic review, scroll and attention patterns | Test where traffic allows, otherwise ship and monitor | Small refinements often need more traffic than the store receives |
| Pricing and policy changes | Behaviour at the payment step, competitor benchmarking | Test, and measure margin alongside conversion | Conversion may improve while contribution falls |
The layout instability row is the clearest case of a usability problem with a mechanical cause. Google’s guidance on Cumulative Layout Shift describes how unexpected movement can make users “click the wrong link or button”, illustrating it with a shopper moved into confirming “a large order they intended to cancel”, and sets the target at “0.1 or less” at the 75th percentile of page loads. That is not a design opinion, and it does not belong in an experiment queue.
Filtering findings deserve similar caution, because some are configuration limits rather than design failures. Shopify’s filter documentation caps how many filters a store may define and notes that “a search that produces more than 100,000 results doesn’t display filters”. A store near either boundary needs a merchandising decision, and an audit recommending a redesign instead has misdiagnosed the cause.
Where UX Audits Fail to Produce Revenue
The most frequent failure is volume. An audit returning 120 findings gives the recipient a research problem rather than a plan, and the practical result is that the team implements whichever items are cheapest rather than whichever items matter. A report of 15 findings, each carrying an estimated value, an implementation cost and an owner, will usually produce more change than one of 120 sorted by severity.
A second failure is auditing the interface without auditing the constraints. Recommendations assuming an unlimited budget, a freely rewritable theme, or a checkout permitting changes the platform disallows will stall at the first technical review. A third is the absence of a measurement plan agreed before anything ships: if nobody defines which metric each fix should move, and over what period, the effect becomes unprovable within weeks, once seasonality and campaign changes have moved as well.
The most damaging commercially is confusing a usability problem with a proposition problem. If shoppers understand the offer perfectly and decline it, interface work is unlikely to change the outcome. Polished sites can still fail on positioning, range or pricing, the pattern behind our work on the Profcentrs.lv rebuild. The same trap appears on product pages, where losses often mix usability, merchandising and expectation setting, as covered in our analysis of why product pages receive traffic but fail to convert.
Who Should Run the Ecommerce UX Audit
There is no single correct answer, and the honest comparison depends on what the store already has rather than on what a supplier prefers to sell. An internal review may be enough when the team includes someone with research experience, the problems sit in one or two known areas, and a developer has capacity to implement findings. Internal reviewers also hold context an outsider has to be taught. The recurring difficulty is that internal reviews rarely question decisions the team made itself, and they compete with operational work that almost always feels more urgent.
An independent researcher often suits a store needing one specific question answered well, such as whether a proposed navigation change will help. The work tends to be sharper and cheaper than a full engagement, but implementation is the limitation: a researcher who does not build depends on someone else prioritising the output, which is where a large share of audit recommendations quietly expire.
An external ecommerce team becomes the more practical choice when findings must be implemented as well as identified, when the store spans several markets or templates, or when previous audits produced recommendations nobody executed. The trade-off is that an external team starts without commercial context, which has to be transferred deliberately. Where the underlying question is about ownership rather than a single study, our comparison of a CRO specialist and an external ecommerce team covers that decision. Before committing to any of the three, it is worth asking:
- Which evidence methods will be used, and what will each one establish?
- How will findings be prioritised, and will the assumptions behind the ranking be open to challenge?
- Will recommendations be checked against what the platform, theme and installed apps actually permit?
- In what format is the output delivered, and can a developer act on it without a translation step?
- Which metrics will judge the work afterwards, and who agrees them before implementation starts?
Internal expert input required: add the number of evidence methods and the typical finding count in WD Market’s own usability audit deliverable, so the article can state the standard rather than describe it generically.
Choosing the Depth That Matches Your Store
The decision is not whether to examine usability but how deeply, and the answer follows from what the store can act on. A smaller catalogue with one template and an available developer may need little more than expert review combined with behavioural data. A multi-market store with several templates and a queue of unimplemented recommendations needs findings that arrive already attached to funnel steps, expressed in margin and ranked against build cost.
What separates an ecommerce UX audit that changes revenue from one that changes nothing is not the quality of the observations. It is whether each observation carries a population, a funnel position and a recoverable share, whether the recommendation respects what the platform permits, and whether a measurement plan was agreed before the first fix shipped.
From Usability Findings to Measurable Improvement
If your store has usability problems that reporting has not been able to isolate, the useful next step is a scoped review rather than a general redesign. Our CRO and growth support team runs ecommerce UX audits that combine expert review, behavioural data and funnel analysis, and return a prioritised set of findings with an estimated value, an implementation cost and a measurement plan for each one.
To discuss what a review would cover for your store, and which of the three audit types is the right purchase, get in touch with the team. Shorter observations from this work are published on WD Market’s LinkedIn page for anyone who would rather follow the thinking first.
Frequently Asked Questions
How long does an ecommerce UX audit usually take?
Duration depends far more on the number of distinct templates and markets than on catalogue size. A single-market store with a handful of page types can often be reviewed in two to three weeks, including time to gather enough behavioural sessions to be representative. Stores running several markets, or separate B2B and retail experiences, take longer because each variant has to be examined on its own terms. Any timeline quoted before someone has seen the store is provisional.
Do we need session recording tools installed before it starts?
It helps considerably, because behavioural tools need time to accumulate sessions before the patterns in them are trustworthy. Installing a recording and heatmap tool several weeks ahead means the review can begin with real data rather than waiting for it. If nothing is installed, the work can still proceed on expert review and existing analytics, though the behavioural component will be thinner or the timeline longer.
Is a usability review worth it for a store with low traffic?
Often yes, and sometimes more so than for a high-traffic store, because low volume rules out the alternative. A store that cannot run experiments at a useful sample size has to rely on expert judgement and qualitative evidence to decide what to change. The caveat is that revenue estimates are less reliable at low volumes, so prioritisation should lean on the severity and certainty of each finding rather than projected values.
How does this differ from a CRO audit?
The difference is the starting point. A CRO audit begins with commercial performance and investigates whatever is depressing it, which may include traffic quality, merchandising, pricing or the offer itself. A usability review begins with the interface and asks where it obstructs a shopper who already wants to buy. The two overlap in the middle, and stores frequently need both, but buying the narrower one to answer a broader question is a reliable way to be disappointed.
How many reviewers should look at the interface?
More than one, and fewer than you might expect. Published guidance on heuristic evaluation puts the lower bound at two, so one person’s preferences do not shape the whole output, and warns that beyond roughly ten the overlapping findings become difficult to reconcile. Two or three experienced reviewers working independently and then comparing notes tends to balance coverage against a report someone can act on.
Should the mobile experience be covered separately?
Yes, and treating mobile as a narrower version of the desktop review tends to hide the problems worth finding. Touch targets, keyboard behaviour on form fields, sticky elements consuming vertical space and layout movement during load all behave differently on a handset. Splitting the funnel by device before the review starts also gives a factual basis for deciding how much effort belongs on mobile, rather than dividing attention by assumption.