Two numbers move in opposite directions the moment a store starts pushing order size, and only one of them usually reaches the report. Suppose a new shipping threshold lifts the typical basket by 8% and quietly costs 5% of orders. Revenue lands close to where it started, the business now absorbs delivery on baskets it used to charge for, and the dashboard still shows a win.
That arithmetic is the whole subject. Levers that raise order value durably tend to help a shopper finish a decision they were already making, while the ones that damage conversion usually interrupt a decision that had already been reached. Useful average order value optimization is therefore a measurement problem before it is a tactics problem, and the number that settles it is margin per session rather than order size.
Why Order Value and Conversion Pull Against Each Other
Every mechanic that raises order size does one of two things to the purchase decision. It either inserts a step between intent and payment, or it changes the terms on which the order can be completed. Cross-sell modules, cart upsells and bundle selectors belong to the first group. Shipping thresholds, order minimums and volume breaks belong to the second, because they attach a condition to the price the shopper believed they were paying.
Both groups can pay for themselves. Both can also fail in ways that stay invisible when order value is the only figure on the report. An offer placed at the moment of commitment competes with the action it interrupts. A threshold set above what most baskets reach turns a share of small orders into no orders at all, and the orders that survive are not always worth what the excluded ones would have been.
Shopify’s own commentary on the metric is blunt about the risk: “A higher AOV doesn’t always equal higher profits”. Margin is the second place the arithmetic hides. A bundle sold at 10% off raises the recorded order value and reduces gross profit on every basket a customer would have assembled anyway. A shipping threshold moves delivery cost from the customer to the business on every qualifying order, including the orders that were always going to qualify.
This is why the guardrail matters more than the headline. Gross margin per session absorbs all three movements at once: how many visitors buy, how much they spend, and what the incentive cost to deliver. Average order value optimization judged on order size alone can report a gain while contribution stays flat, which means money has moved between columns rather than been created.
Average Order Value Optimization Starts With the Distribution, Not the Average
The metric itself is simple. BigCommerce defines it as “a standard ecommerce metric that reflects the average total of every order placed over a defined period of time”, calculated by dividing revenue by the total number of orders, and Salesforce states the same relationship as total revenue divided by the number of orders. The simplicity is also the difficulty. An average compresses a distribution that usually has a long right tail, and the tail is where a small number of unusual orders sit.
A store whose mean sits well above its most common order value is a different commercial problem from one whose orders cluster tightly around the mean. In the first, a threshold placed near the average is placed far above where most baskets land, and the mechanic looks inert in the data because few shoppers are close enough to react to it. In the second, the same threshold may move a large share of baskets by a small amount each. The tactic did not change. The distribution did.
Before choosing a lever, it is worth pulling a small number of views that a standard ecommerce dashboard rarely shows by default:
- Order value in bands rather than as a single average, so that the shape of the distribution and the position of the most common order are both visible.
- Items per order alongside value per order, which separates stores that need a second product from stores that need a more expensive one.
- The same two figures split by new and returning customers, because returning buyers often carry a higher baseline and can flatter a threshold test that did nothing for acquisition.
- Order value by device, since any mechanic that depends on a visible cart summary may behave differently where screen space is limited.
- Gross margin per order rather than revenue per order, so that a lever which raises value by discounting can be recognised for what it is.
These views usually change which lever looks sensible. A catalogue of low-priced consumables with a high items-per-order figure has room for volume mechanics and little need for upsells. A store selling one considered item per visit rarely benefits from a threshold at all, and its order value moves through product tier and attachment instead. Average order value optimization that skips this step tends to default to whatever the platform makes easiest to configure, which is not the same as whatever fits the catalogue.
The Free Shipping Threshold and How It Is Usually Set Wrong
Thresholds are the most widely used order value mechanic because they are the easiest to configure and the easiest to explain to a shopper. On Shopify, a free shipping discount can be limited by a minimum purchase amount, which the documentation describes as a setting that “requires customers to spend a minimum amount to qualify for the discount”, or by a minimum quantity of items, which “requires customers to order a minimum number of products to qualify”.
A merchant can also exclude shipping rates over a chosen amount, an option the same page notes “applies to shipping rates only, and is unrelated to order amounts”. That second control matters for bulky or remote-delivery items, where an uncapped promise can turn a profitable order into a loss.
Shopify’s commercial guidance suggests setting the threshold around 30% above current average order value as a starting point, and cautions that setting it too high risks abandoned carts. The starting point is reasonable, but it inherits the problem described above. A figure 30% above a mean that already sits in the tail is not 30% above where shoppers actually are. Where the distribution is skewed, the most common order value is the more honest anchor to work from.
Whether a threshold earns its cost then depends on three things that have little to do with the number itself. The first is whether shoppers can see the gap: a threshold shown only on a policy page is a rule rather than an incentive, and it has to state the amount remaining rather than the condition. The second is whether the catalogue offers something worth adding at roughly the size of that gap, because where the cheapest sensible addition costs more than the shortfall, the only way to comply is to buy something unwanted.
The third is the delivery cost the business now absorbs on orders that would have cleared the threshold regardless, which is a permanent transfer rather than a promotional one.
Thresholds therefore suit consumable and accessory ranges with frequent multi-item baskets and moderate delivery costs. They rarely suit high-ticket single-item purchases, where the shopper has no realistic way to add value and the offer reads as an irrelevance.
Key takeaway: a shipping threshold is a pricing decision rather than a merchandising one. It changes what every qualifying order is worth to the business, including the orders that would have qualified anyway, so it should be judged on gross margin per session and not on the share of baskets that reach it.
Bundles, Multipacks and Volume Breaks: Where the Discount Goes
Shopify describes a bundle as “a set of two or more related products, commonly offered at a discount” and notes that a bundles app is required to create one, with support across the Online Store, Shop and Shopify POS sales channels. The appeal is that a bundle raises order value and items per order at the same time, and can move slower inventory alongside a product that sells on its own.
The commercial question is narrower than the tactic suggests. A bundle pays when it removes work the customer would otherwise have to do, or when it introduces a product the customer would not have found unaided. It costs money when it discounts a combination the customer was going to assemble anyway, because the discount then applies to revenue that already existed. The difference is measurable. Compare the attach rate of the bundled products before the bundle existed with the bundle’s take-up afterwards, and treat the overlap as cannibalised margin rather than incremental revenue.
Volume breaks follow the same logic in a different shape. They work where consumption is predictable, which is why they are more reliable in consumables, B2B and trade contexts than in considered purchases. Where consumption is unpredictable, a volume break often pulls forward demand that would have arrived later at full price, improving the current period’s order value and weakening the next one.
Bundling can also be built into the storefront rather than bolted on. WD Market’s rebuild of the Evelatus store included bundle deals as a deliberate order value mechanic, and the case study describes the work as integrating bundle deals that allowed the retailer to package related products at better prices, increasing average order value at the point of purchase. Internal expert input required: add the verified before-and-after attach rate for the bundled ranges if the client has approved publication of that figure.
Recommendations, Cross-Sells and Upsells: Placement Decides the Cost
The distinction between the two mechanics is worth keeping precise, because they carry different risks. Omnisend puts it plainly: “Cross-selling recommends complementary products that enhance a purchase. Upselling encourages customers to upgrade to a higher-priced version of the same product”. A cross-sell adds a line to the order and leaves the original decision intact. An upsell asks the shopper to reopen a decision already made, which is why it usually belongs on the product page rather than in the cart. The practical patterns for both are covered in more detail in our guide to upselling and cross-selling in ecommerce.
Shopify’s own recommendation system reflects the same split. Related products are described as “products that are similar to a selected product”, and complementary products as “products often bought in addition to a selected product”. The practical difference is who does the work. Complementary products are chosen manually, up to ten per product, while related products are generated automatically from purchase history, product descriptions and related collections.
Those automatic strategies have conditions attached, and they explain a good deal of underperformance. The documentation states that the purchase-history strategy needs previous sales in order to assess buying behaviour, and that the product-description strategy is available only to merchants with an English storefront. A newly launched product, a new market or a non-English storefront will therefore receive weaker automatic recommendations, and the remedy is usually manual curation on the products that matter most rather than a different app.
Volume of offers is its own risk. Omnisend’s guidance is to keep the count low: “Too many cross-sell offers on a page can overwhelm customers. Less is more; focus on one to three high-value recommendations.” That reads better as a conversion instruction than as a design preference, since each additional offer adds a decision to a page whose job is to produce exactly one.
Placement follows the same reasoning. The product page carries the least risk because the shopper is still comparing, the cart carries more because the decision has been made, and the checkout carries the most. On Shopify the checkout is also the most restricted, since checkout UI extensions rendering on the information, shipping and payment steps “are available only to stores on a Shopify Plus plan”.
Post-Purchase Offers and the Population They Actually Reach
Post-purchase offers are the one order value mechanic that carries no conversion risk by construction. Shopify’s developer documentation defines a post-purchase offer as an additional sales opportunity displayed to customers immediately after they complete checkout, rendered “after the order is confirmed, but before the Thank you page”. Nothing the shopper does with the offer can undo an order that already exists, which is why the mechanic sits so comfortably alongside the wider post-purchase page.
The constraints are where forecasts usually go wrong, and they are specific. The same documentation states that a customer “can accept a maximum of three post-purchase offers for each checkout”, that “orders need to be $0.50 or more to qualify for post-purchase offers”, that “orders need to be placed through the Online Store sales channel to qualify for post-purchase upsells”, and that “only one app can be selected for post-purchase product offers”.
The most consequential restriction concerns payment. The post-purchase page is not surfaced when the customer “chooses to check out with an installment service or a wallet service (such as Klarna, Affirm, AfterPay, Apple Pay, Amazon Pay, or Google Pay)”, nor when “the initial purchase was made with a gift card or any payment method other than a credit card”. For a mobile-heavy store this changes the size of the prize considerably. Where a large share of orders complete through accelerated wallets, the eligible population is a fraction of total orders, and a projection built on total order volume will overstate the return by roughly that share.
The check takes minutes and is worth doing before the app, the integration and the merchandising effort are committed: split recent orders by payment method and size the eligible population directly. Where that population is large enough, the mechanic is unusually clean. Omnisend notes that a one-click add-to-order feature “lets customers add items without re-entering payment details”, which is why the offer converts at all. It succeeds on convenience rather than persuasion, so it suits consumables, accessories and warranty-style attachments far better than a second considered purchase.
Matching the Lever to the Store
No single mechanic suits every catalogue, and several of them interact badly when they run together. The comparison below sets out where each one tends to fit, what it risks and what constrains it.
| Lever | Best suited to | Conversion risk | Effect on margin | Main constraint |
|---|---|---|---|---|
| Free shipping threshold | Consumable and accessory ranges with frequent multi-item baskets | Moderate: may exclude small baskets | Negative on every qualifying order | Needs a visible gap and an affordable item to fill it |
| Product bundles | Catalogues where the combination is obvious but assembly is tedious | Low | Negative where the combination was already being bought | Requires a bundles app and inventory discipline |
| Volume breaks | Predictable consumption, B2B and trade buyers | Low | Negative, and may pull demand forward | Weak fit for considered, infrequent purchases |
| Product page cross-sells | Stores with genuine complements and enough sales history | Low | Neutral to positive | Automatic recommendations depend on purchase history and storefront language |
| Cart and checkout upsells | Stores with a clear tier or upgrade story | High at the checkout stage | Positive where no discount is attached | Pre-purchase checkout extensions require Shopify Plus |
| Post-purchase offers | Repeat-friendly consumables and accessories | None by construction | Positive | Excludes wallet, installment and gift card orders |
| Loyalty and tier incentives | Established repeat-purchase bases | Low | Deferred rather than immediate cost | Slow to show an effect and dependent on retention data |
Two patterns are worth drawing out of that comparison. The mechanics with the lowest conversion risk operate outside the decision itself, whether on a product page the shopper is still evaluating or after the order already exists. The mechanics with the strongest margin profile are the ones that do not rely on a discount, which quietly rules out several of the tactics that are quickest to launch.
Salesforce’s list of order value tactics places loyalty tiers tied to order size in that second category for the same reason: the incentive is deferred and conditional rather than deducted at the point of sale. Most disagreements about average order value optimization are really disagreements about which of these two properties a store is willing to trade away.
A Sequence for Testing Order Value Changes Without Losing Ground
- Establish the baseline on the right unit. Record revenue per session and gross margin per session alongside conversion rate and order value, split by device and by new against returning customers. A change that lifts order value while reducing margin per session is a loss that reports as a win.
- Read the distribution before choosing a mechanic. The position of the most common order relative to the mean determines whether a threshold has anything to work with, and items per order indicates whether the store needs a second product or a better one.
- Fix the guardrails in writing before launch. Decide in advance what fall in conversion rate would make the change unacceptable even if order value rises, and name the person with the authority to stop it. Guardrails agreed afterwards tend to be negotiated rather than applied.
- Change one mechanic at a time. Launching a threshold and a cart upsell in the same week produces a combined result that cannot be attributed, and if the combined result is negative there is no way to know which half to remove.
- Run for a whole number of business cycles. Order value is seasonal at a weekly level in most catalogues, so a test stopped mid-cycle inherits whatever that partial week happened to contain. The wider discipline is covered in our guide to A/B testing for ecommerce.
- Judge on contribution. Compare gross margin per session across variants, subtract the incentive cost, and only then look at what happened to the average order.
- Re-measure by segment after rollout. A mechanic that is neutral overall may be clearly positive for returning customers and negative for first-time buyers, which is an argument for targeting it rather than for abandoning it.
Mistakes That Cost More Than the Lever Adds
- Setting a threshold from the mean in a store with a skewed distribution. The target then sits beyond the reach of most baskets, and a misconfigured mechanic is written off as an ineffective one.
- Discounting a bundle that customers already buy together, which converts existing margin into a promotion while simultaneously producing the rise in order value that appears to justify it.
- Stacking offers through the cart and checkout, where each additional decision competes with the single action those pages exist to produce.
- Forecasting post-purchase revenue from total order volume without checking the payment mix first, since wallet and installment orders are excluded from the offer entirely.
- Reading average order value on its own, which hides both the orders that stopped happening and the margin given away to produce the increase.
- Treating returns as unaffected. Larger baskets assembled to clear a threshold may be returned at a higher rate than baskets a shopper wanted outright, and that cost usually lands in a different report from the one showing the improvement.
Where the Order Value Decision Actually Sits
The choice is rarely between tactics. It is between mechanics that serve a purchase the shopper was already assembling and mechanics that interrupt one they had already settled. The first group tends to hold conversion and improve contribution. The second can raise the recorded average while leaving the business no better off, and occasionally worse off once absorbed delivery, discounts and returns are counted.
The practical starting point is the distribution rather than the average, and the deciding number is gross margin per session rather than order size. With those two in place, most of the levers described here can be tested honestly, and the ones that do not fit the catalogue tend to disqualify themselves quickly. Average order value optimization run this way is slower to produce a headline, and considerably harder to reverse into a loss.
From Order Value Targets to Measurable Contribution
Where order value is the constraint on how much a business can afford to spend acquiring a customer, the first useful deliverable in average order value optimization is not a tactic. It is a measurement: where the order distribution actually sits, what each candidate mechanic would cost in margin, and which of them the catalogue, the platform and the payment mix can realistically support.
WD Market’s CRO and growth support covers that scope directly: an order value and margin baseline by segment and device, an assessment of which levers the platform and catalogue allow, and a prioritised test plan with a guardrail metric and a named owner against each item. To discuss a specific store and its numbers, use the contact page. Shorter analysis of the same problems is published regularly on the WD Market LinkedIn page.
Frequently Asked Questions
How is average order value calculated, and over what period?
Divide total revenue by the number of orders placed in the period being measured. The period should be long enough to cover a full purchasing cycle for the catalogue, which for most stores means at least a month and often a quarter. Shorter windows are dominated by promotions and by the handful of unusually large orders that every store receives. It is also worth agreeing whether the figure runs on gross revenue or on revenue after discounts and returns, because the two can differ enough to change a decision.
What is a sensible free shipping threshold to start from?
A common starting point is around 30% above the current average, which is what Shopify’s guidance suggests, with a warning that a threshold set too high risks abandoned carts. In stores where a few large orders pull the average upward, anchoring on the most frequently occurring order value gives a more realistic target. Whichever anchor is used, the number should be tested rather than assumed, and the result read against margin per session rather than against the proportion of baskets that reach it.
Can raising order value reduce profit?
Yes, and it happens more often than the reporting suggests. Discount-based mechanics such as bundles and volume breaks reduce gross profit on baskets customers were going to buy regardless. Shipping thresholds transfer delivery cost to the business on every qualifying order. Larger assembled baskets may also carry higher return rates. None of those effects appear in the order value figure itself, which is why gross margin per session is the safer number to judge any of these changes on.
Do post-purchase upsells work for every Shopify store?
No. Shopify’s documentation excludes several common situations: orders paid with a wallet or installment service such as Apple Pay, Google Pay or Klarna, orders paid by gift card or any non-credit-card method, orders below $0.50, and orders placed outside the Online Store sales channel. Only one app can serve post-purchase offers, and a customer can accept at most three per checkout. A store with a high share of wallet payments should size the eligible order population before investing in the mechanic.
Should cross-sells appear on the product page or in the cart?
The product page is generally the safer position, because the shopper is still evaluating and an additional suggestion competes with nothing. Cart placement carries more risk, since the decision has already been made and the page exists to move it forward. If cart cross-sells are used, keeping the count to a small number of genuinely complementary items reflects the guidance from vendors in this space, and their effect on completion rate should be watched as closely as their attach rate.
How long should an order value test run before a decision?
Long enough to cover whole business cycles rather than to reach a satisfying number. Order value varies by day of week in most catalogues, so a test ended mid-week carries whatever that partial period contained. Set the duration and the stopping rule before launch, including the fall in conversion rate that would end the test regardless of what happened to order size. Extending a test until it produces the desired answer is the most common way these decisions go wrong.
Where should average order value optimization start in a store that has never tried it?
Start with the order value distribution and the margin figure, not with a tactic. Knowing where the most common order sits relative to the mean, how many items a typical order contains, and what gross margin each order carries usually eliminates most of the candidate levers immediately. A store with single-item orders and thin margins has different options from one with multi-item consumable baskets. Choosing the mechanic first and looking for supporting data afterwards is how misconfigured thresholds and margin-negative bundles get shipped.