Revisiting the use of statistical ranges in transfer pricing

International Tax Review is part of Legal Benchmarking Limited, 1-2 Paris Garden, London, SE1 8ND

Copyright © Legal Benchmarking Limited and its affiliated companies 2026

Accessibility | Terms of Use | Privacy Policy | Modern Slavery Statement


Revisiting the use of statistical ranges in transfer pricing

Sponsored by

Sponsored_Firms_deloitte.png
Graph
autsawin uttisin/Shutterstock

Eric Linge, Vrajesh Dutia, and Ewan Kemsley of Deloitte challenge the routine use of interquartile ranges in transfer pricing, arguing that broader statistical approaches can produce more robust comparability analyses

The number of reported transfer pricing cases has been on the rise for decades, with a sharp increase starting around 2014 (based on reported transfer pricing cases at tpcases.com). This is indicative of the rise in transfer pricing disputes across all fora, and why it is consistently a top issue facing tax heads in multinational enterprises.

Comparability analysis and transfer pricing benchmarks are amongst the most disputed issues. In almost every case there are multiple comparable observations, which must be summarised with some sort of statistical analysis, often maximums, minimums, and quartiles. The use of ranges is frequently at the centre of the dispute: taxpayers claiming they are in the range and tax authorities claiming they are out of the range. This article analyses appropriate use of ranges for summarising comparables data.

Paragraph 3.57 of the OECD Transfer Pricing Guidelines for Multinational Enterprises and Tax Administrations (the Guidelines) refers to the use of statistical measures of “central tendency” when a “sizeable” number of observations exists. However, the Guidelines do not define sizeable. They suggest quartiles and averages to summarise data with a central tendency but do not go into the hundreds of years of statistical science that have given these terms meaning.

Where comparables are considered to show similarity to a controlled transaction, the Guidelines offer that any point in the range may be appropriate. Some recent court judgments have shown sympathy for the full range of comparable results, as opposed to the commonly used interquartile, including recent Supreme Court cases in Sweden and Denmark relating to an alcoholic beverage producer and an electronic parts supplier, respectively.

In practice, however, the interquartile range (IQR) has become the default reference point in many benchmarking analyses, with the median often treated as the desirable point of reference. Other jurisdictions apply alternative statistical ranges, such as India’s 35th to 65th percentile range and Malaysia’s 37.5th to 62.5th percentile range.

This article considers the view that the interquartile should not automatically be concluded as the right range. A statistical analysis that is appropriate for the underlying data, and that better summarises the data-generating process, likely makes a more robust transfer pricing analysis.

The IQR

The IQR is a transfer pricing practitioner’s favoured statistical method. It eliminates outliers by reducing the set of comparable prices or margins to only the middle 50% of observations. It is indeed a statistical method that describes the centre of data, even if it does not describe whether the data has a central tendency.

The de facto explanation as to why the IQR is appropriate is because it corrects for “comparability defects” between the controlled transaction and the comparables. Implicitly, the conclusion must be that the lowest quarter of results and the top quarter of results are less likely to be reliable. In the recent case referenced above, the Danish Supreme Court raised this point. The court ruled the tax authority did not meet its burden of proof to show that limiting the comparables to just the central 50% of observations (the IQR) corrected for the supposed comparability defects.

The Guidelines describe the IQR as one route for summarising comparable pricing data and never say it is the only route for making a comparability analysis. Paragraph 3.57 of the Guidelines implies the use of an analytical hierarchy for summarising quantitative data:

  • If there is only one comparable, its price (or margin) is the arm’s-length price for the controlled transaction;

  • If there are multiple observations, and each is equivalently comparable, there is an arm’s-length range of appropriate prices; and

  • If the observation count is “sizeable”, and comparables have “defects”, then statistical methods taking account of the data’s central tendency may be used.

Despite its popularity in practice, the Guidelines do not mandate the IQR as the only acceptable statistical tool. The Guidelines specifically name medians and averages as types of appropriate statistical tools. While not stated explicitly in the Guidelines, other statistical methods also account for central tendency of data, such as full range (between minimum and maximum), confidence intervals around statistical estimates (e.g., average or median), or conditional statistical methods such as regression. Noteworthy is that the Guidelines say that if the best comparables are already identified, then the full range should suffice as a route to summarise comparables data.

Statistical analysis for transfer pricing

Transfer pricing analysis is inherently concerned with a counterfactual. The arm’s-length price for a related-party transaction can never be observed directly; it must be estimated based on what independent parties would have agreed in comparable circumstances. Statistical science calls this ‘prediction’: finding the pattern amongst all relevant, available data to predict the price or margin that would result from the data-generating process for a set of features similar to those of the controlled parties.

From this perspective, benchmarking is as much an exercise in statistical inference as it is a search exercise, and statistical ranges can be considered as a summary of the information available in the data. The reliability of this summary depends on the quality and size of the sample, as well as the distribution of the data points within it.

A real-world example will showcase this perspective. The analysis draws on a management services benchmarking set with 214 companies covering Europe, APAC, and North America. Table 1 shows summary statistics for this set, using net cost plus (NCP) mark-up as the primary profit level indicator. This type of comparable set, and even the comparables identified, are very common in practice.

Table 1: Summary statistics for management services set

Summary statistic

2024, three-year trailing average NCP

Count

214

Average (mean)

8.5%

Minimum

-0.2%

Lower quartile

2.4%

Median (middle quartile)

5.1%

Upper quartile

11.6%

Maximum

66.0%


At first glance, the median or IQR may appear to provide a clear summary of the comparable set. However, Figure 1 shows that the underlying distribution is right-skewed, meaning a focus on the centre of the range may obscure why certain companies earn higher margins. This lognormal type of distribution is common in corporate profitability data, where downside margins are naturally constrained while upside performance can compound. Looking at the data, evenly spaced as in the histogram below, you might think that the centre should actually be further to the right than predicted by the mean or median.

Figure 1

Histogram of NCP observations.png

By reviewing websites of the companies, two distinct subsets were identified from the broader benchmarking set: back-office services and strategic management services.

The density plots in the figure below show the spread of the data points in each of these functional subsets identified after the website review. The top panel shows all of 214 services companies in the comparable set, and the second and third panels show those 214 companies split into two subsets. Discernibly, the back-office services subset comprises a wider range, while the strategic management services subset has wider quartiles. Like Figure 1, the subsets data has a visible rightward skew, meaning the question remains as to where the central tendency lies. Back-office services have a higher maximum but greater density at the low end, with strategic management services having more weight between approximately 5 to 30%.

Figure 2

Smoothed density of NCP observations.png

Reviewing the websites of potentially comparable companies and narrowing down counts of comparable companies based on business descriptions is common transfer pricing practice. Having done this for services companies, it can be observed that there are some companies performing comparable back-office services that can reasonably achieve margins much higher than the median. By splitting the services companies set into two subsets, each subset now has a weaker claim of containing a “sizeable” number of companies, and fewer high-margin companies are seen in each of the subsets compared with the larger services set.

This discussion shows the nuance on comparability that is missed when the data is summarised only with quartiles. Only when we look at the full distribution do we see that most back-office and strategic management services companies are achieving returns below 10%, but then some can achieve much higher results, leading to the skewed distributions. The reasons are more complex than just their categorisation as back office. There is a full data-generating process behind these NCP results. The question then becomes what is different about these high-margin versus low-margin comparables, because it cannot be fully explained by being back-office versus strategic. This is explored further in the conditional statistics section below.

Ultimately, the inclusion of more comparables provides greater information for understanding the data-generating process that produces margin or price results from comparable transactions. The Guidelines’ “sizeable” criterion is making the point that more information will allow for better prediction.

The figure below presents an incremental reduction in sample sizes using an algorithm that attempts to keep the same distribution in each draw, starting from the full sample of 214 and ending at four observations. As sample sizes became smaller, the upper part of the distribution is where information is lost at the fastest rate. By the time you are down to four observations, the data-generating process simply does not allow for a high margin comparable.

Figure 3

Histogram of NCP at decreasing observation count.png

Quartiles attempt to estimate statistics for the full and complete data-generating process. However, with increasingly less information for estimation, the estimated statistic becomes increasingly uncertain. Confidence intervals, as shown in the figure below, quantify this uncertainty, showing statistical estimates if you resample hundreds or thousands of times and how the estimate tends to bounce around depending on what information is available to calculate it. As the observation count decreases, confidence intervals get wider, and as increasingly more information is lost, especially where there is skew, like near the upper quartile, the estimate becomes less stable. After enough observations are removed, we can no longer even observe the distribution’s characteristic right skew: all that is left are some observations in the densest part of the distribution (towards the bottom, for this high-skewed data). Estimates become increasingly sensitive to whether the top-end outliers are included, reducing the robustness of the quartiles.

As we observe what happens to our quartile estimates as sample sizes reduce, the obvious question is: when we use small sets of comparables, is the median really a ‘true’ median for companies performing that function? Or is the result partly from the data-generating process, partly from statistical error, or partly from something else (an unobserved variable, perhaps)?

Figure 4: Confidence intervals at decreasing observation counts

Confidence intervals at decreasing observation counts.png

We can make our transfer pricing analyses more reliable by keeping more observations in our comparable sets. To do this, we may need to be able to explain why some comparables have higher results (and why some lower) and show why those differences matter or do not matter for our comparability analysis. In comparisons between companies, a tested party may be comparable in function to every company in the set of comparable companies, but it may also have some features that make it more like low-margin (or high-margin) companies. This would be a conditional mean, or regression approach, such as the OECD used in Amount B of Pillar One. Indeed, such conditional statistics will help to strengthen the narrative.

Conditional statistics

The IQR is limited to summarising outcomes conditional only on the observations being included in the set. It asks where the tested party’s result sits relative to the distribution of observed margins, but it does not consider whether the tested party is similar to all parts of that distribution.

Conditional statistics can help address this issue by estimating ranges or expected outcomes conditional on features relevant to the comparison. For example, regression analysis can be used to determine an average of a dependent variable conditional on one or more independent variables. The resulting regression function can be used to predict an average outcome for a given feature set.

The below figure demonstrates this. The data is again drawn from the same comparables, but by adding more data from the data-generating process – i.e., working capital intensity (ratio of working capital to sales) – it can be understood what company features are associated with the different NCP seen in the density plots.

The scatter plot indicates that profitability is correlated with working capital intensity: companies with more working capital intensity tend to have higher NCP. Also, the strength of the correlation (i.e., the line slope) is different between the two functional subsets. The line, in fact, predicts what would be an appropriate NCP for a tested party, based on its features: a tested party with a 0.8 working capital intensity level, and functions classed as more back-office, would on average be expected to have about a 9% NCP. Interestingly, for companies with lower working capital intensity, the regression function, following the same logic, would predict approximately the same NCP, despite the different NCP medians of these comparable sets. Said another way: draw a vertical line, intersecting the horizontal axis where working capital intensity equals 0, and the vertical line intersects the regression lines at essentially the same NCP values. Whether the company is back-office or strategic management, when working capital intensity is zero, the regression lines are almost intersecting, meaning that both regressions would predict NCP of about 5%.

Figure 5

NCP conditional on working capital intensity by group.png

For transfer pricing disputes, this may present a stronger argument than asserting that the median or IQR is the best metric by default. It allows the taxpayer to move from description to prediction. Rather than assuming that all comparable companies should be summarised as if they have the same levels of intensity or functional profile, the analysis asks how profitability changes conditional on those features. This also shows the value of including more information in the analysis and letting the statistics do the work of pulling out the most relevant elements while still taking account of all potentially valuable pieces of information.

Key conclusions

When defending transfer pricing benchmarks, the IQR remains useful but alone may not be sufficient. The goal is to narrate why a transfer price is appropriate, with reference to data from market benchmarks. This narration becomes stronger as more information is included, and statistics can show what is most important. With more information, statistical analysis can be used to describe the data-generating process and focus the prediction on the most useful pieces with higher confidence. With their recommendation of statistical tools, this is what the Guidelines always intended. With too little information, statistics become unreliable, and with more information, statistics can pull signal from the noise.

Deloitte refers to one or more of Deloitte Touche Tohmatsu Limited (DTTL), its global network of member firms, and their related entities (collectively, the “Deloitte organization”). DTTL (also referred to as “Deloitte Global”) and each of its member firms and related entities are legally separate and independent entities, which cannot obligate or bind each other in respect of third parties. DTTL and each DTTL member firm and related entity is liable only for its own acts and omissions, and not those of each other. DTTL does not provide services to clients. Please see www.deloitte.com/about to learn more.  

Deloitte provides leading professional services to nearly 90% of the Fortune Global 500® and thousands of private companies. Our people deliver measurable and lasting results that help reinforce public trust in capital markets and enable clients to transform and thrive. Building on its 180+ year history, Deloitte spans more than 150 countries and territories. Learn how Deloitte’s over 470,000 people worldwide work together every day to make an impact that matters at www.deloitte.com .

This communication contains general information only, and none of Deloitte Touche Tohmatsu Limited (DTTL), its global network of member firms or their related entities (collectively, the “Deloitte organization”) is, by means of this communication, rendering professional advice or services. Before making any decision or taking any action that may affect your finances or your business, you should consult a qualified professional adviser. No representations, warranties or undertakings (express or implied) are given as to the accuracy or completeness of the information in this communication, and none of DTTL, its member firms, related entities, employees or agents shall be liable or responsible for any loss or damage whatsoever arising directly or indirectly in connection with any person relying on this communication. DTTL and each of its member firms, and their related entities, are legally separate and independent entities.  

© 2026. For information, contact Deloitte Global.  

more across site & shared bottom lb ros

More from across our site

One of the two appointments is EY’s Gordon McIntosh, who becomes the big four firm’s second senior tax departure in September
Balson's move from a Tier 1 practice to a Tier 3 competitor looks counterintuitive. The market data suggests it is anything but
Awards
It was another banner year for Deloitte, which picked up more awards than any other firm at a gala ceremony held at The Londoner in Leicester Square
The big four firm has been embroiled in a scandal over partners’ misuse of confidential board papers to pitch for and win corporate audits for Westpac and Dexus
Drawing on lessons from the PepsiCo case, tax lawyer Paul McNab explains why the ATO's latest royalty guidance should concern multinationals well beyond the technology sector
As pillar two exposes the limits of fragmented tax processes, organisations are rethinking their operating models to create the trusted data foundations that AI demands
World Tax data shows Matt Donnelly is moving from a Tier 3 transactional tax practice to a Tier 1 market leader, underlining Kirkland & Ellis’s pull at the top end of the market
Nexdigm's Maulik Doshi and infer360 co-founder Sunil Agarwal dig deeper into their partnership and discuss why the tax technology industry is consolidating
Advisers won’t be short of work in a world of increased valuation disputes, documentation requirements and behavioural responses from clients seeking to protect their wealth
Jaydeep Menon explains how Frazier & Deeter built a specialist practice which helps UK start-ups expand into the US and why private equity backing is accelerating its ambitions
Gift this article