Loading...
Loading...
A practical transfer pricing benchmarking workflow covering database selection, tested-party and PLI decisions, reproducible screening, accept/reject evidence, ranges, and study refreshes.
Borys Ulanenko
CEO of ArmsLength AI

Continue exploring
Browse the full resource library or contact us if you want recommendations for your specific use case.
A transfer pricing benchmarking study tests a controlled transaction against reliable uncontrolled evidence. For a one-sided profit method such as TNMM or the US Comparable Profits Method, the study usually identifies independent companies with sufficiently comparable activities and calculates a profit level indicator for the tested party and the accepted comparables.
The conclusion is only as sound as the route used to reach it. A reviewer should be able to see:
OECD Guidelines paragraphs 3.30-3.34 describe commercial databases as useful sources with real limitations. They should be used objectively, and a database-only search may need to be refined with annual reports, company websites, registries, and other public information. Quantity does not replace comparability.
For the calculation mechanics, see the PLI selection guide, working capital adjustments guide, and interquartile range guide.
A search is not the first step. Start with an accurately delineated transaction and a functional analysis of all parties. Record the following decisions in a short search-planning memo:
| Decision | Question to answer | Evidence to retain |
|---|---|---|
| Controlled transaction | What goods, services, financing, or rights are being tested? | Agreements, invoices, transaction schedules, interviews |
| Parties and period | Which entities and fiscal years are covered? | Legal-entity map, reporting calendar, ledger extracts |
| Method | Why is the selected method more reliable than the practical alternatives? | Method-selection analysis and data-availability review |
| Tested party | Which party can be tested most reliably and has the most reliable comparables? | Functional analysis and tested-party analysis |
| PLI | Which denominator has a sound economic relationship to the tested activity? | Segmented accounts and PLI calculation policy |
| Market | What geography and economic conditions matter? | Market analysis and rationale for local or regional scope |
| Data timing | Is the study being used to set prices, test outcomes, or both? | Pricing policy, true-up process, and filing calendar |
Under , the tested party is generally the party to which the method can be applied most reliably and for which the most reliable comparables can be found. It will often, but not automatically, be the party with the less complex functional profile.
Do not label an entity "routine" simply because the group policy gives it a fixed return. Confirm what its people actually do, which assets it uses, which risks it controls, and whether it makes any unique and valuable contribution.
An internal comparable is an uncontrolled transaction between one party to the controlled transaction and an independent party. OECD paragraphs 3.27-3.28 note that internal comparables can have a closer relationship to the tested transaction and may use the same accounting practices. They are not automatically reliable: volume, contractual terms, geography, product features, risk allocation, and other comparability factors still need to be tested.
Ask these questions before commissioning an external company search:
If a reliable internal comparable exists, an external database search may add little. Document the decision either way.
There is no universally best transfer pricing database. The right source depends on the type of evidence required, the market, the reporting period, and whether the data can be reviewed and reproduced.
| Evidence needed | Typical source category | Useful fields | Main limitations to test |
|---|---|---|---|
| Company profitability for TNMM/CPM | Company accounts and ownership database | Income statement, balance sheet, ownership, activity codes, business descriptions | Uneven private-company disclosure, company-level rather than transactional data, inconsistent accounting classifications |
| Comparable prices or terms | Internal transaction systems, public filings, contract and agreement databases | Price, volume, territory, term, rights, obligations, adjustment clauses | Redactions, incomplete context, and material contractual differences |
| Royalty evidence | Licence agreement databases and public securities filings | Royalty base and rate, licensed rights, exclusivity, geography, term | Selection bias, bundled rights, missing profit potential and negotiation context |
| Financial transaction evidence | Loan, bond, and market-data sources | Currency, tenor, seniority, security, credit quality, issue date, yield | Different credit risk, implicit support, market timing, covenants, and liquidity |
| Qualitative screening | Company websites, annual reports, registries, regulatory filings | Products, functions, business model, group membership, intangibles | Pages change, descriptions may be promotional, and historic evidence may disappear |
The commercial product name matters less than whether the retained evidence answers the comparability questions. Before choosing a source, test:
Record the database name, product or module, release or version, access date, and field definitions in the report. If more than one source is used, explain which source controls when the records conflict.
A search strategy translates the tested party's functional profile into observable criteria. Industry codes are useful for generating a population, but they are not a substitute for manual review.
| Filter | Purpose | Question for the file |
|---|---|---|
| Activity codes and keywords | Find candidates performing potentially similar activities | Why do these codes and terms represent the tested function? |
| Geography | Reflect relevant economic circumstances and available data | Why is the selected market sufficiently comparable? |
| Independence | Remove entities whose results may be affected by controlled dealings | Which ownership fields, thresholds, and manual checks were applied? |
| Operating status | Remove inactive or unsuitable legal entities | What status date was used? |
| Financial-data availability | Ensure the PLI can be calculated consistently | Which years and fields are mandatory, and how is missing data treated? |
| Size | Remove candidates whose scale creates a material comparability difference | What fact supports the chosen measure and threshold? |
| Loss or extreme-result screen | Flag candidates for further analysis | Is the rule investigative, or does local law require a particular treatment? |
Keep a count after each filter. A reviewer should be able to reconstruct the funnel from the initial universe to the manual-review population.
Amount B checkpoint: Before running an ordinary company search for a baseline marketing and distribution transaction, confirm whether the relevant jurisdiction and period apply the OECD's Amount B simplified and streamlined approach, and whether the transaction meets its scope conditions. Domestic implementation and options control. Where Amount B does not apply, use the ordinary method-selection and benchmarking analysis required by the relevant law.
Assume the tested party is a routine distributor of industrial components in a regional market:
| Stage | Candidates | What happened |
|---|---|---|
| Relevant activity codes and keywords | 1,842 | Initial population generated |
| Active entities in the selected geography | 1,106 | Dormant and out-of-scope locations removed |
| Documented independence screen | 284 | Controlled or insufficiently independent entities removed |
| Required financial data available | 96 | Candidates without the fields needed for the PLI removed |
| Quantitative review completed | 47 | Scale and other documented screens applied |
| Qualitative review completed | 11 | Functions, products, risks, and intangibles reviewed manually |
These figures are illustrative. A small accepted set is not inherently weak, and a large set is not inherently strong. The question is whether the retained companies are sufficiently comparable and the exclusions follow the stated criteria.
| Candidate | Evidence reviewed | Decision | Reason recorded in the study |
|---|---|---|---|
| Alpha Distribution | Database description, website, annual report, ownership record | Accept | Distributes similar industrial components, no material manufacturing identified, and passes the documented independence screen |
| Beta Industrial | Website and annual report show manufacturing and distribution; no segmental accounts | Reject | Manufacturing results cannot be separated reliably from distribution results |
| Gamma Components | Comparable activity; losses in two reviewed years | Investigate | Losses are a signal for further analysis, not an automatic rejection; determine whether they reflect normal market conditions and comparable risks |
| Delta Systems | Database code matches, but annual report describes proprietary technology and strategic product development | Reject | Owns and develops valuable technology and controls functions not present in the tested activity |
| Epsilon Trading | Similar products but majority ownership by a larger group | Reject | Does not satisfy the documented independence criterion |
This treatment follows the logic in OECD paragraphs 3.63-3.66. Extreme results and losses call for investigation. A candidate should not be excluded only because its result differs from the rest of the set; the decision should follow the underlying facts.
For every reviewed candidate, retain:
Generic reasons such as "not comparable" do not allow a second reviewer to test the judgment. State the fact that caused the decision.
OECD paragraph 2.82 links PLI selection to the nature of the controlled transaction, the functional analysis, reliable information, comparability, and the reliability of any adjustments. No PLI belongs automatically to a transaction label.
| PLI | Formula | Often considered when | Reliability questions |
|---|---|---|---|
| Operating margin | Operating profit / revenue | Revenue is a relevant indicator of the tested activity, often in distribution | Are revenue recognition, rebates, and operating items classified consistently? |
| Net cost plus | Operating profit / relevant operating costs | Costs bear a relationship to the value of the tested functions, often in services or manufacturing | Which costs belong in the base? Are pass-through items treated consistently? |
| Return on operating assets | Operating profit / relevant operating assets | Operating assets are a meaningful driver of the tested activity | Are asset values, leases, depreciation, and idle assets comparable? |
| Berry ratio | Gross profit / operating expenses | Limited intermediary facts satisfy the conditions in OECD paragraphs 2.106-2.108 | Are cost classifications comparable, and is value unrelated to the value of products sold? |
The Berry ratio is not a general substitute for an operating margin. The OECD states that it is sensitive to cost classification and sets specific conditions for its use, including a relationship between functions and operating expenses and no material relationship between those functions and the value of products distributed.
OECD paragraphs 2.83-2.85 require the net profit calculation to focus on operating items related directly or indirectly to the controlled transaction. A company-wide result can be misleading when the entity conducts different controlled transactions or material uncontrolled activities.
A financial bridge should show:
Apply the same PLI definition to the tested party and comparables. Accounting labels do not guarantee economic consistency.
OECD paragraphs 3.75-3.79 say that multiple-year data are often useful but are not a systematic requirement. The Guidelines do not prescribe a fixed number of years and do not equate examining several years with averaging them.
Multiple-year data may help explain:
State which years are used, whether the analysis uses annual observations or an average, the averaging formula, and why that treatment improves reliability.
An arm's length analysis may produce a single reliable figure or a range. OECD paragraphs 3.55-3.57 provide the relevant sequence:
The interquartile range is therefore a response to residual comparability limitations, not an automatic first step in every OECD analysis. Local law may prescribe a particular range, calculation convention, or adjustment point, so document the jurisdictional rule separately.
Do not remove an observation merely to narrow the range or move the tested result inside it. Tie every exclusion and adjustment to a comparability fact that was not already considered.
A reproducible study allows a qualified reviewer to repeat the search and calculations from the retained record. Use this checklist before sign-off.
The OECD Guidelines do not create a universal three-year validity period for benchmarking studies. Local documentation rules and administrative practice may set their own expectations.
Use an annual facts review to decide whether the prior search can be rolled forward:
| Question | If the answer is yes |
|---|---|
| Did the transaction, value chain, functions, assets, or controlled risks change materially? | Revisit the method, tested party, PLI, and search design |
| Did the tested party acquire or develop valuable intangibles? | Reassess whether a one-sided method remains reliable |
| Did the geography or economic conditions change materially? | Revisit market scope and comparability |
| Did the database change its coverage, fields, or financial restatements? | Reproduce the search using the current release and document the change |
| Did a comparable merge, become controlled, change activity, or lose usable data? | Re-screen the company and update the accepted set |
| Did local law or guidance change? | Apply the current local requirement and record its effective date |
If none of these changes affects the methodology, update the financial data, repeat the qualitative checks, and explain why the original search remains suitable. A roll-forward is a documented conclusion, not simply a new spreadsheet column.
A concise report can still be complete. Include:
The methodology should be understandable without access to the preparer's memory. Keep source data and judgment together so the reviewer does not have to infer why a company was included.
The core references in the 2022 OECD Transfer Pricing Guidelines are:
No database is best for every transaction. Company accounts databases may support TNMM or CPM, agreement databases may support royalty or other CUP analyses, and market-data sources may support financial transactions. Compare coverage, field definitions, ownership data, historic releases, reproducibility, and licensing against the evidence the method requires.
Usually, yes. Commercial databases often provide company-level codes, financials, and short descriptions rather than enough transactional detail to establish comparability. OECD paragraph 3.33 says database searches may need refinement with other public information. Record what was reviewed and the specific reason for every decision.
Not under the OECD Guidelines in every case. Paragraph 3.57 says a statistical tool such as the interquartile range may improve reliability where a sizeable set still contains unidentified or unquantified comparability defects. Local law may impose a specific approach.
The OECD does not prescribe a number. Paragraph 3.75 says multiple-year data should be used when they add value. Choose the period and any averaging convention based on the tested facts, economic cycle, data availability, and local rules.
There is no universal OECD minimum. Reliability depends on the quality of the comparables and the method, not a target count. Explain the search population, each exclusion, and why the final set supports a reliable conclusion.
Not automatically. OECD paragraphs 3.64-3.65 require attention to the facts. Exclude a loss-maker when its losses reflect non-comparable risks or conditions, not solely because the company has a loss.
No. Codes are efficient population filters. Functional comparability still depends on products or services, functions, assets, risks, contractual terms, economic circumstances, and business strategies. Manual review tests those facts.
Review it every year against the current transaction, comparables, data release, and local requirements. Run a new search when changes undermine the earlier methodology. If a roll-forward remains appropriate, record why and repeat the financial and qualitative checks.