Judgement

Building consistent acquisition screening criteria

Turn an acquisition thesis into a written screening rubric. Define disqualifying conditions, scoring rules, and the treatment of missing information, then evaluate the results.

Updated 2026-09-17 · 2,737 words · 10 sources · jump to sources

A written screening rubric helps a team apply its investment criteria consistently across a large target universe. Research on structured judgement offers useful principles for defining the criteria, combining scores, and reviewing the results. The aim is to support investment judgement with a comparable record of the evidence considered for each company.

Research on structured judgement

Since the 1950s, researchers have run head-to-head comparisons between expert judgement and mechanical combination of the same information. A meta-analysis covering a large body of such studies found that mechanical prediction equalled or outperformed clinical judgement in the great majority of comparisons, and was clearly worse in very few [1].

Research on simple linear models suggests that consistent combination of relevant information can be useful even without finely optimised weights [2]. This supports separating the choice of screening criteria from the rules used to combine them.

Define the criteria, assess each separately, and make the basis of the overall score visible.

These studies do not directly establish acquisition outcomes. A practical application to screening is to write down how ownership, size, customer mix, and other criteria will be assessed, then apply those definitions consistently. Partners retain responsibility for selecting the criteria and making the investment judgement.

Reducing variation in screening

Noise is unwanted variability in judgements that should be consistent. It can be difficult to observe when each case is assessed only once [3].

Three sources of variation are relevant to screening:

  • Order effects. A company looks better after three bad ones. The rubric scores each company against fixed anchors instead of against its neighbours.
  • Halo. One strong attribute contaminates the assessment of unrelated ones. The structured approach to strategic decisions addresses this specifically: score the components independently and hold the overall judgement until every component is in [4].
  • Narrative substitution. A good story about a company displaces the base rate for companies like it. Base-rate research describes this risk [6].

Nothing here says the judgement should be automated away. The structured protocol explicitly preserves the final holistic call and simply refuses to let it happen before the parts have been assessed separately. The sequence preserves a role for both structured assessment and final judgement.

Disqualifiers are not scores

The most common rubric design error is making everything a weighted score. Some criteria are not preferences; they are conditions. If your fund cannot buy a subsidiary of a listed parent, then a listed parent is not minus fifteen points, it is the end of the conversation.

Collapsing conditions into scores causes two specific, expensive failures:

  1. Compensation. A company that fails a hard condition accumulates enough points elsewhere to appear on the shortlist. Somebody spends a week on it before noticing.
  2. Cost inversion. Conditions are usually the cheapest questions to answer and the most decisive. Treating them as just another weighted input means you pay for twenty answers before learning the one that ends it.

A practical note on the cost asymmetry, because it is larger than people expect. If a triage step costs a fraction of a deep read and eliminates a meaningful share of a list, running it first is not an optimisation, it is the difference between covering a universe of hundreds of thousands of firms [9] and covering the few hundred somebody had time for.

Building the rubric

A practical sequence for building the rubric:

  1. Write down what a target is, in prose, first. One paragraph, agreed by the people who will vote. If that paragraph cannot be written, the disagreement is about strategy and no rubric will resolve it.
  2. Decompose it into independent dimensions.Independence is the point: two dimensions that measure the same thing double-count it. If “scale” and “revenue” and “headcount” are three dimensions, size is being counted three times.
  3. For each dimension, define the observable.Not “good management” but “years of tenure of the principal.” If you cannot name the observable, the dimension is a feeling and belongs in the judgement stage rather than in the score.
  4. For each observable, define the allowed answers and what each is worth. Anchored, so that two people reading the same evidence assign the same value. This is where most rubrics are under-specified and therefore noisy.
  5. Define how missing answers affect the ranking. Show research completeness alongside the score, with clear rules for records that still need work.

A checklist works for the same reason as the rubric: it converts “did we consider everything” from a memory task into a verification task, and the best-documented evidence for that is not from finance[7].

The missing answer problem

Missing information can affect rankings in ways that are difficult to see from a total score.

A company where you could not find the ownership structure, the customer mix or the real-estate position has no value on those dimensions. What do they score? If the answer is “nothing, and the total is the sum,” then ignorance is being scored as neutral, and a company with two answers and a strong one can outrank a company with fifteen answers and one weakness.

For illustration, consider two companies with the same total score. One has been researched across every criterion; the other has only one supported answer and several default values. Displaying those scores without completeness would suggest a degree of comparability the research does not support.

Treatment of a missing answerEffectConsideration
Score it neutralIncomplete records may rank alongside well-researched companies.Make the default value and its effect visible.
Score it as the worst caseIncomplete records can be deprioritised before they are researched.Maintain a separate research queue.
Impute from similar companiesAdds estimates to the score.Label estimates and test their effect on the ranking.
Score zero, and sort on completeness firstA company that has answered everything sorts above one that has not, whatever the score.A useful policy when the shortlist requires complete research.
Alternative missing-data policies. State the chosen policy and display research completeness.

For a shortlist that requires complete research, sort by completeness before score. Keep incomplete but potentially relevant companies visible in a research queue. This distinguishes readiness for review from potential fit.

Choosing scoring weights

Research on improper linear models provides a reason to begin with simple, consistent weights[2]. For an acquisition rubric, test the weights against reviewed examples and check whether small changes materially alter the shortlist.

Before refining weights, check:

  • That the dimensions are the right dimensions. A missing dimension cannot be fixed by weighting.
  • That the anchors are unambiguous. Two readers must assign the same value to the same evidence.
  • That every company is scored on every dimension. Or the totals are not comparable.
  • That missing answers are handled deliberately. See above.
  • The weights. Assess their effect after reviewing the dimensions and evidence coverage.

How a rubric goes wrong

Four changes that require particular care:

  1. New categories without scoring rules. When extending the accepted answers, define the treatment of each new value at the same time. Check that every category has an explicit scoring rule.
  2. Criteria that live in two places. The rubric in the model and the rubric in the memo drift apart, and both are cited. Derive both from one definition, and add a test that asserts they agree in both directions, so widening one fails as loudly as narrowing it.
  3. Retrofitted criteria. A company everyone likes scores badly, so the rubric is adjusted until it scores well. This risks fitting the rubric to a preferred result. If the rubric is wrong, change it deliberately and re-score the whole list, including everything already reviewed under the old rules.
  4. Questions asked of some companies and not others. A new question gets asked only of companies researched after the change, and the ranking now compares companies assessed under different rules. Apply new questions across the existing list and track the research still pending.

Preparing for committee review

A rubric earns its keep in the meeting, and only if it can answer three questions without a rebuild.

The questionWhat you need ready
Why is this company ranked above that one?The per-dimension scores side by side, and the evidence behind each. Not a total.
What don't we know about it?The list of unanswered questions, by name, and what has been tried. 'Some gaps' is not an answer.
What would change the ranking?Which dimension is closest to a threshold, and what evidence would move it.
Three questions a scored board must answer on its face.

Show each pending question, the evidence already collected, and the reason it remains open. Review completion and research completeness are separate measures: all collected answers may have been reviewed while some required questions still lack evidence.

Note also that the ranking of evidence is itself a committee question. When two sources disagree, the memo should say which outranks which and why, and that ordering should be written down in advance rather than settled per company [10].

A worked rubric

Abstract advice about rubrics is easy to agree with and hard to act on, so here is a complete one for an invented mandate: a buyer rolling up regional operators of a physical-asset service business. The numbers are illustrative. The structure is the point.

Stage one: conditions

Each of these is a verdict with the reason kept, asked before anything expensive runs. Any one of them ends consideration.

  • Listed parent or REIT-owned. No proprietary path; the asset is somebody else’s trade.
  • Majority of operations outside the mandate geography. Out of scope by definition.
  • Wrong business type. A broker or reseller with no owned capacity is a different business wearing the same words.
  • Below the minimum scale. Cheap to test, and it removes the largest part of most universes.

Stage two: scores

DimensionPointsAnchor
Ownership20Founder, family or management owned scores full. Sponsor-backed scores half, because the process is set by somebody else.
Scale20A defined band scores full. Below it is sub-scale; far above it is a platform that will not clear at the mandate's multiple.
Number of sites15Two to six. One site is a single point of failure in diligence; more than eight is a platform, not a tuck-in.
Geography10Secondary markets score highest. Primary markets are already priced.
Customer mix10Diversified end-market majority. Concentration in one large customer is priced down explicitly.
Asset base10Owning the real estate scores full: it is the collateral. A short remaining lease scores zero.
Interconnection or network position10A defensible position in the local network scores full; a single-supplier dependency scores near zero.
Acquisition history5A company that has bought before has a data room, a quality-of-earnings history and a seller who has been through it.
Eight dimensions, 100 points. Each anchor is written so two readers assign the same value to the same evidence.

Stage three: the rules around the score

  1. An unanswered dimension scores zero, and the board sorts on completeness before score.
  2. The bands are stated: a top tier, a second tier, a watch tier and a hold, with the thresholds written down and versioned.
  3. Every dimension’s value is inspectable with the evidence behind it, from the ranked list, in one click.
  4. A change of threshold is a version, and re-scores the whole list rather than only new entrants.

Review the distribution of scores before relying on the tiers. If many companies land in the top band, determine whether the list has strong fit or the thresholds provide too little differentiation. Document any changes and apply them to the full list.

Calibrating against outcomes

A rubric needs periodic comparison with outcomes. Forecasting research examines what improves accuracy: training in base rates and decomposition, working in teams, and above all tracking results and feeding them back [5]. These practices provide a useful basis for reviewing a screening process.

A practical review cycle:

  • Record the score at the time of the decision, immutably. Not the score today.
  • Record the outcome: contacted, responded, met, diligenced, closed, passed and why.
  • Look, once a quarter, at where the rubric and the outcome disagreed. Both directions: high scores that went nowhere, low scores that became deals.
  • Change the rubric deliberately, version it, and re-score the list. A rubric that changes without a version is two rubrics.

This is slow. A fund does not get many deals, so the sample is small for years. That argues for scoring the intermediate outcomes you do get in volume — reply rates, meeting conversion, which dimensions correlate with an owner engaging at all — rather than waiting for closings to accumulate.

One caution if any part of that loop is automated: an automated evaluator has its own biases, including a documented tendency to prefer longer answers and to favour its own outputs [8]. Use it to flag disagreement, and keep a human-scored sample as the anchor.

Who owns the rubric

A rubric with no owner decays into whatever the last person edited. Three roles have to be distinct, and in most teams they are quietly collapsed into one, which is why criteria drift.

RoleDecidesMust not
The mandate ownerWhat a target is: which dimensions exist, what the thresholds are, which conditions are fatal. A business decision, made by whoever answers for the strategy.Change a threshold to make a favoured company score better.
The process ownerHow the questions are asked: the allowed answers, the briefs that make two readers agree, the evidence standard per question, the re-ask cadence.Invent a new allowed answer mid-run, or narrow a definition without re-asking the corpus.
The collectorNothing. It answers from the fixed list and quotes the evidence.Choose a value that is not on the list, or leave a required question silently unanswered.
Three roles. When the collector can extend the vocabulary, you no longer have a rubric.

The third row is the strict one and it is the one that holds the whole thing together. The moment whatever is doing the reading can invent a category, the categories stop being comparable across companies, and everything downstream — the ranking, the coverage statistics, the ability to re-score after a threshold change — silently stops meaning anything. That is why the escape hatch is a measured queue of work for the process owner rather than a convenience for the collector.

When a promising company ranks poorly, review whether the evidence or the rubric explains the result. An exception can be recorded as a judgement call. If the criteria change, record the reason and re-score all companies assessed under the earlier version.

What a rubric is not for

  • It does not make the decision. It orders the queue so that judgement is applied to the right twenty companies instead of to whichever twenty arrived. The investment decision remains with the team.
  • It does not capture everything that matters. A rubric holds what can be observed at scale. Management quality, cultural fit and whether an owner will actually sell are not in it, and pretending otherwise produces false precision.
  • It does not substitute for coverage. Assess the coverage of the underlying target list as well as the quality of its ranking.
  • It does not settle strategy. If two partners disagree about whether you want founder-owned companies, the rubric will encode whoever wrote it. Agree the strategy before encoding it.

A useful rubric applies the same questions to each company, makes answers comparable, and lets reviewers inspect the basis of a ranking. Its value depends on the quality of the criteria, the supporting evidence, and the team’s review of the results.

Sources

References for the research and standards discussed in this guide. Some publications require a subscription or institutional access.

  1. [1]
    Clinical versus mechanical prediction: a meta-analysis
    Psychological Assessment · 2000

    136 studies. Mechanical combination of cues equals or beats expert holistic judgement roughly nine times out of ten.

  2. [2]
    The robust beauty of improper linear models in decision making
    American Psychologist · 1979

    Even crude equal-weight models outperform expert intuition, because consistency is worth more than precision in the weights.

  3. [3]
    Noise: How to Overcome the High, Hidden Cost of Inconsistent Decision Making
    Harvard Business Review · 2016

    Unwanted variability between judges of the same case, and the structured protocols that reduce it.

  4. [4]
    A Structured Approach to Strategic Decisions
    MIT Sloan Management Review · 2019

    The mediating assessments protocol: score the parts independently, hold the overall judgement until the end.

  5. [5]
    Psychological Strategies for Winning a Geopolitical Forecasting Tournament
    Psychological Science · 2014

    Results from the Good Judgment Project: what measurably improved forecasts was training in base rates, decomposition, teaming and frequent small updates, scored against outcomes.

  6. [6]
    Judgment under Uncertainty: Heuristics and Biases
    Science · 1974

    Representativeness, availability and anchoring, including the base-rate neglect that makes a good story beat a good prior.

  7. [7]
    WHO Surgical Safety Checklist
    World Health Organization

    The best-documented case that a short, mandatory checklist changes outcomes in expert hands.

  8. [8]
    Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
    arXiv · 2023

    Documents position, verbosity and self-enhancement bias in model-graded evaluation.

  9. [9]
    Statistics of U.S. Businesses (SUSB)
    U.S. Census Bureau

    Firm counts and employment by enterprise size, which is how you size the population of buyable companies.

  10. [10]
    AS 1105: Audit Evidence
    Public Company Accounting Oversight Board

    Sufficiency and appropriateness of evidence, and why evidence from an independent external source ranks above management's assertion.

Keep reading

This is how Docket works, not just what we think.

Three agents run your criteria across your target list and return a sourced entry on every company, with the page, the sentence and the date behind every field.