Reference

A glossary of deal origination

Definitions of origination, research, and evidence terms, with references to relevant standards and sources.

Updated 2026-09-17 · 3,035 words · 34 sources · jump to sources

Definitions of terms used in deal origination, company research, and evidence review. References link to relevant sources, and Docket-specific terminology is identified where it appears.

Sale process and deal structure

Proprietary deal flow

Used for at least four distinct situations: a genuinely unintermediated approach to a company not for sale; a pre-process conversation before a book is written; a limited process with a handful of invited buyers; and early knowledge of a broad auction. These situations differ in the degree of access and competition they involve. The research on how firms are actually sold shows that competition in a takeover is frequently invisible from the outside, so early access should not be assumed to mean an absence of competing buyers [1]. We discuss the distinction at length on our page about proprietary deal flow.

Auction, limited auction, negotiation

Three points on a spectrum of how many buyers a seller engages. Hand-collected evidence from the pre-public stages of sale processes found roughly half of targets auctioned among multiple bidders and half negotiated with a single bidder, with comparable wealth effects for target shareholders across the two [1]. That last finding is a warning to any buyer whose model assumes a discount for exclusivity.

Merger background section

The narrative in a merger proxy or tender offer document describing how the transaction came about: who was contacted, how many signed confidentiality agreements, what the timetable was. It is a first-person account of a sale process, it is full-text searchable across filings [6], and it is the most under-used origination research source in existence.

Termination fee (break fee)

A payment owed if a signed transaction does not complete for specified reasons, typically because the target accepts a competing offer. Empirically studied as a deal-protection device and associated with the dynamics of competing bids [2]. Its existence is a reminder that the possibility of another bidder rarely disappears entirely, even in a negotiated deal.

Hart-Scott-Rodino threshold

The transaction size above which US antitrust pre-merger notification is required, adjusted annually [3]. Most lower-middle-market acquisitions fall below it and therefore close with no public filing at all. The practical consequence for origination is that a competitive map degrades silently: companies change hands without any public record you can subscribe to.

Add-on, platform, buy-and-build

A platform is the first acquisition in a sector, sized and resourced to acquire further; an add-on (or tuck-in) is a subsequent acquisition folded into it. Add-ons have made up a large and growing share of US private equity deal count for years. The competition agencies have become explicitly more attentive to serial acquisition strategies, which is a reason to know precisely what a platform and its competitors have already bought [4].

Ownership and control

Founder-owned, family-owned, management-owned

Categories of private ownership that determine what kind of conversation is possible. They are not interchangeable: a founder-owned business has one decision-maker, a family-owned business often has several with different time horizons, and a management-owned business usually has a prior transaction in its history that shapes expectations. We treat these as values from a fixed list rather than as prose, because a typed value can be sorted, scored and disagreed with.

Beneficial ownership

The natural person who ultimately owns or controls an entity, as distinct from the entity named on a registry. Most US states publish no shareholder information at all, which is why American ownership research is genuinely hard. The UK publishes a register of persons with significant control for essentially every company, free [8]. Where an entity has a Legal Entity Identifier, direct and ultimate parent relationships may be recorded in structured form [9].

Succession signal

An observable fact that bears on whether a principal is approaching a transition: a stated intention, an adviser engagement, a first finance hire, long tenure with no identified successor. We treat these as a typed answer rather than a narrative, so that “appears to be approaching retirement” is not an allowed output. Owner age alone is much closer to noise than the industry treats it.

Universe, data and measurement

NAICS

The North American Industry Classification System, the coding scheme almost every public dataset keys on [10]. Firms largely self-classify and often never update, so a code is a way of finding candidates and a poor way of defining a market. Write the capability test separately from the codes.

Establishment versus firm

An establishment is a single physical location; a firm is all establishments under common ownership. County Business Patterns counts establishments [11]; Statistics of U.S. Businesses counts firms by enterprise size [12]. Using the first as a count of acquisition targets overstates the universe by whatever the average site count is in your sector, which in multi-site industries is a factor of several.

Denominator, coverage

The denominator is the number of companies that plausibly exist in your defined market, established from statistical sources rather than from your list [11] [14]. Coverage is the share of it you have identified by name — and separately, the share you have researched to standard. These are two different numbers and conflating them is the commonest way coverage gets overstated. The Economic Census adds receipts and concentration data every five years [13].

Entity resolution

Deciding that two records refer to the same real company. The failure mode is not noise but confident duplication: one operating business can appear as a trading name, a holding company, three state registrations and an acquired brand that still answers the phone. Never match on company name; join on a stable key. Confirm candidate matches before merging their evidence.

Coverage bias in commercial databases

Paid company databases do not cover all firms equally, and the gaps vary systematically with firm size and country rather than randomly. This is documented carefully in the literature on constructing nationally representative firm-level data [15]. The practical rule: an aggregator is a candidate generator, never a definition of the universe.

Herfindahl-Hirschman Index (HHI)

The sum of squared market shares of all participants, and the concentration measure the US competition agencies use [5] [4]. Rarely computable exactly for a private-company sector, because shares require revenue you do not have. Approximate it with employment-based shares from your own map, or with published concentration ratios, and state which — “highly fragmented” with no measure behind it is a mood, not a finding.

Evidence and research method

Claim

In our vocabulary, an immutable record: a typed value for one question about one company, the document it came from, the span quoted from that document, the date, and the run that produced it. A newer value is a new claim, not an overwrite. Both stand with their dates, because a file that overwrites cannot answer “what did we know in March?”

Negative (a filed not-found)

The recorded result that a question was asked and the answer is not on the public record: what was asked, where was checked, when, and how long the finding stands before the question is asked again. A blank cell is not a negative. A blank cell is indistinguishable from nobody having looked, which is why the same undisclosed figure gets searched for five times by five people.

Evidence tier

The standard a particular question is held to, decided once for the question rather than per company. Ours are: quotable (a verbatim span must support the value), consistency-checked (no single sentence states it; it must cohere with the rest of the file), and absence-based (the claim is that nothing on the record says otherwise, after checking named places). Auditing standards make the parallel distinction between sufficiency and appropriateness of evidence [16].

Authentication

In evidence law, showing that an item is what its proponent claims it to be [17]. The research analogue is being able to show that the retained page is the page you read and has not changed. It is cheap to build in and impossible to retrofit.

Provenance

The record of what a thing was derived from, by which activity, attributed to which agent, at what time. There is a standard vocabulary for it [18]. A data model with a “source” column and no concept of the run that created the row cannot answer the question that matters after an error, which is what else that run touched.

Machine reading

Grounding / retrieval-augmented generation

Conditioning a language model’s output on documents retrieved at query time rather than on its parameters [20]. It reduces invention substantially and does not eliminate it: a system can retrieve correctly, cite plausibly, and still say something the source does not support.

Attribution (AIS)

Short for “attributable to identified sources”: whether a statement is actually supported by the document cited for it, judged by someone reading both [21]. This is the measurement that matters, and it is different from whether a citation is present. Human evaluation of generative search engines found substantial shares of sentences unsupported by their citations and of citations not supporting their sentences [23].

Hallucination

A blanket term for generated content departing from its source, covering at least two distinct bugs with distinct fixes: unsupported addition and outright contradiction [22]. In company research the more common operational failures are wrong-entity and stale-source, neither of which is a hallucination at all and neither of which is fixed by a better model.

Typed value (enum)

An answer chosen from a fixed list rather than written as prose. Defined values can be compared, filtered, and scored consistently. Each answer should retain its supporting evidence, and missing information should have an explicit state.

The regex trap

Running pattern matching over natural-language output to extract structure. It fails because a pattern matches words and the question is about meaning — negation, subject, and which side of a transaction a company was on. For illustration, matching “subsidiary” in “not a subsidiary” would reverse the meaning. Patterns belong on mechanical formats only: identifiers, dates, currency strings.

Judgement and scoring

Mechanical versus clinical prediction

Combining cues by a fixed rule versus combining them by expert holistic judgement. A meta-analysis of a large body of head-to-head comparisons found mechanical prediction equalled or outperformed expert judgement in the great majority of cases [24]. The expert’s advantage is in knowing which cues matter; their disadvantage is in weighing them the same way twice.

Noise

Unwanted variability between judgements that should be identical: the same company assessed differently on a Tuesday than on a Friday, or after three bad candidates rather than three good ones [25]. Distinct from bias, and harder to notice, because you normally see only one judgement per case.

Base rate neglect

Allowing a compelling story about a specific case to displace the underlying frequency of such cases [26]. In origination it has a signature form: “a large share of owners in this sector are near retirement, therefore this owner is.” Population statistics justify a strategy; they are not evidence about a company.

Disqualifier

A condition rather than a preference: a fact that ends consideration regardless of everything else. Modelling one as a large negative score is a mistake, because a company can accumulate enough points elsewhere to pass, and because conditions are usually the cheapest questions to answer and the most decisive. Ask them first, record the reason, stop.

Completeness sort

Ordering a ranked list by how many questions a record has answered before ordering by score, so that a high score built on thin evidence cannot outrank a fully researched company. Without it, an unanswered question scores as neutral and the board ends up ranking how little you know about each company rather than the companies.

Contact, law and delivery

CAN-SPAM

The US federal statute governing commercial email. Permissive about sending — no prior consent required — and strict about conduct: accurate headers, non-deceptive subject lines, a physical address, a working opt-out honoured promptly, and liability that is not transferred by hiring an agency [27].

CASL

Canada’s regime, which inverts the American default: commercial electronic messages generally require consent before sending, either express or implied [28]. The implied-consent categories relevant to business prospecting depend on facts you must record at the time.

Legitimate interests

The lawful basis most European B2B prospecting relies on [29]. Not a free pass: it requires a documented balancing test and fails where a person would not reasonably expect the processing. Separately, the act of sending is governed by ePrivacy rules transposed differently in each member state; the UK regulator’s guidance is the clearest English-language explanation of how the two interact [30].

SPF, DKIM, DMARC

The three email authentication standards: which hosts may send for a domain [31], a cryptographic signature from the sending domain [32], and the policy and reporting layer that ties them to the visible from address [33]. Receiving providers now require them of anyone sending in volume [34].

Alignment

The DMARC requirement that the domain which passed authentication matches the domain the recipient sees in the from field. Mail can pass SPF and DKIM individually and still fail the policy, which is the single most common reason a technically correct setup is undelivered.

Suppression list

The permanent record of people and companies who must not be contacted. Two properties make it real rather than nominal: it is enforced by the database rather than by application code, so a refactor cannot bypass it, and it survives a change of sending tool. A suppression list that lives only inside one vendor’s account is not a suppression list.

Running the process

Triage

The first and cheapest research stage: resolve identity, establish what kind of business this is, and ask the disqualifying questions before any expensive work runs. Its purpose is not to describe a company but to decide whether describing it is worth paying for. A pipeline without a triage stage pays deep-research prices for companies that fail a condition anyone could have checked in a minute.

Rolling versus wave scheduling

In a wave, a whole cohort completes a stage before any of it moves on. In a rolling schedule, each company advances the moment its own work finishes. Waves idle every lane during each stage’s tail, and the tail is set by the slowest company rather than the median — so a six-stage pipeline pays that tail six times. Rolling schedules can reduce waiting between stages, subject to dependencies and resource limits.

Re-ask

Running a question again across companies that already have an answer, because the standard changed — a new allowed value, a tightened definition of what counts as evidence, a corrected brief. The cost is the whole corpus, and paying it is the price of changing a definition. The alternative is a file where half the answers mean one thing and half mean another with no way to tell them apart.

Coverage area

Docket’s term for applying a research requirement across the relevant company list. When a new question is added, existing records are included in the work queue so that the team can track coverage against the revised criteria.

Projection

The computed current view over an append-only store of observations: given several dated claims about one field, which one is shown. Selection rules account for source quality, date, and review decisions, while the underlying observations remain available.

Fingerprint (of an approved draft)

A hash of the exact message that was approved, re-checked at the moment of sending. It is what makes “the client approved this” a verifiable statement rather than a recollection: edit a sentence after approval and the fingerprint no longer matches, so the message returns to the queue instead of going out.

Terms that need an explicit definition

These terms can describe different standards. The definitions below explain how they are used in Docket’s research workflow.

“Sourced”

An answer linked to retained source material, supporting text where applicable, a collection date, and the research run that produced it. Retaining the material helps preserve the basis of a finding when the live page changes or disappears [19].

“Verified”

An answer reviewed against its supporting evidence. The scope of verification should be stated: confirming that a quotation appears on a page is different from confirming that it supports the recorded answer [23].

“Complete”

Assessed against a named set of requirements. A file may be complete for initial screening while still needing additional information for diligence. State the standard and the treatment of unresolved questions.

“Fragmented”

A description of market structure that should be supported by a stated concentration measure [5]. Where a true index cannot be computed, approximate it and say so. “The four largest operators we identified hold roughly a fifth of mapped employment” is checkable. State the market boundary, measure, and limitations alongside the description.

Sources

References for the research and standards discussed in this guide. Some publications require a subscription or institutional access.

  1. [1]
    How Are Firms Sold?
    The Journal of Finance · 2007

    The canonical study of how takeover targets actually reach a buyer: full auctions, limited auctions and negotiations, hand-collected from SEC filings.

  2. [2]
    Termination fees in mergers and acquisitions
    Journal of Financial Economics · 2003

    On the deal-protection devices that shape what a competing bidder can realistically do after a deal is signed.

  3. [3]
    Premerger Notification Program (Hart-Scott-Rodino)
    U.S. Federal Trade Commission

    Filing thresholds are adjusted annually; below them a deal closes without notifying anyone.

  4. [4]
    2023 Merger Guidelines
    U.S. Department of Justice and Federal Trade Commission · 2023

    How the agencies say they analyse a transaction, including concentration thresholds and serial acquisition.

  5. [5]
    Herfindahl-Hirschman Index
    U.S. Department of Justice

    The concentration measure the agencies use, and a discipline for claiming an industry is 'fragmented'.

  6. [6]
    EDGAR company and filing search
    U.S. Securities and Exchange Commission

    The public search interface. Free, rate-limited, and authoritative for anything a registrant had to disclose.

  7. [7]
    Form D — notice of exempt offering of securities
    U.S. Securities and Exchange Commission

    Private raises leave a public trace here, including the issuer's address and the size of the offering.

  8. [8]
    Companies House
    UK Government

    Free filings, accounts and persons with significant control. Ownership research is trivially easier in the UK than the US.

  9. [9]
    Legal Entity Identifier (LEI) search
    Global Legal Entity Identifier Foundation

    Open, free entity identifiers with parent relationships. The nearest thing to a global primary key for companies.

  10. [10]
    North American Industry Classification System (NAICS)
    U.S. Census Bureau

    The classification every federal dataset keys on, and the first thing a market map has to get right.

  11. [11]
    County Business Patterns (CBP)
    U.S. Census Bureau

    Establishment counts, employment and payroll by NAICS code and county. The spine of any honest market map.

  12. [12]
    Statistics of U.S. Businesses (SUSB)
    U.S. Census Bureau

    Firm counts and employment by enterprise size, which is how you size the population of buyable companies.

  13. [13]
    Economic Census
    U.S. Census Bureau

    Every five years: receipts, establishments and concentration ratios by industry.

  14. [14]
    Quarterly Census of Employment and Wages (QCEW)
    U.S. Bureau of Labor Statistics

    Near-census employment and wage counts by county and industry, monthly granularity, published quarterly.

  15. [15]
    How to Construct Nationally Representative Firm Level Data from the Orbis Global Database
    National Bureau of Economic Research · 2015

    A careful account of what commercial company databases are missing and how their coverage skews by size and country. Read before trusting any vendor's universe count.

  16. [16]
    AS 1105: Audit Evidence
    Public Company Accounting Oversight Board

    Sufficiency and appropriateness of evidence, and why evidence from an independent external source ranks above management's assertion.

  17. [17]
    Federal Rule of Evidence 901 — Authenticating or Identifying Evidence
    Legal Information Institute, Cornell Law School

    What it takes to establish that a record is what its proponent claims it is.

  18. [18]
    PROV-DM: The PROV Data Model
    World Wide Web Consortium (W3C)

    A standard vocabulary for saying which entity was derived from what, by which activity, and when.

  19. [19]
    When Online Content Disappears
    Pew Research Center · 2024

    Measures how much of the web cited a few years ago is already gone. The argument for retaining the page, not the link.

  20. [20]
    Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
    arXiv · 2020

    The paper that named the pattern of grounding a language model's output in retrieved documents.

  21. [21]
    Measuring Attribution in Natural Language Generation Models
    arXiv · 2021

    Defines Attributable to Identified Sources (AIS): whether a statement is actually supported by the document cited for it.

  22. [22]
    Survey of Hallucination in Natural Language Generation
    arXiv · 2022

    Taxonomy of how generated text departs from its source, and of the metrics that try to catch it.

  23. [23]
    Evaluating Verifiability in Generative Search Engines
    arXiv · 2023

    Human evaluation finding that a large share of citations in generative search output do not support the sentence they are attached to.

  24. [24]
    Clinical versus mechanical prediction: a meta-analysis
    Psychological Assessment · 2000

    136 studies. Mechanical combination of cues equals or beats expert holistic judgement roughly nine times out of ten.

  25. [25]
    Noise: How to Overcome the High, Hidden Cost of Inconsistent Decision Making
    Harvard Business Review · 2016

    Unwanted variability between judges of the same case, and the structured protocols that reduce it.

  26. [26]
    Judgment under Uncertainty: Heuristics and Biases
    Science · 1974

    Representativeness, availability and anchoring, including the base-rate neglect that makes a good story beat a good prior.

  27. [27]
    CAN-SPAM Act: A Compliance Guide for Business
    U.S. Federal Trade Commission

    The seven rules, in plain language, from the regulator that enforces them.

  28. [28]
    Canada's Anti-Spam Legislation
    Government of Canada

    Consent-based rather than opt-out, with implied consent categories that matter for B2B prospecting.

  29. [29]
    GDPR Article 6 — Lawfulness of processing
    EUR-Lex / Regulation (EU) 2016/679

    Including legitimate interests, the basis most B2B outreach in Europe relies on, and the balancing test it requires.

  30. [30]
    Direct marketing and privacy and electronic communications
    UK Information Commissioner's Office

    The clearest regulator guidance in English on when a corporate subscriber may be emailed without prior consent.

  31. [31]
    RFC 7208 — Sender Policy Framework (SPF)
    IETF

    Which hosts may send for a domain, published in DNS.

  32. [32]
    RFC 6376 — DomainKeys Identified Mail (DKIM) Signatures
    IETF

    Cryptographic signing of a message by the sending domain.

  33. [33]
    RFC 7489 — Domain-based Message Authentication, Reporting, and Conformance (DMARC)
    IETF

    Alignment, policy and the aggregate reports that tell you who is sending as you.

  34. [34]
    Email sender guidelines
    Google Workspace Admin Help

    The receiving side's published requirements, including authentication and the spam-rate threshold for bulk senders.

Keep reading

This is how Docket works, not just what we think.

Three agents run your criteria across your target list and return a sourced entry on every company, with the page, the sentence and the date behind every field.