Definitions of terms used in deal origination, company research, and evidence review. References link to relevant sources, and Docket-specific terminology is identified where it appears.
Sale process and deal structure
Proprietary deal flow
Used for at least four distinct situations: a genuinely unintermediated approach to a company not for sale; a pre-process conversation before a book is written; a limited process with a handful of invited buyers; and early knowledge of a broad auction. These situations differ in the degree of access and competition they involve. The research on how firms are actually sold shows that competition in a takeover is frequently invisible from the outside, so early access should not be assumed to mean an absence of competing buyers [1]. We discuss the distinction at length on our page about proprietary deal flow.
Auction, limited auction, negotiation
Three points on a spectrum of how many buyers a seller engages. Hand-collected evidence from the pre-public stages of sale processes found roughly half of targets auctioned among multiple bidders and half negotiated with a single bidder, with comparable wealth effects for target shareholders across the two [1]. That last finding is a warning to any buyer whose model assumes a discount for exclusivity.
Merger background section
The narrative in a merger proxy or tender offer document describing how the transaction came about: who was contacted, how many signed confidentiality agreements, what the timetable was. It is a first-person account of a sale process, it is full-text searchable across filings [6], and it is the most under-used origination research source in existence.
Termination fee (break fee)
A payment owed if a signed transaction does not complete for specified reasons, typically because the target accepts a competing offer. Empirically studied as a deal-protection device and associated with the dynamics of competing bids [2]. Its existence is a reminder that the possibility of another bidder rarely disappears entirely, even in a negotiated deal.
Hart-Scott-Rodino threshold
The transaction size above which US antitrust pre-merger notification is required, adjusted annually [3]. Most lower-middle-market acquisitions fall below it and therefore close with no public filing at all. The practical consequence for origination is that a competitive map degrades silently: companies change hands without any public record you can subscribe to.
Add-on, platform, buy-and-build
A platform is the first acquisition in a sector, sized and resourced to acquire further; an add-on (or tuck-in) is a subsequent acquisition folded into it. Add-ons have made up a large and growing share of US private equity deal count for years. The competition agencies have become explicitly more attentive to serial acquisition strategies, which is a reason to know precisely what a platform and its competitors have already bought [4].
Ownership and control
Founder-owned, family-owned, management-owned
Categories of private ownership that determine what kind of conversation is possible. They are not interchangeable: a founder-owned business has one decision-maker, a family-owned business often has several with different time horizons, and a management-owned business usually has a prior transaction in its history that shapes expectations. We treat these as values from a fixed list rather than as prose, because a typed value can be sorted, scored and disagreed with.
Sponsor-backed
Partly or wholly owned by a financial investor. Usually means a process rather than a conversation, and a timetable set by somebody else’s fund life. Detecting it reliably is harder than it sounds: a growth investment from several years ago may or may not still be in place, and the company’s own website rarely mentions it. A Form D filing, where one exists, names the issuer and its officers [7].
Beneficial ownership
The natural person who ultimately owns or controls an entity, as distinct from the entity named on a registry. Most US states publish no shareholder information at all, which is why American ownership research is genuinely hard. The UK publishes a register of persons with significant control for essentially every company, free [8]. Where an entity has a Legal Entity Identifier, direct and ultimate parent relationships may be recorded in structured form [9].
Succession signal
An observable fact that bears on whether a principal is approaching a transition: a stated intention, an adviser engagement, a first finance hire, long tenure with no identified successor. We treat these as a typed answer rather than a narrative, so that “appears to be approaching retirement” is not an allowed output. Owner age alone is much closer to noise than the industry treats it.
Universe, data and measurement
NAICS
The North American Industry Classification System, the coding scheme almost every public dataset keys on [10]. Firms largely self-classify and often never update, so a code is a way of finding candidates and a poor way of defining a market. Write the capability test separately from the codes.
Establishment versus firm
An establishment is a single physical location; a firm is all establishments under common ownership. County Business Patterns counts establishments [11]; Statistics of U.S. Businesses counts firms by enterprise size [12]. Using the first as a count of acquisition targets overstates the universe by whatever the average site count is in your sector, which in multi-site industries is a factor of several.
Denominator, coverage
The denominator is the number of companies that plausibly exist in your defined market, established from statistical sources rather than from your list [11] [14]. Coverage is the share of it you have identified by name — and separately, the share you have researched to standard. These are two different numbers and conflating them is the commonest way coverage gets overstated. The Economic Census adds receipts and concentration data every five years [13].
Entity resolution
Deciding that two records refer to the same real company. The failure mode is not noise but confident duplication: one operating business can appear as a trading name, a holding company, three state registrations and an acquired brand that still answers the phone. Never match on company name; join on a stable key. Confirm candidate matches before merging their evidence.
Coverage bias in commercial databases
Paid company databases do not cover all firms equally, and the gaps vary systematically with firm size and country rather than randomly. This is documented carefully in the literature on constructing nationally representative firm-level data [15]. The practical rule: an aggregator is a candidate generator, never a definition of the universe.
Herfindahl-Hirschman Index (HHI)
The sum of squared market shares of all participants, and the concentration measure the US competition agencies use [5] [4]. Rarely computable exactly for a private-company sector, because shares require revenue you do not have. Approximate it with employment-based shares from your own map, or with published concentration ratios, and state which — “highly fragmented” with no measure behind it is a mood, not a finding.
Evidence and research method
Claim
In our vocabulary, an immutable record: a typed value for one question about one company, the document it came from, the span quoted from that document, the date, and the run that produced it. A newer value is a new claim, not an overwrite. Both stand with their dates, because a file that overwrites cannot answer “what did we know in March?”
Negative (a filed not-found)
The recorded result that a question was asked and the answer is not on the public record: what was asked, where was checked, when, and how long the finding stands before the question is asked again. A blank cell is not a negative. A blank cell is indistinguishable from nobody having looked, which is why the same undisclosed figure gets searched for five times by five people.
Evidence tier
The standard a particular question is held to, decided once for the question rather than per company. Ours are: quotable (a verbatim span must support the value), consistency-checked (no single sentence states it; it must cohere with the rest of the file), and absence-based (the claim is that nothing on the record says otherwise, after checking named places). Auditing standards make the parallel distinction between sufficiency and appropriateness of evidence [16].
Authentication
In evidence law, showing that an item is what its proponent claims it to be [17]. The research analogue is being able to show that the retained page is the page you read and has not changed. It is cheap to build in and impossible to retrofit.
Provenance
The record of what a thing was derived from, by which activity, attributed to which agent, at what time. There is a standard vocabulary for it [18]. A data model with a “source” column and no concept of the run that created the row cannot answer the question that matters after an error, which is what else that run touched.
Link rot
The steady disappearance of cited web content. Measured rates are high enough that a multi-year research asset built on live URLs is being quietly devalued the entire time [19]. A citation to a page that no longer exists is, to a reader, indistinguishable from a fabricated one. The response is to retain the page rather than the address.
Machine reading
Grounding / retrieval-augmented generation
Conditioning a language model’s output on documents retrieved at query time rather than on its parameters [20]. It reduces invention substantially and does not eliminate it: a system can retrieve correctly, cite plausibly, and still say something the source does not support.
Attribution (AIS)
Short for “attributable to identified sources”: whether a statement is actually supported by the document cited for it, judged by someone reading both [21]. This is the measurement that matters, and it is different from whether a citation is present. Human evaluation of generative search engines found substantial shares of sentences unsupported by their citations and of citations not supporting their sentences [23].
Hallucination
A blanket term for generated content departing from its source, covering at least two distinct bugs with distinct fixes: unsupported addition and outright contradiction [22]. In company research the more common operational failures are wrong-entity and stale-source, neither of which is a hallucination at all and neither of which is fixed by a better model.
Typed value (enum)
An answer chosen from a fixed list rather than written as prose. Defined values can be compared, filtered, and scored consistently. Each answer should retain its supporting evidence, and missing information should have an explicit state.
The regex trap
Running pattern matching over natural-language output to extract structure. It fails because a pattern matches words and the question is about meaning — negation, subject, and which side of a transaction a company was on. For illustration, matching “subsidiary” in “not a subsidiary” would reverse the meaning. Patterns belong on mechanical formats only: identifiers, dates, currency strings.
Judgement and scoring
Mechanical versus clinical prediction
Combining cues by a fixed rule versus combining them by expert holistic judgement. A meta-analysis of a large body of head-to-head comparisons found mechanical prediction equalled or outperformed expert judgement in the great majority of cases [24]. The expert’s advantage is in knowing which cues matter; their disadvantage is in weighing them the same way twice.
Noise
Unwanted variability between judgements that should be identical: the same company assessed differently on a Tuesday than on a Friday, or after three bad candidates rather than three good ones [25]. Distinct from bias, and harder to notice, because you normally see only one judgement per case.
Base rate neglect
Allowing a compelling story about a specific case to displace the underlying frequency of such cases [26]. In origination it has a signature form: “a large share of owners in this sector are near retirement, therefore this owner is.” Population statistics justify a strategy; they are not evidence about a company.
Disqualifier
A condition rather than a preference: a fact that ends consideration regardless of everything else. Modelling one as a large negative score is a mistake, because a company can accumulate enough points elsewhere to pass, and because conditions are usually the cheapest questions to answer and the most decisive. Ask them first, record the reason, stop.
Completeness sort
Ordering a ranked list by how many questions a record has answered before ordering by score, so that a high score built on thin evidence cannot outrank a fully researched company. Without it, an unanswered question scores as neutral and the board ends up ranking how little you know about each company rather than the companies.
Contact, law and delivery
CAN-SPAM
The US federal statute governing commercial email. Permissive about sending — no prior consent required — and strict about conduct: accurate headers, non-deceptive subject lines, a physical address, a working opt-out honoured promptly, and liability that is not transferred by hiring an agency [27].
CASL
Canada’s regime, which inverts the American default: commercial electronic messages generally require consent before sending, either express or implied [28]. The implied-consent categories relevant to business prospecting depend on facts you must record at the time.
Legitimate interests
The lawful basis most European B2B prospecting relies on [29]. Not a free pass: it requires a documented balancing test and fails where a person would not reasonably expect the processing. Separately, the act of sending is governed by ePrivacy rules transposed differently in each member state; the UK regulator’s guidance is the clearest English-language explanation of how the two interact [30].
Alignment
The DMARC requirement that the domain which passed authentication matches the domain the recipient sees in the from field. Mail can pass SPF and DKIM individually and still fail the policy, which is the single most common reason a technically correct setup is undelivered.
Suppression list
The permanent record of people and companies who must not be contacted. Two properties make it real rather than nominal: it is enforced by the database rather than by application code, so a refactor cannot bypass it, and it survives a change of sending tool. A suppression list that lives only inside one vendor’s account is not a suppression list.
Running the process
Triage
The first and cheapest research stage: resolve identity, establish what kind of business this is, and ask the disqualifying questions before any expensive work runs. Its purpose is not to describe a company but to decide whether describing it is worth paying for. A pipeline without a triage stage pays deep-research prices for companies that fail a condition anyone could have checked in a minute.
Rolling versus wave scheduling
In a wave, a whole cohort completes a stage before any of it moves on. In a rolling schedule, each company advances the moment its own work finishes. Waves idle every lane during each stage’s tail, and the tail is set by the slowest company rather than the median — so a six-stage pipeline pays that tail six times. Rolling schedules can reduce waiting between stages, subject to dependencies and resource limits.
Re-ask
Running a question again across companies that already have an answer, because the standard changed — a new allowed value, a tightened definition of what counts as evidence, a corrected brief. The cost is the whole corpus, and paying it is the price of changing a definition. The alternative is a file where half the answers mean one thing and half mean another with no way to tell them apart.
Coverage area
Docket’s term for applying a research requirement across the relevant company list. When a new question is added, existing records are included in the work queue so that the team can track coverage against the revised criteria.
Projection
The computed current view over an append-only store of observations: given several dated claims about one field, which one is shown. Selection rules account for source quality, date, and review decisions, while the underlying observations remain available.
Fingerprint (of an approved draft)
A hash of the exact message that was approved, re-checked at the moment of sending. It is what makes “the client approved this” a verifiable statement rather than a recollection: edit a sentence after approval and the fingerprint no longer matches, so the message returns to the queue instead of going out.
Terms that need an explicit definition
These terms can describe different standards. The definitions below explain how they are used in Docket’s research workflow.
“Sourced”
An answer linked to retained source material, supporting text where applicable, a collection date, and the research run that produced it. Retaining the material helps preserve the basis of a finding when the live page changes or disappears [19].
“Verified”
An answer reviewed against its supporting evidence. The scope of verification should be stated: confirming that a quotation appears on a page is different from confirming that it supports the recorded answer [23].
“Complete”
Assessed against a named set of requirements. A file may be complete for initial screening while still needing additional information for diligence. State the standard and the treatment of unresolved questions.
“Fragmented”
A description of market structure that should be supported by a stated concentration measure [5]. Where a true index cannot be computed, approximate it and say so. “The four largest operators we identified hold roughly a fifth of mapped employment” is checkable. State the market boundary, measure, and limitations alongside the description.
Sources
References for the research and standards discussed in this guide. Some publications require a subscription or institutional access.
- [1]How Are Firms Sold?The Journal of Finance · 2007
The canonical study of how takeover targets actually reach a buyer: full auctions, limited auctions and negotiations, hand-collected from SEC filings.
- [2]Termination fees in mergers and acquisitionsJournal of Financial Economics · 2003
On the deal-protection devices that shape what a competing bidder can realistically do after a deal is signed.
- [3]Premerger Notification Program (Hart-Scott-Rodino)U.S. Federal Trade Commission
Filing thresholds are adjusted annually; below them a deal closes without notifying anyone.
- [4]2023 Merger GuidelinesU.S. Department of Justice and Federal Trade Commission · 2023
How the agencies say they analyse a transaction, including concentration thresholds and serial acquisition.
- [5]Herfindahl-Hirschman IndexU.S. Department of Justice
The concentration measure the agencies use, and a discipline for claiming an industry is 'fragmented'.
- [6]EDGAR company and filing searchU.S. Securities and Exchange Commission
The public search interface. Free, rate-limited, and authoritative for anything a registrant had to disclose.
- [7]Form D — notice of exempt offering of securitiesU.S. Securities and Exchange Commission
Private raises leave a public trace here, including the issuer's address and the size of the offering.
- [8]Companies HouseUK Government
Free filings, accounts and persons with significant control. Ownership research is trivially easier in the UK than the US.
- [9]Legal Entity Identifier (LEI) searchGlobal Legal Entity Identifier Foundation
Open, free entity identifiers with parent relationships. The nearest thing to a global primary key for companies.
- [10]North American Industry Classification System (NAICS)U.S. Census Bureau
The classification every federal dataset keys on, and the first thing a market map has to get right.
- [11]County Business Patterns (CBP)U.S. Census Bureau
Establishment counts, employment and payroll by NAICS code and county. The spine of any honest market map.
- [12]Statistics of U.S. Businesses (SUSB)U.S. Census Bureau
Firm counts and employment by enterprise size, which is how you size the population of buyable companies.
- [13]Economic CensusU.S. Census Bureau
Every five years: receipts, establishments and concentration ratios by industry.
- [14]Quarterly Census of Employment and Wages (QCEW)U.S. Bureau of Labor Statistics
Near-census employment and wage counts by county and industry, monthly granularity, published quarterly.
- [15]How to Construct Nationally Representative Firm Level Data from the Orbis Global DatabaseNational Bureau of Economic Research · 2015
A careful account of what commercial company databases are missing and how their coverage skews by size and country. Read before trusting any vendor's universe count.
- [16]AS 1105: Audit EvidencePublic Company Accounting Oversight Board
Sufficiency and appropriateness of evidence, and why evidence from an independent external source ranks above management's assertion.
- [17]Federal Rule of Evidence 901 — Authenticating or Identifying EvidenceLegal Information Institute, Cornell Law School
What it takes to establish that a record is what its proponent claims it is.
- [18]PROV-DM: The PROV Data ModelWorld Wide Web Consortium (W3C)
A standard vocabulary for saying which entity was derived from what, by which activity, and when.
- [19]When Online Content DisappearsPew Research Center · 2024
Measures how much of the web cited a few years ago is already gone. The argument for retaining the page, not the link.
- [20]Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksarXiv · 2020
The paper that named the pattern of grounding a language model's output in retrieved documents.
- [21]Measuring Attribution in Natural Language Generation ModelsarXiv · 2021
Defines Attributable to Identified Sources (AIS): whether a statement is actually supported by the document cited for it.
- [22]Survey of Hallucination in Natural Language GenerationarXiv · 2022
Taxonomy of how generated text departs from its source, and of the metrics that try to catch it.
- [23]Evaluating Verifiability in Generative Search EnginesarXiv · 2023
Human evaluation finding that a large share of citations in generative search output do not support the sentence they are attached to.
- [24]Clinical versus mechanical prediction: a meta-analysisPsychological Assessment · 2000
136 studies. Mechanical combination of cues equals or beats expert holistic judgement roughly nine times out of ten.
- [25]Noise: How to Overcome the High, Hidden Cost of Inconsistent Decision MakingHarvard Business Review · 2016
Unwanted variability between judges of the same case, and the structured protocols that reduce it.
- [26]Judgment under Uncertainty: Heuristics and BiasesScience · 1974
Representativeness, availability and anchoring, including the base-rate neglect that makes a good story beat a good prior.
- [27]CAN-SPAM Act: A Compliance Guide for BusinessU.S. Federal Trade Commission
The seven rules, in plain language, from the regulator that enforces them.
- [28]Canada's Anti-Spam LegislationGovernment of Canada
Consent-based rather than opt-out, with implied consent categories that matter for B2B prospecting.
- [29]GDPR Article 6 — Lawfulness of processingEUR-Lex / Regulation (EU) 2016/679
Including legitimate interests, the basis most B2B outreach in Europe relies on, and the balancing test it requires.
- [30]Direct marketing and privacy and electronic communicationsUK Information Commissioner's Office
The clearest regulator guidance in English on when a corporate subscriber may be emailed without prior consent.
- [31]
- [32]RFC 6376 — DomainKeys Identified Mail (DKIM) SignaturesIETF
Cryptographic signing of a message by the sending domain.
- [33]RFC 7489 — Domain-based Message Authentication, Reporting, and Conformance (DMARC)IETF
Alignment, policy and the aggregate reports that tell you who is sending as you.
- [34]Email sender guidelinesGoogle Workspace Admin Help
The receiving side's published requirements, including authentication and the spam-rate threshold for bulk senders.
Keep reading
Data
Sources of acquisition target data and their limits
A field guide to registries, filings, company websites, and commercial databases: what each source can establish, where coverage is limited, and how to handle conflicting information.
Method
Evidence standards for acquisition research
How principles from auditing, authentication, and archival practice can help teams retain and review the evidence behind company research.
Judgement
Building consistent acquisition screening criteria
Turn an acquisition thesis into a written screening rubric. Define disqualifying conditions, scoring rules, and the treatment of missing information, then evaluate the results.
This is how Docket works, not just what we think.
Three agents run your criteria across your target list and return a sourced entry on every company, with the page, the sentence and the date behind every field.