A useful research file lets a reviewer trace an answer to the material that supports it. Auditing, evidence law, archival practice, and attribution research offer principles for building that record. This guide applies them to acquisition research, with attention to source quality, retained evidence, and the distinction between a quotation and the conclusion drawn from it.
The problem with “sourced”
Research files can provide different levels of source detail:
| What the file contains | What a reviewer can inspect | Limitation |
|---|---|---|
| A value, and a domain name | The named website | The relevant page and supporting passage still need to be located. |
| A value, and a deep URL | The linked page, if it remains available | The page may change or disappear. |
| A value, a URL and the date read | The linked page and the recorded access date | Better. Still requires the page to survive, and still does not say which sentence supports the value. |
| A value, a URL, a date, and the sentence quoted from a retained copy of the page | The retained text and its relationship to the claim | The reviewer must still assess whether the text supports the conclusion. |
Retaining the relevant text reduces the work needed to review a finding. It also preserves what the researcher relied on if the live page subsequently changes.
Principles from auditing
Audit standards are explicit about the two properties evidence must have: sufficiency, which is about quantity, and appropriateness, which is about relevance and reliability. The standards then rank reliability, and the ranking is directly usable in research on private companies [1]:
- Evidence from a knowledgeable independent source outside the entity is more reliable than evidence from the entity itself.
- Evidence obtained directly is more reliable than evidence obtained indirectly or by inference.
- Documentary evidence is more reliable than an oral representation.
- Original documents are more reliable than copies.
For origination research, define source preferences for each field. A relevant filing may provide stronger support than an estimated database value, but date, entity, and scope still matter. Source ranking helps guide a review; it does not automatically settle every disagreement.
Source independence also matters. Ten low-grade sources saying the same thing is often ten copies of one low-grade source, because the internet copies. Counting corroboration without checking independence is a way of manufacturing confidence.
What evidence law adds
Authentication addresses a related question: how do you establish that a record is what it claims to be? The rule on authentication requires evidence sufficient to support a finding that the item is what its proponent claims [2]. The interesting part for our purposes is the self-authentication provisions: a record generated by an electronic process, certified as such, and a copy of data identified by a hash value, both have routes to being accepted without live testimony [3].
That is a remarkably practical template. It says, in effect: if your process is documented, if the record was generated by that process, and if you can show the copy has not changed, the record can stand on its own. The electronic discovery literature builds the same idea into a defensible-process standard rather than a per-document argument [4].
What provenance standards add
The PROV data model describes entities, the activities that generated them, and the agents responsible [5]. Its contribution is not the XML; it is the insistence that derivation is a first-class fact. This value was derived from that document, by this activity, attributed to this agent, at this time.
The media world arrived at the same place from the opposite direction. Content provenance work attaches signed, tamper-evident assertions about how a piece of content was produced [6]. The motivation there is synthetic media; the structure is the same one a research claim needs.
The transferable conclusion: a claim is not a value with a citation glued on. A claim is a small record with a value, a source, an activity that produced it, an agent responsible and a time. If your data model has a column for “source” and no concept of the run that created the row, you cannot answer the question that matters when something turns out to be wrong, which is: what else did that run touch?
What machine attribution research adds
Once a language model is doing the reading, a specific new failure appears: the citation is real, the page exists, and the sentence does not support the claim. This is measurable, and it has been measured.
The concept to know is Attributable to Identified Sources: not “did the system cite something” but “is the statement actually supported by the cited document” [10]. Human evaluation of generative search engines found that a substantial share of generated sentences were not fully supported by their citations, and that a substantial share of citations did not support their associated sentence [11]. A citation therefore needs to be checked for support, not merely presence.
Two further results shape how you check. Factual precision is best measured at the level of atomic claims rather than passages, because a paragraph is usually part right [12]; and the taxonomy of how generated text departs from its source distinguishes contradiction from unsupported addition, which are different bugs with different fixes [13].
A citation identifies a source. Attribution asks whether that source supports the claim.
Retaining the page enables two separate checks. First, verify that the quoted text appears in the retained material. Second, assess whether that text supports the recorded answer. The first check can be automated through text comparison; the second requires interpretation. Passing the first does not establish the second.
A working evidence standard
Assembled from all four disciplines, the working definition we hold to:
Preserving prior entries lets a team review the information available at the time of an earlier decision.
Not every fact deserves the same proof
Demanding a verbatim quotation for everything is a good way to make a system that refuses to answer. Facts divide into at least three kinds, and the standard should be a property of the kind, recorded alongside it, rather than an exception somebody remembers.
| Kind of fact | Standard | Example |
|---|---|---|
| Quotable | A verbatim span from a retained page must support the value. | Ownership structure, services offered, certifications, locations, named leadership. |
| Consistency-checked | No single sentence states it; the value must be consistent with the rest of the file and with its source's own structure. | Employee count from a confirmed company page. The page shows a number; the check is that the page is the right company's. |
| Absence-based | The claim is that nothing on the record says otherwise, after checking a defined list of places. | 'No parent company found.' 'No customer concentration disclosed.' Real answers, provided the places checked are named. |
The reason to make this explicit rather than tacit is that it stops the two failures at either end. Without it, a reviewer either strikes every absence-based claim for lacking a quotation, or waves through a quotable claim that has none.
Recording unsuccessful research
Most research systems treat a blank as nothing. It is not nothing; it is the result of work, and throwing it away is expensive twice over. The first cost is repetition: five analysts search for the same undisclosed revenue figure across a year. The second is interpretive: a blank cell looks identical whether nobody looked, somebody looked and found nothing, or the question does not apply.
A not-found should be a first-class record with four properties:
- What was asked, in the same vocabulary as the answers.
- Where was checked, specifically enough that somebody could repeat it.
- When, because the answer may exist tomorrow.
- How long it stands, after which the question is asked again automatically.
That last property is the one that turns a research file from a snapshot into something alive. Without an expiry, you either re-ask everything constantly and pay for it, or you never re-ask and quietly rot. With one, the cost of keeping a thousand files current is bounded and predictable.
The second reader
A separate review pass needs a defined task and evaluation criteria. Research on variation in judgement supports structuring that review [16].
Three design choices make a second reader worth its cost:
- The reviewer sees the evidence, not just the conclusion. Reviewing a value without its quoted span is not review, it is a vote.
- The reviewer works under a different constraint from the collector. Ours reads with the fetch tools taken away: it can check what was retained, not go and find something new. This separates review of the existing record from additional research.
- The reviewer checks support and records a decision. It can accept a finding, correct it using the retained evidence, or remove it if unsupported. Unresolved conflicts remain visible. Evaluate both missed errors and incorrect rejections against a human-reviewed sample.
Frameworks for managing risk in machine systems converge on the same vocabulary: valid and reliable, accountable and transparent, with measurable characteristics rather than assurances [14]. A second reader whose behaviour is itself unmeasured is an assurance.
Why the ledger is append-only
The single design decision that does the most work, and the one most often argued about, is that a research store records observations and never edits them. A newer value is a new record. The current view is computed.
Preserving the history supports four tasks:
- You can answer “what did we know when we decided?” This is the question asked in every post-mortem, and an overwriting store cannot answer it at all. It can only tell you what it believes now, which is exactly the information you do not need.
- A change becomes visible as a change. Ownership moving from independent to sponsor-backed is a signal. In an overwriting store it is an edit nobody sees.
- A bad run is recoverable. If every record carries the run that produced it, a defective batch can be identified and superseded precisely. Without it, a bad run is mixed irreversibly into the good data and the only honest remedy is to redo everything.
- Disagreement becomes a first-class object. Two conflicting observations can both stand, with their dates and sources, and the conflict becomes visible reconciliation work rather than a silent overwrite by whichever process ran last.
A historical record needs explicit rules for selecting the current answer. Define how source quality, date, evidence type, and review decisions affect that selection. Retain access to the other observations.
Explaining incomplete records
Here is a question that sounds trivial and is not: why is this company’s file incomplete?
A company can have every collected answer reviewed while still lacking evidence for some required questions. Report review status and research completeness separately, against the current mandate.
A coverage view should list the pending questions for each company, the evidence already collected, and any recorded not-found results. It should also explain whether a question needs more research, awaits review, or does not apply under the current criteria.
A useful coverage view has three properties:
- It is per company and per standard.The same file can be complete against a screening standard and incomplete against a diligence one. A single global “complete” flag is therefore always wrong for somebody.
- It names the rule.“This question is not pending because a not-found was filed six weeks ago and stands for six months” ends the conversation. “It looks complete to me” does not.
- It is computed on demand, not stored. The answer depends on the standard you ask about, and materialising it for every company against every standard is both expensive and immediately stale.
Making these distinctions visible helps teams direct additional research to the records that need it.
Objections worth taking seriously
“This is too heavy for a screening exercise.”
Partly fair. The right response is not to lower the standard but to apply it to fewer questions early on. Ask three disqualifying questions to the full standard before asking twenty descriptive ones to any standard. Cheap research done properly on the questions that eliminate candidates is far better value than thorough research done on a list that should have been a fifth of its size.
“Retaining pages has a copyright and storage cost.”
Storage is trivially cheap relative to a single analyst-day. The rights position for retaining copies of public pages for internal research and quoting short extracts is a question for your counsel, and it is the same question every archive, every search engine and every e-discovery process has already had to answer [4]. Public archives exist and are useful [9] [8], and they do not remove the need to keep your own copy of the thing you actually relied on.
“Our sources are stable, so link rot is not our problem.”
Filings on a government system are indeed reasonably stable [15]. Company websites, which is where most private-company evidence lives, are not, and the measured rate at which cited web content disappears is high enough that a multi-year research asset built on live links is being quietly devalued the whole time [7].
“Nobody will ever check.”
Not every finding will be reviewed individually. Retained evidence still allows a team to inspect important claims and sample the wider file for quality. That review can identify errors and inform changes to the research process.
Sources
References for the research and standards discussed in this guide. Some publications require a subscription or institutional access.
- [1]AS 1105: Audit EvidencePublic Company Accounting Oversight Board
Sufficiency and appropriateness of evidence, and why evidence from an independent external source ranks above management's assertion.
- [2]Federal Rule of Evidence 901 — Authenticating or Identifying EvidenceLegal Information Institute, Cornell Law School
What it takes to establish that a record is what its proponent claims it is.
- [3]Federal Rule of Evidence 902 — Evidence That Is Self-AuthenticatingLegal Information Institute, Cornell Law School
Subsections (13) and (14) cover certified records generated by an electronic process and data copies identified by hash.
- [4]The Sedona Conference publicationsThe Sedona Conference
The reference commentary on electronic discovery, preservation and defensible process.
- [5]PROV-DM: The PROV Data ModelWorld Wide Web Consortium (W3C)
A standard vocabulary for saying which entity was derived from what, by which activity, and when.
- [6]Coalition for Content Provenance and AuthenticityC2PA
The emerging technical standard for signed provenance attached to media.
- [7]When Online Content DisappearsPew Research Center · 2024
Measures how much of the web cited a few years ago is already gone. The argument for retaining the page, not the link.
- [8]Perma.cc and the problem of link rot in citationsHarvard Law School Library Innovation Lab
Why legal and academic citation moved to archived copies rather than live URLs.
- [9]Wayback MachineInternet Archive
The public record of what a page said on a date, when you did not retain it yourself.
- [10]Measuring Attribution in Natural Language Generation ModelsarXiv · 2021
Defines Attributable to Identified Sources (AIS): whether a statement is actually supported by the document cited for it.
- [11]Evaluating Verifiability in Generative Search EnginesarXiv · 2023
Human evaluation finding that a large share of citations in generative search output do not support the sentence they are attached to.
- [12]FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationarXiv · 2023
Scores generated text one atomic claim at a time, which is the only granularity at which 'is this true' is answerable.
- [13]Survey of Hallucination in Natural Language GenerationarXiv · 2022
Taxonomy of how generated text departs from its source, and of the metrics that try to catch it.
- [14]AI Risk Management Framework (AI RMF 1.0)U.S. National Institute of Standards and Technology
The vocabulary regulators and enterprise risk committees are converging on: valid and reliable, accountable and transparent.
- [15]EDGAR company and filing searchU.S. Securities and Exchange Commission
The public search interface. Free, rate-limited, and authoritative for anything a registrant had to disclose.
- [16]Noise: How to Overcome the High, Hidden Cost of Inconsistent Decision MakingHarvard Business Review · 2016
Unwanted variability between judges of the same case, and the structured protocols that reduce it.
Keep reading
Data
Sources of acquisition target data and their limits
A field guide to registries, filings, company websites, and commercial databases: what each source can establish, where coverage is limited, and how to handle conflicting information.
Technology
Evaluating AI research agents
How to assess research agents using structured answers, source attribution, separate review, and comparison with human-reviewed examples.
Judgement
Building consistent acquisition screening criteria
Turn an acquisition thesis into a written screening rubric. Define disqualifying conditions, scoring rules, and the treatment of missing information, then evaluate the results.
This is how Docket works, not just what we think.
Three agents run your criteria across your target list and return a sourced entry on every company, with the page, the sentence and the date behind every field.