SCREENING NOTE · METRICS AND RETENTION

A faster result is not yet a benchmark

Why AI performance claims need a protocol, not just a before-and-after number.

In one recent screening, a startup presented a dramatic reduction in document-processing time. A workflow described as taking many hours could reportedly be completed in minutes. The improvement was large enough to attract attention. If reproducible, it could change the economics and usefulness of the product. But the investment package did not include the material needed to interpret the result: no dataset profile, test protocol, hardware configuration, model version, accuracy measures, failure rates, or human-review steps. The claim was specific. The evidence was not.

Speed answers only one part of the performance question

A processing time can show how quickly a system produced an output. It does not show whether the output was complete, accurate, consistent, or useful.

For an AI workflow, the practical result usually depends on several variables:

  • the type, length, language, structure, and quality of the input material;
  • the hardware and software environment;
  • the model and product version;
  • the accuracy threshold used to define completion;
  • the amount of human checking and correction required;
  • the frequency and severity of failures;
  • the user action the output is intended to support.

Without these elements, a faster time is an observation about one reported run. It is not yet a reliable description of product performance.

The baseline must represent the same task

A before-and-after comparison is only useful when both sides perform equivalent work.

The manual baseline may include classification, duplicate removal, verification, exception handling, annotation, and preparation of a final usable output. The automated result may measure only initial extraction.

If the two workflows stop at different points, the time comparison can be arithmetically correct and operationally misleading.

The screening therefore moved beyond asking how long each process took:

What exact task was completed in each workflow, to what quality standard, and at what point was the output considered ready for use?

Accuracy belongs inside the benchmark

Speed and accuracy are not separate claims when the product converts unstructured information into structured outputs.

A system can appear faster by extracting fewer fields, accepting more false matches, overlooking difficult inputs, or transferring unresolved work to the user.

The relevant measures depend on the product, but the benchmark may need to include:

  • precision and recall;
  • false-positive and false-negative rates;
  • duplicate-detection accuracy;
  • field-level completeness;
  • correction frequency;
  • failure and abstention rates;
  • performance across easy, median, and difficult inputs.

The fastest run is rarely the most informative result. The distribution of performance across representative inputs is more useful.

Human review can return the time that automation removed

An automated workflow may finish in minutes and still require hours of review.

That review is not necessarily a weakness. In high-consequence workflows, human approval may be essential. The problem arises when the benchmark reports machine-processing time but excludes the verification needed before the output can be trusted.

A complete comparison should show:

  • machine-processing time;
  • review time;
  • correction time;
  • exception-handling time;
  • total time to a usable output.

This turns an impressive technical result into a measure of workflow improvement.

The deployment environment can change the result

Performance depends on where and how the product runs.

A benchmark produced on high-end development hardware may not describe performance on a customer’s actual infrastructure. Local and offline deployment can introduce different constraints around memory, compute, model size, updates, storage, security controls, and concurrent usage.

The hardware envelope is therefore part of the product claim, not a technical footnote.

The evidence request should identify the processor, memory, accelerators, operating environment, model configuration, concurrency, and any external services used during the test.

Reproducibility changes the status of the claim

A benchmark becomes more decision-useful when another qualified person can reproduce it.

That does not always require disclosure of sensitive data or proprietary code. A controlled evidence package can still provide:

  1. a representative and appropriately redacted dataset profile;
  2. a written test protocol and completion criteria;
  3. hardware, software, and model versions;
  4. accuracy and failure measures;
  5. human-review requirements;
  6. results across repeated runs;
  7. logs or an authorized result summary;
  8. a live demonstration using agreed test inputs.

Each element narrows a different uncertainty. Together, they show whether the result is repeatable, transferable to the customer environment, and relevant to the intended workflow.

The benchmark should end with a user outcome

Even a reproducible technical benchmark does not establish commercial value by itself.

The final question is what changed for the user.

Did the system reduce the time to a decision? Did it allow the same team to process more work? Did it improve completeness, reduce errors, or make a previously impractical workflow possible? Was the result important enough for a customer to adopt, budget for, and continue using the product?

This connects technical performance to operating value and, eventually, to a commercial mechanism.

What the screening changed

The reported speed improvement remained worth investigating. The screening did not treat the absence of evidence as proof that the claim was false.

It changed the next step from repeating the headline number to requesting a reproducible benchmark package.

That distinction matters. A pitch deck can establish what a company claims. A benchmark protocol begins to establish what the product can repeatedly do, under which conditions, at what quality level, and with what practical effect.

The faster result was the reason to look closer.

The protocol was what could make it usable in an investment case.


See how DueCap structures an initial Investment Screening Brief

This note is based on an anonymized initial screening. Company, product, sector, geography, organizations, and identifying benchmark details have been removed or generalized. It does not describe an investment opportunity and does not constitute investment, technical, legal, or other professional advice.

Test the workflow before building the function internally.

Start with a limited Screening Desk pilot. We will discuss your deal flow, agree on the appropriate review capacity, and define the outputs and turnaround expectations.