Speed answers only one part of the performance question
A processing time can show how quickly a system produced an output. It does not show whether the output was complete, accurate, consistent, or useful.
For an AI workflow, the practical result usually depends on several variables:
- the type, length, language, structure, and quality of the input material;
- the hardware and software environment;
- the model and product version;
- the accuracy threshold used to define completion;
- the amount of human checking and correction required;
- the frequency and severity of failures;
- the user action the output is intended to support.
Without these elements, a faster time is an observation about one reported run. It is not yet a reliable description of product performance.
The baseline must represent the same task
A before-and-after comparison is only useful when both sides perform equivalent work.
The manual baseline may include classification, duplicate removal, verification, exception handling, annotation, and preparation of a final usable output. The automated result may measure only initial extraction.
If the two workflows stop at different points, the time comparison can be arithmetically correct and operationally misleading.
The screening therefore moved beyond asking how long each process took:
What exact task was completed in each workflow, to what quality standard, and at what point was the output considered ready for use?
Accuracy belongs inside the benchmark
Speed and accuracy are not separate claims when the product converts unstructured information into structured outputs.
A system can appear faster by extracting fewer fields, accepting more false matches, overlooking difficult inputs, or transferring unresolved work to the user.
The relevant measures depend on the product, but the benchmark may need to include:
- precision and recall;
- false-positive and false-negative rates;
- duplicate-detection accuracy;
- field-level completeness;
- correction frequency;
- failure and abstention rates;
- performance across easy, median, and difficult inputs.
The fastest run is rarely the most informative result. The distribution of performance across representative inputs is more useful.
Human review can return the time that automation removed
An automated workflow may finish in minutes and still require hours of review.
That review is not necessarily a weakness. In high-consequence workflows, human approval may be essential. The problem arises when the benchmark reports machine-processing time but excludes the verification needed before the output can be trusted.
A complete comparison should show:
- machine-processing time;
- review time;
- correction time;
- exception-handling time;
- total time to a usable output.
This turns an impressive technical result into a measure of workflow improvement.
The deployment environment can change the result
Performance depends on where and how the product runs.
A benchmark produced on high-end development hardware may not describe performance on a customer’s actual infrastructure. Local and offline deployment can introduce different constraints around memory, compute, model size, updates, storage, security controls, and concurrent usage.
The hardware envelope is therefore part of the product claim, not a technical footnote.
The evidence request should identify the processor, memory, accelerators, operating environment, model configuration, concurrency, and any external services used during the test.
Reproducibility changes the status of the claim
A benchmark becomes more decision-useful when another qualified person can reproduce it.
That does not always require disclosure of sensitive data or proprietary code. A controlled evidence package can still provide:
- a representative and appropriately redacted dataset profile;
- a written test protocol and completion criteria;
- hardware, software, and model versions;
- accuracy and failure measures;
- human-review requirements;
- results across repeated runs;
- logs or an authorized result summary;
- a live demonstration using agreed test inputs.
Each element narrows a different uncertainty. Together, they show whether the result is repeatable, transferable to the customer environment, and relevant to the intended workflow.
The benchmark should end with a user outcome
Even a reproducible technical benchmark does not establish commercial value by itself.
The final question is what changed for the user.
Did the system reduce the time to a decision? Did it allow the same team to process more work? Did it improve completeness, reduce errors, or make a previously impractical workflow possible? Was the result important enough for a customer to adopt, budget for, and continue using the product?
This connects technical performance to operating value and, eventually, to a commercial mechanism.
What the screening changed
The reported speed improvement remained worth investigating. The screening did not treat the absence of evidence as proof that the claim was false.
It changed the next step from repeating the headline number to requesting a reproducible benchmark package.
That distinction matters. A pitch deck can establish what a company claims. A benchmark protocol begins to establish what the product can repeatedly do, under which conditions, at what quality level, and with what practical effect.
The faster result was the reason to look closer.
The protocol was what could make it usable in an investment case.
See how DueCap structures an initial Investment Screening Brief
This note is based on an anonymized initial screening. Company, product, sector, geography, organizations, and identifying benchmark details have been removed or generalized. It does not describe an investment opportunity and does not constitute investment, technical, legal, or other professional advice.

