Answer quality
Correctness, completeness, relevance and whether the permitted sources support each important claim.
Model evaluation & benchmarking
Assess an existing model or compare candidate approaches using an agreed task, representative examples and evidence your team can inspect.
Discuss your question
An answer can sound right while citing the wrong passage, omitting an exception or adding a condition that the source never stated. A single quality score can conceal these differences.
We assess candidate models or existing systems against defined criteria. The task may be internal question answering, document summarisation or extraction. We separate source retrieval, answer quality and operational trade-offs where the available system access permits it.
This suits teams comparing approaches, investigating failure patterns or deciding whether to invest in further development. It is a technical evaluation, not certification or a guarantee of production suitability.
What gets examined
Correctness, completeness, relevance and whether the permitted sources support each important claim.
Missing evidence, conflicting versions, uncertain questions and cases where asking for clarification is the useful outcome.
Latency, usage cost and review effort, measured under stated conditions rather than assumed from a public benchmark.
We agree the criteria with your domain reviewer and examine scoring disagreements before interpreting the results. Candidates should receive comparable inputs and assessment conditions. Any difference is recorded.
A limited evaluation cannot establish performance on every future input. A result depends on its sample, source coverage and method. Security, governance and production reliability need their own appropriate assessment.
Our worked answer-review example shows the source, illustrative answers and explicit judgments. It contains no model performance claim. For the broader method, read how to evaluate LLM answers.
Before we start
Cost depends on sample preparation, candidate count, access to system traces, repeated runs and specialist review. A narrowly defined comparison is the starting point; additional failure analysis is agreed when the evidence makes it useful.
We agree the question, deliverables, exclusions, fee and review dates in writing before work begins. Data access and reviewer availability affect the schedule. If the question changes, we agree the change before extending the work.
System access or recorded outputs, permitted source material, representative questions and a domain reviewer who can explain acceptable answers. When only outputs are available, causal conclusions will be more limited.
The handover includes the agreed report, methods and shareable project artefacts. Ownership, third-party licences and data permissions are settled in the engagement terms. Production implementation, hosting and ongoing support are separate decisions; the findings can go to your own engineers or an implementation partner.
Yes, within the access and permissions agreed. Findings should be based on the technical evidence, not an assumption that the previous supplier chose the wrong approach.
We can investigate failure patterns. Establishing a cause depends on access to sources, retrieval traces, prompts and settings; an output alone may not be enough.
Yes. Define the task and conditions first, then compare usage costs, latency and review effort alongside quality. A cheaper answer is not useful if it fails a critical requirement.
No. The result supports the defined task under the tested conditions. Production readiness requires the relevant engineering, security, operational and governance work.
A considered next step
Tell us the task, the material available and the decision you need to make. We will discuss whether a focused investigation is the right next step.
Discuss your question hi@manaslabs.co.uk