Research guide · Evaluation
Evaluate the answer, its evidence and its limits.
Fluent text can still be incorrect. A useful evaluation separates the qualities that a polished response tends to blur together.

Define the task before choosing a score.
Evaluating an LLM answer means assessing its behaviour against a defined task and scoring guide. A response can be factually plausible yet unsupported by the allowed sources, or correct in one sentence while omitting an exception that changes its meaning.
Begin with the intended use. Is the answer a draft for expert review, a summary of a supplied document or a response restricted to an approved source collection? The information the model is allowed to use affects what counts as a good answer.
Public benchmarks can inform research, but they assess particular tasks and conditions. For example, Google Research's CURIE publication describes eight scientific understanding and reasoning tasks derived from approximately 250 research papers. A result on that kind of benchmark is evidence about its test design, not proof of performance on your internal documents.
The scoring guide below is a proposed starting method. It is not a claim that Manas Labs has completed an evaluation using it, and its criteria should be adapted to the actual task.
Separate the dimensions of quality.
| Dimension | Reviewer question |
|---|---|
| Correctness | Are the material statements accurate for the task? |
| Evidence support | Do the permitted sources support the important claims? |
| Completeness | Are essential conditions, exceptions or requested elements missing? |
| Relevance | Does the answer address the actual question? |
| Knowing its limits | Does it recognise when the available information is insufficient? |
| Usability | Can the intended reviewer inspect and act on the output? |
A citation is not proof of support. Inspect whether the cited passage actually justifies the claim and whether the answer has silently extended beyond it. Where the system has to find the right passage first, score finding the evidence separately from writing the answer where practical. Finding the wrong passage and misreading the right passage suggest different remedies.
For each dimension, define a small set of scoring categories with examples. A simple scoring guide that reviewers apply consistently can be more useful than a fine-grained scale with unclear distinctions. Preserve material error flags separately from average quality scores.
Build a test set that challenges the approach.
Include representative ordinary cases, then add cases that examine known weaknesses. Useful categories may include ambiguous questions, missing source material, conflicting document versions, long context and questions with an important exception. Keep the reason for including each category visible.
Test what happens when there is no good answer. If every example has a straightforward answer, the evaluation cannot tell you whether the model recognises the limits of its information. Define the expected behaviour when the correct response is a request for clarification or an explanation that the sources are insufficient.
Record where the examples came from, which permissions apply and which intended uses they do not represent. Separate examples used for development from those used for final assessment where possible. If this is not possible, explain the resulting limitation instead of treating the score as fresh evidence.
Calibrate human and automated review
Have reviewers score a shared subset and discuss disagreements. This can expose unclear instructions, multiple acceptable answers or missing domain knowledge. Automated checks are useful for repeatable requirements; model-assisted scoring can help organise review, but it introduces another system whose judgments need scrutiny.
The NIST AI RMF resources include a Generative AI Profile intended to complement the broader framework. That wider risk context is a reminder that output quality is only one part of assessing an AI system.
Report the failures behind the score.
Present results by meaningful category as well as overall. Include the number of tested examples, relevant settings and examples of consequential mistakes. Avoid implying that a small or selectively assembled sample represents every future interaction.
When comparing candidates, keep the task, available context and scoring conditions consistent. State any unavoidable differences. If one approach uses extra source material or a different review process, that is part of the comparison and should be visible.
The final recommendation should explain the trade-offs and the decision the evidence supports. It may identify a preferred candidate, a shared weakness or an unresolved need for better test material. It should not turn an evaluation result into a guarantee of safe operation.
For a scoped comparison, see Model Evaluation & Benchmarking. If you are building an experimental approach first, use proof-of-concept success criteria to establish what the evaluation must resolve.
See a small worked example.
Inspect the source, illustrative answers and review judgments in our answer-evaluation example. The downloadable worksheet shows a possible recording format; it is not a client result or a model benchmark.