How We Test AI Responses with CLEAR

Context. Limits. Expected Answer & Result.

CLEAR is the compact framework behind our Response Quality Tests. The same three elements help define a prompt and then review how well the first answer fulfilled it.

One framework, two jobs

CLEAR is deliberately small. It uses the same three questions while writing a prompt and while reviewing the response. That makes the test criteria visible before a model answers instead of inventing a scoring system afterwards.

C — Context

When prompting: What should the AI know? When reviewing: Did the answer correctly understand and use the relevant context?

L — Limits

When prompting: What rules apply? When reviewing: Did the answer respect the requested rules, constraints and boundaries?

EAR

When prompting: What exactly should the AI return? When reviewing: Did the answer actually deliver the expected answer and result?

Why keep it compact? A rating is useful only when readers can see why it was given. CLEAR keeps the dimensions broad enough to remember, while every score still needs a concrete, visible explanation.

C — Context

Context asks whether the response understood the situation it was given. That can include the user's goal, audience, source material, level of knowledge, facts that must remain unchanged, or other information needed to answer the task correctly.

Example: If a prompt asks for an explanation for a 15-year-old who has never studied statistics, a strong answer should simplify the concept for that audience without replacing the concept with something inaccurate.

A high Context score does not mean every factual claim in the answer has independently been verified. It means the answer correctly used the context supplied for that task.

L — Limits

Limits are the explicit rules that shape the answer. Whenever possible, we make them objectively checkable: word count, number of bullet points, number of paragraphs, required labels, a required closing sentence, or a constraint such as using only free resources.

Measure first

Mechanical checks such as word counts and required item counts are recorded as facts, not guessed from appearance.

Then review

Constraints that require judgment — for example preserving meaning or maintaining a requested tone — are explained in the visible review.

Important: Perfect format compliance does not automatically mean a high-quality result. A response can score 10/10 for Limits and still lose points under Expected Answer & Result if it changes the meaning or misses the real task.

EAR — Expected Answer & Result

EAR asks the most practical question: did the response actually give the user what they needed? This is where we assess completeness, usefulness, preservation of meaning, and whether the result is suitable for the requested purpose.

Good format, weak result

An answer can stay under the word limit and use the right labels but still omit an important dependency, change certainty, or produce something that is not ready for the intended use.

Strong result

The response fulfills the task as a whole: it uses the relevant context, respects the rules, and preserves the meaning and practical outcome the user asked for.

Email example: In our email-rewrite test, a professional email can preserve every required fact and still feel incomplete if it omits a normal sign-off. That is not necessarily a Limits failure when the prompt did not explicitly require one, but it can affect EAR.

How a CLEAR score is produced

Each Response Quality Test follows the same basic sequence. The goal is repeatability and transparency, not a claim that one manual run is a scientific benchmark.

1
Freeze the exact prompt
The same prompt text is used for every assistant in that test.
2
Capture the first complete answer
New chat, no regenerate and no follow-up. The original screenshot is kept as proof.
3
Record the visible conditions
Date, displayed model, account, web/search state and other relevant settings are documented.
4
Check measurable limits
Word counts, bullet counts, labels and other objective constraints are checked directly.
5
Rate C, L and EAR separately
Each score gets a visible reason. We do not hide the rationale behind a total score.

The 1–10 scale

1–3

Poor: the requirement was largely missed.

4–6

Partial: important shortcomings remain.

7–8

Good: the requirement was substantially met, but the issue is meaningful enough to matter.

9

Very good: almost fully met, with only a minor issue.

10

Fully met: no meaningful issue was found for that dimension.

N/A

Used when a dimension genuinely does not apply. We do not turn “not applicable” into an artificial 10/10.

Why no overall winner score? Keeping C, L and EAR separate makes failures easier to see. A perfect format score should not average away a material change in meaning.

A real example: format-perfect, meaning changed

In Test 5, four assistants were asked to summarize a project note for a manager. The source said the website redesign “is planned for launch on 12 September.” ChatGPT summarized that as “Website redesign launches 12 September.”

Why this matters: the date survived, but the status did not. “Planned for launch” is conditional project information; “launches” presents the event as certain. In a management summary, that can materially change decisions, deadlines and expectations based on the project status.

Context 10/10

The response understood that the task was a concise management summary and retained the relevant project topics.

Limits 10/10

It stayed within 75 words, used exactly three bullet points, and used Status:, Risk:, and Next action: as requested.

EAR 5/10

The result is concise and correctly formatted, but it turns a planned launch into a definite launch. In a manager-facing summary, that changes decision-relevant project status and can materially mislead planning.

This is the reason CLEAR separates Limits from EAR: the answer can follow every visible formatting rule and still have a substantive quality problem.

What CLEAR does — and what it does not claim

CLEAR is designed to show

Whether a response understood the task, followed the requested limits, and delivered the expected result — with a visible reason for each score.

CLEAR is not designed to prove

That every factual claim is true, that one assistant is generally “best,” or that a single manual run is a scientific benchmark of a model.

When factual verification matters, sources or claims can be checked separately and documented as evidence. CLEAR itself stays focused on response quality relative to the prompt that was actually tested.

See CLEAR applied to real responses

Open the Response Quality Tests to compare 24 manually captured answers across five practical prompts and one ethical bonus.

Share this CLEAR guide

Add PromptingEasy to your screen

Use your browser menu and choose the option to install this site or add it to your home screen.