One framework, two jobs
CLEAR is deliberately small. It uses the same three questions while writing a prompt and while reviewing the response. That makes the test criteria visible before a model answers instead of inventing a scoring system afterwards.
When prompting: What should the AI know? When reviewing: Did the answer correctly understand and use the relevant context?
When prompting: What rules apply? When reviewing: Did the answer respect the requested rules, constraints and boundaries?
When prompting: What exactly should the AI return? When reviewing: Did the answer actually deliver the expected answer and result?
C — Context
Context asks whether the response understood the situation it was given. That can include the user's goal, audience, source material, level of knowledge, facts that must remain unchanged, or other information needed to answer the task correctly.
A high Context score does not mean every factual claim in the answer has independently been verified. It means the answer correctly used the context supplied for that task.
L — Limits
Limits are the explicit rules that shape the answer. Whenever possible, we make them objectively checkable: word count, number of bullet points, number of paragraphs, required labels, a required closing sentence, or a constraint such as using only free resources.
Mechanical checks such as word counts and required item counts are recorded as facts, not guessed from appearance.
Constraints that require judgment — for example preserving meaning or maintaining a requested tone — are explained in the visible review.
EAR — Expected Answer & Result
EAR asks the most practical question: did the response actually give the user what they needed? This is where we assess completeness, usefulness, preservation of meaning, and whether the result is suitable for the requested purpose.
An answer can stay under the word limit and use the right labels but still omit an important dependency, change certainty, or produce something that is not ready for the intended use.
The response fulfills the task as a whole: it uses the relevant context, respects the rules, and preserves the meaning and practical outcome the user asked for.
How a CLEAR score is produced
Each Response Quality Test follows the same basic sequence. The goal is repeatability and transparency, not a claim that one manual run is a scientific benchmark.
The 1–10 scale
Poor: the requirement was largely missed.
Partial: important shortcomings remain.
Good: the requirement was substantially met, but the issue is meaningful enough to matter.
Very good: almost fully met, with only a minor issue.
Fully met: no meaningful issue was found for that dimension.
Used when a dimension genuinely does not apply. We do not turn “not applicable” into an artificial 10/10.
A real example: format-perfect, meaning changed
In Test 5, four assistants were asked to summarize a project note for a manager. The source said the website redesign “is planned for launch on 12 September.” ChatGPT summarized that as “Website redesign launches 12 September.”
Why this matters: the date survived, but the status did not. “Planned for launch” is conditional project information; “launches” presents the event as certain. In a management summary, that can materially change decisions, deadlines and expectations based on the project status.
The response understood that the task was a concise management summary and retained the relevant project topics.
It stayed within 75 words, used exactly three bullet points, and used Status:, Risk:, and Next action: as requested.
The result is concise and correctly formatted, but it turns a planned launch into a definite launch. In a manager-facing summary, that changes decision-relevant project status and can materially mislead planning.
What CLEAR does — and what it does not claim
Whether a response understood the task, followed the requested limits, and delivered the expected result — with a visible reason for each score.
That every factual claim is true, that one assistant is generally “best,” or that a single manual run is a scientific benchmark of a model.
When factual verification matters, sources or claims can be checked separately and documented as evidence. CLEAR itself stays focused on response quality relative to the prompt that was actually tested.