5 · Documentation
Evidence and the judge
How a second model reads the answers, how every passage is checked and why there is no verdict without evidence.
No evidence, no verdict. That is the rule that separates augenmerk from a mere marketing claim, and it applies to every mention, every strength, every criticism.
The judge
After every answer a second model reads out what it says about the brands: which brands are named, where in the text, in what order, with what verdict, with what source. We call this model the judge. No step shapes the numbers more.
The judge does not form opinions. It receives fixed rules and must deliver the literal passage for every mention. The rules are shown verbatim in the application under "How we measure", with the version in which they apply.
The judge is a different model from the measured systems, with one exception we left deliberately: it comes from one of the providers we also measure, and so assesses that provider's answers itself. That is acceptable because the judge does not give an opinion but extracts passages, which are then checked character by character against the stored answer. A model cannot favour itself that way; it can only read precisely.
The evidence check
The judge states that every statement appears verbatim in the answer. We do not rely on that. After the call we check every passage character by character against the stored raw text. Whatever is not found there we discard entirely rather than soften. The brand then counts as not mentioned. Better one mention too few than one invented.
A mention without evidence is not stored at all. The answers view openly shows the number of discarded passages, and every highlight in the text is a verified passage. If a stored passage can no longer be found after a rule change, we show that as a missing passage rather than silently moving it.
In the prototype from which augenmerk grew, 2,042 of 2,044 passages were found character by character; not one was invented.
Statements, topics, brand image
For strengths and weaknesses a mention is not enough. Questions of the type "What is particularly positive and negative about X?" provoke a verdict the AI rarely gives on its own: of its own accord it voices almost no criticism. Provoked and observed statements are two different measurements. That is why these questions never count towards presence; they name the brand in the question text.
From every answer to such a question the judge extracts statements: one passage, one topic and one verdict each, praise or criticism. What is stored is the topic, not a sentiment number. "Sentiment 0.73" helps no one; "the AI mentions price, baking time and topping for you" does.
Two things we check for every statement:
- Does it really appear in the answer like that? The same evidence check as for mentions.
- Does the answer give a source for it? A statement with a source can be addressed at the source. A statement without a source comes from model knowledge and changes only with new model versions. A negative claim without a source is suspected of being invented — for the brand, the most valuable finding.
Statements come from three question types: strengths/weaknesses ("What is especially positive and negative about X?"), perception (the fixed question about your own brand) and purchase check ("What are the drawbacks of X?", "Who should not buy X?", "Is X worth the price?"). All of them name the brand and do not count towards presence.
The topics view arranges the statements in a matrix of topics and brands and shows where a brand leads, where it alone is criticised and where it is absent. One judgement-provoking question per brand goes in — for competitors their strengths/weaknesses question, for your own brand the perception question unless a strengths/weaknesses question about it exists; otherwise it would count twice. The brand image summarises, per brand, attributes, objections and the company it is named in, and carries the purchase check: what the AI answers when someone double-checks the brand before buying.
Prominence
Beyond praise and criticism it matters how prominently a brand appears in an answer: as the first mention, in a list, in passing. The judge records the position of every mention; from that come the average place and the order in ranking questions. See Presence and position.
Versions
The sentence before every question, the judge's rules and the rules for topics together form a version of the methodology. A version applies from its filing date and is never overwritten: older runs were judged by older rules, and only that keeps the history comparable. Better rules can be applied retroactively to all stored raw answers because those are kept unchanged; the re-assessment then gets a new version.
What this does not mean
We show that an AI said something, not that it is true. Whether a statement is accurate only the brand itself can judge; for statements with a source, the named source helps.
Updated: 2026-09-16