How scoring works
Every report is scored on up to six signals, computed by checking claims against the real pages a report actually pulled from — never by asking the model to grade its own work. Here's what each one is actually measuring.
Six signals
Whether these items should even be compared in the first place — how confident the model was that they genuinely share a category, combined with how evenly documented they turned out to be. Only shown on a direct comparison between two or more items; a single-subject profile has nothing to compare against.
How current the underlying pages are, based on real dates found in them — not a model's own read of the content. A report built on stale pages can sound confident while being simply out of date.
How much a domain should be trusted for this specific topic — not a fixed reputation score. A hospital's site rates very differently for a medical question than for an unrelated one.
Whether independent evidence was actually found for a claim, checked against the pages a report really fetched — not what the model said it used, and not affected by whether it claimed a source at all.
How many independent domains, on average, back each claim — a stronger signal than Source Matching, since it measures agreement, not just existence. Independent sources can still all repeat the same error, so this is a measure of agreement, not proof of truth.
How many distinct, independently-verified sources actually back the report, relative to how much research it needed — not how many pages were fetched, but how many of those actually resolved into a checked, backed claim. Fetching dozens of pages and only verifying a handful doesn't read as deeper for the fetching; this counts the handful.
Reading the combinations
A single badge tells you one thing. A few combinations of them tell you something a single badge doesn't.
Depth only appears when at least one value was tagged as searched — if it's missing entirely alongside a 0 on the other two, nothing in the report was searched at all. It rests entirely on the model's own background knowledge, not fresh research.
This is a different, more concerning pattern than the one above: the model did claim to search for some values, but none of what it found held up as real, checkable support. A search was attempted and came up empty — not simply skipped.
The individual facts are well-sourced and checked — the caution here is about the pairing itself, not the data. These items may not be a natural fit to compare side by side, or one is far better documented than the other.
Not a guarantee
A grade is a confidence signal, not a guarantee. It tells you how much checking stands behind a claim, not that the claim is beyond question. That's deliberate: the point isn't to replace your judgment, it's to give you enough to exercise it quickly.