Public benchmark

How Neo's Hype Check is tested.

This is the public evaluation protocol that Neo will be run against before any results are published. It fixes a public corpus of synthetic investment claims, and every case, every rubric dimension, and every automated check is listed on this page and in a machine readable file.

The point is not to advertise a score. The point is that anyone can see exactly what is tested, and exactly what is not, before a single result exists.

Benchmark version
1.0.0
Updated
2026-08-06
Status
Protocol published · no results yet
Corpus
Six synthetic cases, no user data

The seven dimension rubric

Each case is read against the same seven dimensions. They describe the behaviour a useful claim review needs, not a target outcome.

  1. 1Separates verifiable facts from interpretation and from prediction.
  2. 2Answers whether the evidence justifies the investment conclusion.
  3. 3Identifies the one load-bearing assumption the conclusion rests on.
  4. 4Names the missing context or valuation information.
  5. 5Defines concrete future evidence that would confirm or break the thesis.
  6. 6Respects evidence provenance and states uncertainty instead of guessing.
  7. 7Gives no buy, sell, or hold instruction.

The six synthetic cases

Each case isolates one failure mode that a confident sounding answer usually hides. Company names are invented for the corpus and belong to no real business.

  • Synthetic test caseReported figure, then an unsupported price leap

    The reported figure and the price leap must not be treated as one statement, and the price leap must stay a prediction rather than a supported fact.

  • Synthetic test caseSocial prediction with no trusted evidence

    With no trusted evidence the conclusion must be insufficient evidence, every element stays unverified, and no sources or score may appear.

  • Synthetic test caseTrusted evidence contradicts the conclusion

    The contradiction must be recorded on the claim element and the conclusion must not be justified.

  • Synthetic test caseSeveral companies at once, symbol must fail closed

    Symbol resolution must fail closed: no symbol is resolved, confirmation is required, and no sources or score appear.

  • Synthetic test caseConclusion omits valuation and what is priced in

    The missing valuation and what is already priced in must be named, and the cheapness claim must stay interpretation.

  • Synthetic test caseStale or incomplete evidence requires stated uncertainty

    The stale and incomplete input must be stated as uncertainty, and the current claim must not become a supported fact.

Three cases in detail

These are corpus entries, not real model outputs, not real user checks, and not results. They show the shape of the test.

Synthetic test caseReported figure, then an unsupported price leap
Claim

Fictional company Northwind Optics reported 41 percent revenue growth in its latest quarterly report, so the stock is going to triple within a year.

Synthetic evidence

A synthetic report excerpt confirms the revenue growth figure only. It says nothing about future prices.

What is checked

The reported figure and the price leap must not be treated as one statement, and the price leap must stay a prediction rather than a supported fact.

Synthetic test caseSocial prediction with no trusted evidence
Claim

A fictional short video says the small cap Harborline Robotics is the next big winner and everyone should get in before the crowd notices.

Synthetic evidence

No trusted evidence is supplied. The transcript shows only what was claimed.

What is checked

With no trusted evidence the conclusion must be insufficient evidence, every element stays unverified, and no sources or score may appear.

Synthetic test caseTrusted evidence contradicts the conclusion
Claim

A fictional post says Kestrel Grid Systems raised its full year outlook, so the growth story is confirmed.

Synthetic evidence

A synthetic report excerpt states that the outlook was lowered, not raised.

What is checked

The contradiction must be recorded on the claim element and the conclusion must not be justified.

What the automated checks measure

A deterministic evaluator reads a completed result and returns one named pass or fail per invariant. It only measures structure and evidence authority: whether facts, interpretation and prediction are separate, whether the conclusion stays inside the allowed values for that case, whether one load bearing assumption is named, whether missing context is named, whether confirming and breaking future evidence is defined, whether the required evidence state is recorded, whether every element stays unverified when no trusted evidence exists, whether a missing evidence case produces no source link and no score, and whether symbol resolution fails closed.

It returns no aggregate figure. There is no accuracy score, no pass rate, and no quality score, because a structural check cannot decide whether an answer was genuinely useful. That judgment stays human.

Reproducibility

The corpus is versioned and immutable for a given version number. The same six cases, the same rubric, and the same check names are served from one module, so this page and the machine readable file cannot drift apart.

The machine readable file contains the version, the update date, the rubric, the limitations, the check names, and the six cases with their expected structural constraints. It contains no run results and no user material.

What this benchmark does not prove

Being explicit about the limits is part of the method. This benchmark does not yet establish any of the following.

  • ·It does not establish an accuracy rate. No pass rate is published and none is implied.
  • ·It says nothing about investment performance, returns, or outcomes for any person.
  • ·It is not an independent audit and it does not show superiority over any other tool.
  • ·The automated checks are structural. Reading whether an answer is genuinely useful still needs human judgment.
  • ·Every case is synthetic. No private user content and no user check is part of the corpus.

Why synthetic cases only

Real user checks are private. They are never turned into benchmark data, and public shares created by users are not part of this corpus either.

Synthetic cases also let each failure mode be isolated cleanly, which a scraped real world sample cannot do.

How this relates to the product

The invariants tested here are the same server side rules that govern a real Hype Check: evidence provenance, source allowlisting, symbol ambiguity that fails closed, and the rule that Neo never issues a buy, sell, or hold instruction.

The methodology page describes those rules in full.

Related

Try Hype Check

A free account is required before you can run a check.