Skip to content

Agentic Eyes consensus panels

Any one model has habits — blind spots it always misses, complaints it always invents. A consensus panel removes the single-model variable: the same persona, same site, same mode, run across 2–5 different AI models. Findings are then ranked by how many models independently reported them.

A finding four models independently flagged (“no pricing anywhere”) is about as trustworthy as qualitative feedback gets. A finding exactly one model saw is exploratory — interesting, maybe real, but unverified by peers. The panel output separates these worlds explicitly:

  • Convergent findings — ranked by agreement count, with which models reported each. This is the headline report.
  • Divergence — findings only one model produced, or where models disagreed. Read this section for hypotheses, not conclusions.
  • Panels default to a cross-provider set from the verified model catalog, or you name the models (2–5).
  • Pairwise similarity uses fuzzy matching plus token overlap — models phrase the same complaint differently, and exact-match merging would both hide one model’s finding and inflate another’s agreement count. Area compatibility guards the same way.
  • Runs compose: N models × M trials if you also want within-model stability.
  • Drive mode works with panels — every model drives its own real browser journey.

A panel costs roughly N times a single review, so spend it where the word “confirmed” has value: findings going in front of a customer, driving a roadmap priority, or backing a claim you’ll be held to. For iteration loops — you fixed the header, did Sally’s complaint change? — a single-model rerun is usually enough.

Origin note: the panel design came out of a four-model dogfood where the findings all eight runs agreed on became the report’s spine, and the one-model-only findings — when checked — split evenly between real issues and model folklore. Convergence was the difference.