How we measure
Two lenses. Measured separately. Never blended.
KubeEra reports two independent signals: AI Perception — how frontier models describe a tool — and the Practitioner Verdict — a deterministic rating from verified engineers. This page is the method for both, written to survive a skeptical read. The AI-perception measurement is rolling out; the rules below are how it runs when live.
Lens 1 · AI Perception
How the perception measurement works.
Perception is a measurement of model output, not a judgment of product quality. We say it in the language of visibility — never the language of a verdict.
- A fixed, versioned question set.Buyer-shaped questions a platform engineer actually asks ("best GitOps controller for multi-cluster fleets?"), held constant and version-stamped so results are comparable over time.
- A panel of frontier models, multi-run. Each question is run across several frontier models, multiple times, to average out sampling noise. Model version and run date are logged with every result.
- Perception-native classification. Results map to Highly Visible / Emerging / Overlooked / Contested — never Leader/Laggard. A perception label is not a quality claim, and we refuse to let the two blur.
- Answer-changed vs model-changed, separated.When a score moves we attribute it: did the vendor's footprint change, or did a model ship an update? No false alarms.
- Public attribution stays generic; run logs are retained.Named providers aren't published per-result; raw run logs are kept indefinitely so any claim is replayable — the truth defense.
- Recourse is real. A vendor can contest a characterization with evidence — a lab test, release data, a documented response — shown at equal prominence to the signal it contests.
Lens 2 · Practitioner Verdict
How the verdict is computed.
The verdict is deterministic — one published rule applied identically to every product, computed only from verified practitioner reviews. Money never moves it.
- Verified reviewers only.Reviewers attest to their own production experience, no NDA/EULA breach, and liability for false statements of fact. Vendor employees can't review their own product.
- Seven dimensions. Production Readiness · Operator Experience · Community Health · Security Posture · Lock-in Risk · Enterprise Fit · OSS-vs-Commercial Delta.
- A deterministic rule in the database.The rating is computed by one rule defined once — never reimplemented — so a disputed score's defense is "the same rule, applied to everyone."
- Accumulates honestly. A verdict publishes only at a review threshold per product; until then the state is shown as accumulating, never faked.
- Lab data is a credential, not an input.A "Lab Verified" badge decorates a position reviews earned; it never moves the score. No badge is neutral, never a negative.
Positions money can't move.
The divergence between the two lenses — where the models and the practitioners disagree — is the signal this industry doesn't have anywhere else. Read the manifesto ↗
Perception is measured across a panel of frontier models (identified to customers in the dashboard). Model outputs are measured as-is and may contain inaccuracies; KubeEra measures what models say, not whether it is true. Accuracy probes measure correctness explicitly.